GitHub issue #3431: [Reborn] Add MemoryPromptContextService production adapter - #3471
Conversation
…t in ironclaw_turns
…extService adapter
There was a problem hiding this comment.
Code Review
This pull request introduces significant architectural updates, including the addition of new crates (ironclaw_conversations, ironclaw_event_projections) and infrastructure for runtime policy planning, memory management, and database-backed state stores. It also adds agent documentation across the workspace and implements new boundary tests. The review feedback highlights several efficiency and scalability opportunities: the sanitize_snippet_text function should be optimized using an iterator-based approach to avoid unnecessary allocations, and the full-state replacement patterns in the libSQL and Postgres storage adapters, as well as the redundant I/O in append_file, should be documented as deferred optimizations with a plan for future incremental improvements.
I am having trouble creating individual review comments. Click here to see my feedback.
crates/ironclaw_host_runtime/src/memory_context.rs (152-182)
The current implementation of sanitize_snippet_text is inefficient for large input strings. It collects all non-control characters into a new String before truncating, which can lead to large unnecessary allocations. To optimize performance, use an iterator-based approach (e.g., chars().by_ref().take(N)) instead of calculating the full length or collecting the entire string beforehand.
References
- To optimize performance when truncating strings, use an iterator-based approach (e.g.,
chars().by_ref().take(N)) instead of calculating the full length withchars().count()beforehand.
crates/ironclaw_conversations/src/libsql.rs (535-539)
The save_state_to_conn implementation deletes all rows and re-inserts the entire state. While this snapshot-adapter pattern is acceptable for initial correctness, it should be documented as a deferred optimization. Please add a TODO or documentation noting that incremental updates (e.g., UPSERT) should be implemented as a future task for better scalability.
References
- If a specific design pattern (e.g., snapshot-adapter) necessitates an inefficient implementation (e.g., loading full state), defer optimization and document the requirement for targeted read paths as a follow-up task.
crates/ironclaw_conversations/src/postgres.rs (532-536)
Similar to the libSQL implementation, save_state_to_txn performs a full state replacement. As this is likely for initial correctness, please document the requirement for incremental persistence as a known future task to avoid performance bottlenecks as data volume grows.
References
- When an implementation is intentionally inefficient for the sake of initial correctness, document the deferred optimization (e.g., incremental updates) as a known future task.
crates/ironclaw_memory/src/filesystem.rs (304-318)
The append_file implementation results in the document being read twice. If this redundant I/O is accepted for initial correctness, please document this deferred optimization as a known future task. Note that further performance optimizations should be gated on profiling.
References
- When an implementation is intentionally inefficient for the sake of initial correctness, document the deferred optimization (e.g., incremental updates) as a known future task.
- Performance optimizations, such as adding a caching layer, should be gated on profiling, especially when existing mechanisms already reduce overhead for common cases.
ReviewSummary: Adds Findings:
Verdict: Request changes — (1) and (2) are the asks; (3)–(5) are nits. Scope-filtering, error-redaction, and deterministic ordering are otherwise solid and well-tested. |
|
Addressed the review feedback in
Validation:
Note: attempted workspace |
zmanian
left a comment
There was a problem hiding this comment.
Re-review: approved
All five findings resolved at ed424062:
- Prompt-injection envelope — RESOLVED.
UNTRUSTED_MEMORY_PREFIX = "Untrusted memory content: "applied at the sanitize boundary, plus a 17-marker instruction-hijack denylist with word-boundary matching viacontains_marker_phraseto avoid false positives. Tests cover both rejection and substring safety. - Total-byte budget — RESOLVED.
MAX_TOTAL_SAFE_SUMMARY_BYTES = 4 * 1024enforced incollect_snippets_with_total_budget; test asserts the cap trips beforemax_snippets. - FNV-1a collision docs — RESOLVED. Explicit comment on
snippet_ref_for_path: "unkeyed and not collision-resistant, so callers must never usesnippet_reffor authorization, tenancy checks, or backend lookup." - NaN scores — RESOLVED.
results.retain(|r| ... && r.score.is_finite())before sort, withtotal_cmpordering. Testnon_finite_scores_are_filtered_before_orderingcovers NaN and Infinity. - Magic-string policy — PARTIAL. Now wrapped in
KnownMemoryContextProfileenum withMEMORY_DISABLED_ALIASESconstant and aTODO(reborn/#3333)pointing at the durable context-policy registry. Acknowledged as transitional with a tracked follow-up — acceptable.
Bonus: new error-leak test asserts connection refused / IPs / ports never appear in safe_summary. Good defense in depth.
Follows the foundation slice (see initial commit). Adds the next layer:
1. Capability- and prompt-port middleware (`ironclaw_hooks::middleware`)
* `HookedLoopCapabilityPort` runs `dispatch_before_capability` before
every invocation, translates the composed decision into the existing
`CapabilityOutcome` vocabulary (Deny / PauseApproval / PauseAuth all
map to `Denied` for now; gate-ref plumbing for real pause semantics
lands in the next slice).
* `HookedLoopPromptPort` runs `dispatch_before_prompt` before bundle
construction. Observe-only for snippets in this slice; actual
snippet injection waits for the shared `prompt_envelope::wrap_untrusted`
helper (#3540 / #3471).
2. Declarative predicate evaluator (`ironclaw_hooks::evaluator`)
* `DenyCapability` and `PauseApproval` predicates: stateless, evaluated
directly against `BeforeCapabilityHookContext`.
* `RateOrValueCap` with `InvocationCount` bound: sliding-window counter
keyed by `(hook_id, capability_name)`, in-memory only. Window
parsing supports `s`/`m`/`h`/`d` units; unparseable windows fail
closed.
* `NumericSum` bound: types implemented but evaluation returns Allow
and emits a warn-level audit. Full argument-extraction story is a
follow-up slice once capability arguments become hook-visible.
* `PredicateEvaluator::evaluate_at(...)` test variant accepts an
explicit `Instant` so sliding-window tests don't depend on
real-clock progress.
3. Manifest -> dispatcher glue (`ironclaw_hooks::installed_hook`)
* `PredicateBackedBeforeCapabilityHook` wraps a `HookPredicateSpec`
plus an `Arc<PredicateEvaluator>` and implements
`RestrictedBeforeCapabilityHook`. The registry installer would
construct one of these per `[[hooks]]` entry whose body is
`HookManifestBody::Predicate`.
* Sink reasons are `&'static str`, so the dynamic predicate `reason`
surfaces in audit (via the evaluator's `EvaluatorDecision`) rather
than the model-visible decision. Closed-vocabulary labels carry
through to the sink.
4. Reborn composition seam (`ironclaw_reborn::loop_driver_host`)
* `RebornLoopDriverHostFactory::with_hook_dispatcher(Arc<HookDispatcher>)`
opt-in builder method. When set, the factory wraps the capability
and prompt ports with the hooked middleware. Default behavior
(no dispatcher) is unchanged from the pre-hooks shape, so existing
callers continue to work.
* Added `ironclaw_hooks` as a dep in `ironclaw_reborn`.
Test plan
=========
* `cargo test -p ironclaw_hooks` — 60 tests pass (59 unit + 1
integration smoke; +13 vs the foundation commit covering middleware,
evaluator, installed_hook).
* `cargo test -p ironclaw_reborn` — 118 tests pass; no regressions
from adding the dep.
* `cargo test -p ironclaw_architecture` — 13 tests pass; the
`ironclaw_turns -> ironclaw_hooks` boundary still holds and the new
`ironclaw_hooks` rule (no host_runtime / dispatcher / secrets /
network / wasm / reborn) is unaffected.
* `cargo clippy -p ironclaw_hooks --all-targets --all-features
-- -D warnings` — clean.
* `cargo clippy -p ironclaw_reborn --all-targets -- -D warnings` —
clean.
* `cargo fmt --all -- --check` — clean.
What still defers
==================
* WASM hook execution path.
* Persistent predicate counter (in-memory only for now).
* Argument-extraction so `NumericSum` predicates evaluate against
capability arguments.
* Gate-ref plumbing so PauseApproval / PauseAuth surface real
`CapabilityOutcome::ApprovalRequired` instead of `Denied`.
* Prompt-snippet injection (waits for shared envelope helper).
* Event-triggered hooks.
* Self-authored hooks (#3567).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(reborn): add ironclaw_hooks framework foundation (#3524)
Foundation slice of the Reborn loop hooks framework per nearai/ironclaw#3524.
Lands the trust primitives, sealed decision types, dispatcher contract, and
extension manifest schema; no Reborn middleware composition yet (next slice
wires HookDispatcher into LoopCapabilityPort / LoopPromptPort).
Design comment on #3524:
https://github.com/nearai/ironclaw/issues/3524#issuecomment-4439890144
What this PR ships
==================
* `crates/ironclaw_hooks/` — new crate
* `identity` — content-addressed `HookId` (blake3 of length-prefixed
extension + local + version fields). Same versioning primitive the rest
of Reborn should converge on for replay safety.
* `trust` — `HookTrustClass` enum (Builtin / Trusted / Installed) with
per-kind default attenuation. Trust class is fixed by source, never
declarable.
* `kinds/` — sealed decision DTOs. `BeforeCapabilityHookDecision`,
`HookPatch`, `ObserverFact` all have `pub` outer struct + `pub(crate)`
inner enum + `pub(crate)` constructors. Same #3460 witness pattern.
* `points/` — typed read-only contexts for each hook point.
* `sink` — split sink traits per trust tier. `PrivilegedGateSink` exposes
`allow()`; `RestrictedGateSink` does not. An Installed-tier hook
literally cannot mint Allow at the type level.
* `ordering` — phase → priority → hook id, stable. Phases gated by trust
(Validation/Authorization Builtin-only).
* `failure_policy` — Timeout/Panic/Malformed/AttenuationViolation
categories. Gate/Mutator fail closed, Observer/Effect fail isolated.
Slot poisoning persisted for the rest of the run on any category.
* `registry` — run-profile-sourced bindings; phase-vs-trust gate enforced
at insert; poisoning surface for the dispatcher.
* `dispatch` — HookDispatcher with deterministic ordering, panic
catch-unwind via futures::FutureExt, per-hook tokio::time::timeout,
short-circuit gate composition (Deny > PauseAuth > PauseApproval >
Allow), Telemetry-phase observers always run.
* `manifest` — serde types for the `[[hooks]]` section of extension
manifests. Predicate vs WASM body; same_tenant scope requires explicit
grant; Validation/Authorization phases rejected at parse time because
manifest hooks are always Installed.
* `predicate` — typed predicate language for declarative Installed hooks
(DenyCapability, PauseApproval, RateOrValueCap). Evaluator lives in
the dispatcher follow-up, not here.
* `crates/ironclaw_architecture/tests/reborn_dependency_boundaries.rs`
* Added `ironclaw_turns` -> `ironclaw_hooks` to the forbidden list.
* New BoundaryRule for `ironclaw_hooks` itself (cannot pull host_runtime,
dispatcher, secrets, network, wasm, etc.).
* `Cargo.toml` workspace member registration.
What this PR deliberately does NOT ship
========================================
* Reborn middleware composition wrapping LoopCapabilityPort / LoopPromptPort
with HookDispatcher. Next slice; ironclaw_reborn changes only.
* WASM hook execution path. Programmatic hooks parse and validate from
manifest; the wasmtime integration lands when the WASM dispatcher seam is
built.
* Predicate evaluation. Predicate types serialize and validate; the
evaluator that turns a `RateOrValueCap` spec into a `Deny` decision is in
the next slice alongside Reborn wiring.
* Event-triggered hooks (Phase 5 of the original roadmap).
* Self-authored hooks. Tracked separately at #3567 with monotonic-restriction
+ unforgeable-channel ratification.
Test plan
=========
* `cargo test -p ironclaw_hooks` — 47 tests (46 unit + 1 integration smoke
for the manifest -> binding -> dispatch pipeline).
* `cargo test -p ironclaw_architecture` — 13 tests; new boundary rule
passes, existing rules unaffected.
* `cargo clippy -p ironclaw_hooks --all-targets -- -D warnings` — clean.
* `cargo fmt -p ironclaw_hooks -- --check` — clean.
* `cargo check --workspace` — clean, no regressions in other crates.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(reborn): wire HookDispatcher into LoopCapabilityPort/LoopPromptPort
Follows the foundation slice (see initial commit). Adds the next layer:
1. Capability- and prompt-port middleware (`ironclaw_hooks::middleware`)
* `HookedLoopCapabilityPort` runs `dispatch_before_capability` before
every invocation, translates the composed decision into the existing
`CapabilityOutcome` vocabulary (Deny / PauseApproval / PauseAuth all
map to `Denied` for now; gate-ref plumbing for real pause semantics
lands in the next slice).
* `HookedLoopPromptPort` runs `dispatch_before_prompt` before bundle
construction. Observe-only for snippets in this slice; actual
snippet injection waits for the shared `prompt_envelope::wrap_untrusted`
helper (#3540 / #3471).
2. Declarative predicate evaluator (`ironclaw_hooks::evaluator`)
* `DenyCapability` and `PauseApproval` predicates: stateless, evaluated
directly against `BeforeCapabilityHookContext`.
* `RateOrValueCap` with `InvocationCount` bound: sliding-window counter
keyed by `(hook_id, capability_name)`, in-memory only. Window
parsing supports `s`/`m`/`h`/`d` units; unparseable windows fail
closed.
* `NumericSum` bound: types implemented but evaluation returns Allow
and emits a warn-level audit. Full argument-extraction story is a
follow-up slice once capability arguments become hook-visible.
* `PredicateEvaluator::evaluate_at(...)` test variant accepts an
explicit `Instant` so sliding-window tests don't depend on
real-clock progress.
3. Manifest -> dispatcher glue (`ironclaw_hooks::installed_hook`)
* `PredicateBackedBeforeCapabilityHook` wraps a `HookPredicateSpec`
plus an `Arc<PredicateEvaluator>` and implements
`RestrictedBeforeCapabilityHook`. The registry installer would
construct one of these per `[[hooks]]` entry whose body is
`HookManifestBody::Predicate`.
* Sink reasons are `&'static str`, so the dynamic predicate `reason`
surfaces in audit (via the evaluator's `EvaluatorDecision`) rather
than the model-visible decision. Closed-vocabulary labels carry
through to the sink.
4. Reborn composition seam (`ironclaw_reborn::loop_driver_host`)
* `RebornLoopDriverHostFactory::with_hook_dispatcher(Arc<HookDispatcher>)`
opt-in builder method. When set, the factory wraps the capability
and prompt ports with the hooked middleware. Default behavior
(no dispatcher) is unchanged from the pre-hooks shape, so existing
callers continue to work.
* Added `ironclaw_hooks` as a dep in `ironclaw_reborn`.
Test plan
=========
* `cargo test -p ironclaw_hooks` — 60 tests pass (59 unit + 1
integration smoke; +13 vs the foundation commit covering middleware,
evaluator, installed_hook).
* `cargo test -p ironclaw_reborn` — 118 tests pass; no regressions
from adding the dep.
* `cargo test -p ironclaw_architecture` — 13 tests pass; the
`ironclaw_turns -> ironclaw_hooks` boundary still holds and the new
`ironclaw_hooks` rule (no host_runtime / dispatcher / secrets /
network / wasm / reborn) is unaffected.
* `cargo clippy -p ironclaw_hooks --all-targets --all-features
-- -D warnings` — clean.
* `cargo clippy -p ironclaw_reborn --all-targets -- -D warnings` —
clean.
* `cargo fmt --all -- --check` — clean.
What still defers
==================
* WASM hook execution path.
* Persistent predicate counter (in-memory only for now).
* Argument-extraction so `NumericSum` predicates evaluate against
capability arguments.
* Gate-ref plumbing so PauseApproval / PauseAuth surface real
`CapabilityOutcome::ApprovalRequired` instead of `Denied`.
* Prompt-snippet injection (waits for shared envelope helper).
* Event-triggered hooks.
* Self-authored hooks (#3567).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(reborn): add HookedLoopModelPort/TranscriptPort/CheckpointPort observer middleware
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(reborn): end-to-end hooks integration through RebornLoopDriverHostFactory
Adds crates/ironclaw_reborn/tests/hooks_integration.rs covering the
factory's HookDispatcher wiring seam end-to-end. Tests drive
host.invoke_capability(...) (not dispatcher.dispatch_before_capability(...)
directly) so a regression in RebornLoopDriverHostFactory's wrapping
composition surfaces here.
Scenarios:
- PredicateBackedBeforeCapabilityHook (DenyCapability NameEquals
"cap.blocked") short-circuits invocation; inner port never called;
outcome is Denied(unknown("hook_denied")).
- A privileged selective hook that allows non-matching capabilities
proves the wrapper does not blanket-deny: cap.allowed reaches the
inner port and completes once.
- Factory built without with_hook_dispatcher() lets cap.blocked through
to the inner port, proving the hook plumbing is genuinely opt-in.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(reborn): add pass() + HookRegistrar + self-authored hooks scaffolding
Three additions to ironclaw_hooks:
B. `pass()` on gate sinks — distinguishes "evaluated, no opinion" from
"returned without minting a decision." A passing hook contributes
nothing to the composed decision; a silent hook is still Malformed
and fails closed. `PredicateBackedBeforeCapabilityHook` now routes
the evaluator's `Allow` decision through `sink.pass()` instead of
the previous `deny("hook_predicate_pass")` workaround.
A. `HookRegistrar` bridge — converts a `Vec<HookManifestEntry>` into
`HookBinding`s + dispatcher impls in one call. Predicate bodies are
wired through `PredicateBackedBeforeCapabilityHook`; WASM bodies
return `HookError::RegistryConstruction` for now. Adds
`HookDispatcher::insert_binding` so the registrar can mutate the
registry through the dispatcher rather than reach inside.
I. Self-authored hooks scaffolding — fourth `HookTrustClass` variant
for hooks the agent authors at runtime. Run-scoped only;
monotonic-restriction sink with no `allow`, no trusted-snippet path,
no effect-class constructor. Closed-vocabulary `SelfAuthoredReason`
enum keeps free-text reasons off the audit seam.
`SelfAuthorshipProvenance` captures authoring run/turn, timestamp,
spec digest, optional user ratification, and a generation-trace
pointer. Durable persistence depends on the unforgeable channel
from #3564 and lands separately.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(reborn): real gate-ref plumbing for hook PauseApproval/PauseAuth decisions
Previously, `GateDecisionInner::PauseApproval` and `PauseAuth` returned by
hooks were degraded to `CapabilityOutcome::Denied` at the middleware
boundary because the hook crate had no way to mint a `LoopGateRef` scoped
to the current run. Hooks that wanted to pause the loop for approval or
auth instead failed the call closed, leaving the host's approval-router
machinery unreachable from hook code.
This change introduces a `HookGateRefFactory` trait in
`ironclaw_hooks::middleware::gate_ref` that mints `LoopGateRef`s for
pause-class decisions. `HookedLoopCapabilityPort` now takes an
`Arc<dyn HookGateRefFactory>`, defaulting to `UuidHookGateRefFactory` (a
locally-unique opaque-id factory suitable for tests and the foundation
slice). Production deployments override via `.with_gate_ref_factory(...)`
with a factory bound to the current `LoopRunContext` and the host's
gate-router.
The translation in `decision_to_outcome` is now async so it can await the
factory. `PauseApproval` maps to `CapabilityOutcome::ApprovalRequired
{ gate_ref, safe_summary }` and `PauseAuth` to `AuthRequired`. If the
factory itself errors, the middleware falls back to `Denied` with a
sanitized `hook_gate_ref_unavailable` reason kind so the loop fails
closed rather than routing through an unresolvable suspension. The
underlying error text is dropped to avoid leaking gate-router state into
model-visible output.
Tests:
- `pause_approval_decision_surfaces_as_approval_required`,
`pause_auth_decision_surfaces_as_auth_required`,
`gate_ref_factory_failure_falls_back_to_denied` in
`middleware::capability_port::tests`.
- `pause_approval_hook_surfaces_as_approval_required_with_real_gate_ref`
in `crates/ironclaw_reborn/tests/hooks_integration.rs`, exercising the
full `RebornLoopDriverHostFactory` composition with the default
`UuidHookGateRefFactory`.
- Gate-ref factory unit tests in `gate_ref::tests`.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(reborn): add NumericSum predicate evaluation with capability argument extraction
Wires the missing argument-extraction story for the predicate evaluator so
`ValueOrRateBound::NumericSum` actually enforces a rolling numeric cap
instead of warn-and-allowing.
- Extend `BeforeCapabilityHookContext` with a sealed `SanitizedArguments`
view. Strings truncate to 256 bytes; objects/arrays cap at 8-deep.
`extract_numeric` supports dotted + bracketed paths (`order.amount`,
`items[0].price`) and returns `Option<rust_decimal::Decimal>`. The inner
representation is sealed so external callers can't bypass bounds.
- Introduce `CapabilityInputResolver` + bundled `NullCapabilityInputResolver`
in `middleware/resolver.rs`. The hooks crate intentionally doesn't know
how to dereference a `CapabilityInputRef` — that knowledge belongs to
the production host. Until a real resolver is wired in (follow-up),
arguments are `Unresolved` and `NumericSum` fails closed.
- `HookedLoopCapabilityPort::new` defaults to the null resolver; new
builder `.with_resolver(Arc<dyn CapabilityInputResolver>)` overrides.
- `PredicateEvaluator` gains a tenant-keyed `value_history` map. The
`NumericSum` arm parses `max` + `window`, extracts the numeric value
from sanitized args, accumulates within the rolling window, and applies
`on_exceeded` when the sum exceeds the cap. Unresolved args, missing
field, non-numeric field, unparseable max, and unparseable window all
fail closed via the configured `OnExceededAction`.
- Add `BeforeCapabilityHookContext::new_unresolved(...)` convenience
ctor; existing test sites switch to it instead of churning every call
site through the 4-arg ctor.
Test count: +14 (8 new SanitizedArguments tests, 6 new NumericSum
evaluator tests, 1 null-resolver test; one old NumericSum-stub-related
gap closed).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(reborn): seal hook registration trust boundary + dispatcher hardening
Addresses blocking findings from the security audit of `ironclaw_hooks`:
- C1 (Blocking, Trust Model): "Installed cannot Allow" was not enforced
at the registration boundary. `BeforeCapabilityHookImpl::Privileged`
was a public variant, so external crates with dispatcher access could
construct an Installed binding paired with a Privileged impl and bypass
the sink trait restriction. Sealed `BeforeCapabilityHookImpl`,
`BeforePromptHookImpl`, and `ObserverHookImpl` to `pub(crate)` and
replaced the single generic `install_before_capability` /
`install_before_prompt` / `install_observer` surface with tier-specific
public installers (`install_builtin_*`, `install_trusted_*`,
`install_installed_*`) that build the binding with the matching trust
class internally. Updated registrar, internal middleware tests, the
hooks foundation pipeline test, and the reborn `hooks_integration`
test to drive the new surface. Added regression tests proving the
trust class is set by the installer and that the seal is type-level.
- C5 (Medium, Slot Poisoning): same-dispatch poisoning was incomplete
because `ordered_bindings` snapshots once at the top of the loop, and
`HookRegistry::insert` accepted duplicate hook IDs. Rejected duplicate
hook IDs (any point) in `HookRegistry::insert` and added a poison
re-check before invoking each hook impl in `dispatch_before_capability`,
`dispatch_before_prompt`, and `dispatch_observer_at`. Added regression
tests for both behaviors.
- C6 (Medium, Manifest / Predicate Validation): `parse_window` could
panic on non-ASCII input because `split_at(len - 1)` requires a char
boundary. Rewrote to compute the unit char's UTF-8 byte length and
slice safely, added a public `validate_window` helper, and wired it
into `HookManifestEntry::validate` for both `InvocationCount` and
`NumericSum` bounds. Added tests for non-ASCII, empty, single-char,
and zero-duration windows.
- C2 (High, Tenant Isolation): partial fix only. The
`PredicateEvaluator`'s sliding-window counter was keyed by
`(hook_id, capability)`, so cross-tenant state could leak. Extended
`HistoryKey` to include `tenant_id` and added a regression test
proving counters partition by tenant. Documented the broader
dispatcher-per-build / per-run-fresh-dispatcher pattern as deferred
follow-up in `crates/ironclaw_hooks/CLAUDE.md`.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(reborn): emit hook telemetry milestones for audit/SSE observers
Wires the hook dispatcher into the host's milestone stream so audit
backends and SSE observers can see hook activity. Previously, hook
dispatch was invisible — denies, pauses, failures, and observer fires
left no trace in the host's observability backend.
Changes:
- `ironclaw_turns`: add `HookDispatched`, `HookDecisionEmitted`, and
`HookFailed` variants to `LoopHostMilestoneKind`, with a closed-
vocabulary `HookDecisionSummary` enum (Allow/Deny/PauseApproval/
PauseAuth/Pass/Patch). Introduce a lightweight `HookMilestoneSink`
trait that emits hook-specific *kinds* without requiring a
`LoopRunContext` (the dispatcher is a process-wide singleton that
cannot own a per-run context), plus a `RunScopedHookMilestoneSink`
adapter that injects run context and forwards to the existing
`LoopHostMilestoneSink`. Also add `InMemoryHookMilestoneSink` for
tests.
- `ironclaw_hooks`: add a `telemetry` module that converts hook-crate
types (`HookId`, `HookTrustClass`, `HookPointSpec`, `FailureCategory`,
`FailureDisposition`, `BeforeCapabilityHookDecision`) into the wire-
shape labels and summaries the milestone sink expects. Hook ids cross
the seam as hex strings because the strongly-typed `HookId` cannot be
imported from `ironclaw_turns` (the architecture test enforces
`ironclaw_turns -> ironclaw_hooks` stays absent).
- `ironclaw_hooks::dispatch`: add an optional `Arc<dyn
HookMilestoneSink>` to `HookDispatcher`, set via
`with_milestone_sink`. Emit `HookDispatched` before each hook runs,
`HookDecisionEmitted` after a decision/pass/patch, and `HookFailed`
on timeout/panic/malformed/missing-impl across all three dispatch
paths (before_capability, before_prompt, observer). Default behavior
(no sink attached) emits nothing — preserves the pre-telemetry
observable surface.
- `ironclaw_reborn`: document on `with_hook_dispatcher` that callers
attach the milestone sink to the dispatcher *before* wrapping it in
`Arc` and installing it into the factory, using a
`RunScopedHookMilestoneSink` to inject run-context. The dispatcher
itself is shared across runs, so attaching a fixed run-context inside
it would be wrong. Update `RuntimeEvent` projection in
`milestone_events.rs` to ignore the new hook kinds (no projection
pathway yet; emitted milestones are consumed by SSE observers
directly).
Tests:
- `ironclaw_hooks::dispatch`: 5 new tests covering milestone emission
for deny decisions, panic failures, prompt-mutator patches, observer
pass-throughs, and the no-sink default.
- `ironclaw_reborn` hooks_integration: end-to-end test wiring a
`RunScopedHookMilestoneSink` onto the dispatcher and asserting hook
activity surfaces in the host's `LoopHostMilestoneSink`.
Total: +6 hook telemetry tests; no existing tests modified.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(reborn): extract shared prompt envelope; inject hook patches into prompt bundle
Adds `ironclaw_prompt_envelope`, a leaf crate that owns the single envelope
primitive used by every model-visible untrusted-content path. `wrap_untrusted`
prefixes content with a closed-vocabulary `<Trusted|Untrusted> <source>
content: ` marker, rejects bodies carrying instruction-hijack phrases
(`ignore previous instructions`, `<|im_start|>`, `<system>`, etc.), and
enforces a 4 KiB byte budget by default.
Migrates `ironclaw_host_runtime::memory_context` to delegate envelope
wrapping, marker rejection, and control-character stripping to the new
crate while keeping the `LoopSafeSummary`-specific 512-byte cap and byte
truncation local. Existing memory_context behavior and tests are preserved.
Wires the same envelope into `ironclaw_hooks`:
* `HookPatch::add_enveloped_snippet` now takes a raw body and wraps it
via `wrap_untrusted(EnvelopeSource::Hook, …)`. `Installed` hooks
produce `Untrusted` envelopes; `Builtin`/`Trusted`/`SelfAuthored`
produce `Trusted` envelopes so downstream readers can distinguish the
two paths through a uniform marker.
* `HookedLoopPromptPort::build_prompt_bundle` is no longer observe-only.
After dispatching `before_prompt`, it envelope-wraps every snippet
patch (passing `Enveloped` through, wrapping `Trusted` with the
envelope helper), enforces the 4 KiB aggregate snippet byte budget
across patches, and appends the wrapped snippets to the prompt
bundle's `messages` as `system`-role `LoopModelMessage` entries
carrying deterministic `msg:hook.<ordinal>.<hash>` content refs
(mirroring the skill-snippet ref convention).
The envelope crate is a leaf with no ironclaw dependencies, satisfying
the boundary contract; the existing `ironclaw_hooks` boundary rule in
`reborn_dependency_boundaries` continues to hold because
`ironclaw_prompt_envelope` is not on its forbidden list.
Test count delta:
* `ironclaw_prompt_envelope`: +13 new tests (crate did not exist).
* `ironclaw_hooks`: 84 → 88 tests (+4 prompt-port behavior tests:
`hook_patch_appended_as_envelope_wrapped_message`,
`total_byte_budget_enforced_across_patches`,
`instruction_hijack_in_patch_rejected`,
`trusted_hook_patch_wrapped_with_trust_marker`).
* `ironclaw_host_runtime` memory_context: unchanged (8 tests still pass).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix: align tenant-counter test with SanitizedArguments-extended context ctor
* docs(reborn): document loader contract; pin HookId hex format
Add a "Loader responsibility" section to ironclaw_hooks/CLAUDE.md
explaining that tier-specific installers prevent minting wrong-tier
impls but cannot enforce origin — that's the loader's job — and
recommending registry loaders type-tag extension hooks as
LoadedHook::Installed at the loader seam.
Add tier_specific_installers_are_documented_as_loader_contract as a
regression guard that touches every public install_*_before_capability
and install_*_before_prompt method so any signature change forces the
loader contract to be re-evaluated.
Document HookId::to_hex's 64-char lowercase hex output as part of the
cross-crate contract consumed by LoopHostMilestoneKind::Hook* in
ironclaw_turns; add hook_id_hex_format_is_stable_64_lowercase_chars in
identity::tests and hook_id_string_serialization_matches_to_hex in
telemetry::tests to pin the format and the seam conversion path.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(reborn): pin hook milestone JSON schema + assert pairing invariants
Add L3 schema-snapshot tests for every hook-related LoopHostMilestoneKind
variant (HookDispatched, HookDecisionEmitted per HookDecisionSummary,
HookFailed per FailureCategory) so downstream consumers can rely on the
JSON wire shape and any accidental field rename, enum-tag rename, or type
change fails loudly.
Add L4 pairing-invariant matrix test in the hook dispatcher that drives
every observable outcome (Allow, Deny, PauseApproval, PauseAuth, Pass,
Panic, Timeout, Malformed, MissingImpl) through a recording milestone
sink and asserts the dispatched-then-terminator pairing shape. Document
the MissingImpl path as the one case that emits a sole HookFailed with
no preceding HookDispatched (the dispatcher discovers the protocol
violation before the hook is actually dispatched).
Add a multi-hook dispatch test that installs three hooks with mixed
outcomes (allow/deny/panic) at the same point and asserts each hook
produces its own paired sequence in the deterministic
(phase, priority, hook_id) order taken from the dispatcher's registry.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(reborn): integration tests for observer middleware through RebornLoopDriverHostFactory
Wire the HookedLoopModelPort / HookedLoopTranscriptPort /
HookedLoopCheckpointPort observer wrappers into
RebornLoopDriverHostFactory::build_text_only_host_with_capabilities,
mirroring the existing HookedLoopCapabilityPort / HookedLoopPromptPort
composition. The wrappers are applied only when a HookDispatcher is set
on the factory, so the default factory shape is unchanged.
Add four integration scenarios in crates/ironclaw_reborn/tests/hooks_integration.rs:
- observer_hook_fires_after_model_through_factory
- observer_hook_fires_after_capability_through_factory
- observer_hook_fires_after_checkpoint_through_factory
- observer_panic_does_not_fail_model_call (panic-isolation regression)
Relax the test-fixture model gateway from "panic if invoked" to
returning a stub assistant reply so the AfterModel / panic-isolation
tests can drive stream_model through the wrapped port. The existing
capability-port tests never touch the gateway, so their behavior is
unchanged.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor(reborn): introduce HookDispatcherBuilder for type-enforced sink wiring
Adds a `HookDispatcherBuilder` in `ironclaw_hooks::dispatch` that owns
the dispatcher construction lifecycle: registry -> optional timeout ->
optional milestone sink -> installed hooks -> `.build_arc()`. The
terminal `.build_arc()` wraps in `Arc` and yields an immutable handle.
Tightens the public surface on `HookDispatcher`: `new`, `with_timeout`,
`with_milestone_sink`, and every `install_*_*` method are now
`pub(crate)`. Outside callers route exclusively through the builder, so
"wire the milestone sink before Arc-wrapping" is a compile-time fact
rather than a documentation convention.
`HookRegistrar::install` now takes a `HookDispatcherBuilder` by value
and returns `(HookDispatcherBuilder, Vec<HookId>)`, keeping the builder
chainable through manifest installation.
`RebornLoopDriverHostFactory` gains `with_hook_dispatcher_builder` to
let callers defer `.build_arc()` to the factory — a step toward the
FU8 per-build dispatcher pattern.
Migrates `foundation_pipeline.rs` and `hooks_integration.rs` to the
builder. Internal middleware and dispatch tests continue to use the
crate-private `HookDispatcher::new` directly.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(reborn): production CapabilityInputResolver for NumericSum predicates
Adds HookCapabilityInputResolverAdapter in ironclaw_reborn that bridges
the existing LoopCapabilityInputResolver (already used by
HostRuntimeLoopCapabilityPort for dispatch input resolution) to the
hooks crate's CapabilityInputResolver trait. RebornLoopDriverHostFactory
gains with_capability_input_resolver(...), and when both a hook
dispatcher and resolver are configured the factory threads the adapter
into HookedLoopCapabilityPort::with_resolver — so NumericSum and other
argument-dependent predicates evaluate against real, sanitized inputs
instead of failing closed against the framework's null default.
The adapter also enforces a configurable serialized-byte budget
(default 64 KiB) as defense in depth ahead of the hooks crate's
per-string and depth caps in SanitizedArguments.
Unit tests cover the four adapter branches (resolved JSON,
inner-error → None, non-object pass-through, oversized → None) and a
new end-to-end integration test
(numeric_sum_predicate_caps_total_value_against_real_inputs) drives the
full factory wiring: with a NumericSum cap of 99 over an "amount" field,
two invocations carrying {"amount":"50"} let the first pass through and
deny the second at the hook seam, with the inner port reached exactly
once.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(reborn): per-build HookDispatcher for full per-run isolation (C2)
Introduce `with_hook_dispatcher_factory(F)` on
`RebornLoopDriverHostFactory`. The closure is invoked once per
`build_text_only_host*` call, so dispatcher-owned mutable state — slot
poisoning, registry mutations, predicate-counter siblings — is scoped to
a single host build instead of shared across every host the factory
produces.
The legacy `with_hook_dispatcher(Arc<HookDispatcher>)` adapter is kept as
a thin wrapper that returns clones of the same `Arc` on every build. Its
shared-state behavior is now documented as an explicit opt-in for
backward compat; new wiring should prefer the factory closure.
Adds two regression tests:
- `per_build_dispatcher_state_does_not_leak_across_runs` — installs a
panicking hook, builds two hosts back-to-back, and proves the inner
port is never reached on build 2 (fresh slot still applies the
fail-closed deny). Pins per-run isolation.
- `legacy_with_hook_dispatcher_shares_state_across_builds` — pins the
shared-state semantic of the legacy adapter as the explicit baseline.
Migrates `predicate_deny_hook_short_circuits_inner_port` to the new
factory-closure path so the new wiring is exercised by the existing
suite.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(reborn): project hook telemetry milestones into RuntimeEvent for durable audit
Extend the runtime event substrate with `HookDispatched`, `HookDecisionEmitted`,
and `HookFailed` kinds carrying closed-vocabulary labels and the blake3-hex hook
identity. Project the matching `LoopHostMilestoneKind::Hook*` variants in
`DurableLoopHostMilestoneSink` so hook telemetry now lands in the same durable
event log as model/reply/loop milestones — SSE observers still see live hook
events, and audit replay can reconstruct the full hook trail.
- `ironclaw_events`: add hook variants to `RuntimeEventKind`, optional hook
fields on `RuntimeEvent` (`hook_id`, `hook_point`, `hook_trust_class`,
`hook_decision`, `hook_failure_category`, `hook_failure_disposition`),
typed constructors (`hook_dispatched`, `hook_decision_emitted`,
`hook_failed`), and dedicated sanitizers (`sanitize_hook_label`,
`sanitize_hook_id`) re-run on every wire crossing. No new crate dependency
edges; hook strings cross the boundary opaque.
- `ironclaw_reborn::milestone_events`: project the three hook milestone kinds
via a new `loop.hook` capability id. `HookDecisionSummary` is collapsed to
its closed-vocabulary `kind_name()` so sanitized reasons never enter the
durable substrate.
- `ironclaw_event_projections`: extend `TimelineEntryKind` and the
`RuntimeEventKind -> RunProjectionStatus` mapping so hook events are pure
telemetry — they preserve the current run status rather than changing it.
- Tests: 4 unit tests in `ironclaw_events::runtime_event::tests` (serde
round-trip per variant + unsafe-label collapse), 3 in
`ironclaw_reborn::milestone_events::tests` (projection per variant,
including the assertion that raw `Deny { reason }` text does not reach the
durable wire payload). Existing replay-projection direct-construction
tests updated for the new RuntimeEvent fields.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(reborn): enforce manifest-declared hook scope at dispatch time (C3)
Audit finding C3: extensions could declare `[[hooks]]` with
`scope = "own_capabilities"` in their manifest, but the dispatcher never
enforced it — an Installed hook from ext-A could fire against capabilities
provided by ext-B. Scope was parsed but not load-bearing.
This change makes scope load-bearing end-to-end:
- `BeforeCapabilityHookContext` carries an optional `provider:
ironclaw_host_api::ExtensionId` populated by the middleware. The hook
context is `#[non_exhaustive]` already so this is non-breaking.
- `HookBinding` gains `owning_extension: Option<ExtensionId>` and `scope:
HookBindingScope`. `HookBindingScope` is `Global` / `OwnCapabilities`
/ `SameTenant`. Builtin and Trusted bindings default to `Global` and
carry no `owning_extension`; Installed bindings carry both, sourced
from the manifest.
- `HookDispatcher::install_installed_*` installers now require the
caller to pass `(owning_extension, scope)`. The registrar derives both
from the manifest entry, so manifest authorship is the single source
of truth.
- A new `CapabilityProviderResolver` trait + bundled
`NullCapabilityProviderResolver` lets the middleware lift the
capability id to its provider at invocation time. The middleware
wires the resolved provider into the hook context.
- `dispatch_before_capability` consults `binding.scope.permits(...)`
before invoking each hook. Bindings that don't permit the current
invocation are inert — no sink call, no failure record, no poisoning.
Conservative defaults:
- When the provider resolver returns `None` (no resolver wired, or the
capability has no known provider), `OwnCapabilities`-scoped hooks do
NOT fire. An attacker cannot bypass scope filtering by stripping
provider info from the descriptor.
Tests:
- 5 new dispatcher tests cover OwnCapabilities matching, foreign
provider, unresolved provider, SameTenant, and Builtin Global.
- 1 new registrar test asserts manifest scope and extension propagate
into `HookBinding`.
- 1 new middleware test asserts the provider resolver populates the
hook context.
- 1 new integration test in `ironclaw_reborn` proves an ext-A hook
scoped to `OwnCapabilities` does not intercept invocations that have
no resolved provider (the production composition default).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* style: rustfmt dispatch.rs after FU1 merge
* docs(hooks): prior-art comparison against LSM/eBPF/Envoy/K8s/OPA/CRX/VSC/Tauri
Validates the IronClaw hooks design against 8 established hook/policy
systems across 8 axes (dispatch, trust tiers, attenuation, decision
vocabulary, failure semantics, isolation, manifest, audit).
Surfaces:
- 7 areas where ICLAW stands out vs prior art (type-level trust
enforcement, dispatch-time scope, failure-kind matrix, pause-with-
gate-ref, pairing-invariant audit matrix, tenant-keyed predicates,
phase-ordered dispatch)
- 4 conventional choices we should revisit (in-process Installed-WASM,
sticky poison, no formal dispatch model, no installation rate-limit)
- 3 divergences whose 'why' is weak and need design review
* docs(hooks): STRIDE threat model for v1 framework
Enumerates 7 adversary classes (A1-A7), 6 assets ranked by blast
radius, and ~35 attack vectors across STRIDE categories with mitigations,
existing tests, and residual risk.
Surfaces 7 prioritized follow-ups:
- High: per-extension hook-count cap (D3/D4)
- High: gate-ref unguessability + one-shot test (S1)
- Med: resolver field-level scope (I2)
- Med: per-evaluator state ceiling (D5)
- Med: poison-stickiness operator runbook
- Low: timing side-channel residual acknowledgement (I4)
- Low: instruction-marker denylist periodic review (I5)
Confirms the load-bearing 'Installed cannot Allow' (E1) property holds
via type-level seal + tier-specific installers, backed by
compile_time_seal_test and installed_binding_cannot_be_paired_with_
privileged_impl tests.
Explicit out-of-scope: extension install pipeline (#3492), WASM exec
sandbox (needs separate threat model when it lands), approval gateway
(#3564).
* feat(hooks): close threat-model gaps S1 (gate-ref entropy) and D3/D4 (registration flood)
S1 (gate-ref unguessability, factory side):
- Three new tests on `UuidHookGateRefFactory`:
- `gate_refs_are_v4_uuids` pins the v4 entropy source (122 random
bits per ref per RFC 4122 §4.4); fails if a future change moves to
a counter or weaker UUID version.
- `gate_refs_have_no_collisions_across_many_calls` mints 20k refs
across both namespaces, asserts zero collisions (statistical
proxy for entropy quality).
- `approval_and_auth_namespaces_do_not_overlap` confirms prefix
routing separation.
- Doc comment now documents the security property explicitly and
delineates factory-side vs gateway-side responsibilities for the
one-shot consumption property.
D3/D4 (hook registration flood):
- New `MAX_HOOKS_PER_EXTENSION = 32` and
`MAX_HOOKS_PER_EXTENSION_PER_KIND = 8` consts in `registrar.rs`.
- New `HookRegistrar::enforce_registration_caps` runs pre-flight at
the top of `install()`, before any binding is inserted. Whole-batch
rejection means a partially-installed batch cannot slip past.
- Three regression tests: total-cap rejection, per-kind-cap rejection,
at-cap acceptance.
- Error messages cite the threat-model finding so operators can map
rejection back to the design rationale.
Threat model updated: S1, D3, D4 marked closed in the cross-cutting
properties matrix and the open-follow-ups list.
* test(hooks): three real hooks built against the public API + ergonomics findings
Builds three representative hooks from outside the crate, mimicking
what an extension or system author would actually write:
1. polymarket-daily-cap — Installed predicate hook, InvocationCount
rate-cap with Deny on excess. Canonical 'rate-limit a capability'
use case for the predicate language.
2. large-stake-approval-gate — Installed predicate hook, NumericSum
over amount_usd field, PauseApproval at $1000/24h. Manifest-shape
+ registrar-install coverage from outside Reborn; end-to-end
dispatch lives in ironclaw_reborn integration tests because
NumericSum needs resolved args (a friction finding documented in
the companion doc).
3. pii-redaction-warning — Trusted Rust hook implementing
PrivilegedBeforePromptHook, injects a trusted instruction snippet
reminding the model to redact PII. Demonstrates the path a system
author takes when the predicate language isn't expressive enough.
API change (F1 fix): SanitizedArguments::unresolved() promoted from
pub(crate) to pub. This is the documented safe default — predicates
that need args must fail closed against it — so exposing the
constructor cannot weaken any trust property. The sanitizing
from_json constructor stays sealed; that's the trust boundary.
Without this fix, external hook authors could not construct a
BeforeCapabilityHookContext with both a known provider AND
unresolved args, which made TDD of their own predicate impossible.
Findings documented in docs/real-hooks-findings.md, ranked by
severity. Big-picture observation: writing the Trusted Rust hook
(F4) was easier than writing the declarative predicate hook (F1 +
F2 + F3) — three of seven findings target predicate-authoring
ergonomics. The declarative path needs the most polish before
third-party extension authors will trust it for non-trivial policy.
Tests: 6 new in real_hooks.rs, all pass.
* feat(hooks): close all remaining threat-model and ergonomics gaps
Closes the Med-priority threat-model gaps (I2, D5, poison runbook)
and all real-hooks ergonomics findings (F2, F3, F5, F6, F7) in a
single pass.
Threat model:
- I2 (resolver field-scope): documented in SanitizedArguments rustdoc.
The narrow public surface (only is_resolved + extract_numeric)
enforces field-scope by construction for the current predicate
path. Reassess when Installed-WASM lands.
- D5 (evaluator state ceiling): MAX_HISTORY_KEYS = 8192 per map,
LRU eviction with evictions_observed() metric for operator
monitoring. New regression test
lru_eviction_increments_counter_and_drops_oldest_key.
- Poison-stickiness runbook: new docs/operator-runbook.md with
recovery options ranked by cost.
Ergonomics findings:
- F2 (closed-vocab deny reasons): rustdoc on OnExceededAction
and GateDecisionView::Deny explaining the audit-vs-model split
and why manifest reason text doesn't reach the model.
- F3 (NumericSum can't be TDD'd outside Reborn): new test-support
feature flag with SanitizedArguments::for_tests(value) that
external hook authors can opt into via dev-dep.
- F5 (two ExtensionId types): added
From<&ironclaw_host_api::ExtensionId> impl for
identity::ExtensionId, plus cross-link rustdoc.
- F6 (HookManifestEntry struct-literal fragility): added
#[non_exhaustive] + HookManifestEntry::new(id, kind, body) +
with_scope/with_phase/with_priority/with_description/with_requires_grant
builder methods. Migrated 3 external call sites in tests/.
- F7 (priority guidance): rustdoc on HookPriority with when-to-
deviate guidance, named FIRST/LAST constants documented for
Builtin/Telemetry use cases.
Tests: 151 unit + 1 + 6 integration in ironclaw_hooks all pass
with --all-features. ironclaw_reborn (13 hooks_integration scenarios)
unchanged.
Threat model updated: I2 / D5 / poison runbook marked closed in
both the per-vector table and the cross-cutting properties matrix.
Open follow-ups now down to two Low items (I4 timing side-channel
residual, I5 instruction-marker denylist refresh) plus the deferred
DenyReasonCode enum from F2.
* fix(ci): collapse nested match in hooks_integration test for clippy --all-features
CI runs `cargo clippy --all --tests --examples --all-features -- -D warnings`
which is stricter than the workspace clippy I ran locally and trips
`clippy::collapsible_match` on the nested-if in HookDecisionEmitted
matching. Collapse the inner `if decision.kind_name() == "deny"`
into an arm guard.
* feat(hooks): address henrypark133 review — Critical #1/#2/#3/#5, Concerning #5/#7
Address composition-seam bugs in the Reborn factory wiring + doc tidy.
henrypark133 review findings addressed:
Critical #1 — before_prompt hook messages not materialized.
HookedLoopPromptPort now requires a HookPromptMaterializationSink and
fails closed if patches are emitted without one. The reborn factory
installs an InstructionStoreBackedHookSink adapter that delegates to
the host's InstructionMaterializationStore, so synthetic msg:hook.*
refs are resolvable by the downstream model resolver. New seam trait
(HookPromptMaterializationSink) keeps ironclaw_hooks decoupled from
LoopRunContext.
Critical #2 — OwnCapabilities hooks were inert in production wiring.
Factory now installs SurfaceBackedProviderResolver (consults the
visible-capability surface for capability_id → provider). With this,
ctx.provider is populated and OwnCapabilities-scoped Installed hooks
actually fire against their own provider's capabilities.
Critical #3 — gate refs were unresolvable.
Middleware default switched from UuidHookGateRefFactory to
FailClosedHookGateRefFactory. Tests must explicitly opt into UUID
(via with_gate_ref_factory) to exercise the affirmative ApprovalRequired
path; production deployments must install a router-backed factory.
New factory method RebornLoopDriverHostFactory::with_hook_gate_ref_factory.
Concerning #5 — AfterModel fired twice + before durable finalization.
Removed AfterModel dispatch from HookedLoopModelPort; the transcript
port's finalize_assistant_message is now the sole AfterModel boundary
(the durable one). Model port wrapper is preserved as a no-op shim
for symmetry + future model-response-observed point.
Concerning #7 — doc tidy:
- CLAUDE.md: 3 trust classes → 4 (Builtin/Trusted/Installed/SelfAuthored
with explicit note that SelfAuthored is run-scoped only and not
loadable from an external source).
- operator-runbook.md: "Audit log" → "durable runtime event stream"
where the projection is actually the runtime-event stream, not formal
AuditEnvelope records.
- prior-art.md: poison-lifetime nuance — per-host-build with the
factory pattern, process-lifetime only for the legacy adapter.
- prior-art.md:80: trailing whitespace removed.
Testing gaps from henrypark133 — caller-level tests through
RebornLoopDriverHostFactory:
#1 (before_prompt resolver path):
before_prompt_hook_message_is_resolvable_via_factory_wiring
#2 (OwnCapabilities positive/negative/unknown):
own_capabilities_hook_fires_when_provider_matches
own_capabilities_hook_does_not_fire_when_provider_differs
own_capabilities_hook_does_not_fire_when_provider_unknown
#3 (pause/auth gate lifecycle or fail-closed):
pause_approval_with_default_factory_fails_closed_as_denied
pause_approval_hook_surfaces_as_approval_required_with_real_gate_ref
(updated to require explicit UuidHookGateRefFactory opt-in)
#5 (AfterModel exactly-once at durable boundary):
after_model_fires_exactly_once_at_durable_boundary
Still TODO from review (separate commits):
Critical #4 (telemetry context — two-run attribution) + gap #4
Concerning #6 (TimelineEntry hook metadata projection) + gap #6
Tests: 154 unit + 18 hooks_integration + all other reborn tests pass.
Workspace clippy + fmt + no-panics clean.
* feat(hooks): address remaining henrypark133 review — Critical #4, Concerning #6
Critical #4 — per-run hook telemetry attribution.
New `HookDispatcherBuilderFactory` signature: factory returns a
HookDispatcherBuilder, and `RebornLoopDriverHostFactory` attaches a
`RunScopedHookMilestoneSink` keyed to the CURRENT run's LoopRunContext
inside `build_text_only_host_with_capabilities`, before sealing the
dispatcher. The previous zero-arg signature relied on the closure
capturing run_context — silently misattributed across reuses; new
public API `with_hook_dispatcher_builder_factory` removes that
failure mode entirely. Legacy `with_hook_dispatcher_factory` retained
for back-compat (its sink-wiring contract stays caller-side).
Concerning #6 — TimelineEntry hook metadata.
Added 6 optional fields to `TimelineEntry` (hook_id, hook_point,
hook_trust_class, hook_decision, hook_failure_category,
hook_failure_disposition) and projected them from `RuntimeEvent::Hook*`.
Replay consumers now see which hook fired/failed, not just that some
hook event happened. Each field is closed-vocabulary (no free-form
reason text — that stays in the audit reason payload, not the
product replay DTO).
Testing gaps from henrypark133 — caller-level tests:
#4 (two-run hook telemetry attribution):
hook_telemetry_attribution_is_per_run_not_captured
Builds two hosts from the SAME builder factory closure with two
fresh LoopRunContexts. Asserts each run's hook milestones carry
its OWN run_id (no stale captured one).
#6 (replay projection contract for hook events):
hook_runtime_events_project_with_sanitized_hook_metadata
non_hook_runtime_events_project_with_no_hook_metadata
Constructs RuntimeEvent::Hook{Dispatched,DecisionEmitted,Failed}
and asserts the projection preserves the metadata fields. The
negative test guards against cross-contamination on non-hook
events.
All henrypark133 review items now addressed:
Critical: #1, #2, #3, #4 — done
Concerning: #5, #6, #7 — done
Testing gaps: #1-#6 — done
Tests: 154 unit + 19 hooks_integration in ironclaw_reborn + 61 reborn
unit + 38 + 2 new in ironclaw_event_projections + ... pass.
Workspace clippy + fmt + no-panics clean.
* docs(hooks): scope DenyReasonCode closed-vocabulary enum (successor #6)
Successor PR from #3573 — real-hooks ergonomics finding F2 (deferred).
Adds a curated vocabulary of model-visible denial reasons so hook
authors can communicate why a deny happened without opening a
free-form prompt-injection channel.
* feat(hooks): DenyReasonCode + PauseReasonCode closed-vocabulary enums
Address real-hooks ergonomics finding F2 (deferred from PR #3573). The
prior dispatcher collapsed every Installed-tier deny to the static
label 'hook_predicate_denied', because manifest reason strings are
author-controlled and surfacing them to the model would open a
prompt-injection channel. The cost: the agent couldn't tell *why*
a hook denied.
This PR introduces two closed-vocabulary enums:
- DenyReasonCode: Generic / RateLimit / ValueCap / Blocklist /
RequiresApproval / OutOfPolicy
- PauseReasonCode: Generic / RequiresApproval / OverThreshold /
SensitiveAction
Each variant has an as_label() returning &'static str (so the sink's
&'static str contract is preserved). New OnExceededAction variants
'DenyWithCode { code, reason }' and 'PauseApprovalWithCode { code,
reason }' let manifest authors opt into the richer labels while
keeping reason audit-only.
The legacy Deny { reason } / PauseApproval { reason } variants are
retained for back-compat and map to DenyReasonCode::Generic /
PauseReasonCode::Generic — existing manifests continue to produce
hook_predicate_denied / hook_predicate_pause_requested.
Threat-model regression: a hook author cannot smuggle text into the
model-visible label because the 'code' field is typed as the enum;
there's no String slot exposed model-side. A test
(deny_with_code_only_exposes_enum_variants_to_model) documents this
as a compile-time property.
Tests (+7 new = 161 total):
- deny_reason_code_labels_are_stable: pins the label vocabulary so
rename/relabel is loud.
- pause_reason_code_labels_are_stable: same for PauseReasonCode.
- deny_with_code_round_trips_through_json + pause variant: wire
round-trip + snake_case tag assertion.
- deny_with_code_only_exposes_enum_variants_to_model: compile-time
property check.
- rate_or_value_cap_with_deny_code_routes_to_code_label: end-to-end
affirmative test that the dispatcher emits the code's label.
- rate_or_value_cap_with_pause_code_routes_to_code_label: same for
pause.
Scope doc: crates/ironclaw_hooks/docs/successors/06-deny-reason-code.md
* test(hooks): address codex review on #3636
- Update stale real-hooks-findings.md F2 row to cite this PR's enum
follow-on (was 'deferred').
- Add install_deny_with_code_manifest_surfaces_code_label_on_dispatch:
end-to-end test driving the registrar->dispatcher path for the
new DenyWithCode variant (prior tests covered serde + direct hook
evaluation, but not the manifest install path that downstream
authors actually use).
Codex review on PR #3636: APPROVE with two recommendations; both
addressed.
Tests: 162 unit (+1 new). Clippy/fmt clean.
* fix(hooks): attenuate Installed-tier prompt patches to user role
Installed-tier `before_prompt` patches were injected as role:"system"
messages. Envelope text labels ("[ext-foo says]: ...") do not strip
system-role authority from the model's perspective, so a third-party
extension could inject system-tier instructions through a snippet
patch. This is a prompt-authority escalation against the trust
hierarchy the framework otherwise enforces.
Add `role_for_trust_class()` mapping Installed -> "user" and
Builtin/Trusted/SelfAuthored -> "system". Thread per-patch
trust_class through `wrap_patches_to_messages` and use it for the
emitted `LoopModelMessage.role`.
Tests:
- installed_hook_patch_drops_to_user_role: asserts the role for an
Installed-tier patch is "user"
- trusted_tier_hook_patch_keeps_system_role: regression that Trusted
tier still produces system-role content
* fix(hooks): enforce scope filter on observer dispatch + reject incompatible points
Two related defense-in-depth fixes against silent scope-filter failure:
1. The registry silently accepted Installed bindings with
`HookBindingScope::OwnCapabilities` at points (BeforePrompt,
AfterModel, AfterCheckpoint) whose dispatch context carries no
per-capability provider. The manifest's declared scope had no
effect at all — the hook fired against every dispatch. Reject the
binding at install time so the operator sees the misconfiguration.
2. `dispatch_observer_at` for `AfterCapability` did not consult the
binding's scope, so an Installed observer registered with
`OwnCapabilities` fired against every invocation regardless of
provider. Add `dispatch_observer_at_with_provider` carrying the
resolved capability provider; the capability-port middleware
resolves the provider once per invocation and threads it through
both the BeforeCapability hook context and the AfterCapability
observer dispatch. The dispatcher then enforces
`HookBindingScope::permits` on each observer binding.
`ObserverHookContext` gains a `provider: Option<ExtensionId>` field;
`#[non_exhaustive]` keeps existing authors compiling.
Tests:
- rejects_own_capabilities_at_before_prompt
- rejects_own_capabilities_at_after_model
- accepts_own_capabilities_at_before_capability
- own_capabilities_observer_filters_foreign_providers (covers
foreign / matching / unresolved provider)
* fix(hooks): preserve free-form audit reason alongside closed-vocab model label (serrrfirat #3636)
`PredicateBackedBeforeCapabilityHook::evaluate()` was discarding the
free-form `reason` from `EvaluatorDecision::{Deny, PauseApproval}`
with `..` and only sending `code.as_label()` into the sink. The
`HookDecisionEmitted` milestone therefore carried only the closed-
vocab label, and operator-visible audit/SSE context was silently lost
end-to-end. The fix splits the channels:
- Model sees the closed-vocab label (`hook_rate_limit`,
`hook_pause_over_threshold`, ...) via `sink.deny(label)`. This
channel is unchanged.
- Audit/SSE sees the manifest's free-form `reason` via a new
audit-only sink method `record_audit_reason(reason: String)`. The
recording sink captures it; the dispatcher reads it after the hook
returns and threads it into `LoopHostMilestoneKind::HookDecisionEmitted`.
Surface changes:
- `PrivilegedGateSink` / `RestrictedGateSink` gain
`record_audit_reason(String)` — accepts dynamic `String` (audit-only,
no model-facing seam) unlike the `&'static str` decision reasons.
- `RecordingGateSink` gains an `audit_reason: Option<String>` field.
- `GateHookOutcome::Decision` is now `Decision { decision,
audit_reason }`.
- `HookDispatcher::emit_decision_with_audit` threads the audit reason
into the milestone.
- `LoopHostMilestoneKind::HookDecisionEmitted` gains a
`#[serde(default, skip_serializing_if = "Option::is_none")]`
`audit_reason: Option<String>`. The durable RuntimeEvent projection
intentionally drops this field — audit reasons are operator-facing
in-memory SSE content, never durable cross-process surface.
Tests:
- `deny_with_code_records_audit_reason_separately_from_model_label`:
asserts the recording sink ends with `Deny { reason: "hook_rate_limit" }`
in `state` AND `audit_reason == Some("daily cap of $1000 ...")`.
* fix(hooks): remove unused model_request helper (CI clippy fix)
* fix(hooks): address serrrfirat P1/P2 findings on PR #3573
Three issues from the 5-15 review:
**P1 #1 registrar.rs:70 — `same_tenant` grants not enforced**
`HookManifestEntry::validate` only confirmed `requires_grant` was
present; the registrar then immediately installed the binding with no
host-verified grant context. A manifest could declare
`requires_grant = "anything"` and get a cross-extension binding for
free.
Fix: `HookRegistrar` now carries a `verified_grants: HashSet<String>`
(empty by default — default-deny). Add the host-facing setter
`with_verified_grants(...)`. At `install_one`, if
`entry.requires_grant` is `Some(g)`, require `g ∈ verified_grants` or
reject with a clear error. Tests:
- `install_rejects_same_tenant_without_verified_grant`
- `install_rejects_same_tenant_when_verified_grants_mismatch`
- The existing positive test
`installer_propagates_owning_extension_and_scope_from_manifest` now
wires the verified grant explicitly (proves the API contract).
**P1 #2 prompt_port.rs:150 — zip misalignment**
The materialization loop zipped surviving messages against the
ORIGINAL unfiltered patch list. `wrap_patches_to_messages` skips
metadata patches and over-budget snippets, so the zip silently paired
message[0] with patch[0] even when patch[0] was the skipped metadata
— materializing the wrong content (or none) under the snippet's
synthetic ref.
Fix: `wrap_patches_to_messages` now returns
`Vec<WrappedHookMessage { message, safe_content }>` — surviving
messages paired with their content by construction. The caller
materializes `entry.safe_content` under `entry.message.content_ref`
directly; no zip against unfiltered input. Removed the now-unused
`safe_content_for_patch` helper.
Test:
- `materialization_stays_aligned_when_metadata_patches_are_filtered`:
a hook emits `[metadata, snippet]`; asserts only one model message,
and the materialized content under its ref contains the snippet's
body — proves filtering can no longer desync from materialization.
**P2 #3 loop_driver_host.rs:1343 — `with_hook_dispatcher_builder`**
Docs said it deferred `build_arc()` to let the host factory finalize
wiring; the implementation called `build_arc()` eagerly and routed
through the legacy shared-dispatcher adapter, losing per-run
dispatcher isolation and the run-scoped milestone sink.
Fix: marked `#[deprecated]` with a note pointing callers to
`with_hook_dispatcher_builder_factory(|| ...)` for per-build
isolation, or `with_hook_dispatcher(...)` if they actually meant the
shared adapter. The method body is unchanged so no callers break;
they'll see the deprecation warning. No internal callers exist, so
the deprecation doesn't trip `-D warnings`.
All 162 hooks lib + 19 reborn integration tests pass; clippy clean.
* fix(hooks): address serrrfirat 3573-2026-05-15 review findings
P1 — prompt bundle authority mismatch (prompt_port.rs):
`HookedLoopPromptPort::build_prompt_bundle` called the inner port first,
which caused `HostManagedLoopPromptPort` to issue the prompt-bundle
authority grant against the pre-hook message list. The wrapper then
appended `msg:hook.*` messages to `bundle.messages`, so the downstream
model request hit `grant.messages != messages` and failed closed with
"model request messages do not match the host-built prompt bundle".
Add `with_bundle_authority(authority, run_context)` and re-issue the
grant after appending hook messages so it covers the post-hook bundle.
Reborn wires `prompt_authority.clone()` + `run_context.clone()` into
the wrapper at construction time.
P2 — observer installer accepts non-observer points (dispatch.rs):
`install_observer` accepted any `HookPointSpec` (including
`BeforeCapability` / `BeforePrompt`) and only populated the observer
map. Dispatch later found a binding without a gate/mutator impl and
fail-closed the capability with "binding present without installed
implementation". Reject non-observer points at install time so misuse
fails loudly rather than poisoning bindings at dispatch.
P2 — batch path skipped AfterCapability observers on inner error
(capability_port.rs):
The batch loop used `?` directly on `self.inner.invoke_capability(...)`,
which propagated the error before dispatching `AfterCapability`
observers. Failed batch entries disappeared from telemetry / audit,
while the single-invocation path dispatches observers on error.
Capture the inner result, dispatch observers, then propagate the error.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(hooks): address PR #3573 review feedback round 3
Addresses serrrfirat's CHANGES_REQUESTED review (2026-05-20) by tightening
several install-time / dispatch-time bounds and gating production seams:
- Bound free-form audit reasons crossing telemetry. New
`telemetry::sanitize_audit_reason` strips control characters and caps
length at 512 bytes; `emit_decision_with_audit` routes the manifest-
supplied reason through it before publishing milestones. Manifest
validation also rejects reasons over the same byte limit at install time
so the wire-side cap is a defense-in-depth layer, not the only line.
- Make hot dispatch O(H) instead of O(H^2). The per-binding poison
recheck used to acquire the registry mutex and walk every binding;
`ordered_bindings_with_poison_snapshot` now takes the active bindings
and the poisoned hook-id set under a single lock, and each loop
threads a local `HashSet<HookId>` that absorbs mid-dispatch
poisoning. Removed the redundant `is_poisoned` helper.
- Gate `HookDispatcher::registry_for_test` behind `cfg(any(test,
feature = "test-support"))`. The accessor previously exposed
`&Mutex<HookRegistry>` in production, letting any `Arc<HookDispatcher>`
holder lock and call `HookRegistry::poison` to disable installed
hooks. Added `active_bindings_snapshot(point)` as the read-only
production-safe replacement.
- `#[serde(deny_unknown_fields)]` on every hook-manifest and predicate
DTO (`HookManifestEntry`, `HookManifestBody`, `WasmBudget`,
`HookPredicateSpec`, `CapabilityPredicate`, `ValueOrRateBound`,
`OnExceededAction`). Typoed or unsupported fields (e.g. a
manifest-supplied `trust_class`) now fail loud at install time
instead of being silently dropped.
- Bound predicate trees at install. New
`validate_predicate_tree` enforces `MAX_PREDICATE_DEPTH = 8`,
`MAX_PREDICATE_NODES = 64`, `MAX_PREDICATE_STRING_BYTES = 256`, and
`MAX_MANIFEST_REASON_BYTES = 512`. A hostile registry manifest can no
longer install a deep or huge `All`/`Any` tree that the evaluator
would recursively walk on every match.
- Cap sliding-window samples per key. `MAX_SAMPLES_PER_KEY = 4_096` in
the predicate evaluator. Both the invocation-count and numeric-sum
histories drop the oldest sample once the cap is reached, bounding
memory under attacker-triggered hot capabilities while preserving
rate/value-cap semantics over the most recent window.
- `split_indexer` / `resolve_path` now fail closed on malformed bracket
syntax (`amount[foo]`, `amount[`, trailing garbage). Previously they
silently fell back to the parent field, which could let a typoed
`NumericSum` predicate evaluate against the wrong value and allow
calls the predicate would otherwise have denied.
- Honor `PatchOrdinalHint`. `WrappedHookMessage` carries the source
patch's `ordinal_hint`; `HookedLoopPromptPort` inserts `NearTop`
messages after the bundle's `identity_message_count` and appends
`Last` messages at the end. Safety/policy snippets that need early
placement now get it.
- Update `ironclaw_hooks` top-level docs to reflect the four trust
classes (`Builtin`/`Trusted`/`Installed`/`SelfAuthored`) and the
now-wired Reborn middleware composition.
Tests added:
- `manifest::rejects_unknown_top_level_field`
- `manifest::rejects_unknown_wasm_budget_field`
- `manifest::rejects_predicate_tree_exceeding_max_depth`
- `manifest::rejects_predicate_tree_exceeding_max_nodes`
- `manifest::rejects_predicate_string_exceeding_max_bytes`
- `manifest::rejects_manifest_reason_exceeding_max_bytes`
- `points::capability::malformed_indexer_returns_none_not_parent_value`
- `telemetry::sanitize_audit_reason_*` (truncate / strip control /
preserve / empty)
`cargo fmt`, `cargo clippy --all --benches --tests --examples
--all-features`, and `cargo test -p ironclaw_hooks` all pass clean.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(hooks): batch deferred test coverage from #3573 review (#3914)
* perf(hooks): defer capability input resolution until a predicate needs it (#3913)
* fix(rebase): adapt hooks tests + middleware to upstream API additions
- CapabilityDescriptorView: add parameters_schema field
- LoopModelRequest / LoopPromptBundleRequest: add capability_view field
- TimelineEntry test builder: add hook_id / hook_point / hook_trust_class /
hook_decision / hook_failure_category / hook_failure_disposition fields
- ironclaw_reborn::tests::hooks_integration: switch from
InMemoryLoopCheckpointStore to InMemoryTurnStateStore (which now
impls both LoopCheckpointStore and TurnStateStore), pass TurnActor
in TurnRunState, supply the new turn_state_store factory arg
- ironclaw_reborn lib.rs: drop the pub-use re-exports that upstream
intentionally removed (per the module-directory rationale in the
current ironclaw_reborn lib.rs doc comment); update the
hooks_integration test imports to use module paths
- Cargo.toml: union the hooks-foundation member list with upstream's
new crates (event_streams, auth, first_party_extensions,
reborn_webui_ingress, product_workflow_storage, webui_v2); drop
ironclaw_storage which no longer exists upstream
- crates/ironclaw_architecture/tests/reborn_dependency_boundaries:
keep upstream's removal of ironclaw_filesystem from the ironclaw_turns
forbidden list AND add ironclaw_hooks to that list
- crates/ironclaw_reborn/src/milestone_events.rs: drop dead loop_failure_kind
helper (replaced upstream by loop_failure_kind_name in text_loop_driver.rs);
keep hook_decision_label which is still used
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* perf(hooks): restore batched capability dispatch when hooks acti…
GitHub issue nearai#3431: [Reborn] Add MemoryPromptContextService production adapter
…earai#3573) * feat(reborn): add ironclaw_hooks framework foundation (#3524) Foundation slice of the Reborn loop hooks framework per nearai/ironclaw#3524. Lands the trust primitives, sealed decision types, dispatcher contract, and extension manifest schema; no Reborn middleware composition yet (next slice wires HookDispatcher into LoopCapabilityPort / LoopPromptPort). Design comment on #3524: https://github.com/nearai/ironclaw/issues/3524#issuecomment-4439890144 What this PR ships ================== * `crates/ironclaw_hooks/` — new crate * `identity` — content-addressed `HookId` (blake3 of length-prefixed extension + local + version fields). Same versioning primitive the rest of Reborn should converge on for replay safety. * `trust` — `HookTrustClass` enum (Builtin / Trusted / Installed) with per-kind default attenuation. Trust class is fixed by source, never declarable. * `kinds/` — sealed decision DTOs. `BeforeCapabilityHookDecision`, `HookPatch`, `ObserverFact` all have `pub` outer struct + `pub(crate)` inner enum + `pub(crate)` constructors. Same #3460 witness pattern. * `points/` — typed read-only contexts for each hook point. * `sink` — split sink traits per trust tier. `PrivilegedGateSink` exposes `allow()`; `RestrictedGateSink` does not. An Installed-tier hook literally cannot mint Allow at the type level. * `ordering` — phase → priority → hook id, stable. Phases gated by trust (Validation/Authorization Builtin-only). * `failure_policy` — Timeout/Panic/Malformed/AttenuationViolation categories. Gate/Mutator fail closed, Observer/Effect fail isolated. Slot poisoning persisted for the rest of the run on any category. * `registry` — run-profile-sourced bindings; phase-vs-trust gate enforced at insert; poisoning surface for the dispatcher. * `dispatch` — HookDispatcher with deterministic ordering, panic catch-unwind via futures::FutureExt, per-hook tokio::time::timeout, short-circuit gate composition (Deny > PauseAuth > PauseApproval > Allow), Telemetry-phase observers always run. * `manifest` — serde types for the `[[hooks]]` section of extension manifests. Predicate vs WASM body; same_tenant scope requires explicit grant; Validation/Authorization phases rejected at parse time because manifest hooks are always Installed. * `predicate` — typed predicate language for declarative Installed hooks (DenyCapability, PauseApproval, RateOrValueCap). Evaluator lives in the dispatcher follow-up, not here. * `crates/ironclaw_architecture/tests/reborn_dependency_boundaries.rs` * Added `ironclaw_turns` -> `ironclaw_hooks` to the forbidden list. * New BoundaryRule for `ironclaw_hooks` itself (cannot pull host_runtime, dispatcher, secrets, network, wasm, etc.). * `Cargo.toml` workspace member registration. What this PR deliberately does NOT ship ======================================== * Reborn middleware composition wrapping LoopCapabilityPort / LoopPromptPort with HookDispatcher. Next slice; ironclaw_reborn changes only. * WASM hook execution path. Programmatic hooks parse and validate from manifest; the wasmtime integration lands when the WASM dispatcher seam is built. * Predicate evaluation. Predicate types serialize and validate; the evaluator that turns a `RateOrValueCap` spec into a `Deny` decision is in the next slice alongside Reborn wiring. * Event-triggered hooks (Phase 5 of the original roadmap). * Self-authored hooks. Tracked separately at #3567 with monotonic-restriction + unforgeable-channel ratification. Test plan ========= * `cargo test -p ironclaw_hooks` — 47 tests (46 unit + 1 integration smoke for the manifest -> binding -> dispatch pipeline). * `cargo test -p ironclaw_architecture` — 13 tests; new boundary rule passes, existing rules unaffected. * `cargo clippy -p ironclaw_hooks --all-targets -- -D warnings` — clean. * `cargo fmt -p ironclaw_hooks -- --check` — clean. * `cargo check --workspace` — clean, no regressions in other crates. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(reborn): wire HookDispatcher into LoopCapabilityPort/LoopPromptPort Follows the foundation slice (see initial commit). Adds the next layer: 1. Capability- and prompt-port middleware (`ironclaw_hooks::middleware`) * `HookedLoopCapabilityPort` runs `dispatch_before_capability` before every invocation, translates the composed decision into the existing `CapabilityOutcome` vocabulary (Deny / PauseApproval / PauseAuth all map to `Denied` for now; gate-ref plumbing for real pause semantics lands in the next slice). * `HookedLoopPromptPort` runs `dispatch_before_prompt` before bundle construction. Observe-only for snippets in this slice; actual snippet injection waits for the shared `prompt_envelope::wrap_untrusted` helper (#3540 / #3471). 2. Declarative predicate evaluator (`ironclaw_hooks::evaluator`) * `DenyCapability` and `PauseApproval` predicates: stateless, evaluated directly against `BeforeCapabilityHookContext`. * `RateOrValueCap` with `InvocationCount` bound: sliding-window counter keyed by `(hook_id, capability_name)`, in-memory only. Window parsing supports `s`/`m`/`h`/`d` units; unparseable windows fail closed. * `NumericSum` bound: types implemented but evaluation returns Allow and emits a warn-level audit. Full argument-extraction story is a follow-up slice once capability arguments become hook-visible. * `PredicateEvaluator::evaluate_at(...)` test variant accepts an explicit `Instant` so sliding-window tests don't depend on real-clock progress. 3. Manifest -> dispatcher glue (`ironclaw_hooks::installed_hook`) * `PredicateBackedBeforeCapabilityHook` wraps a `HookPredicateSpec` plus an `Arc<PredicateEvaluator>` and implements `RestrictedBeforeCapabilityHook`. The registry installer would construct one of these per `[[hooks]]` entry whose body is `HookManifestBody::Predicate`. * Sink reasons are `&'static str`, so the dynamic predicate `reason` surfaces in audit (via the evaluator's `EvaluatorDecision`) rather than the model-visible decision. Closed-vocabulary labels carry through to the sink. 4. Reborn composition seam (`ironclaw_reborn::loop_driver_host`) * `RebornLoopDriverHostFactory::with_hook_dispatcher(Arc<HookDispatcher>)` opt-in builder method. When set, the factory wraps the capability and prompt ports with the hooked middleware. Default behavior (no dispatcher) is unchanged from the pre-hooks shape, so existing callers continue to work. * Added `ironclaw_hooks` as a dep in `ironclaw_reborn`. Test plan ========= * `cargo test -p ironclaw_hooks` — 60 tests pass (59 unit + 1 integration smoke; +13 vs the foundation commit covering middleware, evaluator, installed_hook). * `cargo test -p ironclaw_reborn` — 118 tests pass; no regressions from adding the dep. * `cargo test -p ironclaw_architecture` — 13 tests pass; the `ironclaw_turns -> ironclaw_hooks` boundary still holds and the new `ironclaw_hooks` rule (no host_runtime / dispatcher / secrets / network / wasm / reborn) is unaffected. * `cargo clippy -p ironclaw_hooks --all-targets --all-features -- -D warnings` — clean. * `cargo clippy -p ironclaw_reborn --all-targets -- -D warnings` — clean. * `cargo fmt --all -- --check` — clean. What still defers ================== * WASM hook execution path. * Persistent predicate counter (in-memory only for now). * Argument-extraction so `NumericSum` predicates evaluate against capability arguments. * Gate-ref plumbing so PauseApproval / PauseAuth surface real `CapabilityOutcome::ApprovalRequired` instead of `Denied`. * Prompt-snippet injection (waits for shared envelope helper). * Event-triggered hooks. * Self-authored hooks (#3567). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(reborn): add HookedLoopModelPort/TranscriptPort/CheckpointPort observer middleware Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(reborn): end-to-end hooks integration through RebornLoopDriverHostFactory Adds crates/ironclaw_reborn/tests/hooks_integration.rs covering the factory's HookDispatcher wiring seam end-to-end. Tests drive host.invoke_capability(...) (not dispatcher.dispatch_before_capability(...) directly) so a regression in RebornLoopDriverHostFactory's wrapping composition surfaces here. Scenarios: - PredicateBackedBeforeCapabilityHook (DenyCapability NameEquals "cap.blocked") short-circuits invocation; inner port never called; outcome is Denied(unknown("hook_denied")). - A privileged selective hook that allows non-matching capabilities proves the wrapper does not blanket-deny: cap.allowed reaches the inner port and completes once. - Factory built without with_hook_dispatcher() lets cap.blocked through to the inner port, proving the hook plumbing is genuinely opt-in. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(reborn): add pass() + HookRegistrar + self-authored hooks scaffolding Three additions to ironclaw_hooks: B. `pass()` on gate sinks — distinguishes "evaluated, no opinion" from "returned without minting a decision." A passing hook contributes nothing to the composed decision; a silent hook is still Malformed and fails closed. `PredicateBackedBeforeCapabilityHook` now routes the evaluator's `Allow` decision through `sink.pass()` instead of the previous `deny("hook_predicate_pass")` workaround. A. `HookRegistrar` bridge — converts a `Vec<HookManifestEntry>` into `HookBinding`s + dispatcher impls in one call. Predicate bodies are wired through `PredicateBackedBeforeCapabilityHook`; WASM bodies return `HookError::RegistryConstruction` for now. Adds `HookDispatcher::insert_binding` so the registrar can mutate the registry through the dispatcher rather than reach inside. I. Self-authored hooks scaffolding — fourth `HookTrustClass` variant for hooks the agent authors at runtime. Run-scoped only; monotonic-restriction sink with no `allow`, no trusted-snippet path, no effect-class constructor. Closed-vocabulary `SelfAuthoredReason` enum keeps free-text reasons off the audit seam. `SelfAuthorshipProvenance` captures authoring run/turn, timestamp, spec digest, optional user ratification, and a generation-trace pointer. Durable persistence depends on the unforgeable channel from #3564 and lands separately. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(reborn): real gate-ref plumbing for hook PauseApproval/PauseAuth decisions Previously, `GateDecisionInner::PauseApproval` and `PauseAuth` returned by hooks were degraded to `CapabilityOutcome::Denied` at the middleware boundary because the hook crate had no way to mint a `LoopGateRef` scoped to the current run. Hooks that wanted to pause the loop for approval or auth instead failed the call closed, leaving the host's approval-router machinery unreachable from hook code. This change introduces a `HookGateRefFactory` trait in `ironclaw_hooks::middleware::gate_ref` that mints `LoopGateRef`s for pause-class decisions. `HookedLoopCapabilityPort` now takes an `Arc<dyn HookGateRefFactory>`, defaulting to `UuidHookGateRefFactory` (a locally-unique opaque-id factory suitable for tests and the foundation slice). Production deployments override via `.with_gate_ref_factory(...)` with a factory bound to the current `LoopRunContext` and the host's gate-router. The translation in `decision_to_outcome` is now async so it can await the factory. `PauseApproval` maps to `CapabilityOutcome::ApprovalRequired { gate_ref, safe_summary }` and `PauseAuth` to `AuthRequired`. If the factory itself errors, the middleware falls back to `Denied` with a sanitized `hook_gate_ref_unavailable` reason kind so the loop fails closed rather than routing through an unresolvable suspension. The underlying error text is dropped to avoid leaking gate-router state into model-visible output. Tests: - `pause_approval_decision_surfaces_as_approval_required`, `pause_auth_decision_surfaces_as_auth_required`, `gate_ref_factory_failure_falls_back_to_denied` in `middleware::capability_port::tests`. - `pause_approval_hook_surfaces_as_approval_required_with_real_gate_ref` in `crates/ironclaw_reborn/tests/hooks_integration.rs`, exercising the full `RebornLoopDriverHostFactory` composition with the default `UuidHookGateRefFactory`. - Gate-ref factory unit tests in `gate_ref::tests`. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(reborn): add NumericSum predicate evaluation with capability argument extraction Wires the missing argument-extraction story for the predicate evaluator so `ValueOrRateBound::NumericSum` actually enforces a rolling numeric cap instead of warn-and-allowing. - Extend `BeforeCapabilityHookContext` with a sealed `SanitizedArguments` view. Strings truncate to 256 bytes; objects/arrays cap at 8-deep. `extract_numeric` supports dotted + bracketed paths (`order.amount`, `items[0].price`) and returns `Option<rust_decimal::Decimal>`. The inner representation is sealed so external callers can't bypass bounds. - Introduce `CapabilityInputResolver` + bundled `NullCapabilityInputResolver` in `middleware/resolver.rs`. The hooks crate intentionally doesn't know how to dereference a `CapabilityInputRef` — that knowledge belongs to the production host. Until a real resolver is wired in (follow-up), arguments are `Unresolved` and `NumericSum` fails closed. - `HookedLoopCapabilityPort::new` defaults to the null resolver; new builder `.with_resolver(Arc<dyn CapabilityInputResolver>)` overrides. - `PredicateEvaluator` gains a tenant-keyed `value_history` map. The `NumericSum` arm parses `max` + `window`, extracts the numeric value from sanitized args, accumulates within the rolling window, and applies `on_exceeded` when the sum exceeds the cap. Unresolved args, missing field, non-numeric field, unparseable max, and unparseable window all fail closed via the configured `OnExceededAction`. - Add `BeforeCapabilityHookContext::new_unresolved(...)` convenience ctor; existing test sites switch to it instead of churning every call site through the 4-arg ctor. Test count: +14 (8 new SanitizedArguments tests, 6 new NumericSum evaluator tests, 1 null-resolver test; one old NumericSum-stub-related gap closed). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(reborn): seal hook registration trust boundary + dispatcher hardening Addresses blocking findings from the security audit of `ironclaw_hooks`: - C1 (Blocking, Trust Model): "Installed cannot Allow" was not enforced at the registration boundary. `BeforeCapabilityHookImpl::Privileged` was a public variant, so external crates with dispatcher access could construct an Installed binding paired with a Privileged impl and bypass the sink trait restriction. Sealed `BeforeCapabilityHookImpl`, `BeforePromptHookImpl`, and `ObserverHookImpl` to `pub(crate)` and replaced the single generic `install_before_capability` / `install_before_prompt` / `install_observer` surface with tier-specific public installers (`install_builtin_*`, `install_trusted_*`, `install_installed_*`) that build the binding with the matching trust class internally. Updated registrar, internal middleware tests, the hooks foundation pipeline test, and the reborn `hooks_integration` test to drive the new surface. Added regression tests proving the trust class is set by the installer and that the seal is type-level. - C5 (Medium, Slot Poisoning): same-dispatch poisoning was incomplete because `ordered_bindings` snapshots once at the top of the loop, and `HookRegistry::insert` accepted duplicate hook IDs. Rejected duplicate hook IDs (any point) in `HookRegistry::insert` and added a poison re-check before invoking each hook impl in `dispatch_before_capability`, `dispatch_before_prompt`, and `dispatch_observer_at`. Added regression tests for both behaviors. - C6 (Medium, Manifest / Predicate Validation): `parse_window` could panic on non-ASCII input because `split_at(len - 1)` requires a char boundary. Rewrote to compute the unit char's UTF-8 byte length and slice safely, added a public `validate_window` helper, and wired it into `HookManifestEntry::validate` for both `InvocationCount` and `NumericSum` bounds. Added tests for non-ASCII, empty, single-char, and zero-duration windows. - C2 (High, Tenant Isolation): partial fix only. The `PredicateEvaluator`'s sliding-window counter was keyed by `(hook_id, capability)`, so cross-tenant state could leak. Extended `HistoryKey` to include `tenant_id` and added a regression test proving counters partition by tenant. Documented the broader dispatcher-per-build / per-run-fresh-dispatcher pattern as deferred follow-up in `crates/ironclaw_hooks/CLAUDE.md`. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(reborn): emit hook telemetry milestones for audit/SSE observers Wires the hook dispatcher into the host's milestone stream so audit backends and SSE observers can see hook activity. Previously, hook dispatch was invisible — denies, pauses, failures, and observer fires left no trace in the host's observability backend. Changes: - `ironclaw_turns`: add `HookDispatched`, `HookDecisionEmitted`, and `HookFailed` variants to `LoopHostMilestoneKind`, with a closed- vocabulary `HookDecisionSummary` enum (Allow/Deny/PauseApproval/ PauseAuth/Pass/Patch). Introduce a lightweight `HookMilestoneSink` trait that emits hook-specific *kinds* without requiring a `LoopRunContext` (the dispatcher is a process-wide singleton that cannot own a per-run context), plus a `RunScopedHookMilestoneSink` adapter that injects run context and forwards to the existing `LoopHostMilestoneSink`. Also add `InMemoryHookMilestoneSink` for tests. - `ironclaw_hooks`: add a `telemetry` module that converts hook-crate types (`HookId`, `HookTrustClass`, `HookPointSpec`, `FailureCategory`, `FailureDisposition`, `BeforeCapabilityHookDecision`) into the wire- shape labels and summaries the milestone sink expects. Hook ids cross the seam as hex strings because the strongly-typed `HookId` cannot be imported from `ironclaw_turns` (the architecture test enforces `ironclaw_turns -> ironclaw_hooks` stays absent). - `ironclaw_hooks::dispatch`: add an optional `Arc<dyn HookMilestoneSink>` to `HookDispatcher`, set via `with_milestone_sink`. Emit `HookDispatched` before each hook runs, `HookDecisionEmitted` after a decision/pass/patch, and `HookFailed` on timeout/panic/malformed/missing-impl across all three dispatch paths (before_capability, before_prompt, observer). Default behavior (no sink attached) emits nothing — preserves the pre-telemetry observable surface. - `ironclaw_reborn`: document on `with_hook_dispatcher` that callers attach the milestone sink to the dispatcher *before* wrapping it in `Arc` and installing it into the factory, using a `RunScopedHookMilestoneSink` to inject run-context. The dispatcher itself is shared across runs, so attaching a fixed run-context inside it would be wrong. Update `RuntimeEvent` projection in `milestone_events.rs` to ignore the new hook kinds (no projection pathway yet; emitted milestones are consumed by SSE observers directly). Tests: - `ironclaw_hooks::dispatch`: 5 new tests covering milestone emission for deny decisions, panic failures, prompt-mutator patches, observer pass-throughs, and the no-sink default. - `ironclaw_reborn` hooks_integration: end-to-end test wiring a `RunScopedHookMilestoneSink` onto the dispatcher and asserting hook activity surfaces in the host's `LoopHostMilestoneSink`. Total: +6 hook telemetry tests; no existing tests modified. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(reborn): extract shared prompt envelope; inject hook patches into prompt bundle Adds `ironclaw_prompt_envelope`, a leaf crate that owns the single envelope primitive used by every model-visible untrusted-content path. `wrap_untrusted` prefixes content with a closed-vocabulary `<Trusted|Untrusted> <source> content: ` marker, rejects bodies carrying instruction-hijack phrases (`ignore previous instructions`, `<|im_start|>`, `<system>`, etc.), and enforces a 4 KiB byte budget by default. Migrates `ironclaw_host_runtime::memory_context` to delegate envelope wrapping, marker rejection, and control-character stripping to the new crate while keeping the `LoopSafeSummary`-specific 512-byte cap and byte truncation local. Existing memory_context behavior and tests are preserved. Wires the same envelope into `ironclaw_hooks`: * `HookPatch::add_enveloped_snippet` now takes a raw body and wraps it via `wrap_untrusted(EnvelopeSource::Hook, …)`. `Installed` hooks produce `Untrusted` envelopes; `Builtin`/`Trusted`/`SelfAuthored` produce `Trusted` envelopes so downstream readers can distinguish the two paths through a uniform marker. * `HookedLoopPromptPort::build_prompt_bundle` is no longer observe-only. After dispatching `before_prompt`, it envelope-wraps every snippet patch (passing `Enveloped` through, wrapping `Trusted` with the envelope helper), enforces the 4 KiB aggregate snippet byte budget across patches, and appends the wrapped snippets to the prompt bundle's `messages` as `system`-role `LoopModelMessage` entries carrying deterministic `msg:hook.<ordinal>.<hash>` content refs (mirroring the skill-snippet ref convention). The envelope crate is a leaf with no ironclaw dependencies, satisfying the boundary contract; the existing `ironclaw_hooks` boundary rule in `reborn_dependency_boundaries` continues to hold because `ironclaw_prompt_envelope` is not on its forbidden list. Test count delta: * `ironclaw_prompt_envelope`: +13 new tests (crate did not exist). * `ironclaw_hooks`: 84 → 88 tests (+4 prompt-port behavior tests: `hook_patch_appended_as_envelope_wrapped_message`, `total_byte_budget_enforced_across_patches`, `instruction_hijack_in_patch_rejected`, `trusted_hook_patch_wrapped_with_trust_marker`). * `ironclaw_host_runtime` memory_context: unchanged (8 tests still pass). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix: align tenant-counter test with SanitizedArguments-extended context ctor * docs(reborn): document loader contract; pin HookId hex format Add a "Loader responsibility" section to ironclaw_hooks/CLAUDE.md explaining that tier-specific installers prevent minting wrong-tier impls but cannot enforce origin — that's the loader's job — and recommending registry loaders type-tag extension hooks as LoadedHook::Installed at the loader seam. Add tier_specific_installers_are_documented_as_loader_contract as a regression guard that touches every public install_*_before_capability and install_*_before_prompt method so any signature change forces the loader contract to be re-evaluated. Document HookId::to_hex's 64-char lowercase hex output as part of the cross-crate contract consumed by LoopHostMilestoneKind::Hook* in ironclaw_turns; add hook_id_hex_format_is_stable_64_lowercase_chars in identity::tests and hook_id_string_serialization_matches_to_hex in telemetry::tests to pin the format and the seam conversion path. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(reborn): pin hook milestone JSON schema + assert pairing invariants Add L3 schema-snapshot tests for every hook-related LoopHostMilestoneKind variant (HookDispatched, HookDecisionEmitted per HookDecisionSummary, HookFailed per FailureCategory) so downstream consumers can rely on the JSON wire shape and any accidental field rename, enum-tag rename, or type change fails loudly. Add L4 pairing-invariant matrix test in the hook dispatcher that drives every observable outcome (Allow, Deny, PauseApproval, PauseAuth, Pass, Panic, Timeout, Malformed, MissingImpl) through a recording milestone sink and asserts the dispatched-then-terminator pairing shape. Document the MissingImpl path as the one case that emits a sole HookFailed with no preceding HookDispatched (the dispatcher discovers the protocol violation before the hook is actually dispatched). Add a multi-hook dispatch test that installs three hooks with mixed outcomes (allow/deny/panic) at the same point and asserts each hook produces its own paired sequence in the deterministic (phase, priority, hook_id) order taken from the dispatcher's registry. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(reborn): integration tests for observer middleware through RebornLoopDriverHostFactory Wire the HookedLoopModelPort / HookedLoopTranscriptPort / HookedLoopCheckpointPort observer wrappers into RebornLoopDriverHostFactory::build_text_only_host_with_capabilities, mirroring the existing HookedLoopCapabilityPort / HookedLoopPromptPort composition. The wrappers are applied only when a HookDispatcher is set on the factory, so the default factory shape is unchanged. Add four integration scenarios in crates/ironclaw_reborn/tests/hooks_integration.rs: - observer_hook_fires_after_model_through_factory - observer_hook_fires_after_capability_through_factory - observer_hook_fires_after_checkpoint_through_factory - observer_panic_does_not_fail_model_call (panic-isolation regression) Relax the test-fixture model gateway from "panic if invoked" to returning a stub assistant reply so the AfterModel / panic-isolation tests can drive stream_model through the wrapped port. The existing capability-port tests never touch the gateway, so their behavior is unchanged. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(reborn): introduce HookDispatcherBuilder for type-enforced sink wiring Adds a `HookDispatcherBuilder` in `ironclaw_hooks::dispatch` that owns the dispatcher construction lifecycle: registry -> optional timeout -> optional milestone sink -> installed hooks -> `.build_arc()`. The terminal `.build_arc()` wraps in `Arc` and yields an immutable handle. Tightens the public surface on `HookDispatcher`: `new`, `with_timeout`, `with_milestone_sink`, and every `install_*_*` method are now `pub(crate)`. Outside callers route exclusively through the builder, so "wire the milestone sink before Arc-wrapping" is a compile-time fact rather than a documentation convention. `HookRegistrar::install` now takes a `HookDispatcherBuilder` by value and returns `(HookDispatcherBuilder, Vec<HookId>)`, keeping the builder chainable through manifest installation. `RebornLoopDriverHostFactory` gains `with_hook_dispatcher_builder` to let callers defer `.build_arc()` to the factory — a step toward the FU8 per-build dispatcher pattern. Migrates `foundation_pipeline.rs` and `hooks_integration.rs` to the builder. Internal middleware and dispatch tests continue to use the crate-private `HookDispatcher::new` directly. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(reborn): production CapabilityInputResolver for NumericSum predicates Adds HookCapabilityInputResolverAdapter in ironclaw_reborn that bridges the existing LoopCapabilityInputResolver (already used by HostRuntimeLoopCapabilityPort for dispatch input resolution) to the hooks crate's CapabilityInputResolver trait. RebornLoopDriverHostFactory gains with_capability_input_resolver(...), and when both a hook dispatcher and resolver are configured the factory threads the adapter into HookedLoopCapabilityPort::with_resolver — so NumericSum and other argument-dependent predicates evaluate against real, sanitized inputs instead of failing closed against the framework's null default. The adapter also enforces a configurable serialized-byte budget (default 64 KiB) as defense in depth ahead of the hooks crate's per-string and depth caps in SanitizedArguments. Unit tests cover the four adapter branches (resolved JSON, inner-error → None, non-object pass-through, oversized → None) and a new end-to-end integration test (numeric_sum_predicate_caps_total_value_against_real_inputs) drives the full factory wiring: with a NumericSum cap of 99 over an "amount" field, two invocations carrying {"amount":"50"} let the first pass through and deny the second at the hook seam, with the inner port reached exactly once. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(reborn): per-build HookDispatcher for full per-run isolation (C2) Introduce `with_hook_dispatcher_factory(F)` on `RebornLoopDriverHostFactory`. The closure is invoked once per `build_text_only_host*` call, so dispatcher-owned mutable state — slot poisoning, registry mutations, predicate-counter siblings — is scoped to a single host build instead of shared across every host the factory produces. The legacy `with_hook_dispatcher(Arc<HookDispatcher>)` adapter is kept as a thin wrapper that returns clones of the same `Arc` on every build. Its shared-state behavior is now documented as an explicit opt-in for backward compat; new wiring should prefer the factory closure. Adds two regression tests: - `per_build_dispatcher_state_does_not_leak_across_runs` — installs a panicking hook, builds two hosts back-to-back, and proves the inner port is never reached on build 2 (fresh slot still applies the fail-closed deny). Pins per-run isolation. - `legacy_with_hook_dispatcher_shares_state_across_builds` — pins the shared-state semantic of the legacy adapter as the explicit baseline. Migrates `predicate_deny_hook_short_circuits_inner_port` to the new factory-closure path so the new wiring is exercised by the existing suite. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(reborn): project hook telemetry milestones into RuntimeEvent for durable audit Extend the runtime event substrate with `HookDispatched`, `HookDecisionEmitted`, and `HookFailed` kinds carrying closed-vocabulary labels and the blake3-hex hook identity. Project the matching `LoopHostMilestoneKind::Hook*` variants in `DurableLoopHostMilestoneSink` so hook telemetry now lands in the same durable event log as model/reply/loop milestones — SSE observers still see live hook events, and audit replay can reconstruct the full hook trail. - `ironclaw_events`: add hook variants to `RuntimeEventKind`, optional hook fields on `RuntimeEvent` (`hook_id`, `hook_point`, `hook_trust_class`, `hook_decision`, `hook_failure_category`, `hook_failure_disposition`), typed constructors (`hook_dispatched`, `hook_decision_emitted`, `hook_failed`), and dedicated sanitizers (`sanitize_hook_label`, `sanitize_hook_id`) re-run on every wire crossing. No new crate dependency edges; hook strings cross the boundary opaque. - `ironclaw_reborn::milestone_events`: project the three hook milestone kinds via a new `loop.hook` capability id. `HookDecisionSummary` is collapsed to its closed-vocabulary `kind_name()` so sanitized reasons never enter the durable substrate. - `ironclaw_event_projections`: extend `TimelineEntryKind` and the `RuntimeEventKind -> RunProjectionStatus` mapping so hook events are pure telemetry — they preserve the current run status rather than changing it. - Tests: 4 unit tests in `ironclaw_events::runtime_event::tests` (serde round-trip per variant + unsafe-label collapse), 3 in `ironclaw_reborn::milestone_events::tests` (projection per variant, including the assertion that raw `Deny { reason }` text does not reach the durable wire payload). Existing replay-projection direct-construction tests updated for the new RuntimeEvent fields. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(reborn): enforce manifest-declared hook scope at dispatch time (C3) Audit finding C3: extensions could declare `[[hooks]]` with `scope = "own_capabilities"` in their manifest, but the dispatcher never enforced it — an Installed hook from ext-A could fire against capabilities provided by ext-B. Scope was parsed but not load-bearing. This change makes scope load-bearing end-to-end: - `BeforeCapabilityHookContext` carries an optional `provider: ironclaw_host_api::ExtensionId` populated by the middleware. The hook context is `#[non_exhaustive]` already so this is non-breaking. - `HookBinding` gains `owning_extension: Option<ExtensionId>` and `scope: HookBindingScope`. `HookBindingScope` is `Global` / `OwnCapabilities` / `SameTenant`. Builtin and Trusted bindings default to `Global` and carry no `owning_extension`; Installed bindings carry both, sourced from the manifest. - `HookDispatcher::install_installed_*` installers now require the caller to pass `(owning_extension, scope)`. The registrar derives both from the manifest entry, so manifest authorship is the single source of truth. - A new `CapabilityProviderResolver` trait + bundled `NullCapabilityProviderResolver` lets the middleware lift the capability id to its provider at invocation time. The middleware wires the resolved provider into the hook context. - `dispatch_before_capability` consults `binding.scope.permits(...)` before invoking each hook. Bindings that don't permit the current invocation are inert — no sink call, no failure record, no poisoning. Conservative defaults: - When the provider resolver returns `None` (no resolver wired, or the capability has no known provider), `OwnCapabilities`-scoped hooks do NOT fire. An attacker cannot bypass scope filtering by stripping provider info from the descriptor. Tests: - 5 new dispatcher tests cover OwnCapabilities matching, foreign provider, unresolved provider, SameTenant, and Builtin Global. - 1 new registrar test asserts manifest scope and extension propagate into `HookBinding`. - 1 new middleware test asserts the provider resolver populates the hook context. - 1 new integration test in `ironclaw_reborn` proves an ext-A hook scoped to `OwnCapabilities` does not intercept invocations that have no resolved provider (the production composition default). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * style: rustfmt dispatch.rs after FU1 merge * docs(hooks): prior-art comparison against LSM/eBPF/Envoy/K8s/OPA/CRX/VSC/Tauri Validates the IronClaw hooks design against 8 established hook/policy systems across 8 axes (dispatch, trust tiers, attenuation, decision vocabulary, failure semantics, isolation, manifest, audit). Surfaces: - 7 areas where ICLAW stands out vs prior art (type-level trust enforcement, dispatch-time scope, failure-kind matrix, pause-with- gate-ref, pairing-invariant audit matrix, tenant-keyed predicates, phase-ordered dispatch) - 4 conventional choices we should revisit (in-process Installed-WASM, sticky poison, no formal dispatch model, no installation rate-limit) - 3 divergences whose 'why' is weak and need design review * docs(hooks): STRIDE threat model for v1 framework Enumerates 7 adversary classes (A1-A7), 6 assets ranked by blast radius, and ~35 attack vectors across STRIDE categories with mitigations, existing tests, and residual risk. Surfaces 7 prioritized follow-ups: - High: per-extension hook-count cap (D3/D4) - High: gate-ref unguessability + one-shot test (S1) - Med: resolver field-level scope (I2) - Med: per-evaluator state ceiling (D5) - Med: poison-stickiness operator runbook - Low: timing side-channel residual acknowledgement (I4) - Low: instruction-marker denylist periodic review (I5) Confirms the load-bearing 'Installed cannot Allow' (E1) property holds via type-level seal + tier-specific installers, backed by compile_time_seal_test and installed_binding_cannot_be_paired_with_ privileged_impl tests. Explicit out-of-scope: extension install pipeline (#3492), WASM exec sandbox (needs separate threat model when it lands), approval gateway (#3564). * feat(hooks): close threat-model gaps S1 (gate-ref entropy) and D3/D4 (registration flood) S1 (gate-ref unguessability, factory side): - Three new tests on `UuidHookGateRefFactory`: - `gate_refs_are_v4_uuids` pins the v4 entropy source (122 random bits per ref per RFC 4122 §4.4); fails if a future change moves to a counter or weaker UUID version. - `gate_refs_have_no_collisions_across_many_calls` mints 20k refs across both namespaces, asserts zero collisions (statistical proxy for entropy quality). - `approval_and_auth_namespaces_do_not_overlap` confirms prefix routing separation. - Doc comment now documents the security property explicitly and delineates factory-side vs gateway-side responsibilities for the one-shot consumption property. D3/D4 (hook registration flood): - New `MAX_HOOKS_PER_EXTENSION = 32` and `MAX_HOOKS_PER_EXTENSION_PER_KIND = 8` consts in `registrar.rs`. - New `HookRegistrar::enforce_registration_caps` runs pre-flight at the top of `install()`, before any binding is inserted. Whole-batch rejection means a partially-installed batch cannot slip past. - Three regression tests: total-cap rejection, per-kind-cap rejection, at-cap acceptance. - Error messages cite the threat-model finding so operators can map rejection back to the design rationale. Threat model updated: S1, D3, D4 marked closed in the cross-cutting properties matrix and the open-follow-ups list. * test(hooks): three real hooks built against the public API + ergonomics findings Builds three representative hooks from outside the crate, mimicking what an extension or system author would actually write: 1. polymarket-daily-cap — Installed predicate hook, InvocationCount rate-cap with Deny on excess. Canonical 'rate-limit a capability' use case for the predicate language. 2. large-stake-approval-gate — Installed predicate hook, NumericSum over amount_usd field, PauseApproval at $1000/24h. Manifest-shape + registrar-install coverage from outside Reborn; end-to-end dispatch lives in ironclaw_reborn integration tests because NumericSum needs resolved args (a friction finding documented in the companion doc). 3. pii-redaction-warning — Trusted Rust hook implementing PrivilegedBeforePromptHook, injects a trusted instruction snippet reminding the model to redact PII. Demonstrates the path a system author takes when the predicate language isn't expressive enough. API change (F1 fix): SanitizedArguments::unresolved() promoted from pub(crate) to pub. This is the documented safe default — predicates that need args must fail closed against it — so exposing the constructor cannot weaken any trust property. The sanitizing from_json constructor stays sealed; that's the trust boundary. Without this fix, external hook authors could not construct a BeforeCapabilityHookContext with both a known provider AND unresolved args, which made TDD of their own predicate impossible. Findings documented in docs/real-hooks-findings.md, ranked by severity. Big-picture observation: writing the Trusted Rust hook (F4) was easier than writing the declarative predicate hook (F1 + F2 + F3) — three of seven findings target predicate-authoring ergonomics. The declarative path needs the most polish before third-party extension authors will trust it for non-trivial policy. Tests: 6 new in real_hooks.rs, all pass. * feat(hooks): close all remaining threat-model and ergonomics gaps Closes the Med-priority threat-model gaps (I2, D5, poison runbook) and all real-hooks ergonomics findings (F2, F3, F5, F6, F7) in a single pass. Threat model: - I2 (resolver field-scope): documented in SanitizedArguments rustdoc. The narrow public surface (only is_resolved + extract_numeric) enforces field-scope by construction for the current predicate path. Reassess when Installed-WASM lands. - D5 (evaluator state ceiling): MAX_HISTORY_KEYS = 8192 per map, LRU eviction with evictions_observed() metric for operator monitoring. New regression test lru_eviction_increments_counter_and_drops_oldest_key. - Poison-stickiness runbook: new docs/operator-runbook.md with recovery options ranked by cost. Ergonomics findings: - F2 (closed-vocab deny reasons): rustdoc on OnExceededAction and GateDecisionView::Deny explaining the audit-vs-model split and why manifest reason text doesn't reach the model. - F3 (NumericSum can't be TDD'd outside Reborn): new test-support feature flag with SanitizedArguments::for_tests(value) that external hook authors can opt into via dev-dep. - F5 (two ExtensionId types): added From<&ironclaw_host_api::ExtensionId> impl for identity::ExtensionId, plus cross-link rustdoc. - F6 (HookManifestEntry struct-literal fragility): added #[non_exhaustive] + HookManifestEntry::new(id, kind, body) + with_scope/with_phase/with_priority/with_description/with_requires_grant builder methods. Migrated 3 external call sites in tests/. - F7 (priority guidance): rustdoc on HookPriority with when-to- deviate guidance, named FIRST/LAST constants documented for Builtin/Telemetry use cases. Tests: 151 unit + 1 + 6 integration in ironclaw_hooks all pass with --all-features. ironclaw_reborn (13 hooks_integration scenarios) unchanged. Threat model updated: I2 / D5 / poison runbook marked closed in both the per-vector table and the cross-cutting properties matrix. Open follow-ups now down to two Low items (I4 timing side-channel residual, I5 instruction-marker denylist refresh) plus the deferred DenyReasonCode enum from F2. * fix(ci): collapse nested match in hooks_integration test for clippy --all-features CI runs `cargo clippy --all --tests --examples --all-features -- -D warnings` which is stricter than the workspace clippy I ran locally and trips `clippy::collapsible_match` on the nested-if in HookDecisionEmitted matching. Collapse the inner `if decision.kind_name() == "deny"` into an arm guard. * feat(hooks): address henrypark133 review — Critical #1/#2/#3/#5, Concerning #5/#7 Address composition-seam bugs in the Reborn factory wiring + doc tidy. henrypark133 review findings addressed: Critical #1 — before_prompt hook messages not materialized. HookedLoopPromptPort now requires a HookPromptMaterializationSink and fails closed if patches are emitted without one. The reborn factory installs an InstructionStoreBackedHookSink adapter that delegates to the host's InstructionMaterializationStore, so synthetic msg:hook.* refs are resolvable by the downstream model resolver. New seam trait (HookPromptMaterializationSink) keeps ironclaw_hooks decoupled from LoopRunContext. Critical #2 — OwnCapabilities hooks were inert in production wiring. Factory now installs SurfaceBackedProviderResolver (consults the visible-capability surface for capability_id → provider). With this, ctx.provider is populated and OwnCapabilities-scoped Installed hooks actually fire against their own provider's capabilities. Critical #3 — gate refs were unresolvable. Middleware default switched from UuidHookGateRefFactory to FailClosedHookGateRefFactory. Tests must explicitly opt into UUID (via with_gate_ref_factory) to exercise the affirmative ApprovalRequired path; production deployments must install a router-backed factory. New factory method RebornLoopDriverHostFactory::with_hook_gate_ref_factory. Concerning #5 — AfterModel fired twice + before durable finalization. Removed AfterModel dispatch from HookedLoopModelPort; the transcript port's finalize_assistant_message is now the sole AfterModel boundary (the durable one). Model port wrapper is preserved as a no-op shim for symmetry + future model-response-observed point. Concerning #7 — doc tidy: - CLAUDE.md: 3 trust classes → 4 (Builtin/Trusted/Installed/SelfAuthored with explicit note that SelfAuthored is run-scoped only and not loadable from an external source). - operator-runbook.md: "Audit log" → "durable runtime event stream" where the projection is actually the runtime-event stream, not formal AuditEnvelope records. - prior-art.md: poison-lifetime nuance — per-host-build with the factory pattern, process-lifetime only for the legacy adapter. - prior-art.md:80: trailing whitespace removed. Testing gaps from henrypark133 — caller-level tests through RebornLoopDriverHostFactory: #1 (before_prompt resolver path): before_prompt_hook_message_is_resolvable_via_factory_wiring #2 (OwnCapabilities positive/negative/unknown): own_capabilities_hook_fires_when_provider_matches own_capabilities_hook_does_not_fire_when_provider_differs own_capabilities_hook_does_not_fire_when_provider_unknown #3 (pause/auth gate lifecycle or fail-closed): pause_approval_with_default_factory_fails_closed_as_denied pause_approval_hook_surfaces_as_approval_required_with_real_gate_ref (updated to require explicit UuidHookGateRefFactory opt-in) #5 (AfterModel exactly-once at durable boundary): after_model_fires_exactly_once_at_durable_boundary Still TODO from review (separate commits): Critical #4 (telemetry context — two-run attribution) + gap #4 Concerning #6 (TimelineEntry hook metadata projection) + gap #6 Tests: 154 unit + 18 hooks_integration + all other reborn tests pass. Workspace clippy + fmt + no-panics clean. * feat(hooks): address remaining henrypark133 review — Critical #4, Concerning #6 Critical #4 — per-run hook telemetry attribution. New `HookDispatcherBuilderFactory` signature: factory returns a HookDispatcherBuilder, and `RebornLoopDriverHostFactory` attaches a `RunScopedHookMilestoneSink` keyed to the CURRENT run's LoopRunContext inside `build_text_only_host_with_capabilities`, before sealing the dispatcher. The previous zero-arg signature relied on the closure capturing run_context — silently misattributed across reuses; new public API `with_hook_dispatcher_builder_factory` removes that failure mode entirely. Legacy `with_hook_dispatcher_factory` retained for back-compat (its sink-wiring contract stays caller-side). Concerning #6 — TimelineEntry hook metadata. Added 6 optional fields to `TimelineEntry` (hook_id, hook_point, hook_trust_class, hook_decision, hook_failure_category, hook_failure_disposition) and projected them from `RuntimeEvent::Hook*`. Replay consumers now see which hook fired/failed, not just that some hook event happened. Each field is closed-vocabulary (no free-form reason text — that stays in the audit reason payload, not the product replay DTO). Testing gaps from henrypark133 — caller-level tests: #4 (two-run hook telemetry attribution): hook_telemetry_attribution_is_per_run_not_captured Builds two hosts from the SAME builder factory closure with two fresh LoopRunContexts. Asserts each run's hook milestones carry its OWN run_id (no stale captured one). #6 (replay projection contract for hook events): hook_runtime_events_project_with_sanitized_hook_metadata non_hook_runtime_events_project_with_no_hook_metadata Constructs RuntimeEvent::Hook{Dispatched,DecisionEmitted,Failed} and asserts the projection preserves the metadata fields. The negative test guards against cross-contamination on non-hook events. All henrypark133 review items now addressed: Critical: #1, #2, #3, #4 — done Concerning: #5, #6, #7 — done Testing gaps: #1-#6 — done Tests: 154 unit + 19 hooks_integration in ironclaw_reborn + 61 reborn unit + 38 + 2 new in ironclaw_event_projections + ... pass. Workspace clippy + fmt + no-panics clean. * docs(hooks): scope DenyReasonCode closed-vocabulary enum (successor #6) Successor PR from #3573 — real-hooks ergonomics finding F2 (deferred). Adds a curated vocabulary of model-visible denial reasons so hook authors can communicate why a deny happened without opening a free-form prompt-injection channel. * feat(hooks): DenyReasonCode + PauseReasonCode closed-vocabulary enums Address real-hooks ergonomics finding F2 (deferred from PR #3573). The prior dispatcher collapsed every Installed-tier deny to the static label 'hook_predicate_denied', because manifest reason strings are author-controlled and surfacing them to the model would open a prompt-injection channel. The cost: the agent couldn't tell *why* a hook denied. This PR introduces two closed-vocabulary enums: - DenyReasonCode: Generic / RateLimit / ValueCap / Blocklist / RequiresApproval / OutOfPolicy - PauseReasonCode: Generic / RequiresApproval / OverThreshold / SensitiveAction Each variant has an as_label() returning &'static str (so the sink's &'static str contract is preserved). New OnExceededAction variants 'DenyWithCode { code, reason }' and 'PauseApprovalWithCode { code, reason }' let manifest authors opt into the richer labels while keeping reason audit-only. The legacy Deny { reason } / PauseApproval { reason } variants are retained for back-compat and map to DenyReasonCode::Generic / PauseReasonCode::Generic — existing manifests continue to produce hook_predicate_denied / hook_predicate_pause_requested. Threat-model regression: a hook author cannot smuggle text into the model-visible label because the 'code' field is typed as the enum; there's no String slot exposed model-side. A test (deny_with_code_only_exposes_enum_variants_to_model) documents this as a compile-time property. Tests (+7 new = 161 total): - deny_reason_code_labels_are_stable: pins the label vocabulary so rename/relabel is loud. - pause_reason_code_labels_are_stable: same for PauseReasonCode. - deny_with_code_round_trips_through_json + pause variant: wire round-trip + snake_case tag assertion. - deny_with_code_only_exposes_enum_variants_to_model: compile-time property check. - rate_or_value_cap_with_deny_code_routes_to_code_label: end-to-end affirmative test that the dispatcher emits the code's label. - rate_or_value_cap_with_pause_code_routes_to_code_label: same for pause. Scope doc: crates/ironclaw_hooks/docs/successors/06-deny-reason-code.md * test(hooks): address codex review on #3636 - Update stale real-hooks-findings.md F2 row to cite this PR's enum follow-on (was 'deferred'). - Add install_deny_with_code_manifest_surfaces_code_label_on_dispatch: end-to-end test driving the registrar->dispatcher path for the new DenyWithCode variant (prior tests covered serde + direct hook evaluation, but not the manifest install path that downstream authors actually use). Codex review on PR #3636: APPROVE with two recommendations; both addressed. Tests: 162 unit (+1 new). Clippy/fmt clean. * fix(hooks): attenuate Installed-tier prompt patches to user role Installed-tier `before_prompt` patches were injected as role:"system" messages. Envelope text labels ("[ext-foo says]: ...") do not strip system-role authority from the model's perspective, so a third-party extension could inject system-tier instructions through a snippet patch. This is a prompt-authority escalation against the trust hierarchy the framework otherwise enforces. Add `role_for_trust_class()` mapping Installed -> "user" and Builtin/Trusted/SelfAuthored -> "system". Thread per-patch trust_class through `wrap_patches_to_messages` and use it for the emitted `LoopModelMessage.role`. Tests: - installed_hook_patch_drops_to_user_role: asserts the role for an Installed-tier patch is "user" - trusted_tier_hook_patch_keeps_system_role: regression that Trusted tier still produces system-role content * fix(hooks): enforce scope filter on observer dispatch + reject incompatible points Two related defense-in-depth fixes against silent scope-filter failure: 1. The registry silently accepted Installed bindings with `HookBindingScope::OwnCapabilities` at points (BeforePrompt, AfterModel, AfterCheckpoint) whose dispatch context carries no per-capability provider. The manifest's declared scope had no effect at all — the hook fired against every dispatch. Reject the binding at install time so the operator sees the misconfiguration. 2. `dispatch_observer_at` for `AfterCapability` did not consult the binding's scope, so an Installed observer registered with `OwnCapabilities` fired against every invocation regardless of provider. Add `dispatch_observer_at_with_provider` carrying the resolved capability provider; the capability-port middleware resolves the provider once per invocation and threads it through both the BeforeCapability hook context and the AfterCapability observer dispatch. The dispatcher then enforces `HookBindingScope::permits` on each observer binding. `ObserverHookContext` gains a `provider: Option<ExtensionId>` field; `#[non_exhaustive]` keeps existing authors compiling. Tests: - rejects_own_capabilities_at_before_prompt - rejects_own_capabilities_at_after_model - accepts_own_capabilities_at_before_capability - own_capabilities_observer_filters_foreign_providers (covers foreign / matching / unresolved provider) * fix(hooks): preserve free-form audit reason alongside closed-vocab model label (serrrfirat #3636) `PredicateBackedBeforeCapabilityHook::evaluate()` was discarding the free-form `reason` from `EvaluatorDecision::{Deny, PauseApproval}` with `..` and only sending `code.as_label()` into the sink. The `HookDecisionEmitted` milestone therefore carried only the closed- vocab label, and operator-visible audit/SSE context was silently lost end-to-end. The fix splits the channels: - Model sees the closed-vocab label (`hook_rate_limit`, `hook_pause_over_threshold`, ...) via `sink.deny(label)`. This channel is unchanged. - Audit/SSE sees the manifest's free-form `reason` via a new audit-only sink method `record_audit_reason(reason: String)`. The recording sink captures it; the dispatcher reads it after the hook returns and threads it into `LoopHostMilestoneKind::HookDecisionEmitted`. Surface changes: - `PrivilegedGateSink` / `RestrictedGateSink` gain `record_audit_reason(String)` — accepts dynamic `String` (audit-only, no model-facing seam) unlike the `&'static str` decision reasons. - `RecordingGateSink` gains an `audit_reason: Option<String>` field. - `GateHookOutcome::Decision` is now `Decision { decision, audit_reason }`. - `HookDispatcher::emit_decision_with_audit` threads the audit reason into the milestone. - `LoopHostMilestoneKind::HookDecisionEmitted` gains a `#[serde(default, skip_serializing_if = "Option::is_none")]` `audit_reason: Option<String>`. The durable RuntimeEvent projection intentionally drops this field — audit reasons are operator-facing in-memory SSE content, never durable cross-process surface. Tests: - `deny_with_code_records_audit_reason_separately_from_model_label`: asserts the recording sink ends with `Deny { reason: "hook_rate_limit" }` in `state` AND `audit_reason == Some("daily cap of $1000 ...")`. * fix(hooks): remove unused model_request helper (CI clippy fix) * fix(hooks): address serrrfirat P1/P2 findings on PR #3573 Three issues from the 5-15 review: **P1 #1 registrar.rs:70 — `same_tenant` grants not enforced** `HookManifestEntry::validate` only confirmed `requires_grant` was present; the registrar then immediately installed the binding with no host-verified grant context. A manifest could declare `requires_grant = "anything"` and get a cross-extension binding for free. Fix: `HookRegistrar` now carries a `verified_grants: HashSet<String>` (empty by default — default-deny). Add the host-facing setter `with_verified_grants(...)`. At `install_one`, if `entry.requires_grant` is `Some(g)`, require `g ∈ verified_grants` or reject with a clear error. Tests: - `install_rejects_same_tenant_without_verified_grant` - `install_rejects_same_tenant_when_verified_grants_mismatch` - The existing positive test `installer_propagates_owning_extension_and_scope_from_manifest` now wires the verified grant explicitly (proves the API contract). **P1 #2 prompt_port.rs:150 — zip misalignment** The materialization loop zipped surviving messages against the ORIGINAL unfiltered patch list. `wrap_patches_to_messages` skips metadata patches and over-budget snippets, so the zip silently paired message[0] with patch[0] even when patch[0] was the skipped metadata — materializing the wrong content (or none) under the snippet's synthetic ref. Fix: `wrap_patches_to_messages` now returns `Vec<WrappedHookMessage { message, safe_content }>` — surviving messages paired with their content by construction. The caller materializes `entry.safe_content` under `entry.message.content_ref` directly; no zip against unfiltered input. Removed the now-unused `safe_content_for_patch` helper. Test: - `materialization_stays_aligned_when_metadata_patches_are_filtered`: a hook emits `[metadata, snippet]`; asserts only one model message, and the materialized content under its ref contains the snippet's body — proves filtering can no longer desync from materialization. **P2 #3 loop_driver_host.rs:1343 — `with_hook_dispatcher_builder`** Docs said it deferred `build_arc()` to let the host factory finalize wiring; the implementation called `build_arc()` eagerly and routed through the legacy shared-dispatcher adapter, losing per-run dispatcher isolation and the run-scoped milestone sink. Fix: marked `#[deprecated]` with a note pointing callers to `with_hook_dispatcher_builder_factory(|| ...)` for per-build isolation, or `with_hook_dispatcher(...)` if they actually meant the shared adapter. The method body is unchanged so no callers break; they'll see the deprecation warning. No internal callers exist, so the deprecation doesn't trip `-D warnings`. All 162 hooks lib + 19 reborn integration tests pass; clippy clean. * fix(hooks): address serrrfirat 3573-2026-05-15 review findings P1 — prompt bundle authority mismatch (prompt_port.rs): `HookedLoopPromptPort::build_prompt_bundle` called the inner port first, which caused `HostManagedLoopPromptPort` to issue the prompt-bundle authority grant against the pre-hook message list. The wrapper then appended `msg:hook.*` messages to `bundle.messages`, so the downstream model request hit `grant.messages != messages` and failed closed with "model request messages do not match the host-built prompt bundle". Add `with_bundle_authority(authority, run_context)` and re-issue the grant after appending hook messages so it covers the post-hook bundle. Reborn wires `prompt_authority.clone()` + `run_context.clone()` into the wrapper at construction time. P2 — observer installer accepts non-observer points (dispatch.rs): `install_observer` accepted any `HookPointSpec` (including `BeforeCapability` / `BeforePrompt`) and only populated the observer map. Dispatch later found a binding without a gate/mutator impl and fail-closed the capability with "binding present without installed implementation". Reject non-observer points at install time so misuse fails loudly rather than poisoning bindings at dispatch. P2 — batch path skipped AfterCapability observers on inner error (capability_port.rs): The batch loop used `?` directly on `self.inner.invoke_capability(...)`, which propagated the error before dispatching `AfterCapability` observers. Failed batch entries disappeared from telemetry / audit, while the single-invocation path dispatches observers on error. Capture the inner result, dispatch observers, then propagate the error. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(hooks): address PR #3573 review feedback round 3 Addresses serrrfirat's CHANGES_REQUESTED review (2026-05-20) by tightening several install-time / dispatch-time bounds and gating production seams: - Bound free-form audit reasons crossing telemetry. New `telemetry::sanitize_audit_reason` strips control characters and caps length at 512 bytes; `emit_decision_with_audit` routes the manifest- supplied reason through it before publishing milestones. Manifest validation also rejects reasons over the same byte limit at install time so the wire-side cap is a defense-in-depth layer, not the only line. - Make hot dispatch O(H) instead of O(H^2). The per-binding poison recheck used to acquire the registry mutex and walk every binding; `ordered_bindings_with_poison_snapshot` now takes the active bindings and the poisoned hook-id set under a single lock, and each loop threads a local `HashSet<HookId>` that absorbs mid-dispatch poisoning. Removed the redundant `is_poisoned` helper. - Gate `HookDispatcher::registry_for_test` behind `cfg(any(test, feature = "test-support"))`. The accessor previously exposed `&Mutex<HookRegistry>` in production, letting any `Arc<HookDispatcher>` holder lock and call `HookRegistry::poison` to disable installed hooks. Added `active_bindings_snapshot(point)` as the read-only production-safe replacement. - `#[serde(deny_unknown_fields)]` on every hook-manifest and predicate DTO (`HookManifestEntry`, `HookManifestBody`, `WasmBudget`, `HookPredicateSpec`, `CapabilityPredicate`, `ValueOrRateBound`, `OnExceededAction`). Typoed or unsupported fields (e.g. a manifest-supplied `trust_class`) now fail loud at install time instead of being silently dropped. - Bound predicate trees at install. New `validate_predicate_tree` enforces `MAX_PREDICATE_DEPTH = 8`, `MAX_PREDICATE_NODES = 64`, `MAX_PREDICATE_STRING_BYTES = 256`, and `MAX_MANIFEST_REASON_BYTES = 512`. A hostile registry manifest can no longer install a deep or huge `All`/`Any` tree that the evaluator would recursively walk on every match. - Cap sliding-window samples per key. `MAX_SAMPLES_PER_KEY = 4_096` in the predicate evaluator. Both the invocation-count and numeric-sum histories drop the oldest sample once the cap is reached, bounding memory under attacker-triggered hot capabilities while preserving rate/value-cap semantics over the most recent window. - `split_indexer` / `resolve_path` now fail closed on malformed bracket syntax (`amount[foo]`, `amount[`, trailing garbage). Previously they silently fell back to the parent field, which could let a typoed `NumericSum` predicate evaluate against the wrong value and allow calls the predicate would otherwise have denied. - Honor `PatchOrdinalHint`. `WrappedHookMessage` carries the source patch's `ordinal_hint`; `HookedLoopPromptPort` inserts `NearTop` messages after the bundle's `identity_message_count` and appends `Last` messages at the end. Safety/policy snippets that need early placement now get it. - Update `ironclaw_hooks` top-level docs to reflect the four trust classes (`Builtin`/`Trusted`/`Installed`/`SelfAuthored`) and the now-wired Reborn middleware composition. Tests added: - `manifest::rejects_unknown_top_level_field` - `manifest::rejects_unknown_wasm_budget_field` - `manifest::rejects_predicate_tree_exceeding_max_depth` - `manifest::rejects_predicate_tree_exceeding_max_nodes` - `manifest::rejects_predicate_string_exceeding_max_bytes` - `manifest::rejects_manifest_reason_exceeding_max_bytes` - `points::capability::malformed_indexer_returns_none_not_parent_value` - `telemetry::sanitize_audit_reason_*` (truncate / strip control / preserve / empty) `cargo fmt`, `cargo clippy --all --benches --tests --examples --all-features`, and `cargo test -p ironclaw_hooks` all pass clean. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(hooks): batch deferred test coverage from #3573 review (#3914) * perf(hooks): defer capability input resolution until a predicate needs it (#3913) * fix(rebase): adapt hooks tests + middleware to upstream API additions - CapabilityDescriptorView: add parameters_schema field - LoopModelRequest / LoopPromptBundleRequest: add capability_view field - TimelineEntry test builder: add hook_id / hook_point / hook_trust_class / hook_decision / hook_failure_category / hook_failure_disposition fields - ironclaw_reborn::tests::hooks_integration: switch from InMemoryLoopCheckpointStore to InMemoryTurnStateStore (which now impls both LoopCheckpointStore and TurnStateStore), pass TurnActor in TurnRunState, supply the new turn_state_store factory arg - ironclaw_reborn lib.rs: drop the pub-use re-exports that upstream intentionally removed (per the module-directory rationale in the current ironclaw_reborn lib.rs doc comment); update the hooks_integration test imports to use module paths - Cargo.toml: union the hooks-foundation member list with upstream's new crates (event_streams, auth, first_party_extensions, reborn_webui_ingress, product_workflow_storage, webui_v2); drop ironclaw_storage which no longer exists upstream - crates/ironclaw_architecture/tests/reborn_dependency_boundaries: keep upstream's removal of ironclaw_filesystem from the ironclaw_turns forbidden list AND add ironclaw_hooks to that list - crates/ironclaw_reborn/src/milestone_events.rs: drop dead loop_failure_kind helper (replaced upstream by loop_failure_kind_name in text_loop_driver.rs); keep hook_decision_label which is still used Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * perf(hooks): restore batched capability dispatch when hooks acti…
Refs #3431
KB task: KB-008
GitHub issue #3431: [Reborn] Add MemoryPromptContextService production adapter
URL: #3431
Repo: nearai/ironclaw
Labels: reborn
Assignees: nickpismenkov
Scope: IronClaw Reborn issue imported from GitHub.
Before coding: read full issue body/comments, check linked PRs/blockers, work from current origin/reborn-integration unless issue says otherwise, keep PR tightly scoped.
Acceptance: satisfy GitHub issue acceptance criteria; include tests/verification evidence; never merge without explicit user approval.
Auto-opened by kb when task reached In Review. Auto-merge remains disabled; do not merge without explicit operator approval.