feat(hooks): DenyReasonCode closed-vocabulary enum (successor to #3573) - #3636
Conversation
Successor PR from #3573 — real-hooks ergonomics finding F2 (deferred). Adds a curated vocabulary of model-visible denial reasons so hook authors can communicate why a deny happened without opening a free-form prompt-injection channel.
There was a problem hiding this comment.
Code Review
This pull request introduces a design document for a DenyReasonCode closed-vocabulary enum, aimed at providing structured denial reasons to models while mitigating prompt-injection risks. The feedback suggests renaming the enum to avoid ambiguity with host-level authorization codes, removing semantically inconsistent variants like RequiresApproval from the denial path, and extending the PauseApproval action with a structured PauseWithCode variant to maintain consistency across decision paths.
|
|
||
| ## Scope | ||
|
|
||
| 1. Introduce `DenyReasonCode` enum (closed vocabulary) in |
There was a problem hiding this comment.
Consider renaming DenyReasonCode to something more specific, such as HookDenyReason or PolicyDenyCode. This would help distinguish it from the existing DenyReason enum in ironclaw_host_api::decision, which represents host-level authorization failures (e.g., MissingGrant, InvalidPath) rather than hook-level policy decisions.
References
- Documentation for complex logic, such as security policies, must precisely match the code implementation and maintain semantic clarity.
| - `RequiresApproval` (when paired with PauseApproval — different | ||
| decision but same reason space) |
There was a problem hiding this comment.
Including RequiresApproval as a variant of DenyReasonCode is semantically inconsistent. A Deny decision typically represents a terminal rejection of the action. If a hook determines that an action requires approval, it should result in a PauseApproval decision instead. This variant likely belongs exclusively in the proposed PauseReasonCode mentioned in point 4.
References
- Documentation for complex logic, such as security policies, must precisely match the code implementation. Semantic inconsistencies in policy definitions should be resolved.
| pub enum OnExceededAction { | ||
| Deny { reason: String }, // existing | ||
| DenyWithCode { code: DenyReasonCode, reason: String }, // new | ||
| PauseApproval { reason: String }, |
There was a problem hiding this comment.
To maintain consistency with the DenyWithCode improvement and ensure the "Pause" path also avoids prompt-injection risks in model-visible labels, the PauseApproval action should be extended with a structured code variant as well.
| PauseApproval { reason: String }, | |
| PauseApproval { reason: String }, | |
| PauseWithCode { code: PauseReasonCode, reason: String }, |
References
- Documentation for complex logic, such as security policies, must precisely match the code implementation. Consistency across action paths (Deny vs Pause) should be maintained.
b6ac22e to
ac4c109
Compare
Address real-hooks ergonomics finding F2 (deferred from PR #3573). The prior dispatcher collapsed every Installed-tier deny to the static label 'hook_predicate_denied', because manifest reason strings are author-controlled and surfacing them to the model would open a prompt-injection channel. The cost: the agent couldn't tell *why* a hook denied. This PR introduces two closed-vocabulary enums: - DenyReasonCode: Generic / RateLimit / ValueCap / Blocklist / RequiresApproval / OutOfPolicy - PauseReasonCode: Generic / RequiresApproval / OverThreshold / SensitiveAction Each variant has an as_label() returning &'static str (so the sink's &'static str contract is preserved). New OnExceededAction variants 'DenyWithCode { code, reason }' and 'PauseApprovalWithCode { code, reason }' let manifest authors opt into the richer labels while keeping reason audit-only. The legacy Deny { reason } / PauseApproval { reason } variants are retained for back-compat and map to DenyReasonCode::Generic / PauseReasonCode::Generic — existing manifests continue to produce hook_predicate_denied / hook_predicate_pause_requested. Threat-model regression: a hook author cannot smuggle text into the model-visible label because the 'code' field is typed as the enum; there's no String slot exposed model-side. A test (deny_with_code_only_exposes_enum_variants_to_model) documents this as a compile-time property. Tests (+7 new = 161 total): - deny_reason_code_labels_are_stable: pins the label vocabulary so rename/relabel is loud. - pause_reason_code_labels_are_stable: same for PauseReasonCode. - deny_with_code_round_trips_through_json + pause variant: wire round-trip + snake_case tag assertion. - deny_with_code_only_exposes_enum_variants_to_model: compile-time property check. - rate_or_value_cap_with_deny_code_routes_to_code_label: end-to-end affirmative test that the dispatcher emits the code's label. - rate_or_value_cap_with_pause_code_routes_to_code_label: same for pause. Scope doc: crates/ironclaw_hooks/docs/successors/06-deny-reason-code.md
- Update stale real-hooks-findings.md F2 row to cite this PR's enum follow-on (was 'deferred'). - Add install_deny_with_code_manifest_surfaces_code_label_on_dispatch: end-to-end test driving the registrar->dispatcher path for the new DenyWithCode variant (prior tests covered serde + direct hook evaluation, but not the manifest install path that downstream authors actually use). Codex review on PR #3636: APPROVE with two recommendations; both addressed. Tests: 162 unit (+1 new). Clippy/fmt clean.
Codex review addressed (commit c4705aa)Verdict was APPROVE; addressed both recommendations:
162 unit tests pass total. |
SummaryReviewed PR #3636 only. Base PR adds closed-vocabulary deny/pause reason codes. Model-visible labels are safer and legacy mapping looks compatible. Merge stance: Medium correctness/logging gap in promised audit reason preservation. Findings
Security/data-flow notes
Correctness/invariant notes
Missing tests
|
henrypark133
left a comment
There was a problem hiding this comment.
What looks good:
- The model-visible label path is typed and closed-vocabulary through
code.as_label(). - Free-form manifest
reasonremains audit-only in the touched installed-hook path. - There is a registrar-level test for
DenyWithCode, which is the right caller boundary.
No verified findings.
Low-priority notes:
crates/ironclaw_hooks/docs/successors/06-deny-reason-code.md:53: doc examples sayrate_limit-style labels, while runtime emitshook_rate_limit. Worth aligning so downstream consumers do not key on the wrong string.crates/ironclaw_hooks/src/predicate.rs:149:RequiresApprovalexists in both deny and pause vocabularies. It may be intentional, but the docs should make the deny-vs-pause semantics explicit.
Summary:
- Recommended verdict: Approve
- Prior feedback status: Gemini pause-code concern appears addressed by
PauseReasonCode. - Residual risk: no registrar-level pause-code test, though direct hook coverage exists.
…del label (serrrfirat #3636) `PredicateBackedBeforeCapabilityHook::evaluate()` was discarding the free-form `reason` from `EvaluatorDecision::{Deny, PauseApproval}` with `..` and only sending `code.as_label()` into the sink. The `HookDecisionEmitted` milestone therefore carried only the closed- vocab label, and operator-visible audit/SSE context was silently lost end-to-end. The fix splits the channels: - Model sees the closed-vocab label (`hook_rate_limit`, `hook_pause_over_threshold`, ...) via `sink.deny(label)`. This channel is unchanged. - Audit/SSE sees the manifest's free-form `reason` via a new audit-only sink method `record_audit_reason(reason: String)`. The recording sink captures it; the dispatcher reads it after the hook returns and threads it into `LoopHostMilestoneKind::HookDecisionEmitted`. Surface changes: - `PrivilegedGateSink` / `RestrictedGateSink` gain `record_audit_reason(String)` — accepts dynamic `String` (audit-only, no model-facing seam) unlike the `&'static str` decision reasons. - `RecordingGateSink` gains an `audit_reason: Option<String>` field. - `GateHookOutcome::Decision` is now `Decision { decision, audit_reason }`. - `HookDispatcher::emit_decision_with_audit` threads the audit reason into the milestone. - `LoopHostMilestoneKind::HookDecisionEmitted` gains a `#[serde(default, skip_serializing_if = "Option::is_none")]` `audit_reason: Option<String>`. The durable RuntimeEvent projection intentionally drops this field — audit reasons are operator-facing in-memory SSE content, never durable cross-process surface. Tests: - `deny_with_code_records_audit_reason_separately_from_model_label`: asserts the recording sink ends with `Deny { reason: "hook_rate_limit" }` in `state` AND `audit_reason == Some("daily cap of $1000 ...")`.
Code Review — PR #3636 (DenyReasonCode closed-vocab enum)Verdict: APPROVE ✅ — well-executed and thoughtfully designed. The closed-vocabulary approach closes the prompt-injection vector while preserving backward compatibility. OverviewAdds Code quality & style — strong overall
Suggestions (non-blocking)
Test coverage — excellentEight distinct test vectors: backward-compat path, both new affirmative paths, audit-reason separation, label-stability pin, serde round-trip, compile-time type assertion, and an end-to-end registrar test. Security — threat model is soundType enforcement closes the injection channel. The stability + type-assertion tests validate it. One long-term note: if audit logs ever become model-visible in a future feature, the free-form |
serrrfirat
left a comment
There was a problem hiding this comment.
Review rerun against current PR head 5c33c741bdd4f916f0dc596cb6b4b942dbb6e768.
No blocking correctness/security findings. The audit split remains sound: model-visible labels stay closed-vocabulary, the manifest reason flows separately through audit_reason, and durable runtime-event projection does not include that free-form reason.
I found one low-priority docs drift issue and posted it inline.
Verification:
cargo test -p ironclaw_hooks code -- --nocapturepassed: 9 testscargo test -p ironclaw_turns hook_decision_emitted -- --nocapturepassed: 6 testscargo test -p ironclaw_reborn hook_decision_emitted_milestone_projects_to_runtime_event -- --nocapturepassed: 1 test- Refreshed base
origin/hooks-foundation-01to496d0ac121355c3f367282ab2609bfcfd627fc56;git merge-tree origin/hooks-foundation-01 origin/pr/3636produced a clean synthetic merge tree locally.
|
|
||
| ## Required tests | ||
|
|
||
| 1. Manifest with `DenyWithCode { code: RateLimit, .. }` → outcome |
There was a problem hiding this comment.
Low: this required-test text says the outcome carries the rate_limit label, but the implementation pins DenyReasonCode::RateLimit::as_label() to hook_rate_limit. The example enum above also omits the implemented PauseApprovalWithCode variant. Please align this successor doc with the shipped wire labels/shape so downstream consumers do not key on stale labels.
The "No panics in production code" CI check (scripts/check_no_panics.py) only recognizes `// safety:` suppression markers, not `RATIONALE:`. Since std::fmt::Write for String is infallible, just discard the Result with `let _ =` instead of `.expect()` — no panic call, no marker needed. Also merges in latest origin/hooks-foundation-01 (now includes the reborn-integration merge and PR #3636).
Resolves conflicts from the latest base advance: - crates/ironclaw_hooks/src/dispatch.rs: union HEAD's owning_extension plumbing with base's audit_reason field in emit_decision_with_audit - crates/ironclaw_hooks/src/registry.rs: keep both new install-time validations — HEAD's event-kind-filter biconditional (henrypark133 should-fix #1) and base's scope-vs-point coherence check. Extend point_has_capability_context to treat EventTriggered as capability-context-bearing (provider resolved per-event at dispatch time) - crates/ironclaw_turns/src/run_profile/milestones.rs: union HookDecisionEmitted to carry both owning_extension AND audit_reason - crates/ironclaw_reborn/src/milestone_events.rs: project owning_extension into RuntimeEvent::hook_decision_emitted, ignore audit_reason (it stays on the in-memory milestone sink, never crosses the durable boundary) - crates/ironclaw_reborn/src/lib.rs: re-export EventTriggeredHookSubscription* alongside the base's new LoopCapabilityPortFactory - crates/ironclaw_reborn/src/loop_driver_host.rs: keep both the new ProfiledCapabilityHostRuntime/LoopCapabilityPortFactory and the event-subscription handle field; drop unused tokio::sync::Notify - crates/ironclaw_hooks/src/dispatch.rs: reject EventTriggered in the observer install path explicitly (it is installed via the event subscription path, not the generic observer entry point) - propagate event_kind_filter: None to remaining HookBinding test literals exposed by the merge henrypark133 #1 (HIGH): the originally-cited `event.provider` always being `None` for hook milestone events is already addressed on this branch — commit 4bb06c9 / d7f43a3 carry owning_extension end-to-end through both `RuntimeEvent::hook_failed` and `RuntimeEvent::hook_decision_emitted`. The merge preserves and extends that plumbing rather than reverting it. cargo fmt + clippy (-D warnings) green; unit and hooks_integration tests green.
The "No panics in production code" CI check (scripts/check_no_panics.py) only recognizes `// safety:` suppression markers, not `RATIONALE:`. Since std::fmt::Write for String is infallible, just discard the Result with `let _ =` instead of `.expect()` — no panic call, no marker needed. Also merges in latest origin/hooks-foundation-01 (now includes the reborn-integration merge and PR #3636).
* feat(reborn): add ironclaw_hooks framework foundation (#3524)
Foundation slice of the Reborn loop hooks framework per nearai/ironclaw#3524.
Lands the trust primitives, sealed decision types, dispatcher contract, and
extension manifest schema; no Reborn middleware composition yet (next slice
wires HookDispatcher into LoopCapabilityPort / LoopPromptPort).
Design comment on #3524:
https://github.com/nearai/ironclaw/issues/3524#issuecomment-4439890144
What this PR ships
==================
* `crates/ironclaw_hooks/` — new crate
* `identity` — content-addressed `HookId` (blake3 of length-prefixed
extension + local + version fields). Same versioning primitive the rest
of Reborn should converge on for replay safety.
* `trust` — `HookTrustClass` enum (Builtin / Trusted / Installed) with
per-kind default attenuation. Trust class is fixed by source, never
declarable.
* `kinds/` — sealed decision DTOs. `BeforeCapabilityHookDecision`,
`HookPatch`, `ObserverFact` all have `pub` outer struct + `pub(crate)`
inner enum + `pub(crate)` constructors. Same #3460 witness pattern.
* `points/` — typed read-only contexts for each hook point.
* `sink` — split sink traits per trust tier. `PrivilegedGateSink` exposes
`allow()`; `RestrictedGateSink` does not. An Installed-tier hook
literally cannot mint Allow at the type level.
* `ordering` — phase → priority → hook id, stable. Phases gated by trust
(Validation/Authorization Builtin-only).
* `failure_policy` — Timeout/Panic/Malformed/AttenuationViolation
categories. Gate/Mutator fail closed, Observer/Effect fail isolated.
Slot poisoning persisted for the rest of the run on any category.
* `registry` — run-profile-sourced bindings; phase-vs-trust gate enforced
at insert; poisoning surface for the dispatcher.
* `dispatch` — HookDispatcher with deterministic ordering, panic
catch-unwind via futures::FutureExt, per-hook tokio::time::timeout,
short-circuit gate composition (Deny > PauseAuth > PauseApproval >
Allow), Telemetry-phase observers always run.
* `manifest` — serde types for the `[[hooks]]` section of extension
manifests. Predicate vs WASM body; same_tenant scope requires explicit
grant; Validation/Authorization phases rejected at parse time because
manifest hooks are always Installed.
* `predicate` — typed predicate language for declarative Installed hooks
(DenyCapability, PauseApproval, RateOrValueCap). Evaluator lives in
the dispatcher follow-up, not here.
* `crates/ironclaw_architecture/tests/reborn_dependency_boundaries.rs`
* Added `ironclaw_turns` -> `ironclaw_hooks` to the forbidden list.
* New BoundaryRule for `ironclaw_hooks` itself (cannot pull host_runtime,
dispatcher, secrets, network, wasm, etc.).
* `Cargo.toml` workspace member registration.
What this PR deliberately does NOT ship
========================================
* Reborn middleware composition wrapping LoopCapabilityPort / LoopPromptPort
with HookDispatcher. Next slice; ironclaw_reborn changes only.
* WASM hook execution path. Programmatic hooks parse and validate from
manifest; the wasmtime integration lands when the WASM dispatcher seam is
built.
* Predicate evaluation. Predicate types serialize and validate; the
evaluator that turns a `RateOrValueCap` spec into a `Deny` decision is in
the next slice alongside Reborn wiring.
* Event-triggered hooks (Phase 5 of the original roadmap).
* Self-authored hooks. Tracked separately at #3567 with monotonic-restriction
+ unforgeable-channel ratification.
Test plan
=========
* `cargo test -p ironclaw_hooks` — 47 tests (46 unit + 1 integration smoke
for the manifest -> binding -> dispatch pipeline).
* `cargo test -p ironclaw_architecture` — 13 tests; new boundary rule
passes, existing rules unaffected.
* `cargo clippy -p ironclaw_hooks --all-targets -- -D warnings` — clean.
* `cargo fmt -p ironclaw_hooks -- --check` — clean.
* `cargo check --workspace` — clean, no regressions in other crates.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(reborn): wire HookDispatcher into LoopCapabilityPort/LoopPromptPort
Follows the foundation slice (see initial commit). Adds the next layer:
1. Capability- and prompt-port middleware (`ironclaw_hooks::middleware`)
* `HookedLoopCapabilityPort` runs `dispatch_before_capability` before
every invocation, translates the composed decision into the existing
`CapabilityOutcome` vocabulary (Deny / PauseApproval / PauseAuth all
map to `Denied` for now; gate-ref plumbing for real pause semantics
lands in the next slice).
* `HookedLoopPromptPort` runs `dispatch_before_prompt` before bundle
construction. Observe-only for snippets in this slice; actual
snippet injection waits for the shared `prompt_envelope::wrap_untrusted`
helper (#3540 / #3471).
2. Declarative predicate evaluator (`ironclaw_hooks::evaluator`)
* `DenyCapability` and `PauseApproval` predicates: stateless, evaluated
directly against `BeforeCapabilityHookContext`.
* `RateOrValueCap` with `InvocationCount` bound: sliding-window counter
keyed by `(hook_id, capability_name)`, in-memory only. Window
parsing supports `s`/`m`/`h`/`d` units; unparseable windows fail
closed.
* `NumericSum` bound: types implemented but evaluation returns Allow
and emits a warn-level audit. Full argument-extraction story is a
follow-up slice once capability arguments become hook-visible.
* `PredicateEvaluator::evaluate_at(...)` test variant accepts an
explicit `Instant` so sliding-window tests don't depend on
real-clock progress.
3. Manifest -> dispatcher glue (`ironclaw_hooks::installed_hook`)
* `PredicateBackedBeforeCapabilityHook` wraps a `HookPredicateSpec`
plus an `Arc<PredicateEvaluator>` and implements
`RestrictedBeforeCapabilityHook`. The registry installer would
construct one of these per `[[hooks]]` entry whose body is
`HookManifestBody::Predicate`.
* Sink reasons are `&'static str`, so the dynamic predicate `reason`
surfaces in audit (via the evaluator's `EvaluatorDecision`) rather
than the model-visible decision. Closed-vocabulary labels carry
through to the sink.
4. Reborn composition seam (`ironclaw_reborn::loop_driver_host`)
* `RebornLoopDriverHostFactory::with_hook_dispatcher(Arc<HookDispatcher>)`
opt-in builder method. When set, the factory wraps the capability
and prompt ports with the hooked middleware. Default behavior
(no dispatcher) is unchanged from the pre-hooks shape, so existing
callers continue to work.
* Added `ironclaw_hooks` as a dep in `ironclaw_reborn`.
Test plan
=========
* `cargo test -p ironclaw_hooks` — 60 tests pass (59 unit + 1
integration smoke; +13 vs the foundation commit covering middleware,
evaluator, installed_hook).
* `cargo test -p ironclaw_reborn` — 118 tests pass; no regressions
from adding the dep.
* `cargo test -p ironclaw_architecture` — 13 tests pass; the
`ironclaw_turns -> ironclaw_hooks` boundary still holds and the new
`ironclaw_hooks` rule (no host_runtime / dispatcher / secrets /
network / wasm / reborn) is unaffected.
* `cargo clippy -p ironclaw_hooks --all-targets --all-features
-- -D warnings` — clean.
* `cargo clippy -p ironclaw_reborn --all-targets -- -D warnings` —
clean.
* `cargo fmt --all -- --check` — clean.
What still defers
==================
* WASM hook execution path.
* Persistent predicate counter (in-memory only for now).
* Argument-extraction so `NumericSum` predicates evaluate against
capability arguments.
* Gate-ref plumbing so PauseApproval / PauseAuth surface real
`CapabilityOutcome::ApprovalRequired` instead of `Denied`.
* Prompt-snippet injection (waits for shared envelope helper).
* Event-triggered hooks.
* Self-authored hooks (#3567).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(reborn): add HookedLoopModelPort/TranscriptPort/CheckpointPort observer middleware
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(reborn): end-to-end hooks integration through RebornLoopDriverHostFactory
Adds crates/ironclaw_reborn/tests/hooks_integration.rs covering the
factory's HookDispatcher wiring seam end-to-end. Tests drive
host.invoke_capability(...) (not dispatcher.dispatch_before_capability(...)
directly) so a regression in RebornLoopDriverHostFactory's wrapping
composition surfaces here.
Scenarios:
- PredicateBackedBeforeCapabilityHook (DenyCapability NameEquals
"cap.blocked") short-circuits invocation; inner port never called;
outcome is Denied(unknown("hook_denied")).
- A privileged selective hook that allows non-matching capabilities
proves the wrapper does not blanket-deny: cap.allowed reaches the
inner port and completes once.
- Factory built without with_hook_dispatcher() lets cap.blocked through
to the inner port, proving the hook plumbing is genuinely opt-in.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(reborn): add pass() + HookRegistrar + self-authored hooks scaffolding
Three additions to ironclaw_hooks:
B. `pass()` on gate sinks — distinguishes "evaluated, no opinion" from
"returned without minting a decision." A passing hook contributes
nothing to the composed decision; a silent hook is still Malformed
and fails closed. `PredicateBackedBeforeCapabilityHook` now routes
the evaluator's `Allow` decision through `sink.pass()` instead of
the previous `deny("hook_predicate_pass")` workaround.
A. `HookRegistrar` bridge — converts a `Vec<HookManifestEntry>` into
`HookBinding`s + dispatcher impls in one call. Predicate bodies are
wired through `PredicateBackedBeforeCapabilityHook`; WASM bodies
return `HookError::RegistryConstruction` for now. Adds
`HookDispatcher::insert_binding` so the registrar can mutate the
registry through the dispatcher rather than reach inside.
I. Self-authored hooks scaffolding — fourth `HookTrustClass` variant
for hooks the agent authors at runtime. Run-scoped only;
monotonic-restriction sink with no `allow`, no trusted-snippet path,
no effect-class constructor. Closed-vocabulary `SelfAuthoredReason`
enum keeps free-text reasons off the audit seam.
`SelfAuthorshipProvenance` captures authoring run/turn, timestamp,
spec digest, optional user ratification, and a generation-trace
pointer. Durable persistence depends on the unforgeable channel
from #3564 and lands separately.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(reborn): real gate-ref plumbing for hook PauseApproval/PauseAuth decisions
Previously, `GateDecisionInner::PauseApproval` and `PauseAuth` returned by
hooks were degraded to `CapabilityOutcome::Denied` at the middleware
boundary because the hook crate had no way to mint a `LoopGateRef` scoped
to the current run. Hooks that wanted to pause the loop for approval or
auth instead failed the call closed, leaving the host's approval-router
machinery unreachable from hook code.
This change introduces a `HookGateRefFactory` trait in
`ironclaw_hooks::middleware::gate_ref` that mints `LoopGateRef`s for
pause-class decisions. `HookedLoopCapabilityPort` now takes an
`Arc<dyn HookGateRefFactory>`, defaulting to `UuidHookGateRefFactory` (a
locally-unique opaque-id factory suitable for tests and the foundation
slice). Production deployments override via `.with_gate_ref_factory(...)`
with a factory bound to the current `LoopRunContext` and the host's
gate-router.
The translation in `decision_to_outcome` is now async so it can await the
factory. `PauseApproval` maps to `CapabilityOutcome::ApprovalRequired
{ gate_ref, safe_summary }` and `PauseAuth` to `AuthRequired`. If the
factory itself errors, the middleware falls back to `Denied` with a
sanitized `hook_gate_ref_unavailable` reason kind so the loop fails
closed rather than routing through an unresolvable suspension. The
underlying error text is dropped to avoid leaking gate-router state into
model-visible output.
Tests:
- `pause_approval_decision_surfaces_as_approval_required`,
`pause_auth_decision_surfaces_as_auth_required`,
`gate_ref_factory_failure_falls_back_to_denied` in
`middleware::capability_port::tests`.
- `pause_approval_hook_surfaces_as_approval_required_with_real_gate_ref`
in `crates/ironclaw_reborn/tests/hooks_integration.rs`, exercising the
full `RebornLoopDriverHostFactory` composition with the default
`UuidHookGateRefFactory`.
- Gate-ref factory unit tests in `gate_ref::tests`.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(reborn): add NumericSum predicate evaluation with capability argument extraction
Wires the missing argument-extraction story for the predicate evaluator so
`ValueOrRateBound::NumericSum` actually enforces a rolling numeric cap
instead of warn-and-allowing.
- Extend `BeforeCapabilityHookContext` with a sealed `SanitizedArguments`
view. Strings truncate to 256 bytes; objects/arrays cap at 8-deep.
`extract_numeric` supports dotted + bracketed paths (`order.amount`,
`items[0].price`) and returns `Option<rust_decimal::Decimal>`. The inner
representation is sealed so external callers can't bypass bounds.
- Introduce `CapabilityInputResolver` + bundled `NullCapabilityInputResolver`
in `middleware/resolver.rs`. The hooks crate intentionally doesn't know
how to dereference a `CapabilityInputRef` — that knowledge belongs to
the production host. Until a real resolver is wired in (follow-up),
arguments are `Unresolved` and `NumericSum` fails closed.
- `HookedLoopCapabilityPort::new` defaults to the null resolver; new
builder `.with_resolver(Arc<dyn CapabilityInputResolver>)` overrides.
- `PredicateEvaluator` gains a tenant-keyed `value_history` map. The
`NumericSum` arm parses `max` + `window`, extracts the numeric value
from sanitized args, accumulates within the rolling window, and applies
`on_exceeded` when the sum exceeds the cap. Unresolved args, missing
field, non-numeric field, unparseable max, and unparseable window all
fail closed via the configured `OnExceededAction`.
- Add `BeforeCapabilityHookContext::new_unresolved(...)` convenience
ctor; existing test sites switch to it instead of churning every call
site through the 4-arg ctor.
Test count: +14 (8 new SanitizedArguments tests, 6 new NumericSum
evaluator tests, 1 null-resolver test; one old NumericSum-stub-related
gap closed).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(reborn): seal hook registration trust boundary + dispatcher hardening
Addresses blocking findings from the security audit of `ironclaw_hooks`:
- C1 (Blocking, Trust Model): "Installed cannot Allow" was not enforced
at the registration boundary. `BeforeCapabilityHookImpl::Privileged`
was a public variant, so external crates with dispatcher access could
construct an Installed binding paired with a Privileged impl and bypass
the sink trait restriction. Sealed `BeforeCapabilityHookImpl`,
`BeforePromptHookImpl`, and `ObserverHookImpl` to `pub(crate)` and
replaced the single generic `install_before_capability` /
`install_before_prompt` / `install_observer` surface with tier-specific
public installers (`install_builtin_*`, `install_trusted_*`,
`install_installed_*`) that build the binding with the matching trust
class internally. Updated registrar, internal middleware tests, the
hooks foundation pipeline test, and the reborn `hooks_integration`
test to drive the new surface. Added regression tests proving the
trust class is set by the installer and that the seal is type-level.
- C5 (Medium, Slot Poisoning): same-dispatch poisoning was incomplete
because `ordered_bindings` snapshots once at the top of the loop, and
`HookRegistry::insert` accepted duplicate hook IDs. Rejected duplicate
hook IDs (any point) in `HookRegistry::insert` and added a poison
re-check before invoking each hook impl in `dispatch_before_capability`,
`dispatch_before_prompt`, and `dispatch_observer_at`. Added regression
tests for both behaviors.
- C6 (Medium, Manifest / Predicate Validation): `parse_window` could
panic on non-ASCII input because `split_at(len - 1)` requires a char
boundary. Rewrote to compute the unit char's UTF-8 byte length and
slice safely, added a public `validate_window` helper, and wired it
into `HookManifestEntry::validate` for both `InvocationCount` and
`NumericSum` bounds. Added tests for non-ASCII, empty, single-char,
and zero-duration windows.
- C2 (High, Tenant Isolation): partial fix only. The
`PredicateEvaluator`'s sliding-window counter was keyed by
`(hook_id, capability)`, so cross-tenant state could leak. Extended
`HistoryKey` to include `tenant_id` and added a regression test
proving counters partition by tenant. Documented the broader
dispatcher-per-build / per-run-fresh-dispatcher pattern as deferred
follow-up in `crates/ironclaw_hooks/CLAUDE.md`.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(reborn): emit hook telemetry milestones for audit/SSE observers
Wires the hook dispatcher into the host's milestone stream so audit
backends and SSE observers can see hook activity. Previously, hook
dispatch was invisible — denies, pauses, failures, and observer fires
left no trace in the host's observability backend.
Changes:
- `ironclaw_turns`: add `HookDispatched`, `HookDecisionEmitted`, and
`HookFailed` variants to `LoopHostMilestoneKind`, with a closed-
vocabulary `HookDecisionSummary` enum (Allow/Deny/PauseApproval/
PauseAuth/Pass/Patch). Introduce a lightweight `HookMilestoneSink`
trait that emits hook-specific *kinds* without requiring a
`LoopRunContext` (the dispatcher is a process-wide singleton that
cannot own a per-run context), plus a `RunScopedHookMilestoneSink`
adapter that injects run context and forwards to the existing
`LoopHostMilestoneSink`. Also add `InMemoryHookMilestoneSink` for
tests.
- `ironclaw_hooks`: add a `telemetry` module that converts hook-crate
types (`HookId`, `HookTrustClass`, `HookPointSpec`, `FailureCategory`,
`FailureDisposition`, `BeforeCapabilityHookDecision`) into the wire-
shape labels and summaries the milestone sink expects. Hook ids cross
the seam as hex strings because the strongly-typed `HookId` cannot be
imported from `ironclaw_turns` (the architecture test enforces
`ironclaw_turns -> ironclaw_hooks` stays absent).
- `ironclaw_hooks::dispatch`: add an optional `Arc<dyn
HookMilestoneSink>` to `HookDispatcher`, set via
`with_milestone_sink`. Emit `HookDispatched` before each hook runs,
`HookDecisionEmitted` after a decision/pass/patch, and `HookFailed`
on timeout/panic/malformed/missing-impl across all three dispatch
paths (before_capability, before_prompt, observer). Default behavior
(no sink attached) emits nothing — preserves the pre-telemetry
observable surface.
- `ironclaw_reborn`: document on `with_hook_dispatcher` that callers
attach the milestone sink to the dispatcher *before* wrapping it in
`Arc` and installing it into the factory, using a
`RunScopedHookMilestoneSink` to inject run-context. The dispatcher
itself is shared across runs, so attaching a fixed run-context inside
it would be wrong. Update `RuntimeEvent` projection in
`milestone_events.rs` to ignore the new hook kinds (no projection
pathway yet; emitted milestones are consumed by SSE observers
directly).
Tests:
- `ironclaw_hooks::dispatch`: 5 new tests covering milestone emission
for deny decisions, panic failures, prompt-mutator patches, observer
pass-throughs, and the no-sink default.
- `ironclaw_reborn` hooks_integration: end-to-end test wiring a
`RunScopedHookMilestoneSink` onto the dispatcher and asserting hook
activity surfaces in the host's `LoopHostMilestoneSink`.
Total: +6 hook telemetry tests; no existing tests modified.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(reborn): extract shared prompt envelope; inject hook patches into prompt bundle
Adds `ironclaw_prompt_envelope`, a leaf crate that owns the single envelope
primitive used by every model-visible untrusted-content path. `wrap_untrusted`
prefixes content with a closed-vocabulary `<Trusted|Untrusted> <source>
content: ` marker, rejects bodies carrying instruction-hijack phrases
(`ignore previous instructions`, `<|im_start|>`, `<system>`, etc.), and
enforces a 4 KiB byte budget by default.
Migrates `ironclaw_host_runtime::memory_context` to delegate envelope
wrapping, marker rejection, and control-character stripping to the new
crate while keeping the `LoopSafeSummary`-specific 512-byte cap and byte
truncation local. Existing memory_context behavior and tests are preserved.
Wires the same envelope into `ironclaw_hooks`:
* `HookPatch::add_enveloped_snippet` now takes a raw body and wraps it
via `wrap_untrusted(EnvelopeSource::Hook, …)`. `Installed` hooks
produce `Untrusted` envelopes; `Builtin`/`Trusted`/`SelfAuthored`
produce `Trusted` envelopes so downstream readers can distinguish the
two paths through a uniform marker.
* `HookedLoopPromptPort::build_prompt_bundle` is no longer observe-only.
After dispatching `before_prompt`, it envelope-wraps every snippet
patch (passing `Enveloped` through, wrapping `Trusted` with the
envelope helper), enforces the 4 KiB aggregate snippet byte budget
across patches, and appends the wrapped snippets to the prompt
bundle's `messages` as `system`-role `LoopModelMessage` entries
carrying deterministic `msg:hook.<ordinal>.<hash>` content refs
(mirroring the skill-snippet ref convention).
The envelope crate is a leaf with no ironclaw dependencies, satisfying
the boundary contract; the existing `ironclaw_hooks` boundary rule in
`reborn_dependency_boundaries` continues to hold because
`ironclaw_prompt_envelope` is not on its forbidden list.
Test count delta:
* `ironclaw_prompt_envelope`: +13 new tests (crate did not exist).
* `ironclaw_hooks`: 84 → 88 tests (+4 prompt-port behavior tests:
`hook_patch_appended_as_envelope_wrapped_message`,
`total_byte_budget_enforced_across_patches`,
`instruction_hijack_in_patch_rejected`,
`trusted_hook_patch_wrapped_with_trust_marker`).
* `ironclaw_host_runtime` memory_context: unchanged (8 tests still pass).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix: align tenant-counter test with SanitizedArguments-extended context ctor
* docs(reborn): document loader contract; pin HookId hex format
Add a "Loader responsibility" section to ironclaw_hooks/CLAUDE.md
explaining that tier-specific installers prevent minting wrong-tier
impls but cannot enforce origin — that's the loader's job — and
recommending registry loaders type-tag extension hooks as
LoadedHook::Installed at the loader seam.
Add tier_specific_installers_are_documented_as_loader_contract as a
regression guard that touches every public install_*_before_capability
and install_*_before_prompt method so any signature change forces the
loader contract to be re-evaluated.
Document HookId::to_hex's 64-char lowercase hex output as part of the
cross-crate contract consumed by LoopHostMilestoneKind::Hook* in
ironclaw_turns; add hook_id_hex_format_is_stable_64_lowercase_chars in
identity::tests and hook_id_string_serialization_matches_to_hex in
telemetry::tests to pin the format and the seam conversion path.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(reborn): pin hook milestone JSON schema + assert pairing invariants
Add L3 schema-snapshot tests for every hook-related LoopHostMilestoneKind
variant (HookDispatched, HookDecisionEmitted per HookDecisionSummary,
HookFailed per FailureCategory) so downstream consumers can rely on the
JSON wire shape and any accidental field rename, enum-tag rename, or type
change fails loudly.
Add L4 pairing-invariant matrix test in the hook dispatcher that drives
every observable outcome (Allow, Deny, PauseApproval, PauseAuth, Pass,
Panic, Timeout, Malformed, MissingImpl) through a recording milestone
sink and asserts the dispatched-then-terminator pairing shape. Document
the MissingImpl path as the one case that emits a sole HookFailed with
no preceding HookDispatched (the dispatcher discovers the protocol
violation before the hook is actually dispatched).
Add a multi-hook dispatch test that installs three hooks with mixed
outcomes (allow/deny/panic) at the same point and asserts each hook
produces its own paired sequence in the deterministic
(phase, priority, hook_id) order taken from the dispatcher's registry.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(reborn): integration tests for observer middleware through RebornLoopDriverHostFactory
Wire the HookedLoopModelPort / HookedLoopTranscriptPort /
HookedLoopCheckpointPort observer wrappers into
RebornLoopDriverHostFactory::build_text_only_host_with_capabilities,
mirroring the existing HookedLoopCapabilityPort / HookedLoopPromptPort
composition. The wrappers are applied only when a HookDispatcher is set
on the factory, so the default factory shape is unchanged.
Add four integration scenarios in crates/ironclaw_reborn/tests/hooks_integration.rs:
- observer_hook_fires_after_model_through_factory
- observer_hook_fires_after_capability_through_factory
- observer_hook_fires_after_checkpoint_through_factory
- observer_panic_does_not_fail_model_call (panic-isolation regression)
Relax the test-fixture model gateway from "panic if invoked" to
returning a stub assistant reply so the AfterModel / panic-isolation
tests can drive stream_model through the wrapped port. The existing
capability-port tests never touch the gateway, so their behavior is
unchanged.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor(reborn): introduce HookDispatcherBuilder for type-enforced sink wiring
Adds a `HookDispatcherBuilder` in `ironclaw_hooks::dispatch` that owns
the dispatcher construction lifecycle: registry -> optional timeout ->
optional milestone sink -> installed hooks -> `.build_arc()`. The
terminal `.build_arc()` wraps in `Arc` and yields an immutable handle.
Tightens the public surface on `HookDispatcher`: `new`, `with_timeout`,
`with_milestone_sink`, and every `install_*_*` method are now
`pub(crate)`. Outside callers route exclusively through the builder, so
"wire the milestone sink before Arc-wrapping" is a compile-time fact
rather than a documentation convention.
`HookRegistrar::install` now takes a `HookDispatcherBuilder` by value
and returns `(HookDispatcherBuilder, Vec<HookId>)`, keeping the builder
chainable through manifest installation.
`RebornLoopDriverHostFactory` gains `with_hook_dispatcher_builder` to
let callers defer `.build_arc()` to the factory — a step toward the
FU8 per-build dispatcher pattern.
Migrates `foundation_pipeline.rs` and `hooks_integration.rs` to the
builder. Internal middleware and dispatch tests continue to use the
crate-private `HookDispatcher::new` directly.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(reborn): production CapabilityInputResolver for NumericSum predicates
Adds HookCapabilityInputResolverAdapter in ironclaw_reborn that bridges
the existing LoopCapabilityInputResolver (already used by
HostRuntimeLoopCapabilityPort for dispatch input resolution) to the
hooks crate's CapabilityInputResolver trait. RebornLoopDriverHostFactory
gains with_capability_input_resolver(...), and when both a hook
dispatcher and resolver are configured the factory threads the adapter
into HookedLoopCapabilityPort::with_resolver — so NumericSum and other
argument-dependent predicates evaluate against real, sanitized inputs
instead of failing closed against the framework's null default.
The adapter also enforces a configurable serialized-byte budget
(default 64 KiB) as defense in depth ahead of the hooks crate's
per-string and depth caps in SanitizedArguments.
Unit tests cover the four adapter branches (resolved JSON,
inner-error → None, non-object pass-through, oversized → None) and a
new end-to-end integration test
(numeric_sum_predicate_caps_total_value_against_real_inputs) drives the
full factory wiring: with a NumericSum cap of 99 over an "amount" field,
two invocations carrying {"amount":"50"} let the first pass through and
deny the second at the hook seam, with the inner port reached exactly
once.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(reborn): per-build HookDispatcher for full per-run isolation (C2)
Introduce `with_hook_dispatcher_factory(F)` on
`RebornLoopDriverHostFactory`. The closure is invoked once per
`build_text_only_host*` call, so dispatcher-owned mutable state — slot
poisoning, registry mutations, predicate-counter siblings — is scoped to
a single host build instead of shared across every host the factory
produces.
The legacy `with_hook_dispatcher(Arc<HookDispatcher>)` adapter is kept as
a thin wrapper that returns clones of the same `Arc` on every build. Its
shared-state behavior is now documented as an explicit opt-in for
backward compat; new wiring should prefer the factory closure.
Adds two regression tests:
- `per_build_dispatcher_state_does_not_leak_across_runs` — installs a
panicking hook, builds two hosts back-to-back, and proves the inner
port is never reached on build 2 (fresh slot still applies the
fail-closed deny). Pins per-run isolation.
- `legacy_with_hook_dispatcher_shares_state_across_builds` — pins the
shared-state semantic of the legacy adapter as the explicit baseline.
Migrates `predicate_deny_hook_short_circuits_inner_port` to the new
factory-closure path so the new wiring is exercised by the existing
suite.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(reborn): project hook telemetry milestones into RuntimeEvent for durable audit
Extend the runtime event substrate with `HookDispatched`, `HookDecisionEmitted`,
and `HookFailed` kinds carrying closed-vocabulary labels and the blake3-hex hook
identity. Project the matching `LoopHostMilestoneKind::Hook*` variants in
`DurableLoopHostMilestoneSink` so hook telemetry now lands in the same durable
event log as model/reply/loop milestones — SSE observers still see live hook
events, and audit replay can reconstruct the full hook trail.
- `ironclaw_events`: add hook variants to `RuntimeEventKind`, optional hook
fields on `RuntimeEvent` (`hook_id`, `hook_point`, `hook_trust_class`,
`hook_decision`, `hook_failure_category`, `hook_failure_disposition`),
typed constructors (`hook_dispatched`, `hook_decision_emitted`,
`hook_failed`), and dedicated sanitizers (`sanitize_hook_label`,
`sanitize_hook_id`) re-run on every wire crossing. No new crate dependency
edges; hook strings cross the boundary opaque.
- `ironclaw_reborn::milestone_events`: project the three hook milestone kinds
via a new `loop.hook` capability id. `HookDecisionSummary` is collapsed to
its closed-vocabulary `kind_name()` so sanitized reasons never enter the
durable substrate.
- `ironclaw_event_projections`: extend `TimelineEntryKind` and the
`RuntimeEventKind -> RunProjectionStatus` mapping so hook events are pure
telemetry — they preserve the current run status rather than changing it.
- Tests: 4 unit tests in `ironclaw_events::runtime_event::tests` (serde
round-trip per variant + unsafe-label collapse), 3 in
`ironclaw_reborn::milestone_events::tests` (projection per variant,
including the assertion that raw `Deny { reason }` text does not reach the
durable wire payload). Existing replay-projection direct-construction
tests updated for the new RuntimeEvent fields.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(reborn): enforce manifest-declared hook scope at dispatch time (C3)
Audit finding C3: extensions could declare `[[hooks]]` with
`scope = "own_capabilities"` in their manifest, but the dispatcher never
enforced it — an Installed hook from ext-A could fire against capabilities
provided by ext-B. Scope was parsed but not load-bearing.
This change makes scope load-bearing end-to-end:
- `BeforeCapabilityHookContext` carries an optional `provider:
ironclaw_host_api::ExtensionId` populated by the middleware. The hook
context is `#[non_exhaustive]` already so this is non-breaking.
- `HookBinding` gains `owning_extension: Option<ExtensionId>` and `scope:
HookBindingScope`. `HookBindingScope` is `Global` / `OwnCapabilities`
/ `SameTenant`. Builtin and Trusted bindings default to `Global` and
carry no `owning_extension`; Installed bindings carry both, sourced
from the manifest.
- `HookDispatcher::install_installed_*` installers now require the
caller to pass `(owning_extension, scope)`. The registrar derives both
from the manifest entry, so manifest authorship is the single source
of truth.
- A new `CapabilityProviderResolver` trait + bundled
`NullCapabilityProviderResolver` lets the middleware lift the
capability id to its provider at invocation time. The middleware
wires the resolved provider into the hook context.
- `dispatch_before_capability` consults `binding.scope.permits(...)`
before invoking each hook. Bindings that don't permit the current
invocation are inert — no sink call, no failure record, no poisoning.
Conservative defaults:
- When the provider resolver returns `None` (no resolver wired, or the
capability has no known provider), `OwnCapabilities`-scoped hooks do
NOT fire. An attacker cannot bypass scope filtering by stripping
provider info from the descriptor.
Tests:
- 5 new dispatcher tests cover OwnCapabilities matching, foreign
provider, unresolved provider, SameTenant, and Builtin Global.
- 1 new registrar test asserts manifest scope and extension propagate
into `HookBinding`.
- 1 new middleware test asserts the provider resolver populates the
hook context.
- 1 new integration test in `ironclaw_reborn` proves an ext-A hook
scoped to `OwnCapabilities` does not intercept invocations that have
no resolved provider (the production composition default).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* style: rustfmt dispatch.rs after FU1 merge
* docs(hooks): prior-art comparison against LSM/eBPF/Envoy/K8s/OPA/CRX/VSC/Tauri
Validates the IronClaw hooks design against 8 established hook/policy
systems across 8 axes (dispatch, trust tiers, attenuation, decision
vocabulary, failure semantics, isolation, manifest, audit).
Surfaces:
- 7 areas where ICLAW stands out vs prior art (type-level trust
enforcement, dispatch-time scope, failure-kind matrix, pause-with-
gate-ref, pairing-invariant audit matrix, tenant-keyed predicates,
phase-ordered dispatch)
- 4 conventional choices we should revisit (in-process Installed-WASM,
sticky poison, no formal dispatch model, no installation rate-limit)
- 3 divergences whose 'why' is weak and need design review
* docs(hooks): STRIDE threat model for v1 framework
Enumerates 7 adversary classes (A1-A7), 6 assets ranked by blast
radius, and ~35 attack vectors across STRIDE categories with mitigations,
existing tests, and residual risk.
Surfaces 7 prioritized follow-ups:
- High: per-extension hook-count cap (D3/D4)
- High: gate-ref unguessability + one-shot test (S1)
- Med: resolver field-level scope (I2)
- Med: per-evaluator state ceiling (D5)
- Med: poison-stickiness operator runbook
- Low: timing side-channel residual acknowledgement (I4)
- Low: instruction-marker denylist periodic review (I5)
Confirms the load-bearing 'Installed cannot Allow' (E1) property holds
via type-level seal + tier-specific installers, backed by
compile_time_seal_test and installed_binding_cannot_be_paired_with_
privileged_impl tests.
Explicit out-of-scope: extension install pipeline (#3492), WASM exec
sandbox (needs separate threat model when it lands), approval gateway
(#3564).
* feat(hooks): close threat-model gaps S1 (gate-ref entropy) and D3/D4 (registration flood)
S1 (gate-ref unguessability, factory side):
- Three new tests on `UuidHookGateRefFactory`:
- `gate_refs_are_v4_uuids` pins the v4 entropy source (122 random
bits per ref per RFC 4122 §4.4); fails if a future change moves to
a counter or weaker UUID version.
- `gate_refs_have_no_collisions_across_many_calls` mints 20k refs
across both namespaces, asserts zero collisions (statistical
proxy for entropy quality).
- `approval_and_auth_namespaces_do_not_overlap` confirms prefix
routing separation.
- Doc comment now documents the security property explicitly and
delineates factory-side vs gateway-side responsibilities for the
one-shot consumption property.
D3/D4 (hook registration flood):
- New `MAX_HOOKS_PER_EXTENSION = 32` and
`MAX_HOOKS_PER_EXTENSION_PER_KIND = 8` consts in `registrar.rs`.
- New `HookRegistrar::enforce_registration_caps` runs pre-flight at
the top of `install()`, before any binding is inserted. Whole-batch
rejection means a partially-installed batch cannot slip past.
- Three regression tests: total-cap rejection, per-kind-cap rejection,
at-cap acceptance.
- Error messages cite the threat-model finding so operators can map
rejection back to the design rationale.
Threat model updated: S1, D3, D4 marked closed in the cross-cutting
properties matrix and the open-follow-ups list.
* test(hooks): three real hooks built against the public API + ergonomics findings
Builds three representative hooks from outside the crate, mimicking
what an extension or system author would actually write:
1. polymarket-daily-cap — Installed predicate hook, InvocationCount
rate-cap with Deny on excess. Canonical 'rate-limit a capability'
use case for the predicate language.
2. large-stake-approval-gate — Installed predicate hook, NumericSum
over amount_usd field, PauseApproval at $1000/24h. Manifest-shape
+ registrar-install coverage from outside Reborn; end-to-end
dispatch lives in ironclaw_reborn integration tests because
NumericSum needs resolved args (a friction finding documented in
the companion doc).
3. pii-redaction-warning — Trusted Rust hook implementing
PrivilegedBeforePromptHook, injects a trusted instruction snippet
reminding the model to redact PII. Demonstrates the path a system
author takes when the predicate language isn't expressive enough.
API change (F1 fix): SanitizedArguments::unresolved() promoted from
pub(crate) to pub. This is the documented safe default — predicates
that need args must fail closed against it — so exposing the
constructor cannot weaken any trust property. The sanitizing
from_json constructor stays sealed; that's the trust boundary.
Without this fix, external hook authors could not construct a
BeforeCapabilityHookContext with both a known provider AND
unresolved args, which made TDD of their own predicate impossible.
Findings documented in docs/real-hooks-findings.md, ranked by
severity. Big-picture observation: writing the Trusted Rust hook
(F4) was easier than writing the declarative predicate hook (F1 +
F2 + F3) — three of seven findings target predicate-authoring
ergonomics. The declarative path needs the most polish before
third-party extension authors will trust it for non-trivial policy.
Tests: 6 new in real_hooks.rs, all pass.
* feat(hooks): close all remaining threat-model and ergonomics gaps
Closes the Med-priority threat-model gaps (I2, D5, poison runbook)
and all real-hooks ergonomics findings (F2, F3, F5, F6, F7) in a
single pass.
Threat model:
- I2 (resolver field-scope): documented in SanitizedArguments rustdoc.
The narrow public surface (only is_resolved + extract_numeric)
enforces field-scope by construction for the current predicate
path. Reassess when Installed-WASM lands.
- D5 (evaluator state ceiling): MAX_HISTORY_KEYS = 8192 per map,
LRU eviction with evictions_observed() metric for operator
monitoring. New regression test
lru_eviction_increments_counter_and_drops_oldest_key.
- Poison-stickiness runbook: new docs/operator-runbook.md with
recovery options ranked by cost.
Ergonomics findings:
- F2 (closed-vocab deny reasons): rustdoc on OnExceededAction
and GateDecisionView::Deny explaining the audit-vs-model split
and why manifest reason text doesn't reach the model.
- F3 (NumericSum can't be TDD'd outside Reborn): new test-support
feature flag with SanitizedArguments::for_tests(value) that
external hook authors can opt into via dev-dep.
- F5 (two ExtensionId types): added
From<&ironclaw_host_api::ExtensionId> impl for
identity::ExtensionId, plus cross-link rustdoc.
- F6 (HookManifestEntry struct-literal fragility): added
#[non_exhaustive] + HookManifestEntry::new(id, kind, body) +
with_scope/with_phase/with_priority/with_description/with_requires_grant
builder methods. Migrated 3 external call sites in tests/.
- F7 (priority guidance): rustdoc on HookPriority with when-to-
deviate guidance, named FIRST/LAST constants documented for
Builtin/Telemetry use cases.
Tests: 151 unit + 1 + 6 integration in ironclaw_hooks all pass
with --all-features. ironclaw_reborn (13 hooks_integration scenarios)
unchanged.
Threat model updated: I2 / D5 / poison runbook marked closed in
both the per-vector table and the cross-cutting properties matrix.
Open follow-ups now down to two Low items (I4 timing side-channel
residual, I5 instruction-marker denylist refresh) plus the deferred
DenyReasonCode enum from F2.
* fix(ci): collapse nested match in hooks_integration test for clippy --all-features
CI runs `cargo clippy --all --tests --examples --all-features -- -D warnings`
which is stricter than the workspace clippy I ran locally and trips
`clippy::collapsible_match` on the nested-if in HookDecisionEmitted
matching. Collapse the inner `if decision.kind_name() == "deny"`
into an arm guard.
* feat(hooks): address henrypark133 review — Critical #1/#2/#3/#5, Concerning #5/#7
Address composition-seam bugs in the Reborn factory wiring + doc tidy.
henrypark133 review findings addressed:
Critical #1 — before_prompt hook messages not materialized.
HookedLoopPromptPort now requires a HookPromptMaterializationSink and
fails closed if patches are emitted without one. The reborn factory
installs an InstructionStoreBackedHookSink adapter that delegates to
the host's InstructionMaterializationStore, so synthetic msg:hook.*
refs are resolvable by the downstream model resolver. New seam trait
(HookPromptMaterializationSink) keeps ironclaw_hooks decoupled from
LoopRunContext.
Critical #2 — OwnCapabilities hooks were inert in production wiring.
Factory now installs SurfaceBackedProviderResolver (consults the
visible-capability surface for capability_id → provider). With this,
ctx.provider is populated and OwnCapabilities-scoped Installed hooks
actually fire against their own provider's capabilities.
Critical #3 — gate refs were unresolvable.
Middleware default switched from UuidHookGateRefFactory to
FailClosedHookGateRefFactory. Tests must explicitly opt into UUID
(via with_gate_ref_factory) to exercise the affirmative ApprovalRequired
path; production deployments must install a router-backed factory.
New factory method RebornLoopDriverHostFactory::with_hook_gate_ref_factory.
Concerning #5 — AfterModel fired twice + before durable finalization.
Removed AfterModel dispatch from HookedLoopModelPort; the transcript
port's finalize_assistant_message is now the sole AfterModel boundary
(the durable one). Model port wrapper is preserved as a no-op shim
for symmetry + future model-response-observed point.
Concerning #7 — doc tidy:
- CLAUDE.md: 3 trust classes → 4 (Builtin/Trusted/Installed/SelfAuthored
with explicit note that SelfAuthored is run-scoped only and not
loadable from an external source).
- operator-runbook.md: "Audit log" → "durable runtime event stream"
where the projection is actually the runtime-event stream, not formal
AuditEnvelope records.
- prior-art.md: poison-lifetime nuance — per-host-build with the
factory pattern, process-lifetime only for the legacy adapter.
- prior-art.md:80: trailing whitespace removed.
Testing gaps from henrypark133 — caller-level tests through
RebornLoopDriverHostFactory:
#1 (before_prompt resolver path):
before_prompt_hook_message_is_resolvable_via_factory_wiring
#2 (OwnCapabilities positive/negative/unknown):
own_capabilities_hook_fires_when_provider_matches
own_capabilities_hook_does_not_fire_when_provider_differs
own_capabilities_hook_does_not_fire_when_provider_unknown
#3 (pause/auth gate lifecycle or fail-closed):
pause_approval_with_default_factory_fails_closed_as_denied
pause_approval_hook_surfaces_as_approval_required_with_real_gate_ref
(updated to require explicit UuidHookGateRefFactory opt-in)
#5 (AfterModel exactly-once at durable boundary):
after_model_fires_exactly_once_at_durable_boundary
Still TODO from review (separate commits):
Critical #4 (telemetry context — two-run attribution) + gap #4
Concerning #6 (TimelineEntry hook metadata projection) + gap #6
Tests: 154 unit + 18 hooks_integration + all other reborn tests pass.
Workspace clippy + fmt + no-panics clean.
* feat(hooks): address remaining henrypark133 review — Critical #4, Concerning #6
Critical #4 — per-run hook telemetry attribution.
New `HookDispatcherBuilderFactory` signature: factory returns a
HookDispatcherBuilder, and `RebornLoopDriverHostFactory` attaches a
`RunScopedHookMilestoneSink` keyed to the CURRENT run's LoopRunContext
inside `build_text_only_host_with_capabilities`, before sealing the
dispatcher. The previous zero-arg signature relied on the closure
capturing run_context — silently misattributed across reuses; new
public API `with_hook_dispatcher_builder_factory` removes that
failure mode entirely. Legacy `with_hook_dispatcher_factory` retained
for back-compat (its sink-wiring contract stays caller-side).
Concerning #6 — TimelineEntry hook metadata.
Added 6 optional fields to `TimelineEntry` (hook_id, hook_point,
hook_trust_class, hook_decision, hook_failure_category,
hook_failure_disposition) and projected them from `RuntimeEvent::Hook*`.
Replay consumers now see which hook fired/failed, not just that some
hook event happened. Each field is closed-vocabulary (no free-form
reason text — that stays in the audit reason payload, not the
product replay DTO).
Testing gaps from henrypark133 — caller-level tests:
#4 (two-run hook telemetry attribution):
hook_telemetry_attribution_is_per_run_not_captured
Builds two hosts from the SAME builder factory closure with two
fresh LoopRunContexts. Asserts each run's hook milestones carry
its OWN run_id (no stale captured one).
#6 (replay projection contract for hook events):
hook_runtime_events_project_with_sanitized_hook_metadata
non_hook_runtime_events_project_with_no_hook_metadata
Constructs RuntimeEvent::Hook{Dispatched,DecisionEmitted,Failed}
and asserts the projection preserves the metadata fields. The
negative test guards against cross-contamination on non-hook
events.
All henrypark133 review items now addressed:
Critical: #1, #2, #3, #4 — done
Concerning: #5, #6, #7 — done
Testing gaps: #1-#6 — done
Tests: 154 unit + 19 hooks_integration in ironclaw_reborn + 61 reborn
unit + 38 + 2 new in ironclaw_event_projections + ... pass.
Workspace clippy + fmt + no-panics clean.
* docs(hooks): scope DenyReasonCode closed-vocabulary enum (successor #6)
Successor PR from #3573 — real-hooks ergonomics finding F2 (deferred).
Adds a curated vocabulary of model-visible denial reasons so hook
authors can communicate why a deny happened without opening a
free-form prompt-injection channel.
* feat(hooks): DenyReasonCode + PauseReasonCode closed-vocabulary enums
Address real-hooks ergonomics finding F2 (deferred from PR #3573). The
prior dispatcher collapsed every Installed-tier deny to the static
label 'hook_predicate_denied', because manifest reason strings are
author-controlled and surfacing them to the model would open a
prompt-injection channel. The cost: the agent couldn't tell *why*
a hook denied.
This PR introduces two closed-vocabulary enums:
- DenyReasonCode: Generic / RateLimit / ValueCap / Blocklist /
RequiresApproval / OutOfPolicy
- PauseReasonCode: Generic / RequiresApproval / OverThreshold /
SensitiveAction
Each variant has an as_label() returning &'static str (so the sink's
&'static str contract is preserved). New OnExceededAction variants
'DenyWithCode { code, reason }' and 'PauseApprovalWithCode { code,
reason }' let manifest authors opt into the richer labels while
keeping reason audit-only.
The legacy Deny { reason } / PauseApproval { reason } variants are
retained for back-compat and map to DenyReasonCode::Generic /
PauseReasonCode::Generic — existing manifests continue to produce
hook_predicate_denied / hook_predicate_pause_requested.
Threat-model regression: a hook author cannot smuggle text into the
model-visible label because the 'code' field is typed as the enum;
there's no String slot exposed model-side. A test
(deny_with_code_only_exposes_enum_variants_to_model) documents this
as a compile-time property.
Tests (+7 new = 161 total):
- deny_reason_code_labels_are_stable: pins the label vocabulary so
rename/relabel is loud.
- pause_reason_code_labels_are_stable: same for PauseReasonCode.
- deny_with_code_round_trips_through_json + pause variant: wire
round-trip + snake_case tag assertion.
- deny_with_code_only_exposes_enum_variants_to_model: compile-time
property check.
- rate_or_value_cap_with_deny_code_routes_to_code_label: end-to-end
affirmative test that the dispatcher emits the code's label.
- rate_or_value_cap_with_pause_code_routes_to_code_label: same for
pause.
Scope doc: crates/ironclaw_hooks/docs/successors/06-deny-reason-code.md
* test(hooks): address codex review on #3636
- Update stale real-hooks-findings.md F2 row to cite this PR's enum
follow-on (was 'deferred').
- Add install_deny_with_code_manifest_surfaces_code_label_on_dispatch:
end-to-end test driving the registrar->dispatcher path for the
new DenyWithCode variant (prior tests covered serde + direct hook
evaluation, but not the manifest install path that downstream
authors actually use).
Codex review on PR #3636: APPROVE with two recommendations; both
addressed.
Tests: 162 unit (+1 new). Clippy/fmt clean.
* fix(hooks): attenuate Installed-tier prompt patches to user role
Installed-tier `before_prompt` patches were injected as role:"system"
messages. Envelope text labels ("[ext-foo says]: ...") do not strip
system-role authority from the model's perspective, so a third-party
extension could inject system-tier instructions through a snippet
patch. This is a prompt-authority escalation against the trust
hierarchy the framework otherwise enforces.
Add `role_for_trust_class()` mapping Installed -> "user" and
Builtin/Trusted/SelfAuthored -> "system". Thread per-patch
trust_class through `wrap_patches_to_messages` and use it for the
emitted `LoopModelMessage.role`.
Tests:
- installed_hook_patch_drops_to_user_role: asserts the role for an
Installed-tier patch is "user"
- trusted_tier_hook_patch_keeps_system_role: regression that Trusted
tier still produces system-role content
* fix(hooks): enforce scope filter on observer dispatch + reject incompatible points
Two related defense-in-depth fixes against silent scope-filter failure:
1. The registry silently accepted Installed bindings with
`HookBindingScope::OwnCapabilities` at points (BeforePrompt,
AfterModel, AfterCheckpoint) whose dispatch context carries no
per-capability provider. The manifest's declared scope had no
effect at all — the hook fired against every dispatch. Reject the
binding at install time so the operator sees the misconfiguration.
2. `dispatch_observer_at` for `AfterCapability` did not consult the
binding's scope, so an Installed observer registered with
`OwnCapabilities` fired against every invocation regardless of
provider. Add `dispatch_observer_at_with_provider` carrying the
resolved capability provider; the capability-port middleware
resolves the provider once per invocation and threads it through
both the BeforeCapability hook context and the AfterCapability
observer dispatch. The dispatcher then enforces
`HookBindingScope::permits` on each observer binding.
`ObserverHookContext` gains a `provider: Option<ExtensionId>` field;
`#[non_exhaustive]` keeps existing authors compiling.
Tests:
- rejects_own_capabilities_at_before_prompt
- rejects_own_capabilities_at_after_model
- accepts_own_capabilities_at_before_capability
- own_capabilities_observer_filters_foreign_providers (covers
foreign / matching / unresolved provider)
* fix(hooks): preserve free-form audit reason alongside closed-vocab model label (serrrfirat #3636)
`PredicateBackedBeforeCapabilityHook::evaluate()` was discarding the
free-form `reason` from `EvaluatorDecision::{Deny, PauseApproval}`
with `..` and only sending `code.as_label()` into the sink. The
`HookDecisionEmitted` milestone therefore carried only the closed-
vocab label, and operator-visible audit/SSE context was silently lost
end-to-end. The fix splits the channels:
- Model sees the closed-vocab label (`hook_rate_limit`,
`hook_pause_over_threshold`, ...) via `sink.deny(label)`. This
channel is unchanged.
- Audit/SSE sees the manifest's free-form `reason` via a new
audit-only sink method `record_audit_reason(reason: String)`. The
recording sink captures it; the dispatcher reads it after the hook
returns and threads it into `LoopHostMilestoneKind::HookDecisionEmitted`.
Surface changes:
- `PrivilegedGateSink` / `RestrictedGateSink` gain
`record_audit_reason(String)` — accepts dynamic `String` (audit-only,
no model-facing seam) unlike the `&'static str` decision reasons.
- `RecordingGateSink` gains an `audit_reason: Option<String>` field.
- `GateHookOutcome::Decision` is now `Decision { decision,
audit_reason }`.
- `HookDispatcher::emit_decision_with_audit` threads the audit reason
into the milestone.
- `LoopHostMilestoneKind::HookDecisionEmitted` gains a
`#[serde(default, skip_serializing_if = "Option::is_none")]`
`audit_reason: Option<String>`. The durable RuntimeEvent projection
intentionally drops this field — audit reasons are operator-facing
in-memory SSE content, never durable cross-process surface.
Tests:
- `deny_with_code_records_audit_reason_separately_from_model_label`:
asserts the recording sink ends with `Deny { reason: "hook_rate_limit" }`
in `state` AND `audit_reason == Some("daily cap of $1000 ...")`.
* fix(hooks): remove unused model_request helper (CI clippy fix)
* fix(hooks): address serrrfirat P1/P2 findings on PR #3573
Three issues from the 5-15 review:
**P1 #1 registrar.rs:70 — `same_tenant` grants not enforced**
`HookManifestEntry::validate` only confirmed `requires_grant` was
present; the registrar then immediately installed the binding with no
host-verified grant context. A manifest could declare
`requires_grant = "anything"` and get a cross-extension binding for
free.
Fix: `HookRegistrar` now carries a `verified_grants: HashSet<String>`
(empty by default — default-deny). Add the host-facing setter
`with_verified_grants(...)`. At `install_one`, if
`entry.requires_grant` is `Some(g)`, require `g ∈ verified_grants` or
reject with a clear error. Tests:
- `install_rejects_same_tenant_without_verified_grant`
- `install_rejects_same_tenant_when_verified_grants_mismatch`
- The existing positive test
`installer_propagates_owning_extension_and_scope_from_manifest` now
wires the verified grant explicitly (proves the API contract).
**P1 #2 prompt_port.rs:150 — zip misalignment**
The materialization loop zipped surviving messages against the
ORIGINAL unfiltered patch list. `wrap_patches_to_messages` skips
metadata patches and over-budget snippets, so the zip silently paired
message[0] with patch[0] even when patch[0] was the skipped metadata
— materializing the wrong content (or none) under the snippet's
synthetic ref.
Fix: `wrap_patches_to_messages` now returns
`Vec<WrappedHookMessage { message, safe_content }>` — surviving
messages paired with their content by construction. The caller
materializes `entry.safe_content` under `entry.message.content_ref`
directly; no zip against unfiltered input. Removed the now-unused
`safe_content_for_patch` helper.
Test:
- `materialization_stays_aligned_when_metadata_patches_are_filtered`:
a hook emits `[metadata, snippet]`; asserts only one model message,
and the materialized content under its ref contains the snippet's
body — proves filtering can no longer desync from materialization.
**P2 #3 loop_driver_host.rs:1343 — `with_hook_dispatcher_builder`**
Docs said it deferred `build_arc()` to let the host factory finalize
wiring; the implementation called `build_arc()` eagerly and routed
through the legacy shared-dispatcher adapter, losing per-run
dispatcher isolation and the run-scoped milestone sink.
Fix: marked `#[deprecated]` with a note pointing callers to
`with_hook_dispatcher_builder_factory(|| ...)` for per-build
isolation, or `with_hook_dispatcher(...)` if they actually meant the
shared adapter. The method body is unchanged so no callers break;
they'll see the deprecation warning. No internal callers exist, so
the deprecation doesn't trip `-D warnings`.
All 162 hooks lib + 19 reborn integration tests pass; clippy clean.
* fix(hooks): address serrrfirat 3573-2026-05-15 review findings
P1 — prompt bundle authority mismatch (prompt_port.rs):
`HookedLoopPromptPort::build_prompt_bundle` called the inner port first,
which caused `HostManagedLoopPromptPort` to issue the prompt-bundle
authority grant against the pre-hook message list. The wrapper then
appended `msg:hook.*` messages to `bundle.messages`, so the downstream
model request hit `grant.messages != messages` and failed closed with
"model request messages do not match the host-built prompt bundle".
Add `with_bundle_authority(authority, run_context)` and re-issue the
grant after appending hook messages so it covers the post-hook bundle.
Reborn wires `prompt_authority.clone()` + `run_context.clone()` into
the wrapper at construction time.
P2 — observer installer accepts non-observer points (dispatch.rs):
`install_observer` accepted any `HookPointSpec` (including
`BeforeCapability` / `BeforePrompt`) and only populated the observer
map. Dispatch later found a binding without a gate/mutator impl and
fail-closed the capability with "binding present without installed
implementation". Reject non-observer points at install time so misuse
fails loudly rather than poisoning bindings at dispatch.
P2 — batch path skipped AfterCapability observers on inner error
(capability_port.rs):
The batch loop used `?` directly on `self.inner.invoke_capability(...)`,
which propagated the error before dispatching `AfterCapability`
observers. Failed batch entries disappeared from telemetry / audit,
while the single-invocation path dispatches observers on error.
Capture the inner result, dispatch observers, then propagate the error.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(hooks): address PR #3573 review feedback round 3
Addresses serrrfirat's CHANGES_REQUESTED review (2026-05-20) by tightening
several install-time / dispatch-time bounds and gating production seams:
- Bound free-form audit reasons crossing telemetry. New
`telemetry::sanitize_audit_reason` strips control characters and caps
length at 512 bytes; `emit_decision_with_audit` routes the manifest-
supplied reason through it before publishing milestones. Manifest
validation also rejects reasons over the same byte limit at install time
so the wire-side cap is a defense-in-depth layer, not the only line.
- Make hot dispatch O(H) instead of O(H^2). The per-binding poison
recheck used to acquire the registry mutex and walk every binding;
`ordered_bindings_with_poison_snapshot` now takes the active bindings
and the poisoned hook-id set under a single lock, and each loop
threads a local `HashSet<HookId>` that absorbs mid-dispatch
poisoning. Removed the redundant `is_poisoned` helper.
- Gate `HookDispatcher::registry_for_test` behind `cfg(any(test,
feature = "test-support"))`. The accessor previously exposed
`&Mutex<HookRegistry>` in production, letting any `Arc<HookDispatcher>`
holder lock and call `HookRegistry::poison` to disable installed
hooks. Added `active_bindings_snapshot(point)` as the read-only
production-safe replacement.
- `#[serde(deny_unknown_fields)]` on every hook-manifest and predicate
DTO (`HookManifestEntry`, `HookManifestBody`, `WasmBudget`,
`HookPredicateSpec`, `CapabilityPredicate`, `ValueOrRateBound`,
`OnExceededAction`). Typoed or unsupported fields (e.g. a
manifest-supplied `trust_class`) now fail loud at install time
instead of being silently dropped.
- Bound predicate trees at install. New
`validate_predicate_tree` enforces `MAX_PREDICATE_DEPTH = 8`,
`MAX_PREDICATE_NODES = 64`, `MAX_PREDICATE_STRING_BYTES = 256`, and
`MAX_MANIFEST_REASON_BYTES = 512`. A hostile registry manifest can no
longer install a deep or huge `All`/`Any` tree that the evaluator
would recursively walk on every match.
- Cap sliding-window samples per key. `MAX_SAMPLES_PER_KEY = 4_096` in
the predicate evaluator. Both the invocation-count and numeric-sum
histories drop the oldest sample once the cap is reached, bounding
memory under attacker-triggered hot capabilities while preserving
rate/value-cap semantics over the most recent window.
- `split_indexer` / `resolve_path` now fail closed on malformed bracket
syntax (`amount[foo]`, `amount[`, trailing garbage). Previously they
silently fell back to the parent field, which could let a typoed
`NumericSum` predicate evaluate against the wrong value and allow
calls the predicate would otherwise have denied.
- Honor `PatchOrdinalHint`. `WrappedHookMessage` carries the source
patch's `ordinal_hint`; `HookedLoopPromptPort` inserts `NearTop`
messages after the bundle's `identity_message_count` and appends
`Last` messages at the end. Safety/policy snippets that need early
placement now get it.
- Update `ironclaw_hooks` top-level docs to reflect the four trust
classes (`Builtin`/`Trusted`/`Installed`/`SelfAuthored`) and the
now-wired Reborn middleware composition.
Tests added:
- `manifest::rejects_unknown_top_level_field`
- `manifest::rejects_unknown_wasm_budget_field`
- `manifest::rejects_predicate_tree_exceeding_max_depth`
- `manifest::rejects_predicate_tree_exceeding_max_nodes`
- `manifest::rejects_predicate_string_exceeding_max_bytes`
- `manifest::rejects_manifest_reason_exceeding_max_bytes`
- `points::capability::malformed_indexer_returns_none_not_parent_value`
- `telemetry::sanitize_audit_reason_*` (truncate / strip control /
preserve / empty)
`cargo fmt`, `cargo clippy --all --benches --tests --examples
--all-features`, and `cargo test -p ironclaw_hooks` all pass clean.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(hooks): batch deferred test coverage from #3573 review (#3914)
* perf(hooks): defer capability input resolution until a predicate needs it (#3913)
* fix(rebase): adapt hooks tests + middleware to upstream API additions
- CapabilityDescriptorView: add parameters_schema field
- LoopModelRequest / LoopPromptBundleRequest: add capability_view field
- TimelineEntry test builder: add hook_id / hook_point / hook_trust_class /
hook_decision / hook_failure_category / hook_failure_disposition fields
- ironclaw_reborn::tests::hooks_integration: switch from
InMemoryLoopCheckpointStore to InMemoryTurnStateStore (which now
impls both LoopCheckpointStore and TurnStateStore), pass TurnActor
in TurnRunState, supply the new turn_state_store factory arg
- ironclaw_reborn lib.rs: drop the pub-use re-exports that upstream
intentionally removed (per the module-directory rationale in the
current ironclaw_reborn lib.rs doc comment); update the
hooks_integration test imports to use module paths
- Cargo.toml: union the hooks-foundation member list with upstream's
new crates (event_streams, auth, first_party_extensions,
reborn_webui_ingress, product_workflow_storage, webui_v2); drop
ironclaw_storage which no longer exists upstream
- crates/ironclaw_architecture/tests/reborn_dependency_boundaries:
keep upstream's removal of ironclaw_filesystem from the ironclaw_turns
forbidden list AND add ironclaw_hooks to that list
- crates/ironclaw_reborn/src/milestone_events.rs: drop dead loop_failure_kind
helper (replaced upstream by loop_failure_kind_name in text_loop_driver.rs);
keep hook_decision_label which is still used
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* perf(hooks): restore batched capability dispatch when hooks acti…
The "No panics in production code" CI check (scripts/check_no_panics.py) only recognizes `// safety:` suppression markers, not `RATIONALE:`. Since std::fmt::Write for String is infallible, just discard the Result with `let _ =` instead of `.expect()` — no panic call, no marker needed. Also merges in latest origin/hooks-foundation-01 (now includes the reborn-integration merge and PR #3636).
#3635) * docs(hooks): scope persistent predicate counter backend (successor #3) Successor PR from #3573. Current sliding-window state is in-memory and resets on restart. Adds a PredicateStateBackend trait + Postgres/libSQL impls for cross-process and restart-survival semantics. * feat(hooks): extract PredicateStateBackend trait + replay-safe in-memory impl Addresses codex review's three Critical findings on PR #3635: 1. Backend wiring: the trait is now registered (lib.rs:25-26) and PredicateEvaluator delegates to Arc<dyn PredicateStateBackend> via with_backend(...). Default constructor preserves the in-memory behavior so all 154 existing tests pass unchanged. 2. Atomic record-and-read: each record_invocation / record_value call performs the write AND returns the resulting in-window count/sum under a single mutex (in-memory) / transaction (durable backends). Splitting into separate record + read would let two hosts each see 'under cap' and both proceed, drifting past max. 3. Replay refusal: each record call carries a PredicateEventId. Re-emitting the same event_id is a no-op against the count. In-memory backend implements via a per-key bounded set (RECENT_EVENT_ID_CAP = 256); durable backends will use INSERT … ON CONFLICT DO NOTHING. Trait surface (predicate_state.rs): - PredicateEventId(String): opaque dedup key - PredicateBackendError: thiserror enum for fallible durable backends; in-memory backend never returns Err - PredicateStateBackend trait with Result return types - InMemoryPredicateStateBackend default impl - MAX_HISTORY_KEYS const re-exported via evaluator for back-compat Evaluator changes (evaluator.rs): - holds Arc<dyn PredicateStateBackend> (no more inline maps) - evictions_observed() reads through to backend - synth_event_id() generates per-call-unique ids via a process-local atomic counter so tests with identical (hook, ctx, now) still produce distinct ids - LRU helpers + HistoryKey/ValueHistoryKey types moved into predicate_state.rs (as InvocationKey/ValueKey) Tests: - 6 new predicate_state tests: - in_memory_invocation_counts_within_window - in_memory_invocation_trims_outside_window - in_memory_value_sums_within_window - in_memory_tenant_isolation (regression on threat-model C2) - in_memory_duplicate_event_id_is_a_noop_for_invocations - in_memory_duplicate_event_id_is_a_noop_for_values - 160 unit tests pass total. Reborn hooks_integration unchanged at 19 scenarios. Clippy/fmt/no-panics clean. Sync trait + Instant timestamps documented as a v1 choice; durable backends (Postgres, libSQL) will need an async companion trait using SystemTime — tracked in the scope doc as the next slice. Scope doc: crates/ironclaw_hooks/docs/successors/03-persistent-counter.md * fix(hooks): close codex P1 bugs in PredicateStateBackend in-memory impl Addresses codex P1 review on PR #3635: P1 #1 — replay dedup loss under high-throughput keys The prior design used a fixed-size (256) recent_ids ring per bucket decoupled from entries. Under any workload with >256 distinct events in the same window, the first event's id aged out of the ring while its timestamp entry was still live, so a replay silently re-counted. Fix: dedup memory is now intrinsic to entries. Each entry stores (timestamp, event_id), and the dedup check is 'does any in-window entry have this id?'. Dedup memory is therefore exactly the in-window entry set — no fixed cap, no silent loss. P1 #2 — zombie buckets clogging LRU Two-part fix: 1. record_* drops empty buckets eagerly via history.remove(key). This is mostly defense-in-depth — under the new dedup design, the record path can't actually leave a bucket empty (proved in the test rationale comment). 2. evict_lru_* now preferentially targets empty buckets first (find any v.entries.is_empty()), only falling back to the oldest-timestamp scan if no empty bucket exists. Filter-out behavior is gone, so any empty bucket that somehow survives becomes the next eviction victim instead of a permanent zombie. Test changes (+2 new, -0 removed): - dedup_memory_covers_full_window_under_high_throughput: pushes 512 distinct events into one bucket, then replays event-0. Pre-fix this would have counted again (silent dedup loss); post-fix the replay is a no-op. - lru_evicts_empty_buckets_first: crafts an empty bucket alongside a live one, runs LRU eviction, asserts the empty one is evicted and the live one retained. Tests: 162 unit total (+2 new). Clippy/fmt/no-panics clean. * docs(hooks): address gemini review on persistent-counter scope doc Four medium-priority doc nits from gemini-code-assist on the crates/ironclaw_hooks/docs/successors/03-persistent-counter.md scope: 1. run_id in the trait: removed. The trait dedupes on event_id (RuntimeEventId is already run-scoped), not run_id. Replaces the earlier 'backend stores (timestamp, run_id, event_id)' claim. 2. SystemTime vs chrono::DateTime<Utc>: switched to DateTime<Utc> to match project convention (src/db/mod.rs, ironclaw_events). The in-memory backend keeps Instant for monotonic process-local semantics; durable backends require DateTime<Utc> for cross- process serialization. Documented as a clock note. 3. libSQL TEXT column for rust_decimal: per src/db/CLAUDE.md, libSQL can't preserve Decimal precision with numeric/real types. LibSqlPredicateStateBackend serializes value as TEXT via Decimal::to_string() / from_str(). Postgres impl keeps numeric (correct for PG). Documented as the two LibSql-specific schema differences. 4. Batched-writes vs cross-process consistency tension: gemini was right that deferring writes to the tick boundary breaks requirement #1 (two hosts would each see 'under cap' simultaneously). v1 production backend keeps writes synchronous; future optimization batches reads (not writes). * fix(hooks): thread stable caller_event_id through hook context (replay dedup) henrypark133 HIGH on PR #3635 + serrrfirat HIGH #1: the `PredicateBackedBeforeCapabilityHook -> PredicateEvaluator` path always synthesized a fresh `event_id` per evaluation by mixing in a process-local atomic counter, so the same logical invocation retried/replayed always got a different id. The backend's UNIQUE constraint on `event_id` — the load-bearing dedup contract — never engaged on the real production path. Replay dedup was effectively "documented but unused." Plumb a stable per-invocation identity through the public hook surface: - `BeforeCapabilityHookContext` gains a `caller_event_id: Option<PredicateEventId>` field. Middleware that threads through from the calling layer's runtime event identity populates `Some(...)`; older / in-memory-only callers pass `None` and degrade to the current synth path (no behavior change). - New builder method `with_caller_event_id(...)`. - `PredicateEvaluator` resolves the id through a new `resolve_event_id` helper: prefer `ctx.caller_event_id`, fall back to `synth_event_id`. Both `record_invocation` and `record_value` paths use it. - Backend dedup behavior is unchanged — it was already correct on `event_id`. The bug was the caller path never supplying a stable id. Tests (caller-boundary, henrypark133's required regression): - `duplicate_caller_event_id_is_deduped_in_invocation_count`: two evaluations with the same `caller_event_id` count as one invocation; a third with a different id counts as two; a fourth crosses the cap. Sanity branch confirms the no-id synth path still exhibits "every call counts" semantics. This is the API contract slice. Wiring the middleware to actually supply a stable id (e.g. derived from the originating `RuntimeEventId` once that runs through the BeforeCapability path) is the follow-up that lights up the durable backend's end-to-end replay-safety promise. * refactor(hooks): demote PredicateStateBackend to pub(crate) (serrrfirat MED on PR #3635) serrrfirat MED: the `predicate_state` module exposed `PredicateStateBackend` as `pub`, but the trait's `now: Instant` parameter is process-local and not serializable. Any external durable backend impl built against the current trait would have to be rewritten when the durable contract lands with `chrono::DateTime<Utc>` (see successor doc 03-persistent-counter.md). Hold the public surface back until that contract is stable so we don't ship a public API we know we'll break. Demoted to `pub(crate)`: - `PredicateStateBackend` (trait) - `InvocationKey`, `ValueKey` (key types — backend ABI only) - `PredicateBackendError` (error type, with `#[allow(dead_code)]` on the `Unavailable` variant since the in-memory backend is infallible and no durable backend exists yet) - `InMemoryPredicateStateBackend` (the only impl) - `PredicateEvaluator::with_backend` (with `#[allow(dead_code)]` — reserved for future internal injection paths) Kept `pub`: - `PredicateEventId` — it appears on the public hook surface via `BeforeCapabilityHookContext::caller_event_id` (from the #3635 HIGH fix). Hook authors who want stable replay-dedup ids construct one. No behavior change. All 163 hooks lib tests + 19 reborn integration tests still pass. * fix(hooks): address henrypark133 must-fix #1-5 on PR #3635 Five items from the 5-15 review: **#1 (must-fix) O(n) dedup scan** The previous `bucket.entries.iter().any(...)` linear scan held the outer history mutex while walking thousands of in-window entries at high throughput. Add a companion `HashSet<PredicateEventId>` per bucket (`InvocationBucket.dedup_ids` / `ValueBucket.dedup_ids`), maintained alongside the deque via `pop_front`/`push_back` helpers. O(1) dedup, same correctness, same memory bound (one set entry per in-window entry — no fixed ring). **#2 (must-fix) Mutex poison cascade** `.expect("predicate history mutex poisoned")` propagated a panic to every subsequent caller. Replace with `match self.invocation_history.lock() { Ok(g) => g, Err(p) => p.into_inner() }` so a poisoning thread doesn't take down all subsequent evaluations. **#3 (must-fix) `caller_event_id` format validation** `with_caller_event_id` now rejects empty strings and ids containing NUL bytes. Failed validation logs a `tracing::warn!` and leaves `caller_event_id == None` so the synth path takes over — operator sees the warning, predicate dedup still works. Also: `PredicateEventId(pub String)` → `PredicateEventId(String)` with `new()` / `as_str()` (henrypark133 nit #9). Inner field is no longer in-place mutable from outside the crate. **#4 (must-fix) `with_backend` is `#[cfg(test)]`** Previously `#[allow(dead_code)]` — reachable from release builds and inviting future callers to inject backends through an unstable seam. Gated to `cfg(test)`. **#5 (important) `evict_older_than` trait stub** Default-impl no-op added to `PredicateStateBackend` so the trait signature is locked before the first durable-backend PR. Trait-object callers won't break when durable impls override it. **Bonus** (henrypark133 missing-coverage #1): `in_memory_record_invocation_is_atomic_under_concurrent_writers` — 32 threads each record a distinct event id; final count must equal 32, proving the atomic record-and-read contract holds under contention. **Bonus** (henrypark133 nit #10): The third stable id in `duplicate_caller_event_id_is_deduped_in_invocation_count` was 62 chars; bumped to 64 to match the synth output format. * fix(hooks): clippy doc-list-indentation + remove unused with_backend (#3635 CI) * fix(hooks): address serrrfirat HIGH + MEDIUM on PR #3635 (5-15 review) **MEDIUM — `caller_event_id` validation bypass** `with_caller_event_id` validated for empty/NUL but the field on `BeforeCapabilityHookContext` is `pub`, so callers could direct- assign `Some(PredicateEventId::new("..."))` with `new()` permissive and bypass the setter entirely. Move validation INTO the type boundary: - `PredicateEventId::new(...) -> Result<Self, PredicateEventIdError>` validates non-empty + NUL-free at construction. Any value that reaches a downstream backend now satisfies the format invariant by construction. - `PredicateEventId::new_unchecked(...)` for internal synth paths and tests that mint ids from known-good shapes (hex digests). - `with_caller_event_id` drops its now-redundant runtime check; the type already enforces it. - Internal synth in `evaluator.rs` switches to `new_unchecked` (64-char hex output is always valid by construction). Tests: - `predicate_event_id_rejects_empty` - `predicate_event_id_rejects_nul_bytes` - `predicate_event_id_accepts_typical_hex_digest` **HIGH — durable schema: dedup scope mismatch** The successor doc's Postgres schema declared `event_id uuid PRIMARY KEY` (globally unique), but the trait's replay-refusal contract dedupes within the counter `key`. `caller_event_id` is per capability invocation — two predicate-backed hooks observing the same invocation share an id. A global PK lets the first hook's INSERT win and silently undercounts the second hook's bucket. - `docs/successors/03-persistent-counter.md`: PK changes to composite `(tenant_id, hook_id, capability, event_id)` for invocations and `(tenant_id, hook_id, capability, field, event_id)` for values, matching the trait's per-key dedup scope. - `predicate_state.rs` trait doc: replay-refusal section rewritten to spell out the per-key scope and the corresponding `INSERT … ON CONFLICT (tenant, hook, capability[, field], event_id) DO NOTHING` shape durable backends should use. * docs(hooks): document host-assigned trust boundary on PredicateEventId henrypark133 / serrrfirat blocker B4 on PR #3635: the `caller_event_id` threading through `BeforeCapabilityHookContext` partially shipped earlier (commit b4d8a35), but the trust-boundary documentation explaining the host-assigned invariant was still missing. Add rustdoc to `PredicateEventId` and the `PredicateStateBackend` trait clarifying that: - the id MUST be minted by trusted host code from authoritative sources (dispatcher RuntimeEventId, host-side hash, arguments digest) - it MUST NOT pass through unchanged from any tenant-controlled surface (capability arguments, manifest fields, WASM memory, HTTP bodies) - the format invariants in `PredicateEventId::new` (non-empty, NUL-free) are a durability contract for SQL backends, NOT a trust check - a tenant-supplied id can either undercount itself into infinity by replaying a fixed id, or poison adjacent buckets if scoping is ever weakened Doc-only; no behavior change. * test(hooks): add caller-boundary replay-dedup test through wrapper hook henrypark133 HIGH blocker B1 on PR #3635: replay dedup must engage at the caller boundary — `PredicateBackedBeforeCapabilityHook::evaluate` is the production path the dispatcher invokes for installed predicate hooks. A unit test on `PredicateEvaluator::evaluate_at` alone is insufficient regression coverage (repo CLAUDE.md rule "Test through the caller, not just the helper"): the wrapper hook reads `BeforeCapabilityHookContext::caller_event_id` and threads it down to the backend, so the regression test must drive the wrapper itself. The threading work already shipped in commit e6df47d (`caller_event_id` field on the public hook context + evaluator preferring it over the synth path). This commit adds the missing end-to-end test: 1. Two `PredicateBackedBeforeCapabilityHook::evaluate` calls with the same `caller_event_id` and a `RateOrValueCap { max: 1 }` predicate — the second call must stay under cap (dedupe engages at the wrapper boundary, not be re-counted into a deny). 2. A third call with a DISTINCT `caller_event_id` crosses the cap — proving dedup is replay-scoped (same id → no-op), not blanket- suppress (any id → no-op). If the wrapper were synthesizing a fresh id per call (the bug Henry flagged before threading landed), this test would fail at step 2 with the second evaluation being denied. * docs(hooks): D5a + cross-process replay note; add caller-API tests henrypark133 should-fix S8 + S9 on PR #3635. S8 — threat-model expansion: - Add D5a as the correctness-under-attack variant of D5: an attacker flooding high-cardinality keys can LRU-evict legitimate tenants' counters and reset their rate-limit state. Distinct from the memory-only framing of D5; tied back to per-extension caps (D3/D4) and the durable-backend successor (doc 03). - Document the cross-process replay limit on the in-memory backend inside the PredicateStateBackend trait docs, not just in D5 — the process-local dedup is a property callers need at the trait surface, with a pointer to the durable backend as the cross-host story. S9 — three new tests on the in-memory backend public API: - lru_eviction_via_public_api_holds_max_history_keys_cap: drives MAX_HISTORY_KEYS + 1 distinct keys through record_invocation and asserts the map size cap holds + evictions_observed() advances. The previous coverage manually crafted buckets and called the LRU helper directly; this exercises the production path. - in_memory_invocation_retains_entry_at_exact_window_cutoff: pins the `< cutoff` trim semantics so a refactor to `<=` would fail loud. - event_id_dedup_is_isolated_across_invocation_and_value_maps: same event_id used in both record_invocation and record_value must not cross-suppress — the two maps key on disjoint types. The fourth S9 item (concurrent N-thread atomicity) and the caller- boundary replay test on the wrapper hook already landed in earlier commits (f632d22, predicate_state.rs line 840). S2 (evict_older_than stub), S3 (sync-trait docs), and S7 (consistency vs batched-writes) were also already in HEAD; this commit ships the remaining items. Quality gate: cargo fmt clean, cargo clippy -p ironclaw_hooks --all-features --tests -D warnings clean, full hooks test suite green (15 predicate_state unit tests + lib + integration). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * chore(hooks): co-locate synth_event_id with backend; rationale comment; pin synth format henrypark133 nits N1, N2, N5 on PR #3635. N1 — Move `synth_event_id` from `evaluator.rs` to `predicate_state.rs` as `PredicateEventId::synth(...)`. The id format (64-char lowercase hex, no NUL, never empty) is part of the backend's durable contract, so co-locating with `PredicateStateBackend` keeps the format change- surface adjacent to the consumer. To avoid inverting the module dependency (`predicate_state` is a leaf below `points`), the synth helper takes raw bytes / &str rather than a `&BeforeCapabilityHookContext`. The evaluator's `resolve_event_id` fallback unpacks the context and delegates. N2 — `// safety:` comment on a non-`unsafe` block (the `write!(s, "{byte:02x}")` infallibility note) renamed to `// RATIONALE:`. By convention `// SAFETY:` pairs with `unsafe` blocks; using `// safety:` elsewhere conflates the two. N5 — Add `synth_event_id_is_64_char_lowercase_hex` to pin the synth output shape. A refactor that silently changes length or case would break the durable backend's `uuid`-shaped UNIQUE constraint without a test failure today; the new test fails loud. Quality gate: cargo fmt clean, cargo clippy --all --benches --tests --examples --all-features -D warnings clean, full hooks lib test suite green (172 passing including the new pin). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(hooks): drop expect() in hex formatting to satisfy panic CI check The "No panics in production code" CI check (scripts/check_no_panics.py) only recognizes `// safety:` suppression markers, not `RATIONALE:`. Since std::fmt::Write for String is infallible, just discard the Result with `let _ =` instead of `.expect()` — no panic call, no marker needed. Also merges in latest origin/hooks-foundation-01 (now includes the reborn-integration merge and PR #3636). * fix(hooks): narrow caller_event_id visibility to pub(crate) henrypark133 MED on PR #3635 5-19 review. The pub field let external callers bypass with_caller_event_id and assign values the validated PredicateEventId constructor would have rejected. Force every external caller through the typed setter so PredicateEventId::new is the only entry point. * fix(hooks): drop arguments_digest from synth + add in-memory backend warn Two PR #3635 5-19 review findings on the evaluator surface: - henrypark133 LOW (synth oracle): drop arguments_digest from the PredicateEventId synth hash input. The 64-char hex output was an equality oracle for argument shape; replay dedup for durable backends uses the caller-supplied caller_event_id, not the synth path, so synth only needs to be per-call unique, not content-addressed. - henrypark133 HIGH + MED (in-memory production limits): expose PredicateEvaluator::warn_in_memory_backend_active_in_production for hosts to call at startup. Multi-host replay dedup is process-local and the LRU cap is shared across tenants; operators need this surfaced in logs when the durable backend is not wired. * fix(hooks): harden predicate state backend per PR #3635 5-19 review Address five findings on crates/ironclaw_hooks/src/predicate_state.rs: - A1 (henrypark133 HIGH): restrict PredicateEventId::new_unchecked to pub(crate) so external callers cannot bypass the durable UNIQUE-constraint format invariants enforced by ::new. - A4 (henrypark133 MED): per-tenant LRU quota at MAX_HISTORY_KEYS / 4. Without it a noisy tenant could fill the global cap and evict a quiet tenant's bucket, resetting their rate-limit counter. With the quota, a tenant that overflows evicts its OWN oldest-front bucket first. New tests cover single-tenant cap and cross-tenant isolation. - A5 (henrypark133 LOW): drop arguments_digest from the synth hash input (oracle closure mirrored from the evaluator side). Add a thread-local nonce alongside the process-global counter so synth remains per-call unique without relying solely on a contended AtomicU64. New test pins the divergence invariant. - D6 (henrypark133 HIGH): O(1) NumericSum via an incrementally- maintained ValueBucket::running_sum, replacing the O(n) deque walk on every record_value call. New test covers push/trim/replay interactions. - D8 (henrypark133 MED): implement evict_older_than for the in-memory backend (was a no-op Ok(0) default). Drops entries strictly older than the cutoff and removes empty buckets; operator reaper tasks rely on this to reclaim memory from idle keys. - D7 (henrypark133 MED, partial): document the process-global synth COUNTER as a known contention hotspot and add a thread-local nonce so threads can advance without forcing cross-core invalidation in the common path. Tests: 196 passing (+4 new); workspace clippy clean. * fix(hooks): port MAX_SAMPLES_PER_KEY cap into PredicateStateBackend (D5 regression from r3) Round 3 of PR #3573 (already merged into hooks-foundation-01) added an inline per-key sample cap of 4_096 in evaluator.rs to bound memory under attacker-triggered hot capabilities with very large declared windows (threat-model finding D5). The predicate-state extraction in PR #3635 moved that bookkeeping into the PredicateStateBackend trait but missed porting the cap, so the cap would silently disappear from production once this PR rebases onto the foundation branch. This commit moves the cap into the in-memory backend impl next to MAX_HISTORY_KEYS / MAX_KEYS_PER_TENANT and enforces it in both record_invocation and record_value. For the NumericSum path, the bucket helper's pop_front already decrements running_sum, so the incremental sum invariant survives cap-driven eviction. The pre-existing inline copy in evaluator.rs becomes redundant once the trait impl owns the enforcement; the rebase resolution deletes it. Adds two regression tests: - record_invocation_caps_samples_per_key_under_attacker_pressure - record_value_evicts_oldest_keeping_running_sum_consistent * fix(hooks): port predicate_state tests to ::new() after #3912 newtype privatization --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…earai#3573) * feat(reborn): add ironclaw_hooks framework foundation (#3524) Foundation slice of the Reborn loop hooks framework per nearai/ironclaw#3524. Lands the trust primitives, sealed decision types, dispatcher contract, and extension manifest schema; no Reborn middleware composition yet (next slice wires HookDispatcher into LoopCapabilityPort / LoopPromptPort). Design comment on #3524: https://github.com/nearai/ironclaw/issues/3524#issuecomment-4439890144 What this PR ships ================== * `crates/ironclaw_hooks/` — new crate * `identity` — content-addressed `HookId` (blake3 of length-prefixed extension + local + version fields). Same versioning primitive the rest of Reborn should converge on for replay safety. * `trust` — `HookTrustClass` enum (Builtin / Trusted / Installed) with per-kind default attenuation. Trust class is fixed by source, never declarable. * `kinds/` — sealed decision DTOs. `BeforeCapabilityHookDecision`, `HookPatch`, `ObserverFact` all have `pub` outer struct + `pub(crate)` inner enum + `pub(crate)` constructors. Same #3460 witness pattern. * `points/` — typed read-only contexts for each hook point. * `sink` — split sink traits per trust tier. `PrivilegedGateSink` exposes `allow()`; `RestrictedGateSink` does not. An Installed-tier hook literally cannot mint Allow at the type level. * `ordering` — phase → priority → hook id, stable. Phases gated by trust (Validation/Authorization Builtin-only). * `failure_policy` — Timeout/Panic/Malformed/AttenuationViolation categories. Gate/Mutator fail closed, Observer/Effect fail isolated. Slot poisoning persisted for the rest of the run on any category. * `registry` — run-profile-sourced bindings; phase-vs-trust gate enforced at insert; poisoning surface for the dispatcher. * `dispatch` — HookDispatcher with deterministic ordering, panic catch-unwind via futures::FutureExt, per-hook tokio::time::timeout, short-circuit gate composition (Deny > PauseAuth > PauseApproval > Allow), Telemetry-phase observers always run. * `manifest` — serde types for the `[[hooks]]` section of extension manifests. Predicate vs WASM body; same_tenant scope requires explicit grant; Validation/Authorization phases rejected at parse time because manifest hooks are always Installed. * `predicate` — typed predicate language for declarative Installed hooks (DenyCapability, PauseApproval, RateOrValueCap). Evaluator lives in the dispatcher follow-up, not here. * `crates/ironclaw_architecture/tests/reborn_dependency_boundaries.rs` * Added `ironclaw_turns` -> `ironclaw_hooks` to the forbidden list. * New BoundaryRule for `ironclaw_hooks` itself (cannot pull host_runtime, dispatcher, secrets, network, wasm, etc.). * `Cargo.toml` workspace member registration. What this PR deliberately does NOT ship ======================================== * Reborn middleware composition wrapping LoopCapabilityPort / LoopPromptPort with HookDispatcher. Next slice; ironclaw_reborn changes only. * WASM hook execution path. Programmatic hooks parse and validate from manifest; the wasmtime integration lands when the WASM dispatcher seam is built. * Predicate evaluation. Predicate types serialize and validate; the evaluator that turns a `RateOrValueCap` spec into a `Deny` decision is in the next slice alongside Reborn wiring. * Event-triggered hooks (Phase 5 of the original roadmap). * Self-authored hooks. Tracked separately at #3567 with monotonic-restriction + unforgeable-channel ratification. Test plan ========= * `cargo test -p ironclaw_hooks` — 47 tests (46 unit + 1 integration smoke for the manifest -> binding -> dispatch pipeline). * `cargo test -p ironclaw_architecture` — 13 tests; new boundary rule passes, existing rules unaffected. * `cargo clippy -p ironclaw_hooks --all-targets -- -D warnings` — clean. * `cargo fmt -p ironclaw_hooks -- --check` — clean. * `cargo check --workspace` — clean, no regressions in other crates. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(reborn): wire HookDispatcher into LoopCapabilityPort/LoopPromptPort Follows the foundation slice (see initial commit). Adds the next layer: 1. Capability- and prompt-port middleware (`ironclaw_hooks::middleware`) * `HookedLoopCapabilityPort` runs `dispatch_before_capability` before every invocation, translates the composed decision into the existing `CapabilityOutcome` vocabulary (Deny / PauseApproval / PauseAuth all map to `Denied` for now; gate-ref plumbing for real pause semantics lands in the next slice). * `HookedLoopPromptPort` runs `dispatch_before_prompt` before bundle construction. Observe-only for snippets in this slice; actual snippet injection waits for the shared `prompt_envelope::wrap_untrusted` helper (#3540 / #3471). 2. Declarative predicate evaluator (`ironclaw_hooks::evaluator`) * `DenyCapability` and `PauseApproval` predicates: stateless, evaluated directly against `BeforeCapabilityHookContext`. * `RateOrValueCap` with `InvocationCount` bound: sliding-window counter keyed by `(hook_id, capability_name)`, in-memory only. Window parsing supports `s`/`m`/`h`/`d` units; unparseable windows fail closed. * `NumericSum` bound: types implemented but evaluation returns Allow and emits a warn-level audit. Full argument-extraction story is a follow-up slice once capability arguments become hook-visible. * `PredicateEvaluator::evaluate_at(...)` test variant accepts an explicit `Instant` so sliding-window tests don't depend on real-clock progress. 3. Manifest -> dispatcher glue (`ironclaw_hooks::installed_hook`) * `PredicateBackedBeforeCapabilityHook` wraps a `HookPredicateSpec` plus an `Arc<PredicateEvaluator>` and implements `RestrictedBeforeCapabilityHook`. The registry installer would construct one of these per `[[hooks]]` entry whose body is `HookManifestBody::Predicate`. * Sink reasons are `&'static str`, so the dynamic predicate `reason` surfaces in audit (via the evaluator's `EvaluatorDecision`) rather than the model-visible decision. Closed-vocabulary labels carry through to the sink. 4. Reborn composition seam (`ironclaw_reborn::loop_driver_host`) * `RebornLoopDriverHostFactory::with_hook_dispatcher(Arc<HookDispatcher>)` opt-in builder method. When set, the factory wraps the capability and prompt ports with the hooked middleware. Default behavior (no dispatcher) is unchanged from the pre-hooks shape, so existing callers continue to work. * Added `ironclaw_hooks` as a dep in `ironclaw_reborn`. Test plan ========= * `cargo test -p ironclaw_hooks` — 60 tests pass (59 unit + 1 integration smoke; +13 vs the foundation commit covering middleware, evaluator, installed_hook). * `cargo test -p ironclaw_reborn` — 118 tests pass; no regressions from adding the dep. * `cargo test -p ironclaw_architecture` — 13 tests pass; the `ironclaw_turns -> ironclaw_hooks` boundary still holds and the new `ironclaw_hooks` rule (no host_runtime / dispatcher / secrets / network / wasm / reborn) is unaffected. * `cargo clippy -p ironclaw_hooks --all-targets --all-features -- -D warnings` — clean. * `cargo clippy -p ironclaw_reborn --all-targets -- -D warnings` — clean. * `cargo fmt --all -- --check` — clean. What still defers ================== * WASM hook execution path. * Persistent predicate counter (in-memory only for now). * Argument-extraction so `NumericSum` predicates evaluate against capability arguments. * Gate-ref plumbing so PauseApproval / PauseAuth surface real `CapabilityOutcome::ApprovalRequired` instead of `Denied`. * Prompt-snippet injection (waits for shared envelope helper). * Event-triggered hooks. * Self-authored hooks (#3567). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(reborn): add HookedLoopModelPort/TranscriptPort/CheckpointPort observer middleware Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(reborn): end-to-end hooks integration through RebornLoopDriverHostFactory Adds crates/ironclaw_reborn/tests/hooks_integration.rs covering the factory's HookDispatcher wiring seam end-to-end. Tests drive host.invoke_capability(...) (not dispatcher.dispatch_before_capability(...) directly) so a regression in RebornLoopDriverHostFactory's wrapping composition surfaces here. Scenarios: - PredicateBackedBeforeCapabilityHook (DenyCapability NameEquals "cap.blocked") short-circuits invocation; inner port never called; outcome is Denied(unknown("hook_denied")). - A privileged selective hook that allows non-matching capabilities proves the wrapper does not blanket-deny: cap.allowed reaches the inner port and completes once. - Factory built without with_hook_dispatcher() lets cap.blocked through to the inner port, proving the hook plumbing is genuinely opt-in. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(reborn): add pass() + HookRegistrar + self-authored hooks scaffolding Three additions to ironclaw_hooks: B. `pass()` on gate sinks — distinguishes "evaluated, no opinion" from "returned without minting a decision." A passing hook contributes nothing to the composed decision; a silent hook is still Malformed and fails closed. `PredicateBackedBeforeCapabilityHook` now routes the evaluator's `Allow` decision through `sink.pass()` instead of the previous `deny("hook_predicate_pass")` workaround. A. `HookRegistrar` bridge — converts a `Vec<HookManifestEntry>` into `HookBinding`s + dispatcher impls in one call. Predicate bodies are wired through `PredicateBackedBeforeCapabilityHook`; WASM bodies return `HookError::RegistryConstruction` for now. Adds `HookDispatcher::insert_binding` so the registrar can mutate the registry through the dispatcher rather than reach inside. I. Self-authored hooks scaffolding — fourth `HookTrustClass` variant for hooks the agent authors at runtime. Run-scoped only; monotonic-restriction sink with no `allow`, no trusted-snippet path, no effect-class constructor. Closed-vocabulary `SelfAuthoredReason` enum keeps free-text reasons off the audit seam. `SelfAuthorshipProvenance` captures authoring run/turn, timestamp, spec digest, optional user ratification, and a generation-trace pointer. Durable persistence depends on the unforgeable channel from #3564 and lands separately. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(reborn): real gate-ref plumbing for hook PauseApproval/PauseAuth decisions Previously, `GateDecisionInner::PauseApproval` and `PauseAuth` returned by hooks were degraded to `CapabilityOutcome::Denied` at the middleware boundary because the hook crate had no way to mint a `LoopGateRef` scoped to the current run. Hooks that wanted to pause the loop for approval or auth instead failed the call closed, leaving the host's approval-router machinery unreachable from hook code. This change introduces a `HookGateRefFactory` trait in `ironclaw_hooks::middleware::gate_ref` that mints `LoopGateRef`s for pause-class decisions. `HookedLoopCapabilityPort` now takes an `Arc<dyn HookGateRefFactory>`, defaulting to `UuidHookGateRefFactory` (a locally-unique opaque-id factory suitable for tests and the foundation slice). Production deployments override via `.with_gate_ref_factory(...)` with a factory bound to the current `LoopRunContext` and the host's gate-router. The translation in `decision_to_outcome` is now async so it can await the factory. `PauseApproval` maps to `CapabilityOutcome::ApprovalRequired { gate_ref, safe_summary }` and `PauseAuth` to `AuthRequired`. If the factory itself errors, the middleware falls back to `Denied` with a sanitized `hook_gate_ref_unavailable` reason kind so the loop fails closed rather than routing through an unresolvable suspension. The underlying error text is dropped to avoid leaking gate-router state into model-visible output. Tests: - `pause_approval_decision_surfaces_as_approval_required`, `pause_auth_decision_surfaces_as_auth_required`, `gate_ref_factory_failure_falls_back_to_denied` in `middleware::capability_port::tests`. - `pause_approval_hook_surfaces_as_approval_required_with_real_gate_ref` in `crates/ironclaw_reborn/tests/hooks_integration.rs`, exercising the full `RebornLoopDriverHostFactory` composition with the default `UuidHookGateRefFactory`. - Gate-ref factory unit tests in `gate_ref::tests`. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(reborn): add NumericSum predicate evaluation with capability argument extraction Wires the missing argument-extraction story for the predicate evaluator so `ValueOrRateBound::NumericSum` actually enforces a rolling numeric cap instead of warn-and-allowing. - Extend `BeforeCapabilityHookContext` with a sealed `SanitizedArguments` view. Strings truncate to 256 bytes; objects/arrays cap at 8-deep. `extract_numeric` supports dotted + bracketed paths (`order.amount`, `items[0].price`) and returns `Option<rust_decimal::Decimal>`. The inner representation is sealed so external callers can't bypass bounds. - Introduce `CapabilityInputResolver` + bundled `NullCapabilityInputResolver` in `middleware/resolver.rs`. The hooks crate intentionally doesn't know how to dereference a `CapabilityInputRef` — that knowledge belongs to the production host. Until a real resolver is wired in (follow-up), arguments are `Unresolved` and `NumericSum` fails closed. - `HookedLoopCapabilityPort::new` defaults to the null resolver; new builder `.with_resolver(Arc<dyn CapabilityInputResolver>)` overrides. - `PredicateEvaluator` gains a tenant-keyed `value_history` map. The `NumericSum` arm parses `max` + `window`, extracts the numeric value from sanitized args, accumulates within the rolling window, and applies `on_exceeded` when the sum exceeds the cap. Unresolved args, missing field, non-numeric field, unparseable max, and unparseable window all fail closed via the configured `OnExceededAction`. - Add `BeforeCapabilityHookContext::new_unresolved(...)` convenience ctor; existing test sites switch to it instead of churning every call site through the 4-arg ctor. Test count: +14 (8 new SanitizedArguments tests, 6 new NumericSum evaluator tests, 1 null-resolver test; one old NumericSum-stub-related gap closed). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(reborn): seal hook registration trust boundary + dispatcher hardening Addresses blocking findings from the security audit of `ironclaw_hooks`: - C1 (Blocking, Trust Model): "Installed cannot Allow" was not enforced at the registration boundary. `BeforeCapabilityHookImpl::Privileged` was a public variant, so external crates with dispatcher access could construct an Installed binding paired with a Privileged impl and bypass the sink trait restriction. Sealed `BeforeCapabilityHookImpl`, `BeforePromptHookImpl`, and `ObserverHookImpl` to `pub(crate)` and replaced the single generic `install_before_capability` / `install_before_prompt` / `install_observer` surface with tier-specific public installers (`install_builtin_*`, `install_trusted_*`, `install_installed_*`) that build the binding with the matching trust class internally. Updated registrar, internal middleware tests, the hooks foundation pipeline test, and the reborn `hooks_integration` test to drive the new surface. Added regression tests proving the trust class is set by the installer and that the seal is type-level. - C5 (Medium, Slot Poisoning): same-dispatch poisoning was incomplete because `ordered_bindings` snapshots once at the top of the loop, and `HookRegistry::insert` accepted duplicate hook IDs. Rejected duplicate hook IDs (any point) in `HookRegistry::insert` and added a poison re-check before invoking each hook impl in `dispatch_before_capability`, `dispatch_before_prompt`, and `dispatch_observer_at`. Added regression tests for both behaviors. - C6 (Medium, Manifest / Predicate Validation): `parse_window` could panic on non-ASCII input because `split_at(len - 1)` requires a char boundary. Rewrote to compute the unit char's UTF-8 byte length and slice safely, added a public `validate_window` helper, and wired it into `HookManifestEntry::validate` for both `InvocationCount` and `NumericSum` bounds. Added tests for non-ASCII, empty, single-char, and zero-duration windows. - C2 (High, Tenant Isolation): partial fix only. The `PredicateEvaluator`'s sliding-window counter was keyed by `(hook_id, capability)`, so cross-tenant state could leak. Extended `HistoryKey` to include `tenant_id` and added a regression test proving counters partition by tenant. Documented the broader dispatcher-per-build / per-run-fresh-dispatcher pattern as deferred follow-up in `crates/ironclaw_hooks/CLAUDE.md`. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(reborn): emit hook telemetry milestones for audit/SSE observers Wires the hook dispatcher into the host's milestone stream so audit backends and SSE observers can see hook activity. Previously, hook dispatch was invisible — denies, pauses, failures, and observer fires left no trace in the host's observability backend. Changes: - `ironclaw_turns`: add `HookDispatched`, `HookDecisionEmitted`, and `HookFailed` variants to `LoopHostMilestoneKind`, with a closed- vocabulary `HookDecisionSummary` enum (Allow/Deny/PauseApproval/ PauseAuth/Pass/Patch). Introduce a lightweight `HookMilestoneSink` trait that emits hook-specific *kinds* without requiring a `LoopRunContext` (the dispatcher is a process-wide singleton that cannot own a per-run context), plus a `RunScopedHookMilestoneSink` adapter that injects run context and forwards to the existing `LoopHostMilestoneSink`. Also add `InMemoryHookMilestoneSink` for tests. - `ironclaw_hooks`: add a `telemetry` module that converts hook-crate types (`HookId`, `HookTrustClass`, `HookPointSpec`, `FailureCategory`, `FailureDisposition`, `BeforeCapabilityHookDecision`) into the wire- shape labels and summaries the milestone sink expects. Hook ids cross the seam as hex strings because the strongly-typed `HookId` cannot be imported from `ironclaw_turns` (the architecture test enforces `ironclaw_turns -> ironclaw_hooks` stays absent). - `ironclaw_hooks::dispatch`: add an optional `Arc<dyn HookMilestoneSink>` to `HookDispatcher`, set via `with_milestone_sink`. Emit `HookDispatched` before each hook runs, `HookDecisionEmitted` after a decision/pass/patch, and `HookFailed` on timeout/panic/malformed/missing-impl across all three dispatch paths (before_capability, before_prompt, observer). Default behavior (no sink attached) emits nothing — preserves the pre-telemetry observable surface. - `ironclaw_reborn`: document on `with_hook_dispatcher` that callers attach the milestone sink to the dispatcher *before* wrapping it in `Arc` and installing it into the factory, using a `RunScopedHookMilestoneSink` to inject run-context. The dispatcher itself is shared across runs, so attaching a fixed run-context inside it would be wrong. Update `RuntimeEvent` projection in `milestone_events.rs` to ignore the new hook kinds (no projection pathway yet; emitted milestones are consumed by SSE observers directly). Tests: - `ironclaw_hooks::dispatch`: 5 new tests covering milestone emission for deny decisions, panic failures, prompt-mutator patches, observer pass-throughs, and the no-sink default. - `ironclaw_reborn` hooks_integration: end-to-end test wiring a `RunScopedHookMilestoneSink` onto the dispatcher and asserting hook activity surfaces in the host's `LoopHostMilestoneSink`. Total: +6 hook telemetry tests; no existing tests modified. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(reborn): extract shared prompt envelope; inject hook patches into prompt bundle Adds `ironclaw_prompt_envelope`, a leaf crate that owns the single envelope primitive used by every model-visible untrusted-content path. `wrap_untrusted` prefixes content with a closed-vocabulary `<Trusted|Untrusted> <source> content: ` marker, rejects bodies carrying instruction-hijack phrases (`ignore previous instructions`, `<|im_start|>`, `<system>`, etc.), and enforces a 4 KiB byte budget by default. Migrates `ironclaw_host_runtime::memory_context` to delegate envelope wrapping, marker rejection, and control-character stripping to the new crate while keeping the `LoopSafeSummary`-specific 512-byte cap and byte truncation local. Existing memory_context behavior and tests are preserved. Wires the same envelope into `ironclaw_hooks`: * `HookPatch::add_enveloped_snippet` now takes a raw body and wraps it via `wrap_untrusted(EnvelopeSource::Hook, …)`. `Installed` hooks produce `Untrusted` envelopes; `Builtin`/`Trusted`/`SelfAuthored` produce `Trusted` envelopes so downstream readers can distinguish the two paths through a uniform marker. * `HookedLoopPromptPort::build_prompt_bundle` is no longer observe-only. After dispatching `before_prompt`, it envelope-wraps every snippet patch (passing `Enveloped` through, wrapping `Trusted` with the envelope helper), enforces the 4 KiB aggregate snippet byte budget across patches, and appends the wrapped snippets to the prompt bundle's `messages` as `system`-role `LoopModelMessage` entries carrying deterministic `msg:hook.<ordinal>.<hash>` content refs (mirroring the skill-snippet ref convention). The envelope crate is a leaf with no ironclaw dependencies, satisfying the boundary contract; the existing `ironclaw_hooks` boundary rule in `reborn_dependency_boundaries` continues to hold because `ironclaw_prompt_envelope` is not on its forbidden list. Test count delta: * `ironclaw_prompt_envelope`: +13 new tests (crate did not exist). * `ironclaw_hooks`: 84 → 88 tests (+4 prompt-port behavior tests: `hook_patch_appended_as_envelope_wrapped_message`, `total_byte_budget_enforced_across_patches`, `instruction_hijack_in_patch_rejected`, `trusted_hook_patch_wrapped_with_trust_marker`). * `ironclaw_host_runtime` memory_context: unchanged (8 tests still pass). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix: align tenant-counter test with SanitizedArguments-extended context ctor * docs(reborn): document loader contract; pin HookId hex format Add a "Loader responsibility" section to ironclaw_hooks/CLAUDE.md explaining that tier-specific installers prevent minting wrong-tier impls but cannot enforce origin — that's the loader's job — and recommending registry loaders type-tag extension hooks as LoadedHook::Installed at the loader seam. Add tier_specific_installers_are_documented_as_loader_contract as a regression guard that touches every public install_*_before_capability and install_*_before_prompt method so any signature change forces the loader contract to be re-evaluated. Document HookId::to_hex's 64-char lowercase hex output as part of the cross-crate contract consumed by LoopHostMilestoneKind::Hook* in ironclaw_turns; add hook_id_hex_format_is_stable_64_lowercase_chars in identity::tests and hook_id_string_serialization_matches_to_hex in telemetry::tests to pin the format and the seam conversion path. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(reborn): pin hook milestone JSON schema + assert pairing invariants Add L3 schema-snapshot tests for every hook-related LoopHostMilestoneKind variant (HookDispatched, HookDecisionEmitted per HookDecisionSummary, HookFailed per FailureCategory) so downstream consumers can rely on the JSON wire shape and any accidental field rename, enum-tag rename, or type change fails loudly. Add L4 pairing-invariant matrix test in the hook dispatcher that drives every observable outcome (Allow, Deny, PauseApproval, PauseAuth, Pass, Panic, Timeout, Malformed, MissingImpl) through a recording milestone sink and asserts the dispatched-then-terminator pairing shape. Document the MissingImpl path as the one case that emits a sole HookFailed with no preceding HookDispatched (the dispatcher discovers the protocol violation before the hook is actually dispatched). Add a multi-hook dispatch test that installs three hooks with mixed outcomes (allow/deny/panic) at the same point and asserts each hook produces its own paired sequence in the deterministic (phase, priority, hook_id) order taken from the dispatcher's registry. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(reborn): integration tests for observer middleware through RebornLoopDriverHostFactory Wire the HookedLoopModelPort / HookedLoopTranscriptPort / HookedLoopCheckpointPort observer wrappers into RebornLoopDriverHostFactory::build_text_only_host_with_capabilities, mirroring the existing HookedLoopCapabilityPort / HookedLoopPromptPort composition. The wrappers are applied only when a HookDispatcher is set on the factory, so the default factory shape is unchanged. Add four integration scenarios in crates/ironclaw_reborn/tests/hooks_integration.rs: - observer_hook_fires_after_model_through_factory - observer_hook_fires_after_capability_through_factory - observer_hook_fires_after_checkpoint_through_factory - observer_panic_does_not_fail_model_call (panic-isolation regression) Relax the test-fixture model gateway from "panic if invoked" to returning a stub assistant reply so the AfterModel / panic-isolation tests can drive stream_model through the wrapped port. The existing capability-port tests never touch the gateway, so their behavior is unchanged. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(reborn): introduce HookDispatcherBuilder for type-enforced sink wiring Adds a `HookDispatcherBuilder` in `ironclaw_hooks::dispatch` that owns the dispatcher construction lifecycle: registry -> optional timeout -> optional milestone sink -> installed hooks -> `.build_arc()`. The terminal `.build_arc()` wraps in `Arc` and yields an immutable handle. Tightens the public surface on `HookDispatcher`: `new`, `with_timeout`, `with_milestone_sink`, and every `install_*_*` method are now `pub(crate)`. Outside callers route exclusively through the builder, so "wire the milestone sink before Arc-wrapping" is a compile-time fact rather than a documentation convention. `HookRegistrar::install` now takes a `HookDispatcherBuilder` by value and returns `(HookDispatcherBuilder, Vec<HookId>)`, keeping the builder chainable through manifest installation. `RebornLoopDriverHostFactory` gains `with_hook_dispatcher_builder` to let callers defer `.build_arc()` to the factory — a step toward the FU8 per-build dispatcher pattern. Migrates `foundation_pipeline.rs` and `hooks_integration.rs` to the builder. Internal middleware and dispatch tests continue to use the crate-private `HookDispatcher::new` directly. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(reborn): production CapabilityInputResolver for NumericSum predicates Adds HookCapabilityInputResolverAdapter in ironclaw_reborn that bridges the existing LoopCapabilityInputResolver (already used by HostRuntimeLoopCapabilityPort for dispatch input resolution) to the hooks crate's CapabilityInputResolver trait. RebornLoopDriverHostFactory gains with_capability_input_resolver(...), and when both a hook dispatcher and resolver are configured the factory threads the adapter into HookedLoopCapabilityPort::with_resolver — so NumericSum and other argument-dependent predicates evaluate against real, sanitized inputs instead of failing closed against the framework's null default. The adapter also enforces a configurable serialized-byte budget (default 64 KiB) as defense in depth ahead of the hooks crate's per-string and depth caps in SanitizedArguments. Unit tests cover the four adapter branches (resolved JSON, inner-error → None, non-object pass-through, oversized → None) and a new end-to-end integration test (numeric_sum_predicate_caps_total_value_against_real_inputs) drives the full factory wiring: with a NumericSum cap of 99 over an "amount" field, two invocations carrying {"amount":"50"} let the first pass through and deny the second at the hook seam, with the inner port reached exactly once. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(reborn): per-build HookDispatcher for full per-run isolation (C2) Introduce `with_hook_dispatcher_factory(F)` on `RebornLoopDriverHostFactory`. The closure is invoked once per `build_text_only_host*` call, so dispatcher-owned mutable state — slot poisoning, registry mutations, predicate-counter siblings — is scoped to a single host build instead of shared across every host the factory produces. The legacy `with_hook_dispatcher(Arc<HookDispatcher>)` adapter is kept as a thin wrapper that returns clones of the same `Arc` on every build. Its shared-state behavior is now documented as an explicit opt-in for backward compat; new wiring should prefer the factory closure. Adds two regression tests: - `per_build_dispatcher_state_does_not_leak_across_runs` — installs a panicking hook, builds two hosts back-to-back, and proves the inner port is never reached on build 2 (fresh slot still applies the fail-closed deny). Pins per-run isolation. - `legacy_with_hook_dispatcher_shares_state_across_builds` — pins the shared-state semantic of the legacy adapter as the explicit baseline. Migrates `predicate_deny_hook_short_circuits_inner_port` to the new factory-closure path so the new wiring is exercised by the existing suite. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(reborn): project hook telemetry milestones into RuntimeEvent for durable audit Extend the runtime event substrate with `HookDispatched`, `HookDecisionEmitted`, and `HookFailed` kinds carrying closed-vocabulary labels and the blake3-hex hook identity. Project the matching `LoopHostMilestoneKind::Hook*` variants in `DurableLoopHostMilestoneSink` so hook telemetry now lands in the same durable event log as model/reply/loop milestones — SSE observers still see live hook events, and audit replay can reconstruct the full hook trail. - `ironclaw_events`: add hook variants to `RuntimeEventKind`, optional hook fields on `RuntimeEvent` (`hook_id`, `hook_point`, `hook_trust_class`, `hook_decision`, `hook_failure_category`, `hook_failure_disposition`), typed constructors (`hook_dispatched`, `hook_decision_emitted`, `hook_failed`), and dedicated sanitizers (`sanitize_hook_label`, `sanitize_hook_id`) re-run on every wire crossing. No new crate dependency edges; hook strings cross the boundary opaque. - `ironclaw_reborn::milestone_events`: project the three hook milestone kinds via a new `loop.hook` capability id. `HookDecisionSummary` is collapsed to its closed-vocabulary `kind_name()` so sanitized reasons never enter the durable substrate. - `ironclaw_event_projections`: extend `TimelineEntryKind` and the `RuntimeEventKind -> RunProjectionStatus` mapping so hook events are pure telemetry — they preserve the current run status rather than changing it. - Tests: 4 unit tests in `ironclaw_events::runtime_event::tests` (serde round-trip per variant + unsafe-label collapse), 3 in `ironclaw_reborn::milestone_events::tests` (projection per variant, including the assertion that raw `Deny { reason }` text does not reach the durable wire payload). Existing replay-projection direct-construction tests updated for the new RuntimeEvent fields. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * feat(reborn): enforce manifest-declared hook scope at dispatch time (C3) Audit finding C3: extensions could declare `[[hooks]]` with `scope = "own_capabilities"` in their manifest, but the dispatcher never enforced it — an Installed hook from ext-A could fire against capabilities provided by ext-B. Scope was parsed but not load-bearing. This change makes scope load-bearing end-to-end: - `BeforeCapabilityHookContext` carries an optional `provider: ironclaw_host_api::ExtensionId` populated by the middleware. The hook context is `#[non_exhaustive]` already so this is non-breaking. - `HookBinding` gains `owning_extension: Option<ExtensionId>` and `scope: HookBindingScope`. `HookBindingScope` is `Global` / `OwnCapabilities` / `SameTenant`. Builtin and Trusted bindings default to `Global` and carry no `owning_extension`; Installed bindings carry both, sourced from the manifest. - `HookDispatcher::install_installed_*` installers now require the caller to pass `(owning_extension, scope)`. The registrar derives both from the manifest entry, so manifest authorship is the single source of truth. - A new `CapabilityProviderResolver` trait + bundled `NullCapabilityProviderResolver` lets the middleware lift the capability id to its provider at invocation time. The middleware wires the resolved provider into the hook context. - `dispatch_before_capability` consults `binding.scope.permits(...)` before invoking each hook. Bindings that don't permit the current invocation are inert — no sink call, no failure record, no poisoning. Conservative defaults: - When the provider resolver returns `None` (no resolver wired, or the capability has no known provider), `OwnCapabilities`-scoped hooks do NOT fire. An attacker cannot bypass scope filtering by stripping provider info from the descriptor. Tests: - 5 new dispatcher tests cover OwnCapabilities matching, foreign provider, unresolved provider, SameTenant, and Builtin Global. - 1 new registrar test asserts manifest scope and extension propagate into `HookBinding`. - 1 new middleware test asserts the provider resolver populates the hook context. - 1 new integration test in `ironclaw_reborn` proves an ext-A hook scoped to `OwnCapabilities` does not intercept invocations that have no resolved provider (the production composition default). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * style: rustfmt dispatch.rs after FU1 merge * docs(hooks): prior-art comparison against LSM/eBPF/Envoy/K8s/OPA/CRX/VSC/Tauri Validates the IronClaw hooks design against 8 established hook/policy systems across 8 axes (dispatch, trust tiers, attenuation, decision vocabulary, failure semantics, isolation, manifest, audit). Surfaces: - 7 areas where ICLAW stands out vs prior art (type-level trust enforcement, dispatch-time scope, failure-kind matrix, pause-with- gate-ref, pairing-invariant audit matrix, tenant-keyed predicates, phase-ordered dispatch) - 4 conventional choices we should revisit (in-process Installed-WASM, sticky poison, no formal dispatch model, no installation rate-limit) - 3 divergences whose 'why' is weak and need design review * docs(hooks): STRIDE threat model for v1 framework Enumerates 7 adversary classes (A1-A7), 6 assets ranked by blast radius, and ~35 attack vectors across STRIDE categories with mitigations, existing tests, and residual risk. Surfaces 7 prioritized follow-ups: - High: per-extension hook-count cap (D3/D4) - High: gate-ref unguessability + one-shot test (S1) - Med: resolver field-level scope (I2) - Med: per-evaluator state ceiling (D5) - Med: poison-stickiness operator runbook - Low: timing side-channel residual acknowledgement (I4) - Low: instruction-marker denylist periodic review (I5) Confirms the load-bearing 'Installed cannot Allow' (E1) property holds via type-level seal + tier-specific installers, backed by compile_time_seal_test and installed_binding_cannot_be_paired_with_ privileged_impl tests. Explicit out-of-scope: extension install pipeline (#3492), WASM exec sandbox (needs separate threat model when it lands), approval gateway (#3564). * feat(hooks): close threat-model gaps S1 (gate-ref entropy) and D3/D4 (registration flood) S1 (gate-ref unguessability, factory side): - Three new tests on `UuidHookGateRefFactory`: - `gate_refs_are_v4_uuids` pins the v4 entropy source (122 random bits per ref per RFC 4122 §4.4); fails if a future change moves to a counter or weaker UUID version. - `gate_refs_have_no_collisions_across_many_calls` mints 20k refs across both namespaces, asserts zero collisions (statistical proxy for entropy quality). - `approval_and_auth_namespaces_do_not_overlap` confirms prefix routing separation. - Doc comment now documents the security property explicitly and delineates factory-side vs gateway-side responsibilities for the one-shot consumption property. D3/D4 (hook registration flood): - New `MAX_HOOKS_PER_EXTENSION = 32` and `MAX_HOOKS_PER_EXTENSION_PER_KIND = 8` consts in `registrar.rs`. - New `HookRegistrar::enforce_registration_caps` runs pre-flight at the top of `install()`, before any binding is inserted. Whole-batch rejection means a partially-installed batch cannot slip past. - Three regression tests: total-cap rejection, per-kind-cap rejection, at-cap acceptance. - Error messages cite the threat-model finding so operators can map rejection back to the design rationale. Threat model updated: S1, D3, D4 marked closed in the cross-cutting properties matrix and the open-follow-ups list. * test(hooks): three real hooks built against the public API + ergonomics findings Builds three representative hooks from outside the crate, mimicking what an extension or system author would actually write: 1. polymarket-daily-cap — Installed predicate hook, InvocationCount rate-cap with Deny on excess. Canonical 'rate-limit a capability' use case for the predicate language. 2. large-stake-approval-gate — Installed predicate hook, NumericSum over amount_usd field, PauseApproval at $1000/24h. Manifest-shape + registrar-install coverage from outside Reborn; end-to-end dispatch lives in ironclaw_reborn integration tests because NumericSum needs resolved args (a friction finding documented in the companion doc). 3. pii-redaction-warning — Trusted Rust hook implementing PrivilegedBeforePromptHook, injects a trusted instruction snippet reminding the model to redact PII. Demonstrates the path a system author takes when the predicate language isn't expressive enough. API change (F1 fix): SanitizedArguments::unresolved() promoted from pub(crate) to pub. This is the documented safe default — predicates that need args must fail closed against it — so exposing the constructor cannot weaken any trust property. The sanitizing from_json constructor stays sealed; that's the trust boundary. Without this fix, external hook authors could not construct a BeforeCapabilityHookContext with both a known provider AND unresolved args, which made TDD of their own predicate impossible. Findings documented in docs/real-hooks-findings.md, ranked by severity. Big-picture observation: writing the Trusted Rust hook (F4) was easier than writing the declarative predicate hook (F1 + F2 + F3) — three of seven findings target predicate-authoring ergonomics. The declarative path needs the most polish before third-party extension authors will trust it for non-trivial policy. Tests: 6 new in real_hooks.rs, all pass. * feat(hooks): close all remaining threat-model and ergonomics gaps Closes the Med-priority threat-model gaps (I2, D5, poison runbook) and all real-hooks ergonomics findings (F2, F3, F5, F6, F7) in a single pass. Threat model: - I2 (resolver field-scope): documented in SanitizedArguments rustdoc. The narrow public surface (only is_resolved + extract_numeric) enforces field-scope by construction for the current predicate path. Reassess when Installed-WASM lands. - D5 (evaluator state ceiling): MAX_HISTORY_KEYS = 8192 per map, LRU eviction with evictions_observed() metric for operator monitoring. New regression test lru_eviction_increments_counter_and_drops_oldest_key. - Poison-stickiness runbook: new docs/operator-runbook.md with recovery options ranked by cost. Ergonomics findings: - F2 (closed-vocab deny reasons): rustdoc on OnExceededAction and GateDecisionView::Deny explaining the audit-vs-model split and why manifest reason text doesn't reach the model. - F3 (NumericSum can't be TDD'd outside Reborn): new test-support feature flag with SanitizedArguments::for_tests(value) that external hook authors can opt into via dev-dep. - F5 (two ExtensionId types): added From<&ironclaw_host_api::ExtensionId> impl for identity::ExtensionId, plus cross-link rustdoc. - F6 (HookManifestEntry struct-literal fragility): added #[non_exhaustive] + HookManifestEntry::new(id, kind, body) + with_scope/with_phase/with_priority/with_description/with_requires_grant builder methods. Migrated 3 external call sites in tests/. - F7 (priority guidance): rustdoc on HookPriority with when-to- deviate guidance, named FIRST/LAST constants documented for Builtin/Telemetry use cases. Tests: 151 unit + 1 + 6 integration in ironclaw_hooks all pass with --all-features. ironclaw_reborn (13 hooks_integration scenarios) unchanged. Threat model updated: I2 / D5 / poison runbook marked closed in both the per-vector table and the cross-cutting properties matrix. Open follow-ups now down to two Low items (I4 timing side-channel residual, I5 instruction-marker denylist refresh) plus the deferred DenyReasonCode enum from F2. * fix(ci): collapse nested match in hooks_integration test for clippy --all-features CI runs `cargo clippy --all --tests --examples --all-features -- -D warnings` which is stricter than the workspace clippy I ran locally and trips `clippy::collapsible_match` on the nested-if in HookDecisionEmitted matching. Collapse the inner `if decision.kind_name() == "deny"` into an arm guard. * feat(hooks): address henrypark133 review — Critical #1/#2/#3/#5, Concerning #5/#7 Address composition-seam bugs in the Reborn factory wiring + doc tidy. henrypark133 review findings addressed: Critical #1 — before_prompt hook messages not materialized. HookedLoopPromptPort now requires a HookPromptMaterializationSink and fails closed if patches are emitted without one. The reborn factory installs an InstructionStoreBackedHookSink adapter that delegates to the host's InstructionMaterializationStore, so synthetic msg:hook.* refs are resolvable by the downstream model resolver. New seam trait (HookPromptMaterializationSink) keeps ironclaw_hooks decoupled from LoopRunContext. Critical #2 — OwnCapabilities hooks were inert in production wiring. Factory now installs SurfaceBackedProviderResolver (consults the visible-capability surface for capability_id → provider). With this, ctx.provider is populated and OwnCapabilities-scoped Installed hooks actually fire against their own provider's capabilities. Critical #3 — gate refs were unresolvable. Middleware default switched from UuidHookGateRefFactory to FailClosedHookGateRefFactory. Tests must explicitly opt into UUID (via with_gate_ref_factory) to exercise the affirmative ApprovalRequired path; production deployments must install a router-backed factory. New factory method RebornLoopDriverHostFactory::with_hook_gate_ref_factory. Concerning #5 — AfterModel fired twice + before durable finalization. Removed AfterModel dispatch from HookedLoopModelPort; the transcript port's finalize_assistant_message is now the sole AfterModel boundary (the durable one). Model port wrapper is preserved as a no-op shim for symmetry + future model-response-observed point. Concerning #7 — doc tidy: - CLAUDE.md: 3 trust classes → 4 (Builtin/Trusted/Installed/SelfAuthored with explicit note that SelfAuthored is run-scoped only and not loadable from an external source). - operator-runbook.md: "Audit log" → "durable runtime event stream" where the projection is actually the runtime-event stream, not formal AuditEnvelope records. - prior-art.md: poison-lifetime nuance — per-host-build with the factory pattern, process-lifetime only for the legacy adapter. - prior-art.md:80: trailing whitespace removed. Testing gaps from henrypark133 — caller-level tests through RebornLoopDriverHostFactory: #1 (before_prompt resolver path): before_prompt_hook_message_is_resolvable_via_factory_wiring #2 (OwnCapabilities positive/negative/unknown): own_capabilities_hook_fires_when_provider_matches own_capabilities_hook_does_not_fire_when_provider_differs own_capabilities_hook_does_not_fire_when_provider_unknown #3 (pause/auth gate lifecycle or fail-closed): pause_approval_with_default_factory_fails_closed_as_denied pause_approval_hook_surfaces_as_approval_required_with_real_gate_ref (updated to require explicit UuidHookGateRefFactory opt-in) #5 (AfterModel exactly-once at durable boundary): after_model_fires_exactly_once_at_durable_boundary Still TODO from review (separate commits): Critical #4 (telemetry context — two-run attribution) + gap #4 Concerning #6 (TimelineEntry hook metadata projection) + gap #6 Tests: 154 unit + 18 hooks_integration + all other reborn tests pass. Workspace clippy + fmt + no-panics clean. * feat(hooks): address remaining henrypark133 review — Critical #4, Concerning #6 Critical #4 — per-run hook telemetry attribution. New `HookDispatcherBuilderFactory` signature: factory returns a HookDispatcherBuilder, and `RebornLoopDriverHostFactory` attaches a `RunScopedHookMilestoneSink` keyed to the CURRENT run's LoopRunContext inside `build_text_only_host_with_capabilities`, before sealing the dispatcher. The previous zero-arg signature relied on the closure capturing run_context — silently misattributed across reuses; new public API `with_hook_dispatcher_builder_factory` removes that failure mode entirely. Legacy `with_hook_dispatcher_factory` retained for back-compat (its sink-wiring contract stays caller-side). Concerning #6 — TimelineEntry hook metadata. Added 6 optional fields to `TimelineEntry` (hook_id, hook_point, hook_trust_class, hook_decision, hook_failure_category, hook_failure_disposition) and projected them from `RuntimeEvent::Hook*`. Replay consumers now see which hook fired/failed, not just that some hook event happened. Each field is closed-vocabulary (no free-form reason text — that stays in the audit reason payload, not the product replay DTO). Testing gaps from henrypark133 — caller-level tests: #4 (two-run hook telemetry attribution): hook_telemetry_attribution_is_per_run_not_captured Builds two hosts from the SAME builder factory closure with two fresh LoopRunContexts. Asserts each run's hook milestones carry its OWN run_id (no stale captured one). #6 (replay projection contract for hook events): hook_runtime_events_project_with_sanitized_hook_metadata non_hook_runtime_events_project_with_no_hook_metadata Constructs RuntimeEvent::Hook{Dispatched,DecisionEmitted,Failed} and asserts the projection preserves the metadata fields. The negative test guards against cross-contamination on non-hook events. All henrypark133 review items now addressed: Critical: #1, #2, #3, #4 — done Concerning: #5, #6, #7 — done Testing gaps: #1-#6 — done Tests: 154 unit + 19 hooks_integration in ironclaw_reborn + 61 reborn unit + 38 + 2 new in ironclaw_event_projections + ... pass. Workspace clippy + fmt + no-panics clean. * docs(hooks): scope DenyReasonCode closed-vocabulary enum (successor #6) Successor PR from #3573 — real-hooks ergonomics finding F2 (deferred). Adds a curated vocabulary of model-visible denial reasons so hook authors can communicate why a deny happened without opening a free-form prompt-injection channel. * feat(hooks): DenyReasonCode + PauseReasonCode closed-vocabulary enums Address real-hooks ergonomics finding F2 (deferred from PR #3573). The prior dispatcher collapsed every Installed-tier deny to the static label 'hook_predicate_denied', because manifest reason strings are author-controlled and surfacing them to the model would open a prompt-injection channel. The cost: the agent couldn't tell *why* a hook denied. This PR introduces two closed-vocabulary enums: - DenyReasonCode: Generic / RateLimit / ValueCap / Blocklist / RequiresApproval / OutOfPolicy - PauseReasonCode: Generic / RequiresApproval / OverThreshold / SensitiveAction Each variant has an as_label() returning &'static str (so the sink's &'static str contract is preserved). New OnExceededAction variants 'DenyWithCode { code, reason }' and 'PauseApprovalWithCode { code, reason }' let manifest authors opt into the richer labels while keeping reason audit-only. The legacy Deny { reason } / PauseApproval { reason } variants are retained for back-compat and map to DenyReasonCode::Generic / PauseReasonCode::Generic — existing manifests continue to produce hook_predicate_denied / hook_predicate_pause_requested. Threat-model regression: a hook author cannot smuggle text into the model-visible label because the 'code' field is typed as the enum; there's no String slot exposed model-side. A test (deny_with_code_only_exposes_enum_variants_to_model) documents this as a compile-time property. Tests (+7 new = 161 total): - deny_reason_code_labels_are_stable: pins the label vocabulary so rename/relabel is loud. - pause_reason_code_labels_are_stable: same for PauseReasonCode. - deny_with_code_round_trips_through_json + pause variant: wire round-trip + snake_case tag assertion. - deny_with_code_only_exposes_enum_variants_to_model: compile-time property check. - rate_or_value_cap_with_deny_code_routes_to_code_label: end-to-end affirmative test that the dispatcher emits the code's label. - rate_or_value_cap_with_pause_code_routes_to_code_label: same for pause. Scope doc: crates/ironclaw_hooks/docs/successors/06-deny-reason-code.md * test(hooks): address codex review on #3636 - Update stale real-hooks-findings.md F2 row to cite this PR's enum follow-on (was 'deferred'). - Add install_deny_with_code_manifest_surfaces_code_label_on_dispatch: end-to-end test driving the registrar->dispatcher path for the new DenyWithCode variant (prior tests covered serde + direct hook evaluation, but not the manifest install path that downstream authors actually use). Codex review on PR #3636: APPROVE with two recommendations; both addressed. Tests: 162 unit (+1 new). Clippy/fmt clean. * fix(hooks): attenuate Installed-tier prompt patches to user role Installed-tier `before_prompt` patches were injected as role:"system" messages. Envelope text labels ("[ext-foo says]: ...") do not strip system-role authority from the model's perspective, so a third-party extension could inject system-tier instructions through a snippet patch. This is a prompt-authority escalation against the trust hierarchy the framework otherwise enforces. Add `role_for_trust_class()` mapping Installed -> "user" and Builtin/Trusted/SelfAuthored -> "system". Thread per-patch trust_class through `wrap_patches_to_messages` and use it for the emitted `LoopModelMessage.role`. Tests: - installed_hook_patch_drops_to_user_role: asserts the role for an Installed-tier patch is "user" - trusted_tier_hook_patch_keeps_system_role: regression that Trusted tier still produces system-role content * fix(hooks): enforce scope filter on observer dispatch + reject incompatible points Two related defense-in-depth fixes against silent scope-filter failure: 1. The registry silently accepted Installed bindings with `HookBindingScope::OwnCapabilities` at points (BeforePrompt, AfterModel, AfterCheckpoint) whose dispatch context carries no per-capability provider. The manifest's declared scope had no effect at all — the hook fired against every dispatch. Reject the binding at install time so the operator sees the misconfiguration. 2. `dispatch_observer_at` for `AfterCapability` did not consult the binding's scope, so an Installed observer registered with `OwnCapabilities` fired against every invocation regardless of provider. Add `dispatch_observer_at_with_provider` carrying the resolved capability provider; the capability-port middleware resolves the provider once per invocation and threads it through both the BeforeCapability hook context and the AfterCapability observer dispatch. The dispatcher then enforces `HookBindingScope::permits` on each observer binding. `ObserverHookContext` gains a `provider: Option<ExtensionId>` field; `#[non_exhaustive]` keeps existing authors compiling. Tests: - rejects_own_capabilities_at_before_prompt - rejects_own_capabilities_at_after_model - accepts_own_capabilities_at_before_capability - own_capabilities_observer_filters_foreign_providers (covers foreign / matching / unresolved provider) * fix(hooks): preserve free-form audit reason alongside closed-vocab model label (serrrfirat #3636) `PredicateBackedBeforeCapabilityHook::evaluate()` was discarding the free-form `reason` from `EvaluatorDecision::{Deny, PauseApproval}` with `..` and only sending `code.as_label()` into the sink. The `HookDecisionEmitted` milestone therefore carried only the closed- vocab label, and operator-visible audit/SSE context was silently lost end-to-end. The fix splits the channels: - Model sees the closed-vocab label (`hook_rate_limit`, `hook_pause_over_threshold`, ...) via `sink.deny(label)`. This channel is unchanged. - Audit/SSE sees the manifest's free-form `reason` via a new audit-only sink method `record_audit_reason(reason: String)`. The recording sink captures it; the dispatcher reads it after the hook returns and threads it into `LoopHostMilestoneKind::HookDecisionEmitted`. Surface changes: - `PrivilegedGateSink` / `RestrictedGateSink` gain `record_audit_reason(String)` — accepts dynamic `String` (audit-only, no model-facing seam) unlike the `&'static str` decision reasons. - `RecordingGateSink` gains an `audit_reason: Option<String>` field. - `GateHookOutcome::Decision` is now `Decision { decision, audit_reason }`. - `HookDispatcher::emit_decision_with_audit` threads the audit reason into the milestone. - `LoopHostMilestoneKind::HookDecisionEmitted` gains a `#[serde(default, skip_serializing_if = "Option::is_none")]` `audit_reason: Option<String>`. The durable RuntimeEvent projection intentionally drops this field — audit reasons are operator-facing in-memory SSE content, never durable cross-process surface. Tests: - `deny_with_code_records_audit_reason_separately_from_model_label`: asserts the recording sink ends with `Deny { reason: "hook_rate_limit" }` in `state` AND `audit_reason == Some("daily cap of $1000 ...")`. * fix(hooks): remove unused model_request helper (CI clippy fix) * fix(hooks): address serrrfirat P1/P2 findings on PR #3573 Three issues from the 5-15 review: **P1 #1 registrar.rs:70 — `same_tenant` grants not enforced** `HookManifestEntry::validate` only confirmed `requires_grant` was present; the registrar then immediately installed the binding with no host-verified grant context. A manifest could declare `requires_grant = "anything"` and get a cross-extension binding for free. Fix: `HookRegistrar` now carries a `verified_grants: HashSet<String>` (empty by default — default-deny). Add the host-facing setter `with_verified_grants(...)`. At `install_one`, if `entry.requires_grant` is `Some(g)`, require `g ∈ verified_grants` or reject with a clear error. Tests: - `install_rejects_same_tenant_without_verified_grant` - `install_rejects_same_tenant_when_verified_grants_mismatch` - The existing positive test `installer_propagates_owning_extension_and_scope_from_manifest` now wires the verified grant explicitly (proves the API contract). **P1 #2 prompt_port.rs:150 — zip misalignment** The materialization loop zipped surviving messages against the ORIGINAL unfiltered patch list. `wrap_patches_to_messages` skips metadata patches and over-budget snippets, so the zip silently paired message[0] with patch[0] even when patch[0] was the skipped metadata — materializing the wrong content (or none) under the snippet's synthetic ref. Fix: `wrap_patches_to_messages` now returns `Vec<WrappedHookMessage { message, safe_content }>` — surviving messages paired with their content by construction. The caller materializes `entry.safe_content` under `entry.message.content_ref` directly; no zip against unfiltered input. Removed the now-unused `safe_content_for_patch` helper. Test: - `materialization_stays_aligned_when_metadata_patches_are_filtered`: a hook emits `[metadata, snippet]`; asserts only one model message, and the materialized content under its ref contains the snippet's body — proves filtering can no longer desync from materialization. **P2 #3 loop_driver_host.rs:1343 — `with_hook_dispatcher_builder`** Docs said it deferred `build_arc()` to let the host factory finalize wiring; the implementation called `build_arc()` eagerly and routed through the legacy shared-dispatcher adapter, losing per-run dispatcher isolation and the run-scoped milestone sink. Fix: marked `#[deprecated]` with a note pointing callers to `with_hook_dispatcher_builder_factory(|| ...)` for per-build isolation, or `with_hook_dispatcher(...)` if they actually meant the shared adapter. The method body is unchanged so no callers break; they'll see the deprecation warning. No internal callers exist, so the deprecation doesn't trip `-D warnings`. All 162 hooks lib + 19 reborn integration tests pass; clippy clean. * fix(hooks): address serrrfirat 3573-2026-05-15 review findings P1 — prompt bundle authority mismatch (prompt_port.rs): `HookedLoopPromptPort::build_prompt_bundle` called the inner port first, which caused `HostManagedLoopPromptPort` to issue the prompt-bundle authority grant against the pre-hook message list. The wrapper then appended `msg:hook.*` messages to `bundle.messages`, so the downstream model request hit `grant.messages != messages` and failed closed with "model request messages do not match the host-built prompt bundle". Add `with_bundle_authority(authority, run_context)` and re-issue the grant after appending hook messages so it covers the post-hook bundle. Reborn wires `prompt_authority.clone()` + `run_context.clone()` into the wrapper at construction time. P2 — observer installer accepts non-observer points (dispatch.rs): `install_observer` accepted any `HookPointSpec` (including `BeforeCapability` / `BeforePrompt`) and only populated the observer map. Dispatch later found a binding without a gate/mutator impl and fail-closed the capability with "binding present without installed implementation". Reject non-observer points at install time so misuse fails loudly rather than poisoning bindings at dispatch. P2 — batch path skipped AfterCapability observers on inner error (capability_port.rs): The batch loop used `?` directly on `self.inner.invoke_capability(...)`, which propagated the error before dispatching `AfterCapability` observers. Failed batch entries disappeared from telemetry / audit, while the single-invocation path dispatches observers on error. Capture the inner result, dispatch observers, then propagate the error. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(hooks): address PR #3573 review feedback round 3 Addresses serrrfirat's CHANGES_REQUESTED review (2026-05-20) by tightening several install-time / dispatch-time bounds and gating production seams: - Bound free-form audit reasons crossing telemetry. New `telemetry::sanitize_audit_reason` strips control characters and caps length at 512 bytes; `emit_decision_with_audit` routes the manifest- supplied reason through it before publishing milestones. Manifest validation also rejects reasons over the same byte limit at install time so the wire-side cap is a defense-in-depth layer, not the only line. - Make hot dispatch O(H) instead of O(H^2). The per-binding poison recheck used to acquire the registry mutex and walk every binding; `ordered_bindings_with_poison_snapshot` now takes the active bindings and the poisoned hook-id set under a single lock, and each loop threads a local `HashSet<HookId>` that absorbs mid-dispatch poisoning. Removed the redundant `is_poisoned` helper. - Gate `HookDispatcher::registry_for_test` behind `cfg(any(test, feature = "test-support"))`. The accessor previously exposed `&Mutex<HookRegistry>` in production, letting any `Arc<HookDispatcher>` holder lock and call `HookRegistry::poison` to disable installed hooks. Added `active_bindings_snapshot(point)` as the read-only production-safe replacement. - `#[serde(deny_unknown_fields)]` on every hook-manifest and predicate DTO (`HookManifestEntry`, `HookManifestBody`, `WasmBudget`, `HookPredicateSpec`, `CapabilityPredicate`, `ValueOrRateBound`, `OnExceededAction`). Typoed or unsupported fields (e.g. a manifest-supplied `trust_class`) now fail loud at install time instead of being silently dropped. - Bound predicate trees at install. New `validate_predicate_tree` enforces `MAX_PREDICATE_DEPTH = 8`, `MAX_PREDICATE_NODES = 64`, `MAX_PREDICATE_STRING_BYTES = 256`, and `MAX_MANIFEST_REASON_BYTES = 512`. A hostile registry manifest can no longer install a deep or huge `All`/`Any` tree that the evaluator would recursively walk on every match. - Cap sliding-window samples per key. `MAX_SAMPLES_PER_KEY = 4_096` in the predicate evaluator. Both the invocation-count and numeric-sum histories drop the oldest sample once the cap is reached, bounding memory under attacker-triggered hot capabilities while preserving rate/value-cap semantics over the most recent window. - `split_indexer` / `resolve_path` now fail closed on malformed bracket syntax (`amount[foo]`, `amount[`, trailing garbage). Previously they silently fell back to the parent field, which could let a typoed `NumericSum` predicate evaluate against the wrong value and allow calls the predicate would otherwise have denied. - Honor `PatchOrdinalHint`. `WrappedHookMessage` carries the source patch's `ordinal_hint`; `HookedLoopPromptPort` inserts `NearTop` messages after the bundle's `identity_message_count` and appends `Last` messages at the end. Safety/policy snippets that need early placement now get it. - Update `ironclaw_hooks` top-level docs to reflect the four trust classes (`Builtin`/`Trusted`/`Installed`/`SelfAuthored`) and the now-wired Reborn middleware composition. Tests added: - `manifest::rejects_unknown_top_level_field` - `manifest::rejects_unknown_wasm_budget_field` - `manifest::rejects_predicate_tree_exceeding_max_depth` - `manifest::rejects_predicate_tree_exceeding_max_nodes` - `manifest::rejects_predicate_string_exceeding_max_bytes` - `manifest::rejects_manifest_reason_exceeding_max_bytes` - `points::capability::malformed_indexer_returns_none_not_parent_value` - `telemetry::sanitize_audit_reason_*` (truncate / strip control / preserve / empty) `cargo fmt`, `cargo clippy --all --benches --tests --examples --all-features`, and `cargo test -p ironclaw_hooks` all pass clean. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(hooks): batch deferred test coverage from #3573 review (#3914) * perf(hooks): defer capability input resolution until a predicate needs it (#3913) * fix(rebase): adapt hooks tests + middleware to upstream API additions - CapabilityDescriptorView: add parameters_schema field - LoopModelRequest / LoopPromptBundleRequest: add capability_view field - TimelineEntry test builder: add hook_id / hook_point / hook_trust_class / hook_decision / hook_failure_category / hook_failure_disposition fields - ironclaw_reborn::tests::hooks_integration: switch from InMemoryLoopCheckpointStore to InMemoryTurnStateStore (which now impls both LoopCheckpointStore and TurnStateStore), pass TurnActor in TurnRunState, supply the new turn_state_store factory arg - ironclaw_reborn lib.rs: drop the pub-use re-exports that upstream intentionally removed (per the module-directory rationale in the current ironclaw_reborn lib.rs doc comment); update the hooks_integration test imports to use module paths - Cargo.toml: union the hooks-foundation member list with upstream's new crates (event_streams, auth, first_party_extensions, reborn_webui_ingress, product_workflow_storage, webui_v2); drop ironclaw_storage which no longer exists upstream - crates/ironclaw_architecture/tests/reborn_dependency_boundaries: keep upstream's removal of ironclaw_filesystem from the ironclaw_turns forbidden list AND add ironclaw_hooks to that list - crates/ironclaw_reborn/src/milestone_events.rs: drop dead loop_failure_kind helper (replaced upstream by loop_failure_kind_name in text_loop_driver.rs); keep hook_decision_label which is still used Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * perf(hooks): restore batched capability dispatch when hooks acti…
…i#3573) (nearai#3635) * docs(hooks): scope persistent predicate counter backend (successor #3) Successor PR from nearai#3573. Current sliding-window state is in-memory and resets on restart. Adds a PredicateStateBackend trait + Postgres/libSQL impls for cross-process and restart-survival semantics. * feat(hooks): extract PredicateStateBackend trait + replay-safe in-memory impl Addresses codex review's three Critical findings on PR nearai#3635: 1. Backend wiring: the trait is now registered (lib.rs:25-26) and PredicateEvaluator delegates to Arc<dyn PredicateStateBackend> via with_backend(...). Default constructor preserves the in-memory behavior so all 154 existing tests pass unchanged. 2. Atomic record-and-read: each record_invocation / record_value call performs the write AND returns the resulting in-window count/sum under a single mutex (in-memory) / transaction (durable backends). Splitting into separate record + read would let two hosts each see 'under cap' and both proceed, drifting past max. 3. Replay refusal: each record call carries a PredicateEventId. Re-emitting the same event_id is a no-op against the count. In-memory backend implements via a per-key bounded set (RECENT_EVENT_ID_CAP = 256); durable backends will use INSERT … ON CONFLICT DO NOTHING. Trait surface (predicate_state.rs): - PredicateEventId(String): opaque dedup key - PredicateBackendError: thiserror enum for fallible durable backends; in-memory backend never returns Err - PredicateStateBackend trait with Result return types - InMemoryPredicateStateBackend default impl - MAX_HISTORY_KEYS const re-exported via evaluator for back-compat Evaluator changes (evaluator.rs): - holds Arc<dyn PredicateStateBackend> (no more inline maps) - evictions_observed() reads through to backend - synth_event_id() generates per-call-unique ids via a process-local atomic counter so tests with identical (hook, ctx, now) still produce distinct ids - LRU helpers + HistoryKey/ValueHistoryKey types moved into predicate_state.rs (as InvocationKey/ValueKey) Tests: - 6 new predicate_state tests: - in_memory_invocation_counts_within_window - in_memory_invocation_trims_outside_window - in_memory_value_sums_within_window - in_memory_tenant_isolation (regression on threat-model C2) - in_memory_duplicate_event_id_is_a_noop_for_invocations - in_memory_duplicate_event_id_is_a_noop_for_values - 160 unit tests pass total. Reborn hooks_integration unchanged at 19 scenarios. Clippy/fmt/no-panics clean. Sync trait + Instant timestamps documented as a v1 choice; durable backends (Postgres, libSQL) will need an async companion trait using SystemTime — tracked in the scope doc as the next slice. Scope doc: crates/ironclaw_hooks/docs/successors/03-persistent-counter.md * fix(hooks): close codex P1 bugs in PredicateStateBackend in-memory impl Addresses codex P1 review on PR nearai#3635: P1 #1 — replay dedup loss under high-throughput keys The prior design used a fixed-size (256) recent_ids ring per bucket decoupled from entries. Under any workload with >256 distinct events in the same window, the first event's id aged out of the ring while its timestamp entry was still live, so a replay silently re-counted. Fix: dedup memory is now intrinsic to entries. Each entry stores (timestamp, event_id), and the dedup check is 'does any in-window entry have this id?'. Dedup memory is therefore exactly the in-window entry set — no fixed cap, no silent loss. P1 #2 — zombie buckets clogging LRU Two-part fix: 1. record_* drops empty buckets eagerly via history.remove(key). This is mostly defense-in-depth — under the new dedup design, the record path can't actually leave a bucket empty (proved in the test rationale comment). 2. evict_lru_* now preferentially targets empty buckets first (find any v.entries.is_empty()), only falling back to the oldest-timestamp scan if no empty bucket exists. Filter-out behavior is gone, so any empty bucket that somehow survives becomes the next eviction victim instead of a permanent zombie. Test changes (+2 new, -0 removed): - dedup_memory_covers_full_window_under_high_throughput: pushes 512 distinct events into one bucket, then replays event-0. Pre-fix this would have counted again (silent dedup loss); post-fix the replay is a no-op. - lru_evicts_empty_buckets_first: crafts an empty bucket alongside a live one, runs LRU eviction, asserts the empty one is evicted and the live one retained. Tests: 162 unit total (+2 new). Clippy/fmt/no-panics clean. * docs(hooks): address gemini review on persistent-counter scope doc Four medium-priority doc nits from gemini-code-assist on the crates/ironclaw_hooks/docs/successors/03-persistent-counter.md scope: 1. run_id in the trait: removed. The trait dedupes on event_id (RuntimeEventId is already run-scoped), not run_id. Replaces the earlier 'backend stores (timestamp, run_id, event_id)' claim. 2. SystemTime vs chrono::DateTime<Utc>: switched to DateTime<Utc> to match project convention (src/db/mod.rs, ironclaw_events). The in-memory backend keeps Instant for monotonic process-local semantics; durable backends require DateTime<Utc> for cross- process serialization. Documented as a clock note. 3. libSQL TEXT column for rust_decimal: per src/db/CLAUDE.md, libSQL can't preserve Decimal precision with numeric/real types. LibSqlPredicateStateBackend serializes value as TEXT via Decimal::to_string() / from_str(). Postgres impl keeps numeric (correct for PG). Documented as the two LibSql-specific schema differences. 4. Batched-writes vs cross-process consistency tension: gemini was right that deferring writes to the tick boundary breaks requirement #1 (two hosts would each see 'under cap' simultaneously). v1 production backend keeps writes synchronous; future optimization batches reads (not writes). * fix(hooks): thread stable caller_event_id through hook context (replay dedup) henrypark133 HIGH on PR nearai#3635 + serrrfirat HIGH #1: the `PredicateBackedBeforeCapabilityHook -> PredicateEvaluator` path always synthesized a fresh `event_id` per evaluation by mixing in a process-local atomic counter, so the same logical invocation retried/replayed always got a different id. The backend's UNIQUE constraint on `event_id` — the load-bearing dedup contract — never engaged on the real production path. Replay dedup was effectively "documented but unused." Plumb a stable per-invocation identity through the public hook surface: - `BeforeCapabilityHookContext` gains a `caller_event_id: Option<PredicateEventId>` field. Middleware that threads through from the calling layer's runtime event identity populates `Some(...)`; older / in-memory-only callers pass `None` and degrade to the current synth path (no behavior change). - New builder method `with_caller_event_id(...)`. - `PredicateEvaluator` resolves the id through a new `resolve_event_id` helper: prefer `ctx.caller_event_id`, fall back to `synth_event_id`. Both `record_invocation` and `record_value` paths use it. - Backend dedup behavior is unchanged — it was already correct on `event_id`. The bug was the caller path never supplying a stable id. Tests (caller-boundary, henrypark133's required regression): - `duplicate_caller_event_id_is_deduped_in_invocation_count`: two evaluations with the same `caller_event_id` count as one invocation; a third with a different id counts as two; a fourth crosses the cap. Sanity branch confirms the no-id synth path still exhibits "every call counts" semantics. This is the API contract slice. Wiring the middleware to actually supply a stable id (e.g. derived from the originating `RuntimeEventId` once that runs through the BeforeCapability path) is the follow-up that lights up the durable backend's end-to-end replay-safety promise. * refactor(hooks): demote PredicateStateBackend to pub(crate) (serrrfirat MED on PR nearai#3635) serrrfirat MED: the `predicate_state` module exposed `PredicateStateBackend` as `pub`, but the trait's `now: Instant` parameter is process-local and not serializable. Any external durable backend impl built against the current trait would have to be rewritten when the durable contract lands with `chrono::DateTime<Utc>` (see successor doc 03-persistent-counter.md). Hold the public surface back until that contract is stable so we don't ship a public API we know we'll break. Demoted to `pub(crate)`: - `PredicateStateBackend` (trait) - `InvocationKey`, `ValueKey` (key types — backend ABI only) - `PredicateBackendError` (error type, with `#[allow(dead_code)]` on the `Unavailable` variant since the in-memory backend is infallible and no durable backend exists yet) - `InMemoryPredicateStateBackend` (the only impl) - `PredicateEvaluator::with_backend` (with `#[allow(dead_code)]` — reserved for future internal injection paths) Kept `pub`: - `PredicateEventId` — it appears on the public hook surface via `BeforeCapabilityHookContext::caller_event_id` (from the nearai#3635 HIGH fix). Hook authors who want stable replay-dedup ids construct one. No behavior change. All 163 hooks lib tests + 19 reborn integration tests still pass. * fix(hooks): address henrypark133 must-fix #1-5 on PR nearai#3635 Five items from the 5-15 review: **#1 (must-fix) O(n) dedup scan** The previous `bucket.entries.iter().any(...)` linear scan held the outer history mutex while walking thousands of in-window entries at high throughput. Add a companion `HashSet<PredicateEventId>` per bucket (`InvocationBucket.dedup_ids` / `ValueBucket.dedup_ids`), maintained alongside the deque via `pop_front`/`push_back` helpers. O(1) dedup, same correctness, same memory bound (one set entry per in-window entry — no fixed ring). **#2 (must-fix) Mutex poison cascade** `.expect("predicate history mutex poisoned")` propagated a panic to every subsequent caller. Replace with `match self.invocation_history.lock() { Ok(g) => g, Err(p) => p.into_inner() }` so a poisoning thread doesn't take down all subsequent evaluations. **#3 (must-fix) `caller_event_id` format validation** `with_caller_event_id` now rejects empty strings and ids containing NUL bytes. Failed validation logs a `tracing::warn!` and leaves `caller_event_id == None` so the synth path takes over — operator sees the warning, predicate dedup still works. Also: `PredicateEventId(pub String)` → `PredicateEventId(String)` with `new()` / `as_str()` (henrypark133 nit #9). Inner field is no longer in-place mutable from outside the crate. **#4 (must-fix) `with_backend` is `#[cfg(test)]`** Previously `#[allow(dead_code)]` — reachable from release builds and inviting future callers to inject backends through an unstable seam. Gated to `cfg(test)`. **#5 (important) `evict_older_than` trait stub** Default-impl no-op added to `PredicateStateBackend` so the trait signature is locked before the first durable-backend PR. Trait-object callers won't break when durable impls override it. **Bonus** (henrypark133 missing-coverage #1): `in_memory_record_invocation_is_atomic_under_concurrent_writers` — 32 threads each record a distinct event id; final count must equal 32, proving the atomic record-and-read contract holds under contention. **Bonus** (henrypark133 nit #10): The third stable id in `duplicate_caller_event_id_is_deduped_in_invocation_count` was 62 chars; bumped to 64 to match the synth output format. * fix(hooks): clippy doc-list-indentation + remove unused with_backend (nearai#3635 CI) * fix(hooks): address serrrfirat HIGH + MEDIUM on PR nearai#3635 (5-15 review) **MEDIUM — `caller_event_id` validation bypass** `with_caller_event_id` validated for empty/NUL but the field on `BeforeCapabilityHookContext` is `pub`, so callers could direct- assign `Some(PredicateEventId::new("..."))` with `new()` permissive and bypass the setter entirely. Move validation INTO the type boundary: - `PredicateEventId::new(...) -> Result<Self, PredicateEventIdError>` validates non-empty + NUL-free at construction. Any value that reaches a downstream backend now satisfies the format invariant by construction. - `PredicateEventId::new_unchecked(...)` for internal synth paths and tests that mint ids from known-good shapes (hex digests). - `with_caller_event_id` drops its now-redundant runtime check; the type already enforces it. - Internal synth in `evaluator.rs` switches to `new_unchecked` (64-char hex output is always valid by construction). Tests: - `predicate_event_id_rejects_empty` - `predicate_event_id_rejects_nul_bytes` - `predicate_event_id_accepts_typical_hex_digest` **HIGH — durable schema: dedup scope mismatch** The successor doc's Postgres schema declared `event_id uuid PRIMARY KEY` (globally unique), but the trait's replay-refusal contract dedupes within the counter `key`. `caller_event_id` is per capability invocation — two predicate-backed hooks observing the same invocation share an id. A global PK lets the first hook's INSERT win and silently undercounts the second hook's bucket. - `docs/successors/03-persistent-counter.md`: PK changes to composite `(tenant_id, hook_id, capability, event_id)` for invocations and `(tenant_id, hook_id, capability, field, event_id)` for values, matching the trait's per-key dedup scope. - `predicate_state.rs` trait doc: replay-refusal section rewritten to spell out the per-key scope and the corresponding `INSERT … ON CONFLICT (tenant, hook, capability[, field], event_id) DO NOTHING` shape durable backends should use. * docs(hooks): document host-assigned trust boundary on PredicateEventId henrypark133 / serrrfirat blocker B4 on PR nearai#3635: the `caller_event_id` threading through `BeforeCapabilityHookContext` partially shipped earlier (commit b4d8a35), but the trust-boundary documentation explaining the host-assigned invariant was still missing. Add rustdoc to `PredicateEventId` and the `PredicateStateBackend` trait clarifying that: - the id MUST be minted by trusted host code from authoritative sources (dispatcher RuntimeEventId, host-side hash, arguments digest) - it MUST NOT pass through unchanged from any tenant-controlled surface (capability arguments, manifest fields, WASM memory, HTTP bodies) - the format invariants in `PredicateEventId::new` (non-empty, NUL-free) are a durability contract for SQL backends, NOT a trust check - a tenant-supplied id can either undercount itself into infinity by replaying a fixed id, or poison adjacent buckets if scoping is ever weakened Doc-only; no behavior change. * test(hooks): add caller-boundary replay-dedup test through wrapper hook henrypark133 HIGH blocker B1 on PR nearai#3635: replay dedup must engage at the caller boundary — `PredicateBackedBeforeCapabilityHook::evaluate` is the production path the dispatcher invokes for installed predicate hooks. A unit test on `PredicateEvaluator::evaluate_at` alone is insufficient regression coverage (repo CLAUDE.md rule "Test through the caller, not just the helper"): the wrapper hook reads `BeforeCapabilityHookContext::caller_event_id` and threads it down to the backend, so the regression test must drive the wrapper itself. The threading work already shipped in commit e6df47d (`caller_event_id` field on the public hook context + evaluator preferring it over the synth path). This commit adds the missing end-to-end test: 1. Two `PredicateBackedBeforeCapabilityHook::evaluate` calls with the same `caller_event_id` and a `RateOrValueCap { max: 1 }` predicate — the second call must stay under cap (dedupe engages at the wrapper boundary, not be re-counted into a deny). 2. A third call with a DISTINCT `caller_event_id` crosses the cap — proving dedup is replay-scoped (same id → no-op), not blanket- suppress (any id → no-op). If the wrapper were synthesizing a fresh id per call (the bug Henry flagged before threading landed), this test would fail at step 2 with the second evaluation being denied. * docs(hooks): D5a + cross-process replay note; add caller-API tests henrypark133 should-fix S8 + S9 on PR nearai#3635. S8 — threat-model expansion: - Add D5a as the correctness-under-attack variant of D5: an attacker flooding high-cardinality keys can LRU-evict legitimate tenants' counters and reset their rate-limit state. Distinct from the memory-only framing of D5; tied back to per-extension caps (D3/D4) and the durable-backend successor (doc 03). - Document the cross-process replay limit on the in-memory backend inside the PredicateStateBackend trait docs, not just in D5 — the process-local dedup is a property callers need at the trait surface, with a pointer to the durable backend as the cross-host story. S9 — three new tests on the in-memory backend public API: - lru_eviction_via_public_api_holds_max_history_keys_cap: drives MAX_HISTORY_KEYS + 1 distinct keys through record_invocation and asserts the map size cap holds + evictions_observed() advances. The previous coverage manually crafted buckets and called the LRU helper directly; this exercises the production path. - in_memory_invocation_retains_entry_at_exact_window_cutoff: pins the `< cutoff` trim semantics so a refactor to `<=` would fail loud. - event_id_dedup_is_isolated_across_invocation_and_value_maps: same event_id used in both record_invocation and record_value must not cross-suppress — the two maps key on disjoint types. The fourth S9 item (concurrent N-thread atomicity) and the caller- boundary replay test on the wrapper hook already landed in earlier commits (f632d22, predicate_state.rs line 840). S2 (evict_older_than stub), S3 (sync-trait docs), and S7 (consistency vs batched-writes) were also already in HEAD; this commit ships the remaining items. Quality gate: cargo fmt clean, cargo clippy -p ironclaw_hooks --all-features --tests -D warnings clean, full hooks test suite green (15 predicate_state unit tests + lib + integration). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * chore(hooks): co-locate synth_event_id with backend; rationale comment; pin synth format henrypark133 nits N1, N2, N5 on PR nearai#3635. N1 — Move `synth_event_id` from `evaluator.rs` to `predicate_state.rs` as `PredicateEventId::synth(...)`. The id format (64-char lowercase hex, no NUL, never empty) is part of the backend's durable contract, so co-locating with `PredicateStateBackend` keeps the format change- surface adjacent to the consumer. To avoid inverting the module dependency (`predicate_state` is a leaf below `points`), the synth helper takes raw bytes / &str rather than a `&BeforeCapabilityHookContext`. The evaluator's `resolve_event_id` fallback unpacks the context and delegates. N2 — `// safety:` comment on a non-`unsafe` block (the `write!(s, "{byte:02x}")` infallibility note) renamed to `// RATIONALE:`. By convention `// SAFETY:` pairs with `unsafe` blocks; using `// safety:` elsewhere conflates the two. N5 — Add `synth_event_id_is_64_char_lowercase_hex` to pin the synth output shape. A refactor that silently changes length or case would break the durable backend's `uuid`-shaped UNIQUE constraint without a test failure today; the new test fails loud. Quality gate: cargo fmt clean, cargo clippy --all --benches --tests --examples --all-features -D warnings clean, full hooks lib test suite green (172 passing including the new pin). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(hooks): drop expect() in hex formatting to satisfy panic CI check The "No panics in production code" CI check (scripts/check_no_panics.py) only recognizes `// safety:` suppression markers, not `RATIONALE:`. Since std::fmt::Write for String is infallible, just discard the Result with `let _ =` instead of `.expect()` — no panic call, no marker needed. Also merges in latest origin/hooks-foundation-01 (now includes the reborn-integration merge and PR nearai#3636). * fix(hooks): narrow caller_event_id visibility to pub(crate) henrypark133 MED on PR nearai#3635 5-19 review. The pub field let external callers bypass with_caller_event_id and assign values the validated PredicateEventId constructor would have rejected. Force every external caller through the typed setter so PredicateEventId::new is the only entry point. * fix(hooks): drop arguments_digest from synth + add in-memory backend warn Two PR nearai#3635 5-19 review findings on the evaluator surface: - henrypark133 LOW (synth oracle): drop arguments_digest from the PredicateEventId synth hash input. The 64-char hex output was an equality oracle for argument shape; replay dedup for durable backends uses the caller-supplied caller_event_id, not the synth path, so synth only needs to be per-call unique, not content-addressed. - henrypark133 HIGH + MED (in-memory production limits): expose PredicateEvaluator::warn_in_memory_backend_active_in_production for hosts to call at startup. Multi-host replay dedup is process-local and the LRU cap is shared across tenants; operators need this surfaced in logs when the durable backend is not wired. * fix(hooks): harden predicate state backend per PR nearai#3635 5-19 review Address five findings on crates/ironclaw_hooks/src/predicate_state.rs: - A1 (henrypark133 HIGH): restrict PredicateEventId::new_unchecked to pub(crate) so external callers cannot bypass the durable UNIQUE-constraint format invariants enforced by ::new. - A4 (henrypark133 MED): per-tenant LRU quota at MAX_HISTORY_KEYS / 4. Without it a noisy tenant could fill the global cap and evict a quiet tenant's bucket, resetting their rate-limit counter. With the quota, a tenant that overflows evicts its OWN oldest-front bucket first. New tests cover single-tenant cap and cross-tenant isolation. - A5 (henrypark133 LOW): drop arguments_digest from the synth hash input (oracle closure mirrored from the evaluator side). Add a thread-local nonce alongside the process-global counter so synth remains per-call unique without relying solely on a contended AtomicU64. New test pins the divergence invariant. - D6 (henrypark133 HIGH): O(1) NumericSum via an incrementally- maintained ValueBucket::running_sum, replacing the O(n) deque walk on every record_value call. New test covers push/trim/replay interactions. - D8 (henrypark133 MED): implement evict_older_than for the in-memory backend (was a no-op Ok(0) default). Drops entries strictly older than the cutoff and removes empty buckets; operator reaper tasks rely on this to reclaim memory from idle keys. - D7 (henrypark133 MED, partial): document the process-global synth COUNTER as a known contention hotspot and add a thread-local nonce so threads can advance without forcing cross-core invalidation in the common path. Tests: 196 passing (+4 new); workspace clippy clean. * fix(hooks): port MAX_SAMPLES_PER_KEY cap into PredicateStateBackend (D5 regression from r3) Round 3 of PR nearai#3573 (already merged into hooks-foundation-01) added an inline per-key sample cap of 4_096 in evaluator.rs to bound memory under attacker-triggered hot capabilities with very large declared windows (threat-model finding D5). The predicate-state extraction in PR nearai#3635 moved that bookkeeping into the PredicateStateBackend trait but missed porting the cap, so the cap would silently disappear from production once this PR rebases onto the foundation branch. This commit moves the cap into the in-memory backend impl next to MAX_HISTORY_KEYS / MAX_KEYS_PER_TENANT and enforces it in both record_invocation and record_value. For the NumericSum path, the bucket helper's pop_front already decrements running_sum, so the incremental sum invariant survives cap-driven eviction. The pre-existing inline copy in evaluator.rs becomes redundant once the trait impl owns the enforcement; the rebase resolution deletes it. Adds two regression tests: - record_invocation_caps_samples_per_key_under_attacker_pressure - record_value_evicts_oldest_keeping_running_sum_consistent * fix(hooks): port predicate_state tests to ::new() after nearai#3912 newtype privatization --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Successor PR from #3573. Draft — scope doc only, no implementation yet.
Scope
Introduce `DenyReasonCode` enum with curated vocabulary (`RateLimit`, `ValueCap`, `Blocklist`, `RequiresApproval`, `OutOfPolicy`, `Generic`). Extends `OnExceededAction` with `DenyWithCode { code, reason }` so hook authors can communicate why a deny happened in the model-visible label without opening a free-form prompt-injection channel.
Design doc
`crates/ironclaw_hooks/docs/successors/06-deny-reason-code.md`
Threat-model
Initial vocabulary kept coarse-grained — too descriptive labels would let authors smuggle content via label choice.
Status
Draft for design review. Small + bounded; ~3 tests + back-compat path retained.