feat(hooks): event-triggered hooks Phase 5 (successor to #3573) - #3640
Conversation
There was a problem hiding this comment.
Code Review
This pull request introduces a design document for event-triggered hooks, which are asynchronous, observer-only hooks designed to react to runtime events without blocking the main loop. The review feedback highlights several internal inconsistencies and contradictions within the document that need clarification. Specifically, the reviewer noted discrepancies regarding the use of narrowed event projections versus the full RuntimeEvent type, the available methods on the ObserverSink trait, the policy for replaying missed events during downtime, and the current state of crate dependencies.
| ```rust | ||
| // in ironclaw_hooks::points::event_triggered | ||
| pub struct EventHookContext<'a> { | ||
| pub event: &'a RuntimeEvent, |
There was a problem hiding this comment.
The code snippet uses the full RuntimeEvent type, which contradicts the recommendation in the Cross-cutting constraints section (line 105) to pick a narrowed projection (e.g., HookObservableEvent). Using a narrowed projection is preferred to maintain clean crate boundaries. Documentation for complex logic must precisely match the intended implementation.
References
- Documentation for complex logic, such as security policies, must precisely match the code implementation.
There was a problem hiding this comment.
Addressed in 7368139 (refactor) and the doc rewrite in cluster G of #3640 (comment).
The surface sketch in 04-event-triggered-hooks.md now includes a NOTE block at lines 47-53 calling out that the full RuntimeEvent shown is the Phase 5 shipping shape, and the longer-term target is a narrowed HookObservableEvent projection (tracked as issue #3690). The Cross-cutting constraints section (lines 122-129) cross-references the same follow-up. Doc and code now agree: Phase 5 ships RuntimeEvent, narrowing is explicitly deferred.
| pub tenant_id: TenantId, | ||
| } | ||
|
|
||
| // Sink mirrors ObserverSink — `note_fact`, `emit_audit`. No `allow`, |
There was a problem hiding this comment.
The documentation mentions note_fact and emit_audit as methods on the sink, but the existing ObserverSink trait only defines a note method. Align this documentation with the actual trait definition or clearly distinguish these as future or planned states. Documentation must precisely match the code implementation and distinguish between current and future states.
| // Sink mirrors ObserverSink — `note_fact`, `emit_audit`. No `allow`, | |
| // Sink mirrors ObserverSink — note. No allow, |
References
- Documentation for complex logic, such as security policies, must precisely match the code implementation.
- In documentation, clearly distinguish between the current implementation state and future or planned states.
There was a problem hiding this comment.
Addressed in 7368139 (refactor) — EventTriggeredObserverSink was dropped, and the sink shown in the doc sketch (lines 67-78) is now the actual ObserverSink trait with its single note(category, summary) primitive. The "No allow, no deny, no patch" comment matches the trait. The earlier note_fact / emit_audit references that never landed are explicitly called out and corrected in the No re-emission paragraph (lines 138-142).
| (tick-driven or stream-driven, TBD). | ||
| - **Cursor / replay**: subscriptions are cursor-keyed so a restarted | ||
| host can resume from the last-seen `EventCursor`. Lost events | ||
| during downtime is acceptable for observer-only semantics; |
There was a problem hiding this comment.
There is a contradiction between the statement that Lost events during downtime is acceptable and the later requirement for Cursor resume (line 133) which specifies replaying missed events. The design should clarify which semantic is intended to ensure the documentation precisely matches the implementation logic.
References
- Documentation for complex logic, such as security policies, must precisely match the code implementation.
There was a problem hiding this comment.
Addressed in 7368139 cluster G. The at-least-once vs "lost events" wording was a contradiction in the earlier draft; the doc now reads:
Events emitted before a subscription's
start_cursorare not delivered. Concretely: the contract is at-least-once, not at-most-once; on a restart from the same cursor, every event with cursor>= start_cursoris replayed.
"Lost during downtime" only ever meant "events before start_cursor are not replayed," which is consistent with at-least-once for everything >= start_cursor. Hooks set ctx.is_replay = true (also added in 7368139) to dedupe by event.event_id.
|
|
||
| ## Cross-cutting constraints | ||
|
|
||
| - **Cross-crate boundary**: `ironclaw_hooks` already forbids `events` |
There was a problem hiding this comment.
The statement that ironclaw_hooks already forbids a dependency on events is contradicted by the Risk section (line 159), which notes that ironclaw_events is already a dependency. Clarify the constraint to ensure the documentation precisely matches the system architecture.
References
- Documentation for complex logic, such as security policies, must precisely match the code implementation.
There was a problem hiding this comment.
Addressed in 7368139 cluster G. The doc now states (lines 117-121, Phase 5 implementation notes):
ironclaw_hooksdepends onironclaw_eventsas of PR #3640 — the milestone projection in PR #3573 coveredironclaw_reborn's milestone wiring, but the hook crate itself did not gain the dep until the event-triggered consumer landed.
And in Cross-cutting constraints (lines 122-126): ironclaw_hooks already depends on ironclaw_events for the milestone projection (PR #3573)…. The Risk section reference is consistent with this — ironclaw_events is a dep, not forbidden — the forbidden deps remain host_runtime / network / secrets / wasm / reborn internals.
There was a problem hiding this comment.
💡 Codex Review
Here are some automated review suggestions for this pull request.
Reviewed commit: ce610736dc
ℹ️ About Codex in GitHub
Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you
- Open a pull request for review
- Mark a draft as ready
- Comment "@codex review".
If Codex has suggestions, it will comment; otherwise it will react with 👍.
Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".
| if !binding | ||
| .scope | ||
| .permits(binding.owning_extension.as_ref(), event.provider.as_ref()) | ||
| { |
There was a problem hiding this comment.
Preserve OwnCapabilities matching for providerless hook events
Scope filtering for event-triggered hooks is based on event.provider, but RuntimeEvent::hook_failed and RuntimeEvent::hook_decision_emitted populate provider as None (see crates/ironclaw_events/src/runtime_event.rs constructors). As a result, Installed hooks using the default OwnCapabilities scope are always filtered out for those event kinds, so a hook-failure subscription silently never fires unless it is widened to Global/SameTenant. This breaks the expected default behavior for extension-scoped failure observers.
Useful? React with 👍 / 👎.
There was a problem hiding this comment.
Fixed in two commits on this branch:
d7f43a3ff— adds ahook_id-based fallback inscope_provider_for_runtime_eventso that whenevent.providerisNoneonHookDispatched/HookDecisionEmitted/HookFailedevents, the dispatcher resolves the owning extension by looking upevent.hook_idagainst the registry's hex index.OwnCapabilitiesmatching is preserved for the default Installed-hook case.4bb06c93a— the durable fix: plumbsowning_extension: Option<ExtensionId>end-to-end throughLoopHostMilestoneKind::{HookDispatched, HookDecisionEmitted, HookFailed}and theRuntimeEvent::hook_*constructors, soevent.provideris populated at emit time and the primary path resolves without any fallback. Checkpoint payloads round-trip unchanged (#[serde(default, skip_serializing_if = "Option::is_none")]).
Regression tests in crates/ironclaw_reborn/tests/hooks_integration.rs:
event_triggered_own_capabilities_matches_hook_failed_with_carried_provider(primary path)event_triggered_own_capabilities_scope_resolves_hook_failed_owner_from_hook_id(legacy/None-provider fallback)
SummaryReviewed PR #3640 only. Base PR adds event-triggered hooks over durable event replay. Merge stance: blocking security/correctness findings at subscription boundary and replay/reentrancy behavior. Findings
Security/data-flow notes
Correctness/invariant notes
Missing tests
|
henrypark133
left a comment
There was a problem hiding this comment.
What looks good:
- Event-triggered hooks are observer-only at the sink/type boundary.
- Subscription runs from the durable log on a background task, so event emitters are not blocked by hook execution.
- Cursor replay, kind filtering, scope filtering, and task cleanup all have useful coverage.
Findings:
- High -
crates/ironclaw_hooks/src/dispatch.rs:887: event-triggered scope filtering usesevent.provider, butRuntimeEvent::hook_decision_emittedandRuntimeEvent::hook_failedsetprovider: Noneincrates/ironclaw_events/src/runtime_event.rs:550and:576. Why it matters: installed hooks default toOwnCapabilities, so the common/default event-triggered hook silently never fires for the hook-failure/decision events this PR is explicitly meant to observe. Expected fix direction: carry the originating provider into hook milestone runtime events, or resolve provider during replay/dispatch before applyingOwnCapabilities, with a regression test forOwnCapabilities + HookFailed.
Summary:
- Recommended verdict: Request changes
- Prior feedback status: current Codex P1 is confirmed.
- Residual risk: failure is silent non-observation, not loop corruption, but it undermines the Phase 5 default use case.
henrypark133 HIGH + codex P1 on PR #3640: `OwnCapabilities`-scoped event-triggered subscriptions silently never fired for `HookFailed`/`HookDecisionEmitted`/`HookDispatched` events because those `RuntimeEvent` constructors hardcoded `provider: None`. Since Installed hooks default to `OwnCapabilities`, the very events that Phase 5 was designed to observe (hook-failure / decision alerting) never reached their default-configured subscriber. A prior fix added a hook_id-based fallback in `scope_provider_for_runtime_event` that resolves the owning extension through the registry's hex index when `event.provider` is `None`. That covers the case where the failing hook is still registered at replay time, but the durable fix is to stamp the originating provider into the event at emit time so the primary `event.provider` path resolves without any fallback. Plumbed `owning_extension: Option<ExtensionId>` end-to-end: - `LoopHostMilestoneKind::{HookDispatched, HookDecisionEmitted, HookFailed}` gain the field (with `#[serde(default, skip_serializing_if = "Option::is_none")]` so pre-existing checkpoint payloads and the L3 schema-snapshot tests round-trip unchanged when no owner is set). - `RuntimeEvent::hook_{dispatched, decision_emitted, failed}` constructors accept the owner and stamp it into `provider`. - `milestone_events.rs` threads the field through the projection. - `HookDispatcher::emit_dispatched/emit_decision` pass `binding.owning_extension.clone()` directly. - `HookDispatcher::emit_failure` (no binding handy on the failure path) looks the owner up via the registry's existing `owning_extension_for_hook_hex` index. Tests: - `event_triggered_own_capabilities_matches_hook_failed_with_carried_provider`: primary-path regression — two `HookFailed` events with `provider: Some(ext_a|ext_b)` against an `OwnCapabilities` subscription scoped to ext_a; only the own-provider event fires and `event.provider == Some(ext_a)`. - Existing `event_triggered_own_capabilities_scope_resolves_hook_failed_owner_from_hook_id` remains green: passes `None` for the new arg so the fallback path is still exercised for legacy payloads. All other call sites updated to pass `None` (no owner available) or the resolved owner where applicable.
…rfirat HIGH #1 on PR #3640) `EventTriggeredHookSubscription` accepted a caller-supplied `EventStreamKey` + `ReadScope` and used `run_context.scope.tenant_id` as the hook context's tenant — with no validation that the two agreed. A caller wiring tenant A's host with tenant B's stream would cause hooks to observe B's events while the hook context claimed tenant A. Cross-tenant trust-boundary break. Add `EventTriggeredHookSubscription::validate_against_run_scope` and call it from `build_text_only_host_with_capabilities` before spawning. Validation: - Stream `(tenant_id, user_id, agent_id)` must equal `(run_context.scope.tenant_id, thread_scope.owner_user_id, run_context.scope.agent_id)`. - Thread without `owner_user_id` cannot bind any subscription — the user dimension is required to verify stream identity. - Every `Some(want)` in `ReadScope` must equal the corresponding run/thread scope value (project/mission/thread). `None` is permissive (run scope owns the dimension authoritatively). Failures surface as `RebornLoopDriverHostError::ScopeMismatch` with a specific reason naming the offending dimension. Tests: - `event_triggered_subscription_with_foreign_tenant_stream_fails_host_build` - `event_triggered_subscription_with_foreign_user_stream_fails_host_build` The integration fixture's `ThreadScope` now sets `owner_user_id: Some(...)` so it passes validation; previously it was `None`, which the new check (correctly) refuses. Existing tests continue to pass.
…rrfirat MED on PR #3640) When the durable event log returned `EventError::ReplayGap`, the event-triggered subscription's background task previously logged a `tracing::warn!` and broke out of the poll loop — silently killing all future hook event delivery for the run with no operator-visible signal. A scoped audit hook that mattered to compliance would just stop, and nobody downstream would know. Surface the termination through the host's milestone sink: - New `LoopDriverNoteKind::EventSubscriptionTerminated` variant. - The subscription's `spawn`/`run` now takes the host's `Arc<dyn LoopHostMilestoneSink>` and the active `LoopRunContext`. On `ReplayGap`, it constructs a `DriverNote` milestone with that kind plus a `LoopSafeSummary` describing the gap, publishes it through the same sink that carries every other host milestone, and *then* breaks (fail-closed: the at-most-once contract is already broken; resuming from `earliest` would silently lose the gap). - Log level bumped from `warn` to `error` to match the severity. - A best-effort send: failures to publish the milestone are logged but do not stall the subscription teardown. Tests: - `event_triggered_replay_gap_emits_subscription_terminated_milestone`: appends 3 events, `truncate_before_or_at` to cursor 2 to force a replay gap, starts the subscription from cursor origin (now stale), and asserts a `DriverNote { kind: EventSubscriptionTerminated, .. }` shows up on the host's milestone sink within a 2s deadline. Self-emit reentrancy (serrrfirat MED #3 on the same PR) is intentionally not addressed here — that fix needs a design call (task-local re-entry flag vs. removing RuntimeEvent emit capability from event-hook execution contexts) and is a follow-up.
on PR #3640) A hook that subscribes to one of the hook-lifecycle event kinds (`HookDispatched`/`HookDecisionEmitted`/`HookFailed`) with a scope that matches its own provider would otherwise be dispatched for events describing its OWN executions. The dispatcher emits those events itself when running the hook, so a hook subscribing to `HookFailed` with `OwnCapabilities` against its own extension would fail → emit HookFailed → re-dispatch → fail → emit → … storm. `dispatch_event_triggered_at` now skips events whose `event.hook_id` equals the binding's own hook id when the event kind is a hook- lifecycle kind (`is_hook_lifecycle_kind`). The check is intentionally narrow: - It only fires for hook-lifecycle events. Subscriptions to other event kinds are unaffected. - It only suppresses literal self-observation; events about other hooks (even hooks from the same extension) still dispatch. This does NOT cover the broader case of a hook that captures an `Arc<DurableEventLog>` and mints arbitrary `RuntimeEvent`s from inside its `observe()`. That requires architectural restriction on what hook impls can capture — tracked separately as a follow-up. Tests: - `event_triggered_self_lifecycle_event_does_not_redispatch`: appends two `HookFailed` events with the same provider — one targeting the subscriber's own hook id, one targeting a different hook. Asserts only the OTHER hook's failure fires (proves the filter is narrow, not blanket).
Code Review — PR #3640 (event-triggered hooks, Phase 5)Verdict: COMMENT (draft) — design and implementation are sound and well-structured. Several items should be resolved before promotion from draft. OverviewNew IssuesShould fix before promotion from draft:
Non-blocking nits:
Missing test coverage
Security
|
Four items from the 5-15 review (#4 DoS budget and #5 narrowed projection deferred — see below): **#1 (should-fix) Invariant: EventTriggered ↔ event_kind_filter** `HookRegistry::insert` now enforces the biconditional at install time: an `EventTriggered` binding must declare an `event_kind_filter` (otherwise the dispatcher's kind match would silently never fire — a no-op binding), and conversely only `EventTriggered` bindings may declare a filter (other points are kind-agnostic and would ignore the field). Misconfigured bindings fail loud at install. **#2 (should-fix) Remove `Clone` derive on EventTriggeredHookSubscription** `Clone` on a spawn-semantics type was a footgun: external callers cloning + spawning twice would create two consumers reading from the same `start_cursor`, each dispatching every hook. Replace with an explicit `clone_for_independent_spawn(&self)` method named verbosely so the property is visible at the seam. Internal use updated in the factory's host-build path; external callers can no longer accidentally construct a dual-consumer pattern. **#3 (should-fix) catch_unwind around the background `run()` task** The subscription's tokio task body now runs inside `AssertUnwindSafe(...).catch_unwind()`; a panic in `run()` emits the same `EventSubscriptionTerminated` `DriverNote` milestone the `ReplayGap` path already emits, instead of silently terminating with no operator-visible signal. **#6 (should-fix) Replay semantics in rustdoc on public API** Added a "Replay semantics" section to `EventTriggeredHookSubscription` rustdoc: at-least-once, caller-owned cursor persistence, the restart-from-start_cursor replay pattern. Previously only in the design doc; now load-bearing API contract is visible at the type. **#4 (deferred) Per-hook DoS budget for Installed tier** Henry's recommendation was to gate `Installed`-tier event-triggered hooks entirely until the budget design lands, allowing only Builtin/Trusted. That breaks 11 existing tests + the primary use case. Instead: documented the existing first-line throttle (`batch_limit` × `poll_interval`) as the current bound on indirect- recursion fanout, and tracked the full per-hook rate cap with poisoning + milestone-on-overrun as a follow-up. The self-trigger guard (committed earlier in this PR) catches the most common direct pattern; the throttle here bounds the indirect pattern until the proper budget lands. **#5 (deferred) Narrowed `HookObservableEvent` projection** Would prevent full `RuntimeEvent` surface from reaching Installed- tier hooks. Project-wide impact (events crate types, projection glue). Tracked as a follow-up; the existing sanitized-event projection bounds the surface to closed-vocab labels. All 156 hooks lib + 30 reborn integration tests pass.
|
@henrypark133 thanks for the review. Addressing the High finding ( Fixed in two commits, in the order you suggested:
Regression tests in
Full Codex P1 on |
Bundle three nit-tier review items into a single commit: **#9 Replace author-internal tags with NOTE(#3640)** The Phase-5 PR (#3640) had several `serrrfirat HIGH/MED #N on PR #3640` comment tags in this PR's diff. These are review-internal scaffolding, not load-bearing for future readers. Replaced with `NOTE(#3640)` in: - crates/ironclaw_hooks/src/dispatch.rs (self-observation guard) - crates/ironclaw_reborn/src/loop_driver_host.rs (scope validation, replay-gap milestone, subscription binding) - crates/ironclaw_reborn/tests/hooks_integration.rs (three regression tests covering scope validation, self-observation suppression, and replay-gap surfacing) - crates/ironclaw_turns/src/run_profile/host.rs (`EventSubscriptionTerminated` doc) **#10 Replace 10ms spin-poll with tokio::sync::Notify** `wait_for_seen_events` polled the shared `Mutex<Vec<SeenRuntimeEvent>>` every 10 ms until the expected count was reached. Replaced with a `SeenLog` newtype that pairs the events vec with a `Notify`; the hook's `observe()` calls `seen.push(...)` which signals `Notify::notify_one`, and `wait_for_seen_events` parks on `notified().await` under a `tokio::time::timeout`. `notify_one` is a permit-store, so an event landing between snapshot and wait still wakes the waiter immediately. Test latency drops from ~10 ms median to sub-ms and is no longer rate-limited by the polling cadence. All 30 hooks_integration tests still pass. **#11 Remove unused Clone derive on EventTriggeredHookContext** No call site clones the context — it's passed by reference. Dropped the derive to make the borrow contract clearer.
|
@henrypark133 @serrrfirat addressing the 5/15 review batch. Status of each finding below. serrrfirat (MED)MED #2 ( MED #3 (no source tag / self-author suppression beyond hook-id exact match): resolved earlier in this PR — henrypark133should-fix #1 ( should-fix #2 ( should-fix #3 (panic in should-fix #4 (per-hook DoS dispatch budget for Installed-tier event hooks): deferred to follow-up issue #3689. Henry's stricter recommendation was to gate Installed-tier event hooks entirely until the budget lands, which breaks the primary use case + 11 tests. Interim mitigations remain: the subscription throttle ( should-fix #5 (narrowed should-fix #6 (replay semantics on nit #7 (author-internal tags): resolved in nit #8 ( nit #9 (unused nit #10 ( gemini-code-assist (doc fixes)All four resolved in
|
|
Thanks for the update, @zmanian. The resolutions for the identified issues and the follow-up tracking for the DoS budget and narrowed event projection look correct and address the concerns raised in the review. The use of |
|
@henrypark133 Rebased onto latest Re: HIGH finding #1 ( This was already addressed on the branch in commits The merge preserved that plumbing — I had to union it with the new Could you take another look when you have a moment? |
henrypark133
left a comment
There was a problem hiding this comment.
Code Review (multi-agent)
Intent: Add EventTriggered hook point that subscribes to the runtime event bus and reacts to RuntimeEvents asynchronously, outside the inline dispatch tick.
Stats: 2 findings (from 20 raw, 2 after filtering to changed files) across 1 file. Reviewers run: security, bugs, performance, tests, conventions, design. Reviewers failed: 0. Body-only: 2
Security
- Low — Hook dispatcher uses .expect() on mutex in production dispatch paths (
crates/ironclaw_hooks/src/dispatch.rs:224-226, confidence 50) — anchor: crates/ironclaw_hooks/src/dispatch.rs:224
Multiple dispatch-loop methods call .expect() on the registry Mutex. A poisoned mutex would panic the dispatch thread. Fix: Replace .expect() with proper error handling.
Design
- Medium — Duplicated lifecycle-kind match duplicates is_hook_lifecycle_kind helper (
crates/ironclaw_hooks/src/dispatch.rs:1015-1020, confidence 75) — anchor: crates/ironclaw_hooks/src/dispatch.rs:1369
scope_provider_for_runtime_event repeats the same RuntimeEventKind match that is_hook_lifecycle_kind already encapsulates. Fix: Replace inline matches! with !is_hook_lifecycle_kind(event.kind).
Note: This is a docs-only draft PR. Most reviewer findings were on pre-existing code and filtered out. The 2 findings above are the only ones on changed files.
henrypark133
left a comment
There was a problem hiding this comment.
Code Review (multi-agent)
Intent: Add EventTriggered hook point that subscribes to runtime event bus and reacts to durable RuntimeEvents asynchronously outside the inline dispatch tick
Stats: 11 findings (from 15 raw, 11 after dedup) across 6 files. Reviewers run: security, bugs, performance, tests, conventions, design. Reviewers failed: none. Body-only: 11
Security
-
High Full RuntimeEvent exposed to Installed-tier (untrusted) event-triggered hooks (
crates/ironclaw_hooks/src/points/event_triggered.rs:14-18, confidence 100) — anchor: crates/ironclaw_hooks/src/points/event_triggered.rs:15
EventTriggeredHookContext hands the full &RuntimeEvent to all trust tiers including Installed (third-party/untrusted) hooks. The embedded ResourceScope carries tenant_id, user_id, agent_id, project_id, mission_id, thread_id, and invocation_id.
Fix: Introduce a narrowed HookObservableEvent projection for Installed-tier hooks that strips or redacts cross-extension identifiers. -
Medium Indirect mutual recursion between hooks can cause dispatch storm (
crates/ironclaw_hooks/src/dispatch.rs:933-1010, confidence 75) — anchor: crates/ironclaw_hooks/src/dispatch.rs:975
The self-trigger guard only prevents direct self-observation. It does not prevent indirect mutual recursion between hooks.
Fix: Add a per-hook dispatch rate counter with a configurable budget; when exceeded, poison the binding. -
Medium At-least-once replay causes duplicate side effects without API-level guardrails (
crates/ironclaw_reborn/src/loop_driver_host.rs:395-408, confidence 75) — anchor: crates/ironclaw_reborn/src/loop_driver_host.rs:400
Side-effecting hooks will fire again for already-processed events on restart. No idempotency key or replay signal exists.
Fix: Add an is_replay: bool field to EventTriggeredHookContext so hooks can implement idempotent side effects.
Performance
-
High O(n) linear scan of all event-triggered bindings per event — should be indexed by kind (
crates/ironclaw_hooks/src/dispatch.rs:933-1010, confidence 85) — (no diff position — body only)
dispatch_event_triggered_at iterates ALL event-triggered bindings checking event_kind_filter, doing O(k) work per event.
Fix: Add a per-kind index (HashMap<RuntimeEventKind, Vec>) alongside by_point. -
High Subscription poll loop has no backpressure adaptation — dispatch backlog grows unbounded (
crates/ironclaw_reborn/src/loop_driver_host.rs:687-730, confidence 75) — (no diff position — body only)
Fixed 50ms poll interval with batch_limit=64 regardless of dispatch duration. Events accumulate faster than consumed.
Fix: Implement adaptive polling based on wall-clock time spent per dispatch pass. -
Medium Registry mutex held on every event dispatch for scope resolution causes contention under fanout (
crates/ironclaw_hooks/src/dispatch.rs:1012-1048, confidence 70) — (no diff position — body only)
Multiple concurrent dispatch_event_triggered_at calls serialize on Arc<Mutex>.
Fix: Cache owning_extension_for_hook results in a sidecar HashMap populated at install time. -
Medium Cursor advances in-memory with no crash-safe persistence — at-least-once duplicates on restart (
crates/ironclaw_reborn/src/loop_driver_host.rs:634-665, confidence 60) — (no diff position — body only)
Subscription tracks cursor state only in a local variable. Process crash means replay from start_cursor.
Fix: Persist cursor after each successful batch dispatch via a small append-only store.
Tests
-
High dispatch_event_triggered_at missing-impl path not tested (
crates/ironclaw_hooks/src/dispatch.rs:933-1006, confidence 75) — anchor: crates/ironclaw_hooks/src/dispatch.rs:979
No test covers the case where a binding exists but no hook impl is installed (poison slot + Malformed failure).
Fix: Add tests::dispatch::event_triggered_missing_impl_poisons_slot test. -
Medium No test rejects non-event binding that carries an event_kind_filter (
crates/ironclaw_hooks/src/registry.rs:206-212, confidence 75) — anchor: crates/ironclaw_hooks/src/registry.rs:206
New validation rejects bindings at non-event points with event_kind_filter, but no test exercises this path.
Fix: Add tests::registry::rejects_event_kind_filter_on_non_event_point test. -
Medium RuntimeEvent hook constructors untested with Some(owning_extension) (
crates/ironclaw_events/src/runtime_event.rs:510-591, confidence 75) — anchor: crates/ironclaw_events/src/runtime_event.rs:522
All three serde round-trip tests pass None for owning_extension and never assert event.provider.
Fix: Add tests::runtime_event::hook_dispatched_sets_provider_from_owning_extension test. -
Medium scope_provider_for_runtime_event mutex poison fallback never exercised (
crates/ironclaw_hooks/src/dispatch.rs:1024-1032, confidence 75) — anchor: crates/ironclaw_hooks/src/dispatch.rs:1024
Defensive path when registry mutex is poisoned returns None but is never tested.
Fix: Add tests::dispatch::scope_provider_returns_none_on_registry_mutex_poison test. -
Medium No direct unit test for panicking event-triggered hook (
crates/ironclaw_hooks/src/dispatch.rs:1204-1237, confidence 75) — anchor: crates/ironclaw_hooks/src/dispatch.rs:1216
run_event_triggered_hook wraps execution in catch_unwind but no test installs a panicking hook.
Fix: Add tests::dispatch::event_triggered_hook_panic_is_caught_and_poisons test.
Conventions
- Medium PR body claims 'scope doc only' but diff contains ~2100 lines of implementation (
crates/ironclaw_reborn/src/loop_driver_host.rs:1-376, confidence 100) — anchor: Concrete regression risk: reviewers trusting the 'docs only' claim will skip implementation review
The PR body states 'Draft — scope doc only, no implementation yet.' but the diff adds ~2100 lines of implementation code across 23 files.
Fix: Update the PR body to accurately describe the implementation scope of Phase 5 event-triggered hooks.
Design
- Medium EventTriggeredObserverSink duplicates ObserverSink with identical signature (
crates/ironclaw_hooks/src/sink.rs:316-339, confidence 75) — anchor: crates/ironclaw_hooks/docs/successors/04-event-triggered-hooks.md:63
EventTriggeredObserverSink defines note(category, summary) with the exact same signature as ObserverSink, creating two traits that must be kept in sync.
Fix: Use ObserverSink directly in the EventTriggeredHook trait signature instead of defining a duplicate trait.
Cites crates/ironclaw_hooks/docs/successors/04-event-triggered-hooks.md as the scope contract. Adds the EventTriggered observer hook point, durable RuntimeEvent dispatch path, and Reborn pull-driven subscription wiring with caller-level coverage for matching, replay, scope filtering, observer-only authority, and backpressure.
henrypark133 HIGH + codex P1 on PR #3640: `OwnCapabilities`-scoped event-triggered subscriptions silently never fired for `HookFailed`/`HookDecisionEmitted`/`HookDispatched` events because those `RuntimeEvent` constructors hardcoded `provider: None`. Since Installed hooks default to `OwnCapabilities`, the very events that Phase 5 was designed to observe (hook-failure / decision alerting) never reached their default-configured subscriber. A prior fix added a hook_id-based fallback in `scope_provider_for_runtime_event` that resolves the owning extension through the registry's hex index when `event.provider` is `None`. That covers the case where the failing hook is still registered at replay time, but the durable fix is to stamp the originating provider into the event at emit time so the primary `event.provider` path resolves without any fallback. Plumbed `owning_extension: Option<ExtensionId>` end-to-end: - `LoopHostMilestoneKind::{HookDispatched, HookDecisionEmitted, HookFailed}` gain the field (with `#[serde(default, skip_serializing_if = "Option::is_none")]` so pre-existing checkpoint payloads and the L3 schema-snapshot tests round-trip unchanged when no owner is set). - `RuntimeEvent::hook_{dispatched, decision_emitted, failed}` constructors accept the owner and stamp it into `provider`. - `milestone_events.rs` threads the field through the projection. - `HookDispatcher::emit_dispatched/emit_decision` pass `binding.owning_extension.clone()` directly. - `HookDispatcher::emit_failure` (no binding handy on the failure path) looks the owner up via the registry's existing `owning_extension_for_hook_hex` index. Tests: - `event_triggered_own_capabilities_matches_hook_failed_with_carried_provider`: primary-path regression — two `HookFailed` events with `provider: Some(ext_a|ext_b)` against an `OwnCapabilities` subscription scoped to ext_a; only the own-provider event fires and `event.provider == Some(ext_a)`. - Existing `event_triggered_own_capabilities_scope_resolves_hook_failed_owner_from_hook_id` remains green: passes `None` for the new arg so the fallback path is still exercised for legacy payloads. All other call sites updated to pass `None` (no owner available) or the resolved owner where applicable.
…rfirat HIGH #1 on PR #3640) `EventTriggeredHookSubscription` accepted a caller-supplied `EventStreamKey` + `ReadScope` and used `run_context.scope.tenant_id` as the hook context's tenant — with no validation that the two agreed. A caller wiring tenant A's host with tenant B's stream would cause hooks to observe B's events while the hook context claimed tenant A. Cross-tenant trust-boundary break. Add `EventTriggeredHookSubscription::validate_against_run_scope` and call it from `build_text_only_host_with_capabilities` before spawning. Validation: - Stream `(tenant_id, user_id, agent_id)` must equal `(run_context.scope.tenant_id, thread_scope.owner_user_id, run_context.scope.agent_id)`. - Thread without `owner_user_id` cannot bind any subscription — the user dimension is required to verify stream identity. - Every `Some(want)` in `ReadScope` must equal the corresponding run/thread scope value (project/mission/thread). `None` is permissive (run scope owns the dimension authoritatively). Failures surface as `RebornLoopDriverHostError::ScopeMismatch` with a specific reason naming the offending dimension. Tests: - `event_triggered_subscription_with_foreign_tenant_stream_fails_host_build` - `event_triggered_subscription_with_foreign_user_stream_fails_host_build` The integration fixture's `ThreadScope` now sets `owner_user_id: Some(...)` so it passes validation; previously it was `None`, which the new check (correctly) refuses. Existing tests continue to pass.
…rrfirat MED on PR #3640) When the durable event log returned `EventError::ReplayGap`, the event-triggered subscription's background task previously logged a `tracing::warn!` and broke out of the poll loop — silently killing all future hook event delivery for the run with no operator-visible signal. A scoped audit hook that mattered to compliance would just stop, and nobody downstream would know. Surface the termination through the host's milestone sink: - New `LoopDriverNoteKind::EventSubscriptionTerminated` variant. - The subscription's `spawn`/`run` now takes the host's `Arc<dyn LoopHostMilestoneSink>` and the active `LoopRunContext`. On `ReplayGap`, it constructs a `DriverNote` milestone with that kind plus a `LoopSafeSummary` describing the gap, publishes it through the same sink that carries every other host milestone, and *then* breaks (fail-closed: the at-most-once contract is already broken; resuming from `earliest` would silently lose the gap). - Log level bumped from `warn` to `error` to match the severity. - A best-effort send: failures to publish the milestone are logged but do not stall the subscription teardown. Tests: - `event_triggered_replay_gap_emits_subscription_terminated_milestone`: appends 3 events, `truncate_before_or_at` to cursor 2 to force a replay gap, starts the subscription from cursor origin (now stale), and asserts a `DriverNote { kind: EventSubscriptionTerminated, .. }` shows up on the host's milestone sink within a 2s deadline. Self-emit reentrancy (serrrfirat MED #3 on the same PR) is intentionally not addressed here — that fix needs a design call (task-local re-entry flag vs. removing RuntimeEvent emit capability from event-hook execution contexts) and is a follow-up.
on PR #3640) A hook that subscribes to one of the hook-lifecycle event kinds (`HookDispatched`/`HookDecisionEmitted`/`HookFailed`) with a scope that matches its own provider would otherwise be dispatched for events describing its OWN executions. The dispatcher emits those events itself when running the hook, so a hook subscribing to `HookFailed` with `OwnCapabilities` against its own extension would fail → emit HookFailed → re-dispatch → fail → emit → … storm. `dispatch_event_triggered_at` now skips events whose `event.hook_id` equals the binding's own hook id when the event kind is a hook- lifecycle kind (`is_hook_lifecycle_kind`). The check is intentionally narrow: - It only fires for hook-lifecycle events. Subscriptions to other event kinds are unaffected. - It only suppresses literal self-observation; events about other hooks (even hooks from the same extension) still dispatch. This does NOT cover the broader case of a hook that captures an `Arc<DurableEventLog>` and mints arbitrary `RuntimeEvent`s from inside its `observe()`. That requires architectural restriction on what hook impls can capture — tracked separately as a follow-up. Tests: - `event_triggered_self_lifecycle_event_does_not_redispatch`: appends two `HookFailed` events with the same provider — one targeting the subscriber's own hook id, one targeting a different hook. Asserts only the OTHER hook's failure fires (proves the filter is narrow, not blanket).
Four items from the 5-15 review (#4 DoS budget and #5 narrowed projection deferred — see below): **#1 (should-fix) Invariant: EventTriggered ↔ event_kind_filter** `HookRegistry::insert` now enforces the biconditional at install time: an `EventTriggered` binding must declare an `event_kind_filter` (otherwise the dispatcher's kind match would silently never fire — a no-op binding), and conversely only `EventTriggered` bindings may declare a filter (other points are kind-agnostic and would ignore the field). Misconfigured bindings fail loud at install. **#2 (should-fix) Remove `Clone` derive on EventTriggeredHookSubscription** `Clone` on a spawn-semantics type was a footgun: external callers cloning + spawning twice would create two consumers reading from the same `start_cursor`, each dispatching every hook. Replace with an explicit `clone_for_independent_spawn(&self)` method named verbosely so the property is visible at the seam. Internal use updated in the factory's host-build path; external callers can no longer accidentally construct a dual-consumer pattern. **#3 (should-fix) catch_unwind around the background `run()` task** The subscription's tokio task body now runs inside `AssertUnwindSafe(...).catch_unwind()`; a panic in `run()` emits the same `EventSubscriptionTerminated` `DriverNote` milestone the `ReplayGap` path already emits, instead of silently terminating with no operator-visible signal. **#6 (should-fix) Replay semantics in rustdoc on public API** Added a "Replay semantics" section to `EventTriggeredHookSubscription` rustdoc: at-least-once, caller-owned cursor persistence, the restart-from-start_cursor replay pattern. Previously only in the design doc; now load-bearing API contract is visible at the type. **#4 (deferred) Per-hook DoS budget for Installed tier** Henry's recommendation was to gate `Installed`-tier event-triggered hooks entirely until the budget design lands, allowing only Builtin/Trusted. That breaks 11 existing tests + the primary use case. Instead: documented the existing first-line throttle (`batch_limit` × `poll_interval`) as the current bound on indirect- recursion fanout, and tracked the full per-hook rate cap with poisoning + milestone-on-overrun as a follow-up. The self-trigger guard (committed earlier in this PR) catches the most common direct pattern; the throttle here bounds the indirect pattern until the proper budget lands. **#5 (deferred) Narrowed `HookObservableEvent` projection** Would prevent full `RuntimeEvent` surface from reaching Installed- tier hooks. Project-wide impact (events crate types, projection glue). Tracked as a follow-up; the existing sanitized-event projection bounds the surface to closed-vocab labels. All 156 hooks lib + 30 reborn integration tests pass.
Bundle three nit-tier review items into a single commit: **#9 Replace author-internal tags with NOTE(#3640)** The Phase-5 PR (#3640) had several `serrrfirat HIGH/MED #N on PR #3640` comment tags in this PR's diff. These are review-internal scaffolding, not load-bearing for future readers. Replaced with `NOTE(#3640)` in: - crates/ironclaw_hooks/src/dispatch.rs (self-observation guard) - crates/ironclaw_reborn/src/loop_driver_host.rs (scope validation, replay-gap milestone, subscription binding) - crates/ironclaw_reborn/tests/hooks_integration.rs (three regression tests covering scope validation, self-observation suppression, and replay-gap surfacing) - crates/ironclaw_turns/src/run_profile/host.rs (`EventSubscriptionTerminated` doc) **#10 Replace 10ms spin-poll with tokio::sync::Notify** `wait_for_seen_events` polled the shared `Mutex<Vec<SeenRuntimeEvent>>` every 10 ms until the expected count was reached. Replaced with a `SeenLog` newtype that pairs the events vec with a `Notify`; the hook's `observe()` calls `seen.push(...)` which signals `Notify::notify_one`, and `wait_for_seen_events` parks on `notified().await` under a `tokio::time::timeout`. `notify_one` is a permit-store, so an event landing between snapshot and wait still wakes the waiter immediately. Test latency drops from ~10 ms median to sub-ms and is no longer rate-limited by the polling cadence. All 30 hooks_integration tests still pass. **#11 Remove unused Clone derive on EventTriggeredHookContext** No call site clones the context — it's passed by reference. Dropped the derive to make the borrow contract clearer.
…reality Address gemini-code-assist review on `04-event-triggered-hooks.md`: - L50 (Likely surface): annotated the sketch's full `RuntimeEvent` use with a pointer to the narrowed-projection follow-up so the snippet no longer reads as a recommendation contradicting L119–121. - L55 (sink methods): replaced `note_fact` / `emit_audit` (which never shipped on `ObserverSink`) with the actual `note(category, summary)` primitive and cross-referenced Reborn's `EventTriggeredObserverSink`. - L95 (cursor / replay): "lost events during downtime acceptable" contradicted the at-least-once replay semantics described in the Phase 5 implementation notes. Rewrote the bullet to say replay is at-least-once from the persisted cursor and to spell out the operator obligation around cursor persistence before shutdown. - L100/115 (forbids events dep): the original doc claimed `ironclaw_hooks` forbids an `ironclaw_events` dep, but the Risk section noted the dep is already established via PR #3573. Updated both passages to reflect that the dep direction is set; Phase 5 adds the *consumer* side. The narrowed `HookObservableEvent` projection is now framed as a follow-up tracked in #3690.
Address PR #3640 review findings A3, C4, F14, and cluster G: - F14: drop duplicate `EventTriggeredObserverSink` trait and reuse `ObserverSink` directly in the `EventTriggeredHook` trait. The two surfaces were signature-identical; keeping them separate let them drift, and a future gate/mutator method added to one would not surface as a compile error on the other. - A3: add `is_replay: bool` to `EventTriggeredHookContext` and a dedicated `dispatch_event_triggered_replay_at` entry point. The subscription contract is at-least-once, so side-effecting hooks need to dedupe by `event.event_id` on restart-driven replay. - C4: index event-triggered bindings by `RuntimeEventKind` at install time so dispatch is O(matches) instead of scanning every event-triggered binding for every event. - Cluster G: doc/04-event-triggered-hooks.md updated to reflect the unified sink, the explicit at-least-once semantics + `is_replay` signal, the actual `note(category, summary)` primitive (not the speculative `note_fact` / `emit_audit`), the corrected `ironclaw_events` dep status, and the issue #3690 reference for the narrowed `HookObservableEvent` projection. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Address PR #3640 review findings C5, A1, A2: - C5: empty-poll backoff for `EventTriggeredHookSubscription`. The previous loop hammered the durable log at a fixed 50ms cadence under sustained idle, even when no events had arrived for minutes. The subscription now tracks consecutive empty polls and sleeps for `min(poll_interval << streak, max_poll_interval)` before the next poll, defaulting to a 1s cap; a non-empty batch resets the streak so producer bursts restore low-latency dispatch immediately. Exposed via `with_max_poll_interval` so callers can tune. - A1 / A2: explicit issue references for the deferred narrowed `HookObservableEvent` projection (#3690) and the per-hook DoS dispatch budget (#3689). The current self-trigger guard catches direct-recursion storms; the backoff bounds indirect ones until the proper budget design lands. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Address PR #3640 review findings D8, D9, D10, D11, D12, and B: - D8: dispatching an event-triggered binding that has no installed hook impl must poison the slot and surface a Malformed failure rather than silently no-op. A follow-up dispatch on the same kind must skip the poisoned slot. - D9: registry validation rejects non-event-point bindings that carry an `event_kind_filter`, mirroring the existing reverse-direction check. - D10: the existing hook-meta serde round-trip tests always passed `None` for `owning_extension` and never asserted `event.provider`. Add `hook_meta_events_round_trip_owning_extension_as_provider` to pin the projection that scope filtering depends on. - D11: `scope_provider_for_runtime_event` falls back to `None` when the registry mutex is poisoned. Force a poison on a spawned thread and assert the resolver remains fail-closed. - D12: `run_event_triggered_hook` catches panics from the hook impl via `AssertUnwindSafe::catch_unwind`. Drive it with a deliberately panicking impl and assert `FailureCategory::Panic`. - Cluster B: when a hook-meta event has `provider: None`, the dispatcher recovers the owning extension from the registry's hex-keyed index so `OwnCapabilities` watchers still fire. Add a full end-to-end test exercising that path through `dispatch_event_triggered_at`. Also pin C4 indexing: a registry-level test that `active_for_event_kind` returns only bindings whose declared filter matches. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…#3912/#3913) - Add event_kind_filter: None to HookBinding test constructions (foundation added new field) - Replace tuple-struct construction of ExtensionId/HookLocalId with ::new() per #3912 newtype privatization - Lowercase RuntimeEventKind debug repr for HookLocalId validation (lowercase-only identifiers) - Replace pub-use re-exports with module-path imports per foundation cleanup - Rename fixture.user_id to fixture.actor_id per #3633 final naming
- Add event_kind_filter: None to WASM HookBinding constructions in dispatch.rs - Extend HookManifestKind match arms in registrar.rs and wasm/runtime.rs to handle EventTriggered (rejected: WASM-bodied event-triggered hooks are not yet supported)
0a2574e to
2728471
Compare
Post-merge Codex reviewVerdict: REQUEST CHANGES — event-triggered hooks are correctly modeled as a parallel observer path, but subscription scoping and replay contracts have real holes. Follow-up PR needed before tenant-safe audit automation can rely on this. Critical1. Cross-thread/project event leakage via permissive Fix: build an effective read scope from the run/thread scope, or require 2. Replay signal is effectively unused. Fix: track cursor progress and mark replayed records, or remove the flag until it is actually wired. 3. Hook-lifecycle provider spoofing/conflict. Fix: for hook lifecycle kinds, resolve the owner from Recommendations
|
…provider claim Bug 3 (CRITICAL) from Codex review of nearai#3640: scope_provider_for_runtime_event trusted event.provider before consulting the registry. A hook could synthesize a HookFailed event for hook B (owned by ext-B) while claiming provider = ext-A, causing ext-A's OwnCapabilities hooks to fire on an event about another extension's hook. For hook-lifecycle event kinds (HookDispatched/HookDecisionEmitted/HookFailed) the authoritative owner is now resolved from event.hook_id via HookRegistry::owning_extension_for_hook_hex. If the resolved owner disagrees with the payload claim, the discrepancy is logged and the resolved owner is used (never the claim). Lifecycle events with no hook_id, or whose hook_id does not resolve, fall back to None (inert) — fail-closed. Non-lifecycle events, which carry no hook_id anchor, continue to use the host-established provider. TDD: hook_failed_with_spoofed_provider_does_not_fire_target_extension_hooks. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Bug 2 (CRITICAL) from Codex review of nearai#3640: the event-triggered hook subscription always called dispatch_event_triggered_at (live), never dispatch_event_triggered_replay_at, so EventTriggeredHookContext::is_replay was always false. On reconnect/restart, the backlog caught up from the resume cursor replayed duplicate side effects with no dedupe signal, despite the public is_replay contract. The subscription now tracks the replay/live boundary without new durable state: events caught up from start_cursor (the caller's committed resume point) to the stream head at subscription time are dispatched through the replay path (is_replay = true). The first empty poll means the backlog is drained and head is reached; every event after that boundary is live (is_replay = false). Over-marking a live event as replay is the fail-safe direction; the previous bug marked replayed events as live. TDD: event_subscription_marks_is_replay_true_for_replayed_events (integration). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
…ive any() Bug 1 (CRITICAL) from Codex review of nearai#3640: validate_against_run_scope accepted a permissive read_scope (including ReadScope::any()) and passed it unchanged to read_after_cursor. Durable event streams are keyed only by (tenant, user, agent), so within one stream a run could observe events from other projects/threads/missions — a cross-thread/project trust-boundary break. Replaced with effective_read_scope: it rejects ReadScope::any() with a typed RebornLoopDriverHostError::ScopeMismatch (fail-closed), keeps the existing stream-key (tenant/user/agent) checks, requires any caller-supplied Some(want) to equal the run/thread value (tighten, never widen), and then *derives* the filter actually used from the authoritative run/thread scope: thread_id pinned to the run thread, project_id/mission_id pinned when present. The host build applies the derived scope via with_read_scope before spawning the subscription. TDD: event_subscription_rejects_read_scope_any, event_subscription_does_not_observe_other_project_events, effective_read_scope_pins_thread_and_project_from_run_scope. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
…provider claim Bug 3 (CRITICAL) from Codex review of nearai#3640: scope_provider_for_runtime_event trusted event.provider before consulting the registry. A hook could synthesize a HookFailed event for hook B (owned by ext-B) while claiming provider = ext-A, causing ext-A's OwnCapabilities hooks to fire on an event about another extension's hook. For hook-lifecycle event kinds (HookDispatched/HookDecisionEmitted/HookFailed) the authoritative owner is now resolved from event.hook_id via HookRegistry::owning_extension_for_hook_hex. If the resolved owner disagrees with the payload claim, the discrepancy is logged and the resolved owner is used (never the claim). Lifecycle events with no hook_id, or whose hook_id does not resolve, fall back to None (inert) — fail-closed. Non-lifecycle events, which carry no hook_id anchor, continue to use the host-established provider. TDD: hook_failed_with_spoofed_provider_does_not_fire_target_extension_hooks. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
Bug 2 (CRITICAL) from Codex review of nearai#3640: the event-triggered hook subscription always called dispatch_event_triggered_at (live), never dispatch_event_triggered_replay_at, so EventTriggeredHookContext::is_replay was always false. On reconnect/restart, the backlog caught up from the resume cursor replayed duplicate side effects with no dedupe signal, despite the public is_replay contract. The subscription now tracks the replay/live boundary without new durable state: events caught up from start_cursor (the caller's committed resume point) to the stream head at subscription time are dispatched through the replay path (is_replay = true). The first empty poll means the backlog is drained and head is reached; every event after that boundary is live (is_replay = false). Over-marking a live event as replay is the fail-safe direction; the previous bug marked replayed events as live. TDD: event_subscription_marks_is_replay_true_for_replayed_events (integration). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
…ive any() Bug 1 (CRITICAL) from Codex review of nearai#3640: validate_against_run_scope accepted a permissive read_scope (including ReadScope::any()) and passed it unchanged to read_after_cursor. Durable event streams are keyed only by (tenant, user, agent), so within one stream a run could observe events from other projects/threads/missions — a cross-thread/project trust-boundary break. Replaced with effective_read_scope: it rejects ReadScope::any() with a typed RebornLoopDriverHostError::ScopeMismatch (fail-closed), keeps the existing stream-key (tenant/user/agent) checks, requires any caller-supplied Some(want) to equal the run/thread value (tighten, never widen), and then *derives* the filter actually used from the authoritative run/thread scope: thread_id pinned to the run thread, project_id/mission_id pinned when present. The host build applies the derived scope via with_read_scope before spawning the subscription. TDD: event_subscription_rejects_read_scope_any, event_subscription_does_not_observe_other_project_events, effective_read_scope_pins_thread_and_project_from_run_scope. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com>
…n event-triggered (#3640 followup) (#3931) * fix(hooks): resolve hook-lifecycle owner from hook_id, not spoofable provider claim Bug 3 (CRITICAL) from Codex review of #3640: scope_provider_for_runtime_event trusted event.provider before consulting the registry. A hook could synthesize a HookFailed event for hook B (owned by ext-B) while claiming provider = ext-A, causing ext-A's OwnCapabilities hooks to fire on an event about another extension's hook. For hook-lifecycle event kinds (HookDispatched/HookDecisionEmitted/HookFailed) the authoritative owner is now resolved from event.hook_id via HookRegistry::owning_extension_for_hook_hex. If the resolved owner disagrees with the payload claim, the discrepancy is logged and the resolved owner is used (never the claim). Lifecycle events with no hook_id, or whose hook_id does not resolve, fall back to None (inert) — fail-closed. Non-lifecycle events, which carry no hook_id anchor, continue to use the host-established provider. TDD: hook_failed_with_spoofed_provider_does_not_fire_target_extension_hooks. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * fix(hooks): wire is_replay through event-triggered replay dispatch path Bug 2 (CRITICAL) from Codex review of #3640: the event-triggered hook subscription always called dispatch_event_triggered_at (live), never dispatch_event_triggered_replay_at, so EventTriggeredHookContext::is_replay was always false. On reconnect/restart, the backlog caught up from the resume cursor replayed duplicate side effects with no dedupe signal, despite the public is_replay contract. The subscription now tracks the replay/live boundary without new durable state: events caught up from start_cursor (the caller's committed resume point) to the stream head at subscription time are dispatched through the replay path (is_replay = true). The first empty poll means the backlog is drained and head is reached; every event after that boundary is live (is_replay = false). Over-marking a live event as replay is the fail-safe direction; the previous bug marked replayed events as live. TDD: event_subscription_marks_is_replay_true_for_replayed_events (integration). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * fix(hooks): derive effective ReadScope from run scope, reject permissive any() Bug 1 (CRITICAL) from Codex review of #3640: validate_against_run_scope accepted a permissive read_scope (including ReadScope::any()) and passed it unchanged to read_after_cursor. Durable event streams are keyed only by (tenant, user, agent), so within one stream a run could observe events from other projects/threads/missions — a cross-thread/project trust-boundary break. Replaced with effective_read_scope: it rejects ReadScope::any() with a typed RebornLoopDriverHostError::ScopeMismatch (fail-closed), keeps the existing stream-key (tenant/user/agent) checks, requires any caller-supplied Some(want) to equal the run/thread value (tighten, never widen), and then *derives* the filter actually used from the authoritative run/thread scope: thread_id pinned to the run thread, project_id/mission_id pinned when present. The host build applies the derived scope via with_read_scope before spawning the subscription. TDD: event_subscription_rejects_read_scope_any, event_subscription_does_not_observe_other_project_events, effective_read_scope_pins_thread_and_project_from_run_scope. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * fix(reborn): snapshot stream head at subscription start for replay/live boundary PR #3931 Hole 1: the event-triggered hook subscription treated "the first poll that returns no entries" as the replay→live boundary. That races: a live event appended after subscription start but before the first empty poll (or while a continuous backlog drains past the true head) was dispatched via the replay path and marked is_replay=true, violating the "gap from start_cursor to head-at-startup" contract and risking dedupe/skip of legitimate live side effects. Fix: add DurableEventLog::head_cursor (default impl drains unfiltered; in-memory backend overrides O(1)) and snapshot startup_head atomically in EventTriggeredHookSubscription::spawn — synchronously, before tokio::spawn — so "head-at-startup" means the head at host-build time, not at the spawned task's first lazy poll. Records classify per-cursor: cursor <= startup_head is replay, everything beyond is live; mixed batches split at the boundary. Fail-closed: if the head snapshot fails, the subscription does not dispatch and emits an operator-visible terminated milestone. Tests: live event before first empty poll is LIVE; event during continuous backlog drain past startup_head is LIVE; head_cursor contract (latest cursor, future-cursor ReplayGap). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * fix(hooks): fail closed on unknown/poison lifecycle owner, never trust carried provider PR #3931 Hole 2: scope_provider_for_runtime_event resolved the owner of a hook-lifecycle event (HookFailed/HookDispatched/HookDecisionEmitted) from the registry by hook_id, but the `(None, claimed) => claimed` arm fell back to the event's carried `provider` payload whenever the registry lookup returned None. That payload field is forgeable — it is a plain serialized field, not an unforgeable host stamp — so two attacks slipped through: 1. Unknown hook_id + provider=ext-a: a synthesized lifecycle event naming a hook_id absent from the registry could fire ext-a's OwnCapabilities hooks. 2. Poisoned registry + provider=target: registry poison collapsed to None and then trusted the carried provider, defeating fail-closed. Fix: for lifecycle events, the owner is the registry-resolved owner of hook_id and ONLY that. Poisoned registry and unknown hook_id are kept as distinct branches that both resolve to None (hook inert); neither falls back to the carried provider. A registry hit still wins, logging a spoof warning when the payload disagrees. Design tension resolved (Phase 5 alerting): two integration tests previously asserted that an OwnCapabilities watcher fires for a lifecycle event whose *subject* hook_id was UNREGISTERED, relying on the carried provider. That is the exact spoof shape the fix rejects. In production the subject hook ran in this same host and so resolves via the registry — the carried-provider fallback was never needed for a legitimate flow. The tests are reworked to register the subject hooks (so ownership resolves from the registry, the production reality): the ext-A-owned event fires, the ext-B-owned event stays inert, and the self-suppression control case still proves the filter is narrow. There is no unforgeable host-stamped owner source distinct from the payload today; if a future extension-emitted or cross-host lifecycle path needs one, it must add an authenticated owner field rather than trusting the payload. Tests: unknown lifecycle hook_id + provider=target stays inert (resolves None); poisoned registry + provider=target resolves None; reworked Phase 5 and self-redispatch integration tests to resolve ownership from the registry. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * refactor(hooks): address serrrfirat review — atomic head_cursor, drop any() ceremony, extract lifecycle owner resolver P1: make DurableEventLog::head_cursor a required backend operation (no default). The previous default drained the stream page-by-page until an empty read, which is non-atomic (a record appended during the probe folds into the observed head and gets mis-marked replay) and moved an unbounded scan into host construction. Each backend now reads its own tail atomically: - JsonlDurableEventLog: read_last_jsonl_cursor (O(1) tail seek) - FilesystemDurableEventLog: single tail(after-1) snapshot, max seq - InMemoryDurableEventLog: already O(1) via next_cursor All production + test impls updated. Added filesystem head_cursor contract test (latest cursor, mid-stream probe, future-cursor ReplayGap). P2: drop the ReadScope::any() rejection in EventTriggeredHookSubscription::effective_read_scope. Verified the read uses the DERIVED authoritative scope (installed via with_read_scope before the task spawns), not the caller's filter — the derivation always pins thread/project/mission from the run/thread scope, so any() cannot widen the read or reopen the cross-thread/project leak. The rejection was redundant ceremony. Reverted the forced hooks_integration.rs fixture churn (event_log_subscription back to ReadScope::any(), no &fixture param). Test renamed to assert any() is accepted and constrained by the derived scope. P2: extract scope_provider_for_runtime_event into dispatch/lifecycle_owner.rs (resolve_event_owner + LifecycleOwnerLookup), converting dispatch.rs to dispatch/mod.rs. Pure, table-tested resolver with 6 cases: registered owner, spoofed provider override, unknown hook id, poisoned registry, missing hook id, non-lifecycle passthrough. Behavior identical to the security-fixed inline version. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * test(events): prove head_cursor atomicity — post-snapshot append classified live Adds head_cursor_snapshot_classifies_later_appends_as_live to the durable log contract suite (PR #3931 P1): a record appended after head_cursor returns receives a cursor strictly greater than the observed startup_head, so it is classified live and never folded into the replay window. Locks in the atomic-head contract that the required (no-default) head_cursor method and its per-backend O(1)/single-snapshot overrides must satisfy. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * fix(hooks): address henrypark133 review on #3931 (head_cursor perf + tests + must_use) - Add `RootFilesystem::head_seq(path, from)` returning `Option<SeqNo>` with a correctness-preserving `tail`-backed default; override in Postgres/libSQL with O(1) `MAX(seq)` so a fresh subscription (`after = 0`) no longer materializes the whole stream gap into memory just to find its head. Route `FilesystemDurableEventLog::head_cursor` through `head_seq`. (Perf finding) - Add JSONL backend head_cursor coverage: `jsonl_head_cursor_reports_latest_ and_rejects_future` exercises empty-stream head, post-append head, mid-stream probe, and future-cursor ReplayGap rejection. (Tests finding) - Add `event_triggered_head_cursor_failure_emits_terminated_milestone`: a stub log whose head_cursor fails at subscription spawn must emit exactly one EventSubscriptionTerminated milestone (fail-closed). (Tests finding) - Add `#[must_use]` to `EventTriggeredHookSubscription::with_read_scope` to match its four sibling builder methods. (Local-patterns finding) The in_memory doc-comment nit is declined: the head_cursor doc records the atomicity contract at the impl site, which is intentional. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * fix(hooks,reborn): close exhaustiveness + scope-derivation review gaps (#3931) M1: replace the non-exhaustive `matches!` in `is_lifecycle_kind` with an exhaustive `match` enumerating every `RuntimeEventKind` variant (no `_` wildcard). A future lifecycle variant now fails to compile until it is explicitly classified, instead of silently falling through to the non-lifecycle branch that trusts the forgeable `event.provider` (Hole 2). Adds a test asserting every variant classifies on the correct side. M2: narrow `EventTriggeredHookSubscription::with_read_scope` from `pub(crate)` to module-private. It has no cross-module caller (only the in-module host build path at line ~1351 invokes it), so the narrower visibility removes the latent surface for other `ironclaw_reborn` code to install an arbitrary `ReadScope::any()` and bypass `effective_read_scope`. M3: emit a `tracing::warn!` in `effective_read_scope` when `run_scope.project_id` is `None`, documenting that the subscription is not project-pinned (thread_id still pins it). Not fail-closed: tenant/agent-level project-less runs are legitimate. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
…nearai#3640) * docs(hooks): scope event-triggered hooks (Phase 5, successor #4) Successor PR from nearai#3573. Adds a new EventTriggered hook point that subscribes to RuntimeEvents asynchronously, outside the loop's inline tick. Observer-only by construction (no Allow/Deny/Patch); typed against a narrowed HookObservableEvent projection to keep the cross-crate boundary clean. Scope doc only; design questions about cursor/replay semantics and per-extension event-rate caps need design review before implementation. * Implement Phase 5 event-triggered hooks Cites crates/ironclaw_hooks/docs/successors/04-event-triggered-hooks.md as the scope contract. Adds the EventTriggered observer hook point, durable RuntimeEvent dispatch path, and Reborn pull-driven subscription wiring with caller-level coverage for matching, replay, scope filtering, observer-only authority, and backpressure. * Fix hook event OwnCapabilities owner lookup * fix(hooks): carry owning extension into hook milestone runtime events henrypark133 HIGH + codex P1 on PR nearai#3640: `OwnCapabilities`-scoped event-triggered subscriptions silently never fired for `HookFailed`/`HookDecisionEmitted`/`HookDispatched` events because those `RuntimeEvent` constructors hardcoded `provider: None`. Since Installed hooks default to `OwnCapabilities`, the very events that Phase 5 was designed to observe (hook-failure / decision alerting) never reached their default-configured subscriber. A prior fix added a hook_id-based fallback in `scope_provider_for_runtime_event` that resolves the owning extension through the registry's hex index when `event.provider` is `None`. That covers the case where the failing hook is still registered at replay time, but the durable fix is to stamp the originating provider into the event at emit time so the primary `event.provider` path resolves without any fallback. Plumbed `owning_extension: Option<ExtensionId>` end-to-end: - `LoopHostMilestoneKind::{HookDispatched, HookDecisionEmitted, HookFailed}` gain the field (with `#[serde(default, skip_serializing_if = "Option::is_none")]` so pre-existing checkpoint payloads and the L3 schema-snapshot tests round-trip unchanged when no owner is set). - `RuntimeEvent::hook_{dispatched, decision_emitted, failed}` constructors accept the owner and stamp it into `provider`. - `milestone_events.rs` threads the field through the projection. - `HookDispatcher::emit_dispatched/emit_decision` pass `binding.owning_extension.clone()` directly. - `HookDispatcher::emit_failure` (no binding handy on the failure path) looks the owner up via the registry's existing `owning_extension_for_hook_hex` index. Tests: - `event_triggered_own_capabilities_matches_hook_failed_with_carried_provider`: primary-path regression — two `HookFailed` events with `provider: Some(ext_a|ext_b)` against an `OwnCapabilities` subscription scoped to ext_a; only the own-provider event fires and `event.provider == Some(ext_a)`. - Existing `event_triggered_own_capabilities_scope_resolves_hook_failed_owner_from_hook_id` remains green: passes `None` for the new arg so the fallback path is still exercised for legacy payloads. All other call sites updated to pass `None` (no owner available) or the resolved owner where applicable. * fix(hooks): validate event subscription scope against run scope (serrrfirat HIGH #1 on PR nearai#3640) `EventTriggeredHookSubscription` accepted a caller-supplied `EventStreamKey` + `ReadScope` and used `run_context.scope.tenant_id` as the hook context's tenant — with no validation that the two agreed. A caller wiring tenant A's host with tenant B's stream would cause hooks to observe B's events while the hook context claimed tenant A. Cross-tenant trust-boundary break. Add `EventTriggeredHookSubscription::validate_against_run_scope` and call it from `build_text_only_host_with_capabilities` before spawning. Validation: - Stream `(tenant_id, user_id, agent_id)` must equal `(run_context.scope.tenant_id, thread_scope.owner_user_id, run_context.scope.agent_id)`. - Thread without `owner_user_id` cannot bind any subscription — the user dimension is required to verify stream identity. - Every `Some(want)` in `ReadScope` must equal the corresponding run/thread scope value (project/mission/thread). `None` is permissive (run scope owns the dimension authoritatively). Failures surface as `RebornLoopDriverHostError::ScopeMismatch` with a specific reason naming the offending dimension. Tests: - `event_triggered_subscription_with_foreign_tenant_stream_fails_host_build` - `event_triggered_subscription_with_foreign_user_stream_fails_host_build` The integration fixture's `ThreadScope` now sets `owner_user_id: Some(...)` so it passes validation; previously it was `None`, which the new check (correctly) refuses. Existing tests continue to pass. * fix(hooks): surface event subscription replay-gap as a milestone (serrrfirat MED on PR nearai#3640) When the durable event log returned `EventError::ReplayGap`, the event-triggered subscription's background task previously logged a `tracing::warn!` and broke out of the poll loop — silently killing all future hook event delivery for the run with no operator-visible signal. A scoped audit hook that mattered to compliance would just stop, and nobody downstream would know. Surface the termination through the host's milestone sink: - New `LoopDriverNoteKind::EventSubscriptionTerminated` variant. - The subscription's `spawn`/`run` now takes the host's `Arc<dyn LoopHostMilestoneSink>` and the active `LoopRunContext`. On `ReplayGap`, it constructs a `DriverNote` milestone with that kind plus a `LoopSafeSummary` describing the gap, publishes it through the same sink that carries every other host milestone, and *then* breaks (fail-closed: the at-most-once contract is already broken; resuming from `earliest` would silently lose the gap). - Log level bumped from `warn` to `error` to match the severity. - A best-effort send: failures to publish the milestone are logged but do not stall the subscription teardown. Tests: - `event_triggered_replay_gap_emits_subscription_terminated_milestone`: appends 3 events, `truncate_before_or_at` to cursor 2 to force a replay gap, starts the subscription from cursor origin (now stale), and asserts a `DriverNote { kind: EventSubscriptionTerminated, .. }` shows up on the host's milestone sink within a 2s deadline. Self-emit reentrancy (serrrfirat MED #3 on the same PR) is intentionally not addressed here — that fix needs a design call (task-local re-entry flag vs. removing RuntimeEvent emit capability from event-hook execution contexts) and is a follow-up. * fix(hooks): suppress event-triggered self-observation (serrrfirat MED #3 on PR nearai#3640) A hook that subscribes to one of the hook-lifecycle event kinds (`HookDispatched`/`HookDecisionEmitted`/`HookFailed`) with a scope that matches its own provider would otherwise be dispatched for events describing its OWN executions. The dispatcher emits those events itself when running the hook, so a hook subscribing to `HookFailed` with `OwnCapabilities` against its own extension would fail → emit HookFailed → re-dispatch → fail → emit → … storm. `dispatch_event_triggered_at` now skips events whose `event.hook_id` equals the binding's own hook id when the event kind is a hook- lifecycle kind (`is_hook_lifecycle_kind`). The check is intentionally narrow: - It only fires for hook-lifecycle events. Subscriptions to other event kinds are unaffected. - It only suppresses literal self-observation; events about other hooks (even hooks from the same extension) still dispatch. This does NOT cover the broader case of a hook that captures an `Arc<DurableEventLog>` and mints arbitrary `RuntimeEvent`s from inside its `observe()`. That requires architectural restriction on what hook impls can capture — tracked separately as a follow-up. Tests: - `event_triggered_self_lifecycle_event_does_not_redispatch`: appends two `HookFailed` events with the same provider — one targeting the subscriber's own hook id, one targeting a different hook. Asserts only the OTHER hook's failure fires (proves the filter is narrow, not blanket). * fix(hooks): address henrypark133 should-fix #1, #2, #3, #6 on PR nearai#3640 Four items from the 5-15 review (#4 DoS budget and #5 narrowed projection deferred — see below): **#1 (should-fix) Invariant: EventTriggered ↔ event_kind_filter** `HookRegistry::insert` now enforces the biconditional at install time: an `EventTriggered` binding must declare an `event_kind_filter` (otherwise the dispatcher's kind match would silently never fire — a no-op binding), and conversely only `EventTriggered` bindings may declare a filter (other points are kind-agnostic and would ignore the field). Misconfigured bindings fail loud at install. **#2 (should-fix) Remove `Clone` derive on EventTriggeredHookSubscription** `Clone` on a spawn-semantics type was a footgun: external callers cloning + spawning twice would create two consumers reading from the same `start_cursor`, each dispatching every hook. Replace with an explicit `clone_for_independent_spawn(&self)` method named verbosely so the property is visible at the seam. Internal use updated in the factory's host-build path; external callers can no longer accidentally construct a dual-consumer pattern. **#3 (should-fix) catch_unwind around the background `run()` task** The subscription's tokio task body now runs inside `AssertUnwindSafe(...).catch_unwind()`; a panic in `run()` emits the same `EventSubscriptionTerminated` `DriverNote` milestone the `ReplayGap` path already emits, instead of silently terminating with no operator-visible signal. **#6 (should-fix) Replay semantics in rustdoc on public API** Added a "Replay semantics" section to `EventTriggeredHookSubscription` rustdoc: at-least-once, caller-owned cursor persistence, the restart-from-start_cursor replay pattern. Previously only in the design doc; now load-bearing API contract is visible at the type. **#4 (deferred) Per-hook DoS budget for Installed tier** Henry's recommendation was to gate `Installed`-tier event-triggered hooks entirely until the budget design lands, allowing only Builtin/Trusted. That breaks 11 existing tests + the primary use case. Instead: documented the existing first-line throttle (`batch_limit` × `poll_interval`) as the current bound on indirect- recursion fanout, and tracked the full per-hook rate cap with poisoning + milestone-on-overrun as a follow-up. The self-trigger guard (committed earlier in this PR) catches the most common direct pattern; the throttle here bounds the indirect pattern until the proper budget lands. **#5 (deferred) Narrowed `HookObservableEvent` projection** Would prevent full `RuntimeEvent` surface from reaching Installed- tier hooks. Project-wide impact (events crate types, projection glue). Tracked as a follow-up; the existing sanitized-event projection bounds the surface to closed-vocab labels. All 156 hooks lib + 30 reborn integration tests pass. * chore(hooks): address nits from PR nearai#3640 review Bundle three nit-tier review items into a single commit: **#9 Replace author-internal tags with NOTE(nearai#3640)** The Phase-5 PR (nearai#3640) had several `serrrfirat HIGH/MED #N on PR nearai#3640` comment tags in this PR's diff. These are review-internal scaffolding, not load-bearing for future readers. Replaced with `NOTE(nearai#3640)` in: - crates/ironclaw_hooks/src/dispatch.rs (self-observation guard) - crates/ironclaw_reborn/src/loop_driver_host.rs (scope validation, replay-gap milestone, subscription binding) - crates/ironclaw_reborn/tests/hooks_integration.rs (three regression tests covering scope validation, self-observation suppression, and replay-gap surfacing) - crates/ironclaw_turns/src/run_profile/host.rs (`EventSubscriptionTerminated` doc) **#10 Replace 10ms spin-poll with tokio::sync::Notify** `wait_for_seen_events` polled the shared `Mutex<Vec<SeenRuntimeEvent>>` every 10 ms until the expected count was reached. Replaced with a `SeenLog` newtype that pairs the events vec with a `Notify`; the hook's `observe()` calls `seen.push(...)` which signals `Notify::notify_one`, and `wait_for_seen_events` parks on `notified().await` under a `tokio::time::timeout`. `notify_one` is a permit-store, so an event landing between snapshot and wait still wakes the waiter immediately. Test latency drops from ~10 ms median to sub-ms and is no longer rate-limited by the polling cadence. All 30 hooks_integration tests still pass. **#11 Remove unused Clone derive on EventTriggeredHookContext** No call site clones the context — it's passed by reference. Dropped the derive to make the borrow contract clearer. * docs(hooks): reconcile event-triggered hooks design doc with Phase 5 reality Address gemini-code-assist review on `04-event-triggered-hooks.md`: - L50 (Likely surface): annotated the sketch's full `RuntimeEvent` use with a pointer to the narrowed-projection follow-up so the snippet no longer reads as a recommendation contradicting L119–121. - L55 (sink methods): replaced `note_fact` / `emit_audit` (which never shipped on `ObserverSink`) with the actual `note(category, summary)` primitive and cross-referenced Reborn's `EventTriggeredObserverSink`. - L95 (cursor / replay): "lost events during downtime acceptable" contradicted the at-least-once replay semantics described in the Phase 5 implementation notes. Rewrote the bullet to say replay is at-least-once from the persisted cursor and to spell out the operator obligation around cursor persistence before shutdown. - L100/115 (forbids events dep): the original doc claimed `ironclaw_hooks` forbids an `ironclaw_events` dep, but the Risk section noted the dep is already established via PR nearai#3573. Updated both passages to reflect that the dep direction is set; Phase 5 adds the *consumer* side. The narrowed `HookObservableEvent` projection is now framed as a follow-up tracked in nearai#3690. * refactor(hooks): unify event-triggered sink with ObserverSink Address PR nearai#3640 review findings A3, C4, F14, and cluster G: - F14: drop duplicate `EventTriggeredObserverSink` trait and reuse `ObserverSink` directly in the `EventTriggeredHook` trait. The two surfaces were signature-identical; keeping them separate let them drift, and a future gate/mutator method added to one would not surface as a compile error on the other. - A3: add `is_replay: bool` to `EventTriggeredHookContext` and a dedicated `dispatch_event_triggered_replay_at` entry point. The subscription contract is at-least-once, so side-effecting hooks need to dedupe by `event.event_id` on restart-driven replay. - C4: index event-triggered bindings by `RuntimeEventKind` at install time so dispatch is O(matches) instead of scanning every event-triggered binding for every event. - Cluster G: doc/04-event-triggered-hooks.md updated to reflect the unified sink, the explicit at-least-once semantics + `is_replay` signal, the actual `note(category, summary)` primitive (not the speculative `note_fact` / `emit_audit`), the corrected `ironclaw_events` dep status, and the issue nearai#3690 reference for the narrowed `HookObservableEvent` projection. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * perf(hooks): adaptive backoff for event-triggered subscription Address PR nearai#3640 review findings C5, A1, A2: - C5: empty-poll backoff for `EventTriggeredHookSubscription`. The previous loop hammered the durable log at a fixed 50ms cadence under sustained idle, even when no events had arrived for minutes. The subscription now tracks consecutive empty polls and sleeps for `min(poll_interval << streak, max_poll_interval)` before the next poll, defaulting to a 1s cap; a non-empty batch resets the streak so producer bursts restore low-latency dispatch immediately. Exposed via `with_max_poll_interval` so callers can tune. - A1 / A2: explicit issue references for the deferred narrowed `HookObservableEvent` projection (nearai#3690) and the per-hook DoS dispatch budget (nearai#3689). The current self-trigger guard catches direct-recursion storms; the backoff bounds indirect ones until the proper budget design lands. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(hooks): cover event-triggered dispatch edge cases Address PR nearai#3640 review findings D8, D9, D10, D11, D12, and B: - D8: dispatching an event-triggered binding that has no installed hook impl must poison the slot and surface a Malformed failure rather than silently no-op. A follow-up dispatch on the same kind must skip the poisoned slot. - D9: registry validation rejects non-event-point bindings that carry an `event_kind_filter`, mirroring the existing reverse-direction check. - D10: the existing hook-meta serde round-trip tests always passed `None` for `owning_extension` and never asserted `event.provider`. Add `hook_meta_events_round_trip_owning_extension_as_provider` to pin the projection that scope filtering depends on. - D11: `scope_provider_for_runtime_event` falls back to `None` when the registry mutex is poisoned. Force a poison on a spawned thread and assert the resolver remains fail-closed. - D12: `run_event_triggered_hook` catches panics from the hook impl via `AssertUnwindSafe::catch_unwind`. Drive it with a deliberately panicking impl and assert `FailureCategory::Panic`. - Cluster B: when a hook-meta event has `provider: None`, the dispatcher recovers the owning extension from the registry's hex-keyed index so `OwnCapabilities` watchers still fire. Add a full end-to-end test exercising that path through `dispatch_event_triggered_at`. Also pin C4 indexing: a registry-level test that `active_for_event_kind` returns only bindings whose declared filter matches. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(hooks): port event-triggered tests after foundation rebases (nearai#3911/nearai#3912/nearai#3913) - Add event_kind_filter: None to HookBinding test constructions (foundation added new field) - Replace tuple-struct construction of ExtensionId/HookLocalId with ::new() per nearai#3912 newtype privatization - Lowercase RuntimeEventKind debug repr for HookLocalId validation (lowercase-only identifiers) - Replace pub-use re-exports with module-path imports per foundation cleanup - Rename fixture.user_id to fixture.actor_id per nearai#3633 final naming * fix(hooks): adapt event-triggered to WASM hook runtime (nearai#3920) - Add event_kind_filter: None to WASM HookBinding constructions in dispatch.rs - Extend HookManifestKind match arms in registrar.rs and wasm/runtime.rs to handle EventTriggered (rejected: WASM-bodied event-triggered hooks are not yet supported) --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…n event-triggered (nearai#3640 followup) (nearai#3931) * fix(hooks): resolve hook-lifecycle owner from hook_id, not spoofable provider claim Bug 3 (CRITICAL) from Codex review of nearai#3640: scope_provider_for_runtime_event trusted event.provider before consulting the registry. A hook could synthesize a HookFailed event for hook B (owned by ext-B) while claiming provider = ext-A, causing ext-A's OwnCapabilities hooks to fire on an event about another extension's hook. For hook-lifecycle event kinds (HookDispatched/HookDecisionEmitted/HookFailed) the authoritative owner is now resolved from event.hook_id via HookRegistry::owning_extension_for_hook_hex. If the resolved owner disagrees with the payload claim, the discrepancy is logged and the resolved owner is used (never the claim). Lifecycle events with no hook_id, or whose hook_id does not resolve, fall back to None (inert) — fail-closed. Non-lifecycle events, which carry no hook_id anchor, continue to use the host-established provider. TDD: hook_failed_with_spoofed_provider_does_not_fire_target_extension_hooks. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * fix(hooks): wire is_replay through event-triggered replay dispatch path Bug 2 (CRITICAL) from Codex review of nearai#3640: the event-triggered hook subscription always called dispatch_event_triggered_at (live), never dispatch_event_triggered_replay_at, so EventTriggeredHookContext::is_replay was always false. On reconnect/restart, the backlog caught up from the resume cursor replayed duplicate side effects with no dedupe signal, despite the public is_replay contract. The subscription now tracks the replay/live boundary without new durable state: events caught up from start_cursor (the caller's committed resume point) to the stream head at subscription time are dispatched through the replay path (is_replay = true). The first empty poll means the backlog is drained and head is reached; every event after that boundary is live (is_replay = false). Over-marking a live event as replay is the fail-safe direction; the previous bug marked replayed events as live. TDD: event_subscription_marks_is_replay_true_for_replayed_events (integration). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * fix(hooks): derive effective ReadScope from run scope, reject permissive any() Bug 1 (CRITICAL) from Codex review of nearai#3640: validate_against_run_scope accepted a permissive read_scope (including ReadScope::any()) and passed it unchanged to read_after_cursor. Durable event streams are keyed only by (tenant, user, agent), so within one stream a run could observe events from other projects/threads/missions — a cross-thread/project trust-boundary break. Replaced with effective_read_scope: it rejects ReadScope::any() with a typed RebornLoopDriverHostError::ScopeMismatch (fail-closed), keeps the existing stream-key (tenant/user/agent) checks, requires any caller-supplied Some(want) to equal the run/thread value (tighten, never widen), and then *derives* the filter actually used from the authoritative run/thread scope: thread_id pinned to the run thread, project_id/mission_id pinned when present. The host build applies the derived scope via with_read_scope before spawning the subscription. TDD: event_subscription_rejects_read_scope_any, event_subscription_does_not_observe_other_project_events, effective_read_scope_pins_thread_and_project_from_run_scope. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * fix(reborn): snapshot stream head at subscription start for replay/live boundary PR nearai#3931 Hole 1: the event-triggered hook subscription treated "the first poll that returns no entries" as the replay→live boundary. That races: a live event appended after subscription start but before the first empty poll (or while a continuous backlog drains past the true head) was dispatched via the replay path and marked is_replay=true, violating the "gap from start_cursor to head-at-startup" contract and risking dedupe/skip of legitimate live side effects. Fix: add DurableEventLog::head_cursor (default impl drains unfiltered; in-memory backend overrides O(1)) and snapshot startup_head atomically in EventTriggeredHookSubscription::spawn — synchronously, before tokio::spawn — so "head-at-startup" means the head at host-build time, not at the spawned task's first lazy poll. Records classify per-cursor: cursor <= startup_head is replay, everything beyond is live; mixed batches split at the boundary. Fail-closed: if the head snapshot fails, the subscription does not dispatch and emits an operator-visible terminated milestone. Tests: live event before first empty poll is LIVE; event during continuous backlog drain past startup_head is LIVE; head_cursor contract (latest cursor, future-cursor ReplayGap). Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * fix(hooks): fail closed on unknown/poison lifecycle owner, never trust carried provider PR nearai#3931 Hole 2: scope_provider_for_runtime_event resolved the owner of a hook-lifecycle event (HookFailed/HookDispatched/HookDecisionEmitted) from the registry by hook_id, but the `(None, claimed) => claimed` arm fell back to the event's carried `provider` payload whenever the registry lookup returned None. That payload field is forgeable — it is a plain serialized field, not an unforgeable host stamp — so two attacks slipped through: 1. Unknown hook_id + provider=ext-a: a synthesized lifecycle event naming a hook_id absent from the registry could fire ext-a's OwnCapabilities hooks. 2. Poisoned registry + provider=target: registry poison collapsed to None and then trusted the carried provider, defeating fail-closed. Fix: for lifecycle events, the owner is the registry-resolved owner of hook_id and ONLY that. Poisoned registry and unknown hook_id are kept as distinct branches that both resolve to None (hook inert); neither falls back to the carried provider. A registry hit still wins, logging a spoof warning when the payload disagrees. Design tension resolved (Phase 5 alerting): two integration tests previously asserted that an OwnCapabilities watcher fires for a lifecycle event whose *subject* hook_id was UNREGISTERED, relying on the carried provider. That is the exact spoof shape the fix rejects. In production the subject hook ran in this same host and so resolves via the registry — the carried-provider fallback was never needed for a legitimate flow. The tests are reworked to register the subject hooks (so ownership resolves from the registry, the production reality): the ext-A-owned event fires, the ext-B-owned event stays inert, and the self-suppression control case still proves the filter is narrow. There is no unforgeable host-stamped owner source distinct from the payload today; if a future extension-emitted or cross-host lifecycle path needs one, it must add an authenticated owner field rather than trusting the payload. Tests: unknown lifecycle hook_id + provider=target stays inert (resolves None); poisoned registry + provider=target resolves None; reworked Phase 5 and self-redispatch integration tests to resolve ownership from the registry. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * refactor(hooks): address serrrfirat review — atomic head_cursor, drop any() ceremony, extract lifecycle owner resolver P1: make DurableEventLog::head_cursor a required backend operation (no default). The previous default drained the stream page-by-page until an empty read, which is non-atomic (a record appended during the probe folds into the observed head and gets mis-marked replay) and moved an unbounded scan into host construction. Each backend now reads its own tail atomically: - JsonlDurableEventLog: read_last_jsonl_cursor (O(1) tail seek) - FilesystemDurableEventLog: single tail(after-1) snapshot, max seq - InMemoryDurableEventLog: already O(1) via next_cursor All production + test impls updated. Added filesystem head_cursor contract test (latest cursor, mid-stream probe, future-cursor ReplayGap). P2: drop the ReadScope::any() rejection in EventTriggeredHookSubscription::effective_read_scope. Verified the read uses the DERIVED authoritative scope (installed via with_read_scope before the task spawns), not the caller's filter — the derivation always pins thread/project/mission from the run/thread scope, so any() cannot widen the read or reopen the cross-thread/project leak. The rejection was redundant ceremony. Reverted the forced hooks_integration.rs fixture churn (event_log_subscription back to ReadScope::any(), no &fixture param). Test renamed to assert any() is accepted and constrained by the derived scope. P2: extract scope_provider_for_runtime_event into dispatch/lifecycle_owner.rs (resolve_event_owner + LifecycleOwnerLookup), converting dispatch.rs to dispatch/mod.rs. Pure, table-tested resolver with 6 cases: registered owner, spoofed provider override, unknown hook id, poisoned registry, missing hook id, non-lifecycle passthrough. Behavior identical to the security-fixed inline version. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * test(events): prove head_cursor atomicity — post-snapshot append classified live Adds head_cursor_snapshot_classifies_later_appends_as_live to the durable log contract suite (PR nearai#3931 P1): a record appended after head_cursor returns receives a cursor strictly greater than the observed startup_head, so it is classified live and never folded into the replay window. Locks in the atomic-head contract that the required (no-default) head_cursor method and its per-backend O(1)/single-snapshot overrides must satisfy. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * fix(hooks): address henrypark133 review on nearai#3931 (head_cursor perf + tests + must_use) - Add `RootFilesystem::head_seq(path, from)` returning `Option<SeqNo>` with a correctness-preserving `tail`-backed default; override in Postgres/libSQL with O(1) `MAX(seq)` so a fresh subscription (`after = 0`) no longer materializes the whole stream gap into memory just to find its head. Route `FilesystemDurableEventLog::head_cursor` through `head_seq`. (Perf finding) - Add JSONL backend head_cursor coverage: `jsonl_head_cursor_reports_latest_ and_rejects_future` exercises empty-stream head, post-append head, mid-stream probe, and future-cursor ReplayGap rejection. (Tests finding) - Add `event_triggered_head_cursor_failure_emits_terminated_milestone`: a stub log whose head_cursor fails at subscription spawn must emit exactly one EventSubscriptionTerminated milestone (fail-closed). (Tests finding) - Add `#[must_use]` to `EventTriggeredHookSubscription::with_read_scope` to match its four sibling builder methods. (Local-patterns finding) The in_memory doc-comment nit is declined: the head_cursor doc records the atomicity contract at the impl site, which is intentional. Co-Authored-By: Claude Opus 4.7 <noreply@anthropic.com> * fix(hooks,reborn): close exhaustiveness + scope-derivation review gaps (nearai#3931) M1: replace the non-exhaustive `matches!` in `is_lifecycle_kind` with an exhaustive `match` enumerating every `RuntimeEventKind` variant (no `_` wildcard). A future lifecycle variant now fails to compile until it is explicitly classified, instead of silently falling through to the non-lifecycle branch that trusts the forgeable `event.provider` (Hole 2). Adds a test asserting every variant classifies on the correct side. M2: narrow `EventTriggeredHookSubscription::with_read_scope` from `pub(crate)` to module-private. It has no cross-module caller (only the in-module host build path at line ~1351 invokes it), so the narrower visibility removes the latent surface for other `ironclaw_reborn` code to install an arbitrary `ReadScope::any()` and bypass `effective_read_scope`. M3: emit a `tracing::warn!` in `effective_read_scope` when `run_scope.project_id` is `None`, documenting that the subscription is not project-pinned (thread_id still pins it). Not fail-closed: tenant/agent-level project-less runs are legitimate. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 <noreply@anthropic.com>
Successor PR from #3573. Phase 5 implementation: new
EventTriggeredhook point that subscribes to the runtime event bus and reacts to durableRuntimeEvents asynchronously, outside the inline dispatch tick. Observer-only by construction (no Allow/Deny/Patch decisions; sink isObserverSink).Scope
HookPointSpec::EventTriggeredpoint +EventTriggeredHooktrait that reusesObserverSink(no duplicate sink trait — see review finding F14).HookDispatcher::dispatch_event_triggered_at+dispatch_event_triggered_replay_atentry points.HookRegistryso dispatch is O(matches) rather than O(all event-triggered bindings).EventTriggeredHookSubscriptioninironclaw_reborn— pull-driven durable-log consumer with adaptive backoff (50ms base → 1s cap) so an idle stream backs off instead of hammering the log at a fixed cadence.EventTriggeredHookContext::is_replaylets side-effecting hooks dedupe byevent.event_idon at-least-once restart replay.providerfield resolve the owning extension via the registry's hex-keyed index soOwnCapabilitieswatchers still fire.Motivation
Inline hook points are synchronous against the loop. Some legitimate hook use cases don't fit:
HookFailedshouldn't block the loop)Design doc
crates/ironclaw_hooks/docs/successors/04-event-triggered-hooks.mdCursor / replay
The subscription is at-least-once. Restarting from the same cursor replays every event with cursor
>= start_cursor, including events already dispatched before the prior shutdown. Hooks treatctx.is_replay = trueas a signal to dedupe byevent.event_id. Cursor persistence is caller-owned for this slice; exact-once acknowledgement is a future ratification slice.Follow-ups tracked as issues
HookObservableEventnarrowed projection for Installed-tier hooks (stripResourceScopefields from durable events handed to Installed hooks): hooks: narrow RuntimeEvent to HookObservableEvent projection for Installed-tier event hooks #3690Coordination
Status
Ready for re-review. The 14 findings from the 2026-05-19 multi-agent review are addressed in three focused commits (refactor + perf + tests); see the consolidated PR comment for the per-finding map.