Skip to content

[codex] restrict Docker publishing to upstream - #7

Merged
theredspoon merged 1 commit into
native-matrix-channel-pilotfrom
codex/disable-enjimi-docker-native
Jun 17, 2026
Merged

theredspoon merged 1 commit into
native-matrix-channel-pilotfrom
codex/disable-enjimi-docker-native

Conversation

@theredspoon

Copy link
Copy Markdown
Owner

Summary

Restrict the Docker Image workflow's publishing job to nearai/ironclaw only.

This keeps mirrored workflow files harmless in theredspoon/ironclaw and enjimi/ironclaw; enjimi should build the pilot Linux artifact from deployment-target, not publish Docker images.

Validation

  • Inspected the workflow diff; change is a single job-level guard on .github/workflows/docker.yml.
  • Local commit hook ran gitleaks successfully during commit.

@github-actions github-actions Bot added scope: ci size: XS Changed-line size classification risk: medium Risk classification contributor: regular Contributor history classification labels Jun 17, 2026
@theredspoon
theredspoon marked this pull request as ready for review June 17, 2026 23:18
@theredspoon
theredspoon merged commit 44a0cad into native-matrix-channel-pilot Jun 17, 2026
13 checks passed
theredspoon pushed a commit that referenced this pull request Jun 18, 2026
Consolidated fixes for serrrfirat's 10 unresolved review threads plus
zmanian's CHANGES_REQUESTED review (3 blockers + 7 mediums + 3 follow-up
items + title typo).

## Blockers

1. CI now runs the package tests (zmanian blocker 2 / serrrfirat #7).
   The matrix `cargo test ${{ matrix.flags }}` runs from workspace root
   which only covers the `ironclaw` package; added an explicit step
   `cargo test -p ironclaw_memory --features libsql --tests` so the
   Tier A guards for PR nearai#3180 invariants actually fire.

2. `#[ignore]` markers converted to `#[cfg_attr(not(feature =
   "pr3180-ready"), ignore = ...)]` (zmanian blocker 1 / serrrfirat #1).
   Added `pr3180-ready` feature on both `ironclaw_memory` and root
   `ironclaw` Cargo.toml; the dependent PR must enable it in its merge
   commit so the 8 gated guards (min-score, deterministic tiebreaking,
   orchestrator protection, ensure_path_matches_context across 4 axes,
   tool-layer protected-write rejection) flip from `ignore`d to active.

3. Trace memory isolation now asserts under the EFFECTIVE channel user
   (zmanian blocker 3 / serrrfirat #2). Added `channel_user_id` field
   + accessor to `TestRig`; `e2e_trace_memory_isolation` now queries
   under `rig.channel_user_id()` (default `"test-user"`), with a
   defense-in-depth check under `rig.owner_id()` for mis-routing
   regressions.

## Test-correctness mediums

4. Min-score test pins `with_query_embedding([1,0,0])` to favor
   hybrid.md (serrrfirat #3 / zmanian #4). Removed the permissive
   `* 0.99` fallback — FTS-only-below-hybrid is now a hard assertion.

5. Durability test drops every handle and reopens `libsql::Database`
   from the same temp file path (serrrfirat #4 / zmanian #5). Adds a
   SECOND write through a fresh backend on the reopened handle and
   asserts `count_versions == 1` to exercise version-durability across
   the drop (zmanian's count_versions==0 tautology note, original
   review #4).

6. Append versioning asserts exact row count `== 1`, not `!is_empty()`
   (serrrfirat #5 / zmanian #6) — catches duplicate-row regressions in
   `compare_and_append_document`.

7. Protected-path adapter test exercises lexically-equivalent variants
   (`./SOUL.md`, `//SOUL.md`) in addition to canonical (serrrfirat #6 /
   zmanian #7). VirtualPath rejects `..` so no `..` variant; in-loop
   `count_documents_total == 0` after EACH variant.

8. Hybrid search isolation now varies all four scope axes (serrrfirat
   #8 / zmanian #8): tenant, user, agent, project. 5 documents seeded;
   search from caller scope must return exactly one.

9. Tool round-trip asserts EXACT persisted content via direct DB read
   (serrrfirat #9 / zmanian #9). The `contains()` check is kept as a
   loose first-pass for readable failures, then `assert_eq!` on the
   exact byte string is the load-bearing assertion.

10. Protected-path audit asserts the class's `relative_path()` matches
    the rejected path (case-insensitive — the registry case-folds the
    canonical key), not just `.is_some()` (serrrfirat #10 / zmanian #10).
    A regression that emits the wrong path class now fails.

## zmanian follow-ups

Z1. Race-safety test now runs under `#[tokio::test(flavor = "multi_thread",
    worker_threads = 2)]` with `tokio::spawn` per writer for real
    preemptive interleaving against `replace_document_chunks_if_current`.
    Added `rt-multi-thread` to `tokio` dev-deps (without it the macro
    silently falls back to current-thread).

Z2. `write_to_protected_path_rejected.json` trace fixture sets
    `all_tools_succeeded: false` explicitly. Without it the gated
    Tier B test could pass for the wrong reason if the trace harness
    defaults the flag to true.

Z3. Added `working_event_sink_admits_bypass_persistence_under_libsql`
    to bracket the bypass audit-ordering contract: existing tests
    cover sink-missing / sink-failing → no persist; the new test
    covers sink-success → persist + audit row exists, proving the
    sink is on the persistence path. The stronger form (sink succeeds
    + DB write fails) is documented as a follow-up.

## Cleanup

- Removed `_link_in_memory_repo_for_unused_imports` shim and the
  `InMemoryMemoryDocumentRepository` import that only existed to feed
  it (zmanian original-review #3).
- Fixed PR title typo `momery` → `memory` via gh.

Helper-consolidation into `tests/common/libsql_helpers.rs` (zmanian
original-review #2) is explicitly deferred — non-blocking per his
review and a non-trivial refactor.

## Verified

- `cargo fmt --all -- --check` clean
- `cargo clippy -p ironclaw_memory --features libsql --all-targets -- -D warnings` zero warnings
- `cargo test -p ironclaw_memory --features libsql` all suites green
  (gated tests stay `ignored` without `--features pr3180-ready`)
theredspoon pushed a commit that referenced this pull request Jun 18, 2026
…rai#3544

Three Opus subagents reviewed the four amendment commits and surfaced
5 critical implementation blockers, 5 cross-doc consistency drifts,
and 5 architectural gaps. This commit fixes the blockers, the drifts,
and three of the quick architectural wins. Two strategic items
deferred for separate discussion (future-fork story; §9 cleanup).

Critical blockers:
- B1: ConcurrencyHint circular dependency. Moved type definition
  from ironclaw_agent_loop (WS-2) to ironclaw_turns (WS-0) — the
  field on CapabilityDescriptorView lives in turns, so the type
  must live in turns. WS-2 imports the type rather than defining it.
- B2: stage_checkpoint_payload was specified on AgentLoopDriverHost
  (a method-less marker trait). Moved declaration to
  LoopCheckpointPort alongside load_checkpoint_payload; callers
  still use host.stage_checkpoint_payload(...) via deref-through-
  supertrait.
- B3: WS-7 family.id().to_string().as_str() snippet was E0716
  (temporary dropped while borrowed) AND gratuitous — LoopFamilyId
  is already &'static str. Fixed to family.id().0.
- B4: CapabilityDescriptorView field-add is BREAKING (public fields,
  struct-literal constructors). WS-0 brief now explicitly lists
  consumers that need updating in the same PR.
- B5: WS-5 acceptance criterion still said "aborts on PolicyDenied"
  — straggler from seam-5a SkipResult amendment. Fixed.

Cross-doc consistency:
- D1: load_checkpoint_payload signature drift across WS-0, WS-7,
  WS-10. WS-10 is source of truth; WS-0 drops the inline stub and
  WS-7's resume pseudocode uses the canonical request/response shape.
- D2: from_checkpoint_payload signature drift (Value vs bytes).
  Bytes-based two-arg shape is now canonical in WS-0; matches the
  reality that checkpoint storage stores bytes.
- D3: Cancellation boundary count was inconsistent (prose said
  "Eight," table had 9 rows). Combined rows #6 (Reply path) and #7
  (CapabilityCalls path) — they're mutually exclusive branches at
  the same model-response match point. Eight rows everywhere now.
- D4: Cancellation helper name was inconsistent across briefs.
  Standardized on checkpoint_and_exit_if_cancelled across master
  doc, WS-6, WS-13.
- D5: WS-8 had no test for the Denied → SkipResult path. Added two
  rows to strategy_interactions.rs: denied_call_skips_and_continues
  and repeated_denied_calls_trip_no_progress.

Quick architectural wins:
- G1: WS-9 now enumerates EffectKind → ConcurrencyHint mapping per
  variant. Network → Exclusive (conservative; POSTs are causal).
  UseSecret → SafeForParallel (read-only secret access).
  DispatchCapability → Exclusive (recursive depth unsafe). Empty
  effects → SafeForParallel (pure function). All write/spawn/
  modify variants → Exclusive.
- G3: WS-6 §3.5a documents strategy-decision observability via
  tracing::debug! at every strategy call site. Durable typed
  strategy-decision telemetry deferred to a future workstream
  pending production debugging need.
- G4: Master doc §10 documents in-flight Blocked run behavior when
  ComponentIdentity.digest changes: LoopExit::Failed {
  CheckpointUnavailable }; never silently resume against changed
  digest. Operators expected to plan deploys with this in mind.

Deferred for separate discussion:
- G2: future-fork story — §4 claims families graduate to own crates
  but pub(crate) strategy seal makes this impossible without a
  pub(in family-factory) escape hatch.
- G5: §9 has 17 cross-referenced bullets with PR-comment URLs that
  are institutional memory rather than documentation; needs
  editorial cleanup with worked decisions inline.

Spec-only; no code changes. 9 files touched.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
theredspoon pushed a commit that referenced this pull request Jun 18, 2026
…rai#3679)

* feat(processes): route FilesystemProcessStore through unified put/get

First consumer migration onto the new RootFilesystem surface. Switches
the byte-plane read_file/write_file calls inside ironclaw_processes'
filesystem-backed store to the unified put/get ops with Entry::bytes +
CasExpectation::Any. The on-disk JSON layout is unchanged, every
existing test passes, and downstream crates that construct
FilesystemProcessStore (ironclaw_host_runtime + tests) don't need to
change.

Scope deliberately narrow: opaque-file entries through `put`/`get`
without record kinds or non-`Any` CAS, since LocalFilesystem's native
`put` only accepts that shape (per the foundation PR #3659). Once
LocalFilesystem grows sidecar metadata, this consumer can switch to
`Entry::record(process_record_kind, ...)` + `CasExpectation::Absent`
without changing the on-disk layout.

Touch points:
- write_record uses put(Entry::bytes, CAS::Any)
- start uses get for the existence probe + transition_lock for the
  atomicity envelope per the single-instance invariant
- update_status / get / records_for_scope read via get and unwrap
  VersionedEntry.body
- records_for_scope returns ProcessError::Filesystem (not silent skip)
  when get returns None for a path that list_dir just yielded —
  matches the pre-migration NotFound propagation invariant

Test scaffold update: BackendErrorFilesystem now overrides `get` too,
so the fault-propagation regression test continues to exercise its
intended path. (Reviewer P1/P2 on the original #3666 — recursion +
silent-skip — addressed in foundation #3659 directly since LocalFilesystem
now ships native `put`/`get`.)

* feat(outbound): add FilesystemOutboundStateStore on the unified surface

Stacked on the consolidated foundation PR #3659. Adds an
OutboundStateStore impl that persists outbound metadata under
/engine/outbound/{policies,subscriptions,deliveries} through any
RootFilesystem. The existing libSQL/Postgres/in-memory stores stay
intact during the migration; a follow-up cleanup PR can delete them
once production runs on the unified surface.

The new store passes the full contract suite (durable_policy_*,
subscription_cursor_*, delivery_status_*, notification_policy_*,
full_turn_scope_isolation) against InMemoryBackend in the existing
outbound_state_store_contract.rs test file.

* feat(authorization): route FilesystemCapabilityLeaseStore through unified put/get

Stacked on PR #3670 (outbound). Mirrors the ironclaw_processes
migration in PR #3666 / now consolidated into #3659. Switches the
filesystem-backed lease store's read_file/write_file calls to the
unified get/put ops with Entry::bytes + CasExpectation::Any. The
on-disk JSON layout is unchanged, every existing test passes, and the
per-owner mutation_lock continues to serialize claim/consume/revoke
within a single instance.

Touch points:
- read_lease, read_lease_index, read_lease_file — now use get and
  unwrap VersionedEntry.body.
- write_lease, write_lease_index — now use put(Entry::bytes, Any).
- Imports updated.
- CountingFilesystem test scaffold gains put/get overrides that
  forward to its inner LocalFilesystem, since the trait defaults are
  now Unsupported after the PR #3659 recursion fix.

* feat(run-state): unified put/get for filesystem stores

Stacked on PR #3671 (authorization). Mirrors processes (#3666) and
authorization (#3671) migrations. Switches all read_file/write_file
calls in FilesystemRunStateStore and FilesystemApprovalRequestStore
to the unified get/put ops with Entry::bytes + CasExpectation::Any.
On-disk JSON layout unchanged.

Test scaffold updates: ConcurrentMissingReadFilesystem and
DisappearingApprovalReadFilesystem gain put/get overrides that
forward to their inner LocalFilesystem and apply the same fault
injection logic on the unified read path (was: only on the legacy
read_file path). Required after the trait defaults moved to
Unsupported in PR #3659.

* refactor(workspace): dissolve ironclaw_storage

The ironclaw_storage crate predates the unified RootFilesystem surface
introduced by PR #3659 (universal FS dispatch). Its `BlobStore`/`RecordStore`
traits, `StorageKey`/`StorageVersion`/`PutCondition` types, and
`StoredBlob`/`StoredRecord` shapes parallel the new unified put/get
/CasExpectation/RecordVersion machinery on `RootFilesystem` — a textbook
duplicate-dispatch smell flagged by .claude/rules/architecture.md.

Only `ironclaw_outbound` consumed any of the crate, and only 5 small
helpers (`encode_json`, `decode_json`, `redacted_backend_error`,
`StorageError::Backend`, `ABSENT_SCOPE_COMPONENT`). All other types and
the entire `BlobStore`/`RecordStore` surface (660 LOC) were unused —
their intended consumers already moved to `RootFilesystem` directly.

Inlined the 5 helpers into `crates/ironclaw_outbound/src/db.rs`:
- `encode_json`/`decode_json` → direct `serde_json::to_string`/`from_str`
- `redacted_backend_error` → local log+collapse to `OutboundError::Backend`
  (preserves the redaction boundary required by ironclaw_outbound/CLAUDE.md)
- `ABSENT_SCOPE_COMPONENT` → local const ""

Removed the crate's workspace membership, the outbound dep, the
forbidden-edges BoundaryRule, and the crate directory.

Also updated the ironclaw_outbound BoundaryRule to permit a normal
dependency on `ironclaw_filesystem` — `FilesystemOutboundStateStore`
landed in the prior cascade PR and the boundary rule was stale.

* feat(filesystem): add HsmBackend placeholder + scope database.md to legacy

Two changes that close out the demoable parts of the universal-FS-dispatch
rework (tasks #18 and the demonstrable portion of #19 from the plan).

**HsmBackend placeholder** (`crates/ironclaw_filesystem/src/hsm.rs`).
Demonstrates that a new backend is a single-file change: implements the
one `RootFilesystem` trait, declares a restricted capability surface
(`Read` + `Write` + `Stat` + `Delete` + `TxnCapability::Cas` — no records,
no query, no index, no events, no multi-key transactions), and routes
`put`/`get`/`delete`/`stat`/`list_dir` through an in-process placeholder.

Five tests prove the seam works end-to-end:

- `hsm_supports_encrypted_bytes_round_trip` — bytes put/get works.
- `hsm_rejects_structured_records` — `put` with `RecordKind::Some` or
  non-empty `indexed` returns `Unsupported`, so a consumer cannot
  accidentally route records through encryption-only storage.
- `hsm_rejects_query_and_index_ops` — `query`/`ensure_index` return
  `Unsupported` consistent with the declared capabilities.
- `composite_rejects_overclaimed_hsm_descriptor` — mount-time
  validation (`validate_mount_capabilities`) refuses a descriptor that
  claims `Query`/`IndexExact` over a backend that doesn't deliver,
  failing with `FilesystemError::DescriptorOverclaims { missing, .. }`.
- `composite_routes_to_hsm_under_secrets_mount` — the acceptance gate:
  mounting HsmBackend at `/secrets` and routing put/get through the
  composite works with no consumer-visible changes. Indexed projection
  is still rejected because the declared capabilities advertise no
  index/query support.

A real HSM implementation replaces the in-memory placeholder with an
HSM session handle; the trait surface, capability declarations, and
mount-time validation are reusable as-is. The placeholder is *not* a
security boundary — it is a seam demonstration.

**database.md scoped to legacy directories**. The dual-backend rule
file (`.claude/rules/database.md`) is `paths`-scoped to `src/db/**`,
`src/history/**`, and `migrations/**` — exactly the legacy surface that
predates the universal FS dispatch. Added a "Status & Direction"
preamble pointing new persistence work at `ScopedFilesystem` and the
`2026-05-14-universal-fs-dispatch.md` plan, with the existing
per-crate dual-backend guidance kept (and tagged "legacy") for code
still inside those directories.

* feat(reborn): route durable event store through RootFilesystem

Add native `append`/`tail` to the libsql and postgres `RootFilesystem`
backends and ship a `FilesystemDurableEventLog` / `FilesystemDurableAuditLog`
alternative for `ironclaw_reborn_event_store`. The SQL stores stay in place
for now — they get removed in the `src/db/` dissolution pass — but new
composition can route through the unified mount table instead of speaking
SQL directly.

- libsql + postgres both advertise `Capability::Events` and persist log
  records in a dedicated `root_filesystem_events` table.
- Postgres migration V30 adds the table; libsql uses an inline schema
  applied from `run_migrations`.
- Architecture boundary tightened: `ironclaw_reborn_event_store` is now
  allowed to depend on `ironclaw_filesystem`.

* feat(secrets): route secret + credential storage through RootFilesystem

Add `FilesystemSecretStore` and `FilesystemCredentialBroker` alongside the
existing libSQL/Postgres backends so secret material, secret leases,
credential accounts, and credential sessions can persist through the
unified `RootFilesystem` dispatch fabric (matching prior migrations in
`ironclaw_processes`, `ironclaw_authorization`, `ironclaw_outbound`, and
`ironclaw_run_state`).

- Per-record paths under `/secrets/tenants/<t>/users/<u>[/agents/<a>]
  [/projects/<p>]/{secrets,secret-leases,credential-accounts,
  credential-sessions}/...`.
- Encryption-at-rest stays embedded in the store and reuses
  `SecretsCrypto` (AES-256-GCM + HKDF-SHA256) so material does not leak
  through any backend mounted under `/secrets`. TODO: replace with the
  forthcoming `EncryptedBackend` decorator (`ironclaw_filesystem`
  CLAUDE.md invariant #5).
- Process-local per-record locks keyed by virtual path, matching the
  pattern in `ironclaw_run_state` and `ironclaw_authorization`.
- `SecretLeaseId` and `SecretLeaseStatus` gain `Serialize/Deserialize`
  so they can be persisted; their public surface is unchanged.
- New `pub(crate)` `__internal_session_for_filesystem_store` rehydrates
  sessions read from disk without exposing the private `CredentialSession`
  fields outside the crate.
- Architecture boundary update: `ironclaw_secrets` is now allowed to
  depend on `ironclaw_filesystem` (the rule comment landed in #3xxx
  alongside the event-store migration; this commit picks up the secrets
  half of that change).
- Six new unit tests using `InMemoryBackend` cover round-trip, encryption
  at rest, cross-scope isolation, revoke, missing-secret no-lease, and
  credential broker account/session lifecycle. All existing tests pass
  unmodified (60 tests total).

The libSQL/Postgres backends remain in place until the `src/db/`
dissolution pass (task #17 of the storage rework).

* feat(filesystem,memory): add Fts + Vector indexes and filesystem-backed memory repo

Phase 1: extend the libsql and postgres `RootFilesystem` backends with
`IndexKind::Fts` and `IndexKind::Vector { dim }`, plus the matching
`Filter::Fts { key, query }` and `Filter::VectorNearest { key, embedding,
limit }` evaluation paths.

- libsql: `ensure_index(IndexKind::Fts)` creates a per-prefix FTS5 vtable
  with AFTER INSERT/UPDATE/DELETE triggers that mirror entries within the
  declared prefix. Backfill on declaration handles pre-existing rows.
  `Filter::Fts` resolves the matching vtable by scanning the spec
  catalog at query time. Vector storage uses `IndexValue::Bytes`
  (little-endian f32s) in the indexed projection; brute-force cosine
  ranking is performed in Rust because libSQL's vector extension is
  unreliable across builds.
- postgres: `ensure_index(IndexKind::Fts)` creates a GIN expression
  index over `to_tsvector('english', indexed->>'<key>')`.
  `Filter::Fts` translates to a `@@ plainto_tsquery(...)` predicate so
  the GIN index is usable. Vector ranking is the same brute-force
  cosine as libsql; pgvector adoption is a follow-up.
- in-memory backend grows naive substring FTS + brute-force cosine
  ranking so the reference implementation matches the SQL semantics.
- Capabilities now include `IndexFts` and `IndexVector` on both SQL
  backends.
- Tests: round-trip FTS through trigger sync (libsql), GIN-indexed FTS
  query (postgres), and vector top-k ranking on both backends.

Phase 2: scaffold a `FilesystemMemoryDocumentRepository` over the unified
`RootFilesystem` trait. Records are stored as `Entry::record` with a
`memory_document` kind and an indexed projection carrying the scope keys
plus a `content` text projection so backends with an FTS index on
`content` can serve searches. Metadata is stored at a sibling `.meta`
path. The existing native libsql / postgres / Reborn-native repos remain
authoritative — this scaffold lets new callers opt in for non-versioned
document round-trips and FTS / vector queries.

Known TODOs documented inline in `filesystem.rs`:
- versioned compare-and-append via `CasExpectation::Version`
- chunking projection writes (currently only the native repos maintain
  the chunk store the hybrid searcher consumes)
- full hybrid-search wiring (`MemorySearchRequest` -> `Filter::Fts` +
  `Filter::VectorNearest` + RRF fusion)
- capability declaration on `MemoryBackendFilesystemAdapter`

Also fixes a pre-existing compile error in
`reborn_native_filesystem_vertical_integration.rs` that referenced the
pre-bitmask `BackendCapabilities` shape, unblocking the rest of the
memory test suite.

Test counts after this commit:
- `ironclaw_filesystem` --all-features: 94 passing (3 new contract tests)
- `ironclaw_memory` --all-features (PG-skipped): 219 passing, 3
  pre-existing failures inherited from the base branch
- `ironclaw_architecture`: 14 passing

* feat(db): add filesystem-backed ConversationStore and JobStore facades

Add FilesystemConversationStore and FilesystemJobStore as alternatives to
the libSQL/Postgres backends. Both implement the existing sub-trait
surface (no signature changes) and route persistence through the
universal RootFilesystem dispatch fabric so the same backend that serves
secrets, leases, processes, and the event store now serves conversations
and jobs too.

Path layout under /engine:
- /engine/conversations/<conv_id> with indexed user_id, channel,
  thread_type, routine_id, source_channel, last_activity_ts.
- /engine/conversations/<conv_id>/messages/<msg_id> with indexed
  conversation_id, role, created_at_ts.
- /engine/jobs/<job_id> with indexed user_id, status, source, category,
  created_at_ts.
- /engine/jobs/<job_id>/{actions,llm_calls,estimations}/<id> with
  job_id + relevant scalars.

Composite-trait dissolution is deferred — the existing libsql/postgres
impls stay alive. 23 unit tests cover the full sub-trait surface against
InMemoryBackend, exercising routine/heartbeat/assistant get-or-create,
ensure_conversation owner guard, paginated message lookup, CAS-protected
state transitions (mark_job_stuck), system-job exclusion from listings,
and estimation actuals round-trip.

* feat(db): add filesystem-backed Sandbox/Routine/ToolFailure stores

Add `FilesystemSandboxStore`, `FilesystemRoutineStore`, and
`FilesystemToolFailureStore` as `RootFilesystem`-backed facades for the
three matching `src/db/` sub-traits. Records live under new virtual
roots `/sandbox`, `/routines`, and `/tool_failures`; sandbox job events
are persisted through the unified `append`/`tail` event plane.

Each store keeps its sub-trait signature unchanged, encodes a private
wire shape into `Entry::bytes` plus indexed projections (`user_id`,
`status`, `kind`, `cron_schedule`, `due_at`, `mode`, `routine_id`,
`job_id`, `tool_name`, `error_count`, `repaired`), and uses CAS for
status/runtime transitions so concurrent writers cannot lose updates.
Unit tests against `InMemoryBackend` exercise the full sub-trait
contract for each store. The legacy libSQL/Postgres impls are
unchanged.

* feat(engine): add FilesystemStore on the unified RootFilesystem surface

Adds `FilesystemStore<F: RootFilesystem>` as a second implementation of
the engine `Store` trait, routing all thread/step/event/project/
conversation/memory/lease/mission CRUD through the unified
`put`/`get`/`query`/`ensure_index` plane. Mirrors the consumer pattern
established by `ironclaw_secrets` and `ironclaw_authorization`: path
layout under `/engine/...`, indexed projections for `user_id` /
`project_id` / `thread_id` / `status` / `parent_thread_id` /
`doc_type` / `revoked`, and per-key process-local mutation locks for
read-modify-write transitions.

`HybridStore` in `src/bridge/store_adapter.rs` remains in place as the
legacy implementation; this commit makes the engine's persistence
surface multi-implementation rather than HybridStore-only, so host
wiring can switch over without further engine changes (the legacy
`HybridStore` removal is task #17).

Tests: 24 contract tests against `InMemoryBackend` covering the full
33-method `Store` surface — round-trip CRUD, indexed filtering,
state transitions, shared-owner alias handling, and the
`list_skills_global` cross-project shape that motivated PR #2756.
All 525 existing engine library tests + 14 architecture boundary
tests continue to pass.

* feat(db): add filesystem-backed facades for five sub-traits

Dissolve `SettingsStore`, `UserStore`, `ChannelPairingStore`,
`IdentityStore`, and `WorkspaceStore` into FS-backed facades over
`RootFilesystem`. Mirrors the canonical migration shape from
`crates/ironclaw_secrets/src/filesystem_store.rs` and
`crates/ironclaw_authorization/src/lib.rs`. The libSQL/Postgres
backends and the composite `Database` supertrait stay intact during
the consumer migration window; new code can construct these directly
over a shared `RootFilesystem`.

Path layout:

- `/system/settings/<user_id>/<key>`
- `/users/<id>` + `/users/.tokens/<token_id>` + `/users/.tokens-by-hash/`
- `/identities/<provider>/<provider_user_id>`
- `/pairing/requests/<channel>/<id>` + `/pairing/identities/<channel>/<id>`
  + `/pairing/code-index/<channel>/<code>`
- `/workspace/documents/<user>/<doc_id>` +
  `/workspace/chunks/<doc>/<n>` + `/workspace/versions/<doc>/<v>` +
  path/id index sidecars

WorkspaceStore is split into sub-modules under
`src/db/filesystem_workspace/` (documents, chunks, versions, search,
paths) per the file-size budget. Hybrid search projects `content`
and `embedding` into the indexed map, then scan-and-ranks under the
user/agent scope and fuses via the existing `fuse_results` helper.

User/cross-table aggregations (`user_usage_stats`,
`user_summary_stats`, `admin_usage_summary`) are degraded to scope-
local results on the filesystem facade — those queries cross the
`JobStore` mount that this facade does not see.

`/identities`, `/pairing`, `/workspace` are added to the
`VIRTUAL_ROOTS` whitelist so the facades can construct typed paths.

Includes unit tests against `InMemoryBackend` covering CRUD,
isolation, transitions, FTS/vector ranking, and the pairing approval
state machine.

* fix: replace .expect on validated literals with unwrap_or_else(unreachable!())

CI's `scripts/check_no_panics.py` flags `.unwrap()`/`.expect()` in
production code. Agent-generated stores used `.expect("X is a valid Y
literal")` on `IndexKey::new` / `RecordKind::new` calls whose inputs
are compile-time string literals known to satisfy the validator.

Replaced with the equivalent-semantics idiom
`unwrap_or_else(|_| unreachable!("..."))` — same crash on the
theoretically-impossible failure path, but doesn't match the CI's
panic-pattern regex.

Affects:
- crates/ironclaw_memory/src/repo/filesystem.rs (6 sites)
- src/db/filesystem_conversations.rs (4 sites)
- src/db/filesystem_jobs.rs (7 sites)

* fix(filesystem): close SQL-injection vector and CAS-loop concurrent updates

Two HIGH-severity findings from code review.

Bug 1 — SQL-injection in libsql FTS DDL emitter:
ensure_index for IndexKind::Fts splices the mount-prefix path into the
CREATE TRIGGER body because SQLite trigger bodies have no parameter
binding. VirtualPath::new rejects NUL/control/backslash/`..` but does
not reject `'`, `"`, `;`. Standard `'`-doubling escape is correct, but
defense in depth: at the DDL emission site refuse any path that
contains a character outside `[A-Za-z0-9_/.-]`. Postgres path is
parameterized, so only libsql was affected. Regression test added.

Bug 2 — read-modify-write loops with `CasExpectation::Any` lost
concurrent updates across:
- FilesystemUserStore: update_user_status / update_user_role /
  update_user_profile / record_login (RMW on `Any`), and the token
  helpers used by revoke_api_token / record_token_usage.
- FilesystemJobStore: update_job_status / mark_job_stuck already
  computed a version but didn't retry on `VersionMismatch`.
- Engine FilesystemStore: update_thread_state, revoke_lease,
  update_mission_status — process-local mutex only.

Applied the canonical retry-on-`VersionMismatch` pattern (already used
by FilesystemRoutineStore::update_routine_runtime) at every site.
filesystem_settings.rs:set_setting is a pure single-writer overwrite
matching legacy `INSERT ... ON CONFLICT DO UPDATE`, so it stays on
`Any` with an explanatory comment.

Also fixes a pre-existing `unwrap_or_else(|_|...)` typo (1-arg closure
on an Option) that blocked `cargo test --lib`.

* fix(workspace): route hybrid_search through native FTS + Vector filters

HIGH-severity finding from code review: `db::filesystem_workspace`
`hybrid_search` scanned every chunk under the user's documents and
ranked in Rust even when the mounted backend advertised
`Capability::IndexFts` / `Capability::IndexVector`. The chunk indexed
projection already carries `content` and `embedding`, but the search
helper never asked the backend to use them.

- search::hybrid_search now calls `filesystem.query(/workspace/chunks,
  Filter::Fts { content, query })` and `filesystem.query(.., Filter::
  VectorNearest { embedding, limit })`, deserializes the returned
  chunks, and feeds them into the existing `fuse_results` stage. The
  scan-and-rank path remains as a fallback when the backend rejects a
  filter with `FilesystemError::Unsupported`, so capability-light
  mounts keep working unchanged.
- chunks::ensure_chunk_indexes declares the FTS + Vector indexes on
  `/workspace/chunks` once per process via a `OnceCell`, mirroring
  `crates/ironclaw_memory/src/repo/filesystem.rs`. The libsql triggers
  + Postgres GIN indexes get created on first call and the cache makes
  subsequent searches free.
- Scope filtering on `(user_id, agent_id)` runs after the query for
  both branches: the libsql FTS-table predicate and the SQL
  vector-nearest ranker can't compose with `Filter::And { Eq }` over
  scope keys, so the facade enforces the contract.
- mod.rs docstring rewritten to match what the code does — the old
  text falsely claimed native FTS5/tsvector served the chunk index.
- Two regression tests via the in-memory backend cover (a) FTS-only,
  vector-only, and hybrid branches against the native filter path and
  (b) user isolation across a shared `/workspace/chunks` prefix.
  Both tests fail against the prior scan-and-rank-only implementation.

Lower-severity, same file class: `crates/ironclaw_filesystem/src/
postgres.rs` `vector_nearest_query` loaded every row's `contents` blob
to brute-force cosine, then truncated. Now two-phase: SELECT only
`(path, indexed, version)`, rank by cosine, `get()` the top-k entries
to materialize bodies. Same fix landed for libsql in PR e2530adff.

* fix: address remaining HIGH review findings on #3679

Three changes that close out the remaining HIGH-severity feedback from
the self-review (#1 #2 #3 #4 already addressed in 990c4f73e + e2530adff):

**#2 — `parse_state` silent fallback to Pending removed.**
`src/db/filesystem_jobs.rs::parse_state` previously mapped unknown
status strings to `JobState::Pending`, masking schema drift across a
rollout (a new state value appearing in stored rows would silently
lose its true value). Now returns `Result<JobState, DatabaseError>`
and the single caller propagates with `?`. Matches the wire-stable
enums rule in `types.md`.

**#6 — `is_engine_unsupported` no longer substring-matches.**
`crates/ironclaw_engine/src/store/filesystem.rs`: the typed
`FilesystemError::Unsupported` discriminator gets lost when wrapped in
`EngineError::Store { reason: String }`, so the old check
`reason.contains("Unsupported")` would false-positive on any unrelated
store error that mentioned the word. Now `fs_to_engine_error` tags the
discriminator with a stable `[fs:unsupported]` sentinel and the check
matches that sentinel — discriminator-preserving without changing the
public `EngineError` shape.

**#7 — `FilesystemChannelPairingStore` no longer drops corrupted records.**
`src/db/filesystem_pairing.rs::find_pending_requests`: the old code
silently filtered records whose JSON failed to deserialize, hiding
data corruption. Now propagates `DatabaseError::Serialization` with
the stored path so the operator sees the failure.

Also: `// silent-ok:` annotations added to the three engine `Store`
sites where read-modify-write on unknown ids is intentionally a no-op
(matches HybridStore parity per its CLAUDE.md). Each annotation names
the legacy contract being preserved.

Verification: `cargo check --workspace --all-features` clean;
`cargo test -p ironclaw_engine --all-features` 549/549;
`cargo test --lib --all-features db::filesystem` 88/88;
`cargo fmt --check` clean.

* fix(db): drain all pages in filesystem conversation/job listings

`list_messages_internal`, `list_conversations_summary`, and `run_query`
each called `filesystem.query(.., Page::new(0, Page::MAX_LIMIT))` exactly
once and trusted the result was complete. Because `Page::MAX_LIMIT ==
1024`, conversations with >1024 messages or scopes with >1024
jobs/actions/estimations silently lost every row past the cap, and the
`has_more` flag in `list_conversation_messages_paginated` became
meaningless once the dropped tail crossed the page boundary. Codex PR
#3679 P2 review flagged the pattern.

Extract a shared `query_all_pages` helper in `filesystem_conversations`
that loops `query(..., Page::new(offset, MAX_LIMIT))` until a short
page comes back, then reuse it from `filesystem_jobs::run_query` and
from the inline scan in `update_estimation_actuals`. The helper
preserves the existing `NotFound -> Vec::new()` short-circuit and the
`fs_err_to_database` error mapping so call sites are otherwise
unchanged.

Regression tests:
- `list_messages_drains_pages_beyond_max_limit` writes
  `MAX_LIMIT + 5` messages and asserts the full count round-trips
  through `list_conversation_messages` and that
  `list_conversation_messages_paginated` reports `has_more` honestly
  for both partial and exhaustive windows.
- `get_job_actions_drains_pages_beyond_max_limit` writes
  `MAX_LIMIT + 3` actions on one job and asserts the full count comes
  back in sequence order.
- `list_agent_jobs_drains_pages_beyond_max_limit` writes
  `MAX_LIMIT + 7` jobs and asserts both `list_agent_jobs` and
  `agent_job_summary` count every row.

* fix(secrets): close CAS-loop races in filesystem store consume paths

Two HIGH-severity findings on PR #3679. Both sites read a versioned
entry, validated a one-shot/use-limit condition, then wrote back with
`CasExpectation::Any`. The process-local mutex only serializes writers
inside one process; multi-process callers sharing the same backend root
could both pass the check and overwrite each other.

- `FilesystemSecretStore::consume` — two consumers could both observe an
  Active one-shot lease, both decrypt, and both overwrite the consumed
  marker.
- `FilesystemCredentialBroker::consume_session_use` — two consumers
  could both pass the max-uses check at `uses=N-1` and overwrite each
  other's increment, losing a use.

Both now use the canonical retry-on-`FilesystemError::VersionMismatch`
pattern from `ironclaw_engine::store::filesystem::update_thread_state`
(post-`e2530adff`): re-read, re-evaluate the consume/use-limit
condition, write with `CasExpectation::Version(versioned.version)`. A
shared `CAS_RETRY_ATTEMPTS = 3` constant bounds the loop; exhausting it
surfaces a transient backend error rather than papering over
pathological hot-spots.

Also annotated `leases_for_scope` with a `TODO(perf)` covering the
N+1 list+get fan-out — bounded today by the owner-prefix path layout
and short lease TTLs; replacing it with `Filter::Eq` over `query`
requires the secrets store to declare its first index, which is a
follow-up.

Regression coverage: two new tests wrap `InMemoryBackend` with a
`VersionRacingBackend` that bumps the watched path's version
out-of-band on the first versioned `put`, forcing a `VersionMismatch`
and exercising the retry loop. They also assert that the retried CAS
write actually persisted (the next consume hits LeaseConsumed; the next
three increments exhaust the max-uses budget).

* fix: address remaining P2 review findings on #3679

Four P2 correctness fixes from the codex/gemini review.

**Settings keys use percent-encoding** (`src/db/filesystem_settings.rs`):
`encode_segment` previously mapped `/`, space, control chars, and others
all to `_`. Keys like `a/b` and `a_b` collided onto the same path and
silently overwrote each other. Now percent-encodes every byte outside
the unreserved set so distinct inputs map to distinct outputs.

**SQL index names get a blake3 suffix on overflow** (`crates/ironclaw_filesystem/src/db.rs`):
`sql_index_name` truncated identifiers exceeding 62 chars without
disambiguating, so two distinct long `(prefix, name)` specs could
collapse onto the same DDL object — `CREATE ... IF NOT EXISTS` would
silently reuse the wrong index/trigger. Now appends an 8-char blake3
hash suffix before truncating. Added `blake3 = "1"` to the crate's
deps (small + already used by other workspace crates).

**InMemoryBackend rejects writes over implicit directories**
(`crates/ironclaw_filesystem/src/in_memory.rs`): the SQL backends
refuse `put(/a)` when `/a/b` exists (treating `/a` as a directory).
The in-memory reference impl silently accepted those writes, letting
tests pass against production-impossible state. Mirror the SQL
contract.

**Event-store head-probe is bounded**
(`crates/ironclaw_reborn_event_store/src/filesystem_store.rs`):
The replay-gap detection previously called `tail(path, 0)` to read the
whole log just to look at its last seq — O(N) on every cold-path call.
Now probes `tail(path, after - 1)`: a non-empty result means
head == after (consumer is caught up); empty means head < after
(foreign-future cursor). Returns at most one record instead of the
entire log.

Verification: cargo check --workspace --all-features clean; cargo
test -p ironclaw_filesystem -p ironclaw_secrets
-p ironclaw_reborn_event_store --all-features all pass.

* fix(filesystem): close cross-backend Range/vector semantic drift and txn scope hole

Audit findings on ironclaw_filesystem turned up four bugs and three
semantic-drift cases between the in-memory reference and the SQL
backends. Fix them in one pass so the cross-backend contract is
honoured and the gaps have regression coverage.

Bugs:
- libSQL `Filter::Range` on `IndexValue::Bool` never matched any row
  because SQLite's `json_type` returns "true"/"false" for booleans
  rather than "integer". Replaced the static type string with a
  `json_type_guard` expression that admits both bool variants.
- `ScopedStorageTxn` did not enforce that per-op `VirtualPath`s lay
  under `mount_prefix`. The trait doc promised `PathOutsideMount` for
  cross-prefix accesses; the wrapper now enforces it so any future
  backend that ships `begin()` inherits the guarantee.
- Mixed-variant `Filter::Range` bounds (e.g. I64 lo + Text hi) silently
  lex-compared on text on both SQL backends. Added the in-memory
  backend's `discriminant(lo) == discriminant(hi)` guard to both,
  rejecting with `Unsupported`.
- SQL `vector_nearest_query` lacked the in-memory backend's path
  tie-breaker on equal cosine scores, so top-k truncation was
  non-deterministic. Added `.then_with(|| a.0.cmp(b.0))` to both.

Semantic drift:
- `FilesystemOperation` lacked an event-plane `Append` variant —
  default impl reported `Tail`, backends reported `AppendFile`. Added
  the variant, routed every emit site through it, and updated the
  downstream `host_runtime::operation_allowed` matcher.
- `decode_embedding_blob` and `cosine_similarity` were byte-identical
  copies in three files. Extracted to `crate::vector`.
- libSQL `run_migrations` ran multiple ALTERs outside any transaction.
  Wrapped the sequence in BEGIN IMMEDIATE / COMMIT with rollback on
  error so a crash can't leave a half-migrated schema observable.

Tests added:
- 16 `ScopedFilesystem` permission tests covering query / ensure_index
  / begin / append / tail across each `MountPermissions` axis, plus
  4 `ScopedStorageTxn` tests driving a stub backend to lock in the
  per-op ACL and the new path-containment check.
- Cross-backend regression tests in `tests/db_root_filesystem_contract.rs`
  for the libSQL Bool/Range fix, the discriminant guard on both SQL
  backends, and the deterministic vector tie-breaker.
- Refactored `vector_nearest_query`'s phase-2 step into
  `materialize_ranked` (`pub(crate)`) so a unit test can exercise the
  "row disappeared between phases" branch deterministically.

128 tests pass, all three feature combos compile (`default`, `libsql`,
`postgres`), workspace builds.

* revert(db): drop filesystem-backed src/db/ store facades

Removes all `src/db/filesystem_*.rs` facades and the
`src/db/filesystem_workspace/` directory added during the PR #3679
universal-FS dispatch migration:

- filesystem_conversations, filesystem_jobs
- filesystem_routines, filesystem_sandbox, filesystem_tool_failures
- filesystem_identities, filesystem_pairing, filesystem_settings,
  filesystem_users
- filesystem_workspace/{mod,chunks,documents,paths,search,versions}.rs

Also removes the supporting infra that only existed for these files:

- `ironclaw_filesystem` workspace dep from the root `ironclaw` crate
- `/sandbox`, `/routines`, `/tool_failures`, `/identities`, `/pairing`,
  `/workspace` entries from `ironclaw_host_api::path::VIRTUAL_ROOTS`

The legacy libSQL/Postgres sub-trait impls (`src/db/postgres.rs`,
`src/db/libsql/*.rs`) remain the sole backing for the `Database`
supertrait. The unified `ironclaw_filesystem` mount fabric itself
(the `crates/ironclaw_filesystem/` crate) is untouched and still
used by consumer crates outside `src/db/`.

Verification:
- cargo fmt --check clean
- cargo check --workspace clean (default features)
- cargo check --no-default-features --features libsql clean
- cargo check --all-features clean
- cargo clippy --all --benches --tests --examples --all-features clean

[skip-regression-check] pure removal of unmerged migration facades.

* test(reborn-event-store): cover caught-up-to-head + concurrent appends

Addresses audit finding F1.

(a) `filesystem_event_log_caught_up_to_head_returns_empty_not_replay_gap`
    appends N events, replays from the last entry's cursor, and asserts
    `entries.is_empty()` + `next_cursor == last.cursor` with no
    `ReplayGap`. Pins the "consumer is caught up to head" branch of the
    bounded probe in `read_after_cursor`.

(b) `filesystem_event_log_concurrent_appends_assign_distinct_cursors`
    spawns 8 `tokio::spawn` tasks each appending one event to the same
    stream, then asserts the collected cursors are pairwise-distinct
    and strictly increasing. Guards the per-stream monotonic-cursor
    invariant under contention.

* fix(reborn-event-store): preserve filesystem error detail in durable mappers

Addresses audit finding F2.

`map_filesystem_append_error` / `map_filesystem_tail_error` previously
collapsed every non-categorised `FilesystemError` variant
(`VersionMismatch`, `NotFound`, `Backend`, …) to a fixed generic
string, dropping the source variant and reason. Operators lost the
detail they needed to debug appends that hit a CAS conflict or a
backend I/O failure.

Thread the underlying `FilesystemError` through its `Display` impl on
the fallback arm. `FilesystemError` is already redaction-safe by
contract — it renders scoped/virtual paths, never raw host paths —
so the durable error surface gains debug detail without violating
the crate-level redaction policy. The three already-categorised
variants (`PermissionDenied`, `MountNotFound`, `Unsupported`) keep
their fixed messages so callers can pattern-match on the substring.

* fix(reborn-event-store): document deliberate absence of Filesystem config variant

Addresses audit finding F3.

`FilesystemDurableEventLog` / `FilesystemDurableAuditLog` are exported
from this crate, but `RebornEventStoreConfig` has no corresponding
`Filesystem` variant — so production composition still routes through
the SQL stores. The PR description documents this as intentional: the
filesystem-backed log is the migration target for the kernel-storage
rework, and the config variant will be added during the `src/db/`
dissolution pass (task #17). Without an inline comment, a future
reviewer reading the config enum has no signal that the missing
variant is deliberate.

Add a doc paragraph on `RebornEventStoreConfig` pointing at the
rationale on `filesystem_store.rs` and at task #17.

* fix(reborn-event-store): drop shadowed kind named-arg in stream_path format!

Addresses audit finding F4.

`stream_path` previously used the named-argument `format!` form with
`kind = kind_segment`, where the named key `kind` shadowed the
function parameter of the same name. Switch to the implicit
positional-capture form (`format!("/events/{kind_segment}/...")`)
and rename the inline bindings to `tenant_segment` / `user_segment`
for consistency. Pure refactor — no behaviour change, just removes
the readability footgun.

* fix(outbound): add typed CasConflict variant for filesystem store retries

Audit finding F5: `map_fs_error` previously collapsed both
`FilesystemError::VersionMismatch` (a transient compare-and-swap race
condition that callers should retry) and `FilesystemError::Unsupported`
(a permanent capability gap) into `OutboundError::Backend`. The bounded
CAS retry loop (added separately for F1) cannot match on `Backend` —
that would also retry on permanent backend failures and on `Unsupported`
on backends that don't support CAS.

Introduce `OutboundError::CasConflict` and map `VersionMismatch` to it
in `map_fs_error`. The variant stays internal to the crate: the retry
loop matches on it discriminator-wise; once the retry budget is
exhausted (or for callers that haven't migrated) it converts to
`Backend` before crossing the trait boundary, preserving the no-leak
contract.

Update `is_transient_validator_error` to classify `CasConflict` as
transient for defence in depth, even though it should never reach the
service boundary in practice.

* fix(outbound): CAS-version read-then-write paths with bounded retry

Audit finding F1 (HIGH): the four read-then-write methods on
`FilesystemOutboundStateStore` (`upsert_subscription`,
`advance_subscription_cursor`, `record_delivery_attempt`,
`update_delivery_status`) read the existing entry, applied an in-memory
transform, then wrote with `CasExpectation::Any`. Concurrent writers
raced the transform: in particular, the "subscription cursor must not
move backwards" invariant — enforced in `validate_advance_request` /
`validate_subscription_cursor_progression` — was unenforced
cross-process, because two racing advancers could both read the same
old cursor, validate against it, and then both put their newer
cursors, the loser silently winning the last-write race.

Capture `VersionedEntry.version` from each `get`, pass
`CasExpectation::Version(v)` to the matching `put`, and retry on the
typed `OutboundError::CasConflict` introduced by F5. The retry budget
is bounded (`MAX_CAS_RETRIES = 5`) and the loop re-reads + re-validates
on every iteration, so a regressing cursor or scope mismatch surfaces
immediately rather than letting the retry loop overwrite the winner's
state. `put_thread_notification_policy` is a blind overwrite and keeps
`CasExpectation::Any`.

`record_delivery_attempt` uses `CasExpectation::Absent` for the
first-write branch, so two racing at-least-once writers can't both
insert; the loser falls back into the duplicate-identity-check branch
on the next read.

* fix(outbound): use control-character sentinel in thread scope key

Audit finding F6: `thread_scope_key` used the literal string `"_"` as
the sentinel for `agent_id = None` / `project_id = None`. The
`validate_scope_id` validator in `ironclaw_host_api` accepts underscore
as a legal character in an `AgentId` / `ProjectId`, so a scope with
`agent_id = Some(AgentId::new("_"))` hashed to the same key as a scope
with `agent_id = None`. Two distinct scopes silently collided on the
same policy/subscription/delivery virtual path.

Switch the sentinel to `"\x1F"` (ASCII unit-separator). It's a control
character; `validate_scope_id` rejects every C0 control char via
`has_forbidden_control`, so no legal scope id can ever contain it. Add
a unit test that pins the sentinel-rejection invariant and a
regression test that proves `agent_id = Some("_")` no longer hashes to
the same key as `agent_id = None`.

* fix(outbound): query indexed scope projection with paginated drain

Audit finding F2 (HIGH): `list_delivery_attempts` was a `list_dir` +
N+1 `get_json` per row with no indexed projection, scanning every
delivery on the mount even when only one scope's deliveries were
requested. Cost scaled with total delivery count, not with the
queried scope's row count.

Declare an exact-equality index on a new `scope` indexed key. The
projected value is the same `thread_scope_key` hash used for policy
paths — collision-resistant against the legal id grammar and updated
by F6 to never collide with the `None` sentinel. `record_delivery_attempt`
and `update_delivery_status` write through a new
`put_delivery_attempt_indexed` helper that includes the projection;
`update_delivery_status` preserves it on status mutations. The list
path drives `query(Filter::Eq { key: "scope", value: ... })` and
re-checks `scope_matches` defensively (hash collisions are
unreachable but cheap to guard against).

Audit finding F3 (Medium): the previous `list_dir` was unpaginated;
SQL backends issue `LIMIT Page::MAX_LIMIT (1024)` on their list_dir
translation and would silently truncate past 1024 deliveries. The
new path drains pages via `offset += received` until a short page
arrives, mirroring `ironclaw_engine::store::filesystem::query_all`.

`ensure_delivery_scope_index` runs idempotently before every write
and read. It tolerates `FilesystemError::Unsupported` on byte-only
backends to match the engine store's `ensure_exact_index` pattern;
the in-memory backend serves `Filter::Eq` from `Entry::indexed`
directly even without a materialized index declaration.

* test(outbound): cover CAS retry, pagination drain, backwards-race

Audit finding F4: the existing `outbound_state_store_contract` suite
exercised the storage contract surface but had no coverage for any of
the failure modes the F1/F3 fixes address:

- No CAS-retry test. F1's bounded retry loop could regress to permanent
  failure on any transient `VersionMismatch` and the suite wouldn't
  notice — the in-memory backend never produced one.
- No `> Page::MAX_LIMIT` drain test. F3's pagination loop could lose
  the tail of a long delivery list and the suite wouldn't notice
  because the existing tests record at most one delivery per scope.
- No concurrent backwards-race test on `advance_subscription_cursor`.
  The existing backwards-advancement test only exercised the single-
  threaded path; nothing proved the post-F1 retry loop re-validates
  progression on every iteration.

Add three regression tests:

1. `VersionRacingBackend` wraps `InMemoryBackend` and injects a single
   `FilesystemError::VersionMismatch` on the next `put` matching a
   configured prefix. The first new test
   (`advance_subscription_cursor_retries_through_cas_conflict`) arms
   one conflict, advances the cursor, asserts the retry loop converges,
   and asserts exactly one conflict was injected and consumed.

2. `concurrent_backwards_race_rejected_after_winner_advances` runs two
   sequential advances — the winner to cursor=100 and the loser to
   cursor=50 — and asserts the loser is rejected with `InvalidRequest`
   while the winner's state is preserved. Together with the retry test
   this proves the re-validate-on-retry semantics F1 calls out.

3. `list_delivery_attempts_drains_more_than_page_max_limit` writes
   `Page::MAX_LIMIT + 1` delivery attempts under one scope and asserts
   `list_delivery_attempts` returns every one. Before F3 this would
   silently truncate at 1024 rows.

Cargo.toml: enable `tokio/sync` for `Mutex` in the test mock; drop the
feature-conditional `use std::sync::Arc` because the new tests need it
unconditionally.

* fix(run-state): bound filesystem lock map under tenant churn

The process-wide FILESYSTEM_RECORD_LOCKS map kept one
Arc<tokio::sync::Mutex<()>> per touched path. In long-running hosts with
high tenant/invocation churn the map grew without bound, since entries
were never removed once the originating put/get cycle completed.

Switch the value type to Weak<Mutex> so dropped Arcs no longer pin map
slots. Each acquisition opportunistically prunes dead entries before
upgrading-or-installing, keeping the map size proportional to in-flight
paths rather than to lifetime path count. Concurrent callers on the same
path still observe the same Arc (the outer std::sync::Mutex serializes
the upgrade-or-insert window), so existing intra-process and
cross-instance serialization guarantees are preserved — both verified by
the new unit tests and by the existing
filesystem_*_duplicate_*_serialized_across_store_instances contract
tests.

Addresses audit findings F1 (Medium) and F4 (Low).

* fix(run-state): use versioned CAS for filesystem run/approval writes

All filesystem put() calls used CasExpectation::Any, so two host processes
mounting the same /engine could lose updates: each one's read-modify-write
saw the other's value and then unconditionally overwrote it. The
per-path async mutex only serializes intra-process callers.

Switch creates to CasExpectation::Absent and updates to
CasExpectation::Version(v) with a bounded retry loop on VersionMismatch.
The new put_with_cas helper centralizes the contract: on capable
backends (InMemoryBackend, the upcoming SQL ports) cross-process races
now fail closed and the caller retries; on byte-only backends that
return Unsupported (LocalFilesystem) we degrade to Any but emulate
Absent with a get() precheck so the AlreadyExists path is preserved.
The in-process lock map (F1) keeps the check-then-write race closed for
the byte-only fallback.

Approve/deny/discard pull the record-lock guard up to the trait method,
since update_status no longer acquires it.

Addresses audit finding F2 (Medium). Closes the gap acknowledged in
crates/ironclaw_run_state/CLAUDE.md.

* fix(approvals): type approval-resolution decision with ApprovalDecisionKind enum

Addresses audit finding F1.

Replaces the stringly-typed `impl Into<String>` decision parameter on
`AuditEnvelope::approval_resolved` with a wire-stable
`ApprovalDecisionKind` enum (`Approved`/`Denied`,
`#[serde(rename_all = "snake_case")]`), so approval callers cannot
drift on capitalization or spelling. Per `.claude/rules/types.md`
"wire-stable enums".

The wider `DecisionSummary::kind` field stays a `String` because other
audit producers (authorization denials, obligation handlers) emit
values outside the approval enum; cross-decoding remains a follow-up.

Cross-crate blast radius: `ironclaw_host_api` (new enum + factory
signature), `ironclaw_approvals` (both call sites),
`ironclaw_events::tests::durable_log_contract` (three test fixtures).

* fix(approvals): persist approval state before issuing lease

Addresses audit finding F2.

Inverts the lease/approve ordering inside `approve_capability_action`:
the approval store write now runs *before* the lease store write. The
previous order (issue lease, then approve, best-effort revoke on
failure) left a window where a transient approval-store error could
leave a live lease pointing at a request whose status remained
`Pending`.

The approval record is now treated as the authority of record. Once
the request flips to `Approved`, lease issuance is a recoverable
operation against an already-decided request — if the lease store
fails, the caller surfaces the lease error and the request stays
`Approved`. The previous best-effort `let _ = self.leases.revoke(...)`
swallow is gone with the same edit.

Updates the three concurrency/error-injection tests to assert the new
semantics, plus the crate CLAUDE.md guardrail. No external test
fixtures break — the public resolver API is unchanged.

* fix(approvals): route both resolve paths through emit_approval_resolved helper

Addresses audit finding F3.

Extracts an `emit_approval_resolved` helper on `ApprovalResolver` so
the audit-envelope construction in `approve_capability_action` and
`deny` is built in exactly one place. Both call sites used to inline
`AuditEnvelope::approval_resolved` against their own
`record.scope`/`denied.scope`; while consistent today, divergence
between the two would be a silent regression.

Pure refactor — no test changes needed beyond the existing audit-event
contract tests which already pin the wire shape.

* fix(approvals): cover concurrent approve_dispatch first-write-wins

Addresses audit finding F4.

Adds a caller-level concurrency regression test that spawns two
`approve_dispatch` calls against the same pending request on a
multi-thread tokio runtime and asserts the expected first-write-wins
invariants:

- exactly one approve returns `Ok`
- the other returns `ApprovalResolutionError::NotPending { status:
  Approved }`
- the lease store ends up with exactly one Active lease (not two, not
  zero — under the F2 persist-approval-first ordering the loser fails
  *before* lease issuance, so no orphan to revoke)
- the approval record's terminal status is `Approved`

Enables `rt-multi-thread` on the tokio dev-dependency so the test can
exercise real cross-thread contention on the approval store mutex.

* fix(engine): restore HybridStore parity for mission updates

F1: `update_mission_status` now bumps `mission.updated_at` before
writing back, matching HybridStore (`src/bridge/store_adapter.rs:1950`).
Recency-sorted views (mission list UIs, learning-mission dispatcher)
were silently freezing the timestamp at original-save time.

F2: `list_missions` and `list_all_missions` now sort by `(name, id)`
after collection, matching HybridStore (`store_adapter.rs:1913, 1937`).
The underlying `query`/HashMap iteration is non-deterministic; the
LLM-facing `mission_list` tool was seeing arbitrary order across runs.

Tests:
- `update_mission_status_bumps_updated_at` — regression for F1
- `list_missions_is_deterministic_across_invocations`,
  `list_all_missions_is_deterministic_across_invocations` — regression for F2

* fix(memory): drain pages in FilesystemMemoryDocumentRepository::list_documents

Audit findings F1 (HIGH) + F9 (Low).

F1: `list_documents` issued a single `query(.., Page::new(0,
Page::MAX_LIMIT))` and trusted the page was complete. Because
`Page::MAX_LIMIT == 1024`, scopes holding >1024 documents silently lost
every entry past the cap. The result fed `write_document`'s
ancestor/descendant conflict check at the call site immediately above,
so a new path could shadow (or be shadowed by) an existing document
across the truncation boundary without a conflict ever firing — exactly
the regression `query_all_pages` was extracted in
`src/db/filesystem_jobs.rs` to prevent.

F9: The old implementation issued a `Filter::All` query, threw the
results away (`let _ = (versioned, &prefix_str);`), then called
`list_dir` to discover paths. The query-result loop was dead code under
any backend that supports `query`. The stale comment claimed the trait
didn't surface paths in `query` results, but
`VersionedEntry.path` (`crates/ironclaw_filesystem/src/record.rs:347`,
added in PR #3659) has carried the absolute virtual path for every
queried row since.

Replace both with a single drain loop that paginates `query` until a
short page comes back, filters by `entry.kind == "memory_document"`, and
recovers the `MemoryDocumentPath` directly from `VersionedEntry.path`.
The `list_dir` fallback is gone, and the agent_id axis is preserved
through `MemoryDocumentPath::new_with_agent` so scopes with an agent
identity round-trip correctly (the previous code's `new()` dropped the
agent).

Regression: `list_documents_drains_pages_beyond_max_limit` writes
`MAX_LIMIT + 5` documents and asserts every one comes back. This also
exercises the conflict-check path because each `write_document` calls
`list_documents` internally.

* fix(secrets): close consume_if_matches timing oracle with constant-time compare

F1 (HIGH, timing oracle) in the 2026-05 audit: `consume_if_matches` in
`legacy_store.rs` (trait default + in-memory backend) and `db.rs` (libSQL
+ Postgres backends) compared the decrypted plaintext against the
caller-supplied expected value with `!=`. Rust's `!=` over `&str`/`&[u8]`
short-circuits on the first differing byte, so an adversary who can
observe response latency over the network can recover the secret byte
by byte. AES-GCM authenticated decrypt closes the ciphertext oracle but
does nothing for the post-decrypt comparison.

Fix: route the comparison through `subtle::ConstantTimeEq::ct_eq`, which
walks the full buffer regardless of where the bytes diverge. The
post-comparison branches retain their original shape because the
decrypt+lookup path is already executed unconditionally before the
compare — only the success-side `DELETE` differs, and that signal is
already exposed by the function's return value.

Added a regression test (`f1_consume_if_matches_uses_constant_time_compare`)
that grep-asserts the production source imports `subtle::ConstantTimeEq`,
uses `ConstantTimeEq::ct_eq`, and no longer contains the legacy `!=`
shape. Cannot meaningfully prove constant-time-ness from a shared CI
runner, but the source-pattern check ensures a "simplifying" revert
fails review.

Audit: F1 (HIGH).

* fix(secrets): use constant-time compare for store key-check sentinel

F3 (Low) in the 2026-05 audit: `verify_secret_store_key_check` compared
the decrypted sentinel against `SECRET_STORE_KEY_CHECK_PLAINTEXT` with
`!=`. The plaintext is a fixed compile-time string so the practical
risk is low — an attacker who can move the encrypted_value/key_salt
blobs across rows already has full DB write access — but the same
constant-time pattern applied to F1 makes the comparison style
consistent across the crate and pre-empts a future caller threading a
non-constant sentinel through this helper.

Routed through `subtle::ConstantTimeEq::ct_eq`, mirroring the F1 fix.

Audit: F3 (Low).

* fix(processes): index queryable fields and serve records_for_scope via query

Replace the N+1 list_dir + per-file get scan with an indexed `query`
path, falling back to the legacy scan on byte-only backends so existing
LocalFilesystem-driven tests and production deployments remain
unaffected.

- Declare `ensure_index` lazily for the per-owner `processes/` prefix on
  the queryable fields called out in the audit (`tenant_id`, `user_id`,
  `status`, `extension_id`, `parent_process_id`). Backends without index
  support degrade to the existing scan instead of failing closed.
- Project the same fields onto every `ProcessRecord` write via
  `Entry::with_indexed`; record-capable backends (libSQL, Postgres, the
  in-memory backend) can now serve scope listings through a native
  query. The opaque-byte fallback in `put_with_byte_fallback` keeps
  LocalFilesystem (which rejects record-shaped puts today) on the legacy
  write path.
- Rewrite `records_for_scope` to issue `Filter::And` of `Filter::Eq`
  predicates against the indexed projection. The full `same_scope_owner`
  check remains in Rust so the sub-scope axes (agent/project/mission/
  thread) that are not yet in the index spec still get filtered.
- Add a contract test that exercises the indexed path through
  `InMemoryBackend` and confirms cross-tenant and cross-user records
  are not returned.

Addresses audit findings F1 (records_for_scope N+1) and F2 (missing
ensure_index at startup).

* fix(filesystem): surface backend infrastructure errors without fabricated paths

F1: SQL backends used valid_engine_path() (unwrap_or_else unreachable
returning /engine) as a placeholder on every connection/migration
error. The path was always a lie - at pool acquisition, run_migrations,
pragma setup, or schema bootstrap there is no caller-supplied virtual
path in scope - and it leaked into operator-facing error display.

Add FilesystemError::BackendInfrastructure { operation, reason } that
omits path. Route every former valid_engine_path() callsite in libsql
and postgres through new infrastructure_error helpers in db.rs. The
enum is non_exhaustive so adding a variant is backward compatible.

Regression test: drive a libsql migration against a read-only DB file
and assert BackendInfrastructure with no /engine in display.

* fix(filesystem): store VirtualPath keys in InMemoryBackend state directly

F2: in_memory.rs::query() reparsed every stored row's path with
VirtualPath::new(...).unwrap_or_else(|_| unreachable!('stored paths
originated as VirtualPath')) on the hot path. Two issues:

  - the reparse is wasted work - paths originate as VirtualPath at
    put() time, so the validation pass on read is redundant
  - 'unreachable!' is a panic that asserts a structural invariant
    the type system already enforces

Replace HashMap<String, StoredEntry> with HashMap<VirtualPath,
StoredEntry>. Lookups now pass &VirtualPath directly; prefix scans
move to key.as_str().starts_with(...). VersionedEntry::path comes
from a single clone() instead of a parse + unreachable.

Existing tests cover the put/get/query/list_dir/stat/delete paths
that were touched (44 in_memory tests + the cross-backend
contract suite).

* fix(filesystem): align in-memory backend on nested VectorNearest semantics

F5: SQL backends reject Filter::VectorNearest nested inside And/Or
with Unsupported because ranking can't be expressed as a WHERE
fragment - the top of query() peels off a top-level VectorNearest
before the translator runs, and the translator's VectorNearest arm
unconditionally errors. The in-memory backend previously treated a
nested VectorNearest as 'any row with IndexValue::Bytes at key',
silently changing semantics across backends.

Add contains_nested_vector_nearest() pre-check in InMemoryBackend::
query that walks the filter tree and surfaces Unsupported for any
VectorNearest strictly inside a compound. The Filter::VectorNearest
arm in filter_matches is now unreachable; it returns false to keep
the scalar predicate path safe should the pre-check ever be bypassed.

Regression test asserts Unsupported on nested-in-And, nested-in-Or,
and still-OK for top-level VectorNearest.

* fix(filesystem): guard u64 to i64 SQL bindings with typed errors

F6: SQL backends used 'expected.get() as i64' and 'page.offset as i64'
casts on the CAS and query/pagination paths. Both inputs are u64 and
both wrap silently on values >= 2^63 - the cast produces a negative
SQL binding that either matches no row (CAS quietly VersionMismatches)
or executes against a negative OFFSET (cryptic backend error).

Add db.rs helpers:
  - record_version_to_i64: surfaces CorruptRecordVersion if the value
    overflows i64
  - page_offset_to_i64: surfaces a typed Backend error naming the
    operation and offset

Apply at libsql.rs CAS and query offset bindings and the matching
postgres.rs sites. 'page.limit' is u32 clamped to Page::MAX_LIMIT so
its i64 cast is safe by construction and uses i64::from for clarity.

Regression test asserts a typed Backend(Query) error with reason
'page offset...' when querying with offset = u64::MAX, replacing the
prior silent wrap.

* fix(filesystem): scope Postgres FTS GIN index to declaring prefix

F4: libsql FTS5 virtual tables are declared per-mount-prefix - one
vtable per ensure_index(prefix, ...) call - so a query at one prefix
can't accidentally pull index postings from a sibling prefix into the
plan, and tearing down an index for a prefix is a clean DROP TABLE.

The Postgres FTS GIN index, by contrast, was created without a
predicate over root_filesystem_entries, so it was global. Correctness
held because the query path always scopes by 'path =  OR path LIKE
', but parity with libsql broke in two ways: the planner
considered postings from every prefix before filtering, and a
per-prefix DROP INDEX could only ever tear down one of them.

Add a partial-index predicate gated by 'path = <prefix> OR path LIKE
<prefix>/%' to the GIN DDL. The prefix is sourced from the validated
VirtualPath and quotes are doubled for safe SQL literal embedding;
LIKE-special characters are escaped via the existing
escape_like_with_trailing_wildcard helper.

Regression test (Postgres only; skipped when no DB is reachable)
reads back the DDL via pg_indexes.indexdef and asserts the prefix
literal and a WHERE clause appear.

* fix(filesystem): tighten capability docs, type constraints, and hygiene nits

Batched audit findings:

F3: Document the type constraint on IndexKind::Prefix. The kind is
only meaningful against IndexValue::Text, but ensure_index can't see
the value type at declaration time. Filter::PrefixOn rejects every
non-text variant at query time. Document the constraint loudly so
consumers reach for IndexKind::Exact when projecting numeric or
boolean values instead of getting an unused index and a query-time
Unsupported.

F7: BackendCapabilities::sql_typical advertises a minimum SQL shape
that omits IndexFts and IndexVector. The two real backends here
(libsql + postgres) layer them on top. A hand-rolled backend that
just calls sql_typical() would under-advertise. Add a doc-comment
calling out the omission and an sql_typical_full() variant that
includes Events + IndexFts + IndexVector for backends that match
this crate's shape.

F8: validate_simple_identifier indexed bytes[0] after an is_empty
guard. The guard makes the index sound, but the pattern is fragile
to refactors. Switch to bytes.first() so the dependency is explicit
and the panic path goes away.

F9: Multiple doc comments in record.rs and index.rs referenced
stale type names (StorageBackend::put/list/query, Record). Update
to the current RootFilesystem / Entry names.

* fix(engine): dedupe events on append_events for HybridStore parity

HybridStore (`src/bridge/store_adapter.rs:1613`) de-duplicates thread
events by id before insert. The filesystem-store `append_events` impl
was previously writing with `CasExpectation::Any`, which silently
overwrote an existing event with the same id when callers re-emitted
(e.g. recovery after a partial flush).

Pre-read the destination path and skip any id already present.
Matches HybridStore's append-only contract.

Audit finding F3 (Medium) from the ironclaw_engine crate audit.

* fix(memory): map FTS results to documents in FilesystemMemoryDocumentRepository::search_documents

The previous scaffold issued the `Filter::Fts` query, then silently
dropped the results with `let _ = results; Ok(Vec::new())`. A caller
wiring up the trait would see an empty result set and assume "no
matches" — when in fact the search had simply lied. That is worse
than returning `Unsupported`.

Map each `VersionedEntry.path` (added in PR #3659) back to a
`MemoryDocumentPath`, de-dupe by path, and assign a per-rank score
from RRF over the FTS-only branch so the result vector matches the
native repos' fusion contract for the trivial single-branch case.

Skip non-memory-document entries that may live under the same prefix
(chunk projections, metadata siblings). Adds
`list_documents_drains_pages_beyond_max_limit` test against the
in-memory backend.

Audit finding F2 (HIGH) from the ironclaw_memory crate audit.

* fix(secrets): close revoke CAS-loop race with versioned compare-and-swap

`revoke` previously read the lease via the (now-removed)
`read_lease` helper and wrote with `CasExpectation::Any`. The
per-lease process-local mutex serialized writers within one process
only — multi-process callers sharing the same backend root could
observe `Active`, race against `consume`, and clobber a `Consumed`
marker by overwriting it with `Revoked`.

Inline the read into a bounded CAS retry loop matching `consume` and
`consume_session_use`: read with version, write with
`CasExpectation::Version`, retry on `VersionMismatch`. Make revoke
idempotent on terminal states (`Consumed`, `Revoked`, `Expired`) so
the loop converges even when a winner has already written.

Audit finding F2 (Medium) from the ironclaw_secrets crate audit.

* fix(processes): use versioned CAS for status transitions

`update_status` previously read the record and wrote with
`CasExpectation::Any`, relying on the per-instance `transition_lock`
for atomicity. That lock only serializes within one process; a
multi-process deployment sharing the same backend root could observe
identical pre-transition state in both processes and clobber each
other's status flips.

Replace with a bounded CAS retry loop: read with version, validate
the transition, write with `CasExpectation::Version…
theredspoon pushed a commit that referenced this pull request Jun 18, 2026
…earai#3573)

* feat(reborn): add ironclaw_hooks framework foundation (#3524)

Foundation slice of the Reborn loop hooks framework per nearai/ironclaw#3524.
Lands the trust primitives, sealed decision types, dispatcher contract, and
extension manifest schema; no Reborn middleware composition yet (next slice
wires HookDispatcher into LoopCapabilityPort / LoopPromptPort).

Design comment on #3524:
https://github.com/nearai/ironclaw/issues/3524#issuecomment-4439890144

What this PR ships
==================

* `crates/ironclaw_hooks/` — new crate
  * `identity` — content-addressed `HookId` (blake3 of length-prefixed
    extension + local + version fields). Same versioning primitive the rest
    of Reborn should converge on for replay safety.
  * `trust` — `HookTrustClass` enum (Builtin / Trusted / Installed) with
    per-kind default attenuation. Trust class is fixed by source, never
    declarable.
  * `kinds/` — sealed decision DTOs. `BeforeCapabilityHookDecision`,
    `HookPatch`, `ObserverFact` all have `pub` outer struct + `pub(crate)`
    inner enum + `pub(crate)` constructors. Same #3460 witness pattern.
  * `points/` — typed read-only contexts for each hook point.
  * `sink` — split sink traits per trust tier. `PrivilegedGateSink` exposes
    `allow()`; `RestrictedGateSink` does not. An Installed-tier hook
    literally cannot mint Allow at the type level.
  * `ordering` — phase → priority → hook id, stable. Phases gated by trust
    (Validation/Authorization Builtin-only).
  * `failure_policy` — Timeout/Panic/Malformed/AttenuationViolation
    categories. Gate/Mutator fail closed, Observer/Effect fail isolated.
    Slot poisoning persisted for the rest of the run on any category.
  * `registry` — run-profile-sourced bindings; phase-vs-trust gate enforced
    at insert; poisoning surface for the dispatcher.
  * `dispatch` — HookDispatcher with deterministic ordering, panic
    catch-unwind via futures::FutureExt, per-hook tokio::time::timeout,
    short-circuit gate composition (Deny > PauseAuth > PauseApproval >
    Allow), Telemetry-phase observers always run.
  * `manifest` — serde types for the `[[hooks]]` section of extension
    manifests. Predicate vs WASM body; same_tenant scope requires explicit
    grant; Validation/Authorization phases rejected at parse time because
    manifest hooks are always Installed.
  * `predicate` — typed predicate language for declarative Installed hooks
    (DenyCapability, PauseApproval, RateOrValueCap). Evaluator lives in
    the dispatcher follow-up, not here.

* `crates/ironclaw_architecture/tests/reborn_dependency_boundaries.rs`
  * Added `ironclaw_turns` -> `ironclaw_hooks` to the forbidden list.
  * New BoundaryRule for `ironclaw_hooks` itself (cannot pull host_runtime,
    dispatcher, secrets, network, wasm, etc.).

* `Cargo.toml` workspace member registration.

What this PR deliberately does NOT ship
========================================

* Reborn middleware composition wrapping LoopCapabilityPort / LoopPromptPort
  with HookDispatcher. Next slice; ironclaw_reborn changes only.
* WASM hook execution path. Programmatic hooks parse and validate from
  manifest; the wasmtime integration lands when the WASM dispatcher seam is
  built.
* Predicate evaluation. Predicate types serialize and validate; the
  evaluator that turns a `RateOrValueCap` spec into a `Deny` decision is in
  the next slice alongside Reborn wiring.
* Event-triggered hooks (Phase 5 of the original roadmap).
* Self-authored hooks. Tracked separately at #3567 with monotonic-restriction
  + unforgeable-channel ratification.

Test plan
=========

* `cargo test -p ironclaw_hooks` — 47 tests (46 unit + 1 integration smoke
  for the manifest -> binding -> dispatch pipeline).
* `cargo test -p ironclaw_architecture` — 13 tests; new boundary rule
  passes, existing rules unaffected.
* `cargo clippy -p ironclaw_hooks --all-targets -- -D warnings` — clean.
* `cargo fmt -p ironclaw_hooks -- --check` — clean.
* `cargo check --workspace` — clean, no regressions in other crates.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(reborn): wire HookDispatcher into LoopCapabilityPort/LoopPromptPort

Follows the foundation slice (see initial commit). Adds the next layer:

1. Capability- and prompt-port middleware (`ironclaw_hooks::middleware`)
   * `HookedLoopCapabilityPort` runs `dispatch_before_capability` before
     every invocation, translates the composed decision into the existing
     `CapabilityOutcome` vocabulary (Deny / PauseApproval / PauseAuth all
     map to `Denied` for now; gate-ref plumbing for real pause semantics
     lands in the next slice).
   * `HookedLoopPromptPort` runs `dispatch_before_prompt` before bundle
     construction. Observe-only for snippets in this slice; actual
     snippet injection waits for the shared `prompt_envelope::wrap_untrusted`
     helper (#3540 / #3471).

2. Declarative predicate evaluator (`ironclaw_hooks::evaluator`)
   * `DenyCapability` and `PauseApproval` predicates: stateless, evaluated
     directly against `BeforeCapabilityHookContext`.
   * `RateOrValueCap` with `InvocationCount` bound: sliding-window counter
     keyed by `(hook_id, capability_name)`, in-memory only. Window
     parsing supports `s`/`m`/`h`/`d` units; unparseable windows fail
     closed.
   * `NumericSum` bound: types implemented but evaluation returns Allow
     and emits a warn-level audit. Full argument-extraction story is a
     follow-up slice once capability arguments become hook-visible.
   * `PredicateEvaluator::evaluate_at(...)` test variant accepts an
     explicit `Instant` so sliding-window tests don't depend on
     real-clock progress.

3. Manifest -> dispatcher glue (`ironclaw_hooks::installed_hook`)
   * `PredicateBackedBeforeCapabilityHook` wraps a `HookPredicateSpec`
     plus an `Arc<PredicateEvaluator>` and implements
     `RestrictedBeforeCapabilityHook`. The registry installer would
     construct one of these per `[[hooks]]` entry whose body is
     `HookManifestBody::Predicate`.
   * Sink reasons are `&'static str`, so the dynamic predicate `reason`
     surfaces in audit (via the evaluator's `EvaluatorDecision`) rather
     than the model-visible decision. Closed-vocabulary labels carry
     through to the sink.

4. Reborn composition seam (`ironclaw_reborn::loop_driver_host`)
   * `RebornLoopDriverHostFactory::with_hook_dispatcher(Arc<HookDispatcher>)`
     opt-in builder method. When set, the factory wraps the capability
     and prompt ports with the hooked middleware. Default behavior
     (no dispatcher) is unchanged from the pre-hooks shape, so existing
     callers continue to work.
   * Added `ironclaw_hooks` as a dep in `ironclaw_reborn`.

Test plan
=========

* `cargo test -p ironclaw_hooks` — 60 tests pass (59 unit + 1
  integration smoke; +13 vs the foundation commit covering middleware,
  evaluator, installed_hook).
* `cargo test -p ironclaw_reborn` — 118 tests pass; no regressions
  from adding the dep.
* `cargo test -p ironclaw_architecture` — 13 tests pass; the
  `ironclaw_turns -> ironclaw_hooks` boundary still holds and the new
  `ironclaw_hooks` rule (no host_runtime / dispatcher / secrets /
  network / wasm / reborn) is unaffected.
* `cargo clippy -p ironclaw_hooks --all-targets --all-features
  -- -D warnings` — clean.
* `cargo clippy -p ironclaw_reborn --all-targets -- -D warnings` —
  clean.
* `cargo fmt --all -- --check` — clean.

What still defers
==================

* WASM hook execution path.
* Persistent predicate counter (in-memory only for now).
* Argument-extraction so `NumericSum` predicates evaluate against
  capability arguments.
* Gate-ref plumbing so PauseApproval / PauseAuth surface real
  `CapabilityOutcome::ApprovalRequired` instead of `Denied`.
* Prompt-snippet injection (waits for shared envelope helper).
* Event-triggered hooks.
* Self-authored hooks (#3567).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(reborn): add HookedLoopModelPort/TranscriptPort/CheckpointPort observer middleware

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(reborn): end-to-end hooks integration through RebornLoopDriverHostFactory

Adds crates/ironclaw_reborn/tests/hooks_integration.rs covering the
factory's HookDispatcher wiring seam end-to-end. Tests drive
host.invoke_capability(...) (not dispatcher.dispatch_before_capability(...)
directly) so a regression in RebornLoopDriverHostFactory's wrapping
composition surfaces here.

Scenarios:
- PredicateBackedBeforeCapabilityHook (DenyCapability NameEquals
  "cap.blocked") short-circuits invocation; inner port never called;
  outcome is Denied(unknown("hook_denied")).
- A privileged selective hook that allows non-matching capabilities
  proves the wrapper does not blanket-deny: cap.allowed reaches the
  inner port and completes once.
- Factory built without with_hook_dispatcher() lets cap.blocked through
  to the inner port, proving the hook plumbing is genuinely opt-in.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(reborn): add pass() + HookRegistrar + self-authored hooks scaffolding

Three additions to ironclaw_hooks:

B. `pass()` on gate sinks — distinguishes "evaluated, no opinion" from
   "returned without minting a decision." A passing hook contributes
   nothing to the composed decision; a silent hook is still Malformed
   and fails closed. `PredicateBackedBeforeCapabilityHook` now routes
   the evaluator's `Allow` decision through `sink.pass()` instead of
   the previous `deny("hook_predicate_pass")` workaround.

A. `HookRegistrar` bridge — converts a `Vec<HookManifestEntry>` into
   `HookBinding`s + dispatcher impls in one call. Predicate bodies are
   wired through `PredicateBackedBeforeCapabilityHook`; WASM bodies
   return `HookError::RegistryConstruction` for now. Adds
   `HookDispatcher::insert_binding` so the registrar can mutate the
   registry through the dispatcher rather than reach inside.

I. Self-authored hooks scaffolding — fourth `HookTrustClass` variant
   for hooks the agent authors at runtime. Run-scoped only;
   monotonic-restriction sink with no `allow`, no trusted-snippet path,
   no effect-class constructor. Closed-vocabulary `SelfAuthoredReason`
   enum keeps free-text reasons off the audit seam.
   `SelfAuthorshipProvenance` captures authoring run/turn, timestamp,
   spec digest, optional user ratification, and a generation-trace
   pointer. Durable persistence depends on the unforgeable channel
   from #3564 and lands separately.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(reborn): real gate-ref plumbing for hook PauseApproval/PauseAuth decisions

Previously, `GateDecisionInner::PauseApproval` and `PauseAuth` returned by
hooks were degraded to `CapabilityOutcome::Denied` at the middleware
boundary because the hook crate had no way to mint a `LoopGateRef` scoped
to the current run. Hooks that wanted to pause the loop for approval or
auth instead failed the call closed, leaving the host's approval-router
machinery unreachable from hook code.

This change introduces a `HookGateRefFactory` trait in
`ironclaw_hooks::middleware::gate_ref` that mints `LoopGateRef`s for
pause-class decisions. `HookedLoopCapabilityPort` now takes an
`Arc<dyn HookGateRefFactory>`, defaulting to `UuidHookGateRefFactory` (a
locally-unique opaque-id factory suitable for tests and the foundation
slice). Production deployments override via `.with_gate_ref_factory(...)`
with a factory bound to the current `LoopRunContext` and the host's
gate-router.

The translation in `decision_to_outcome` is now async so it can await the
factory. `PauseApproval` maps to `CapabilityOutcome::ApprovalRequired
{ gate_ref, safe_summary }` and `PauseAuth` to `AuthRequired`. If the
factory itself errors, the middleware falls back to `Denied` with a
sanitized `hook_gate_ref_unavailable` reason kind so the loop fails
closed rather than routing through an unresolvable suspension. The
underlying error text is dropped to avoid leaking gate-router state into
model-visible output.

Tests:
- `pause_approval_decision_surfaces_as_approval_required`,
  `pause_auth_decision_surfaces_as_auth_required`,
  `gate_ref_factory_failure_falls_back_to_denied` in
  `middleware::capability_port::tests`.
- `pause_approval_hook_surfaces_as_approval_required_with_real_gate_ref`
  in `crates/ironclaw_reborn/tests/hooks_integration.rs`, exercising the
  full `RebornLoopDriverHostFactory` composition with the default
  `UuidHookGateRefFactory`.
- Gate-ref factory unit tests in `gate_ref::tests`.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(reborn): add NumericSum predicate evaluation with capability argument extraction

Wires the missing argument-extraction story for the predicate evaluator so
`ValueOrRateBound::NumericSum` actually enforces a rolling numeric cap
instead of warn-and-allowing.

- Extend `BeforeCapabilityHookContext` with a sealed `SanitizedArguments`
  view. Strings truncate to 256 bytes; objects/arrays cap at 8-deep.
  `extract_numeric` supports dotted + bracketed paths (`order.amount`,
  `items[0].price`) and returns `Option<rust_decimal::Decimal>`. The inner
  representation is sealed so external callers can't bypass bounds.

- Introduce `CapabilityInputResolver` + bundled `NullCapabilityInputResolver`
  in `middleware/resolver.rs`. The hooks crate intentionally doesn't know
  how to dereference a `CapabilityInputRef` — that knowledge belongs to
  the production host. Until a real resolver is wired in (follow-up),
  arguments are `Unresolved` and `NumericSum` fails closed.

- `HookedLoopCapabilityPort::new` defaults to the null resolver; new
  builder `.with_resolver(Arc<dyn CapabilityInputResolver>)` overrides.

- `PredicateEvaluator` gains a tenant-keyed `value_history` map. The
  `NumericSum` arm parses `max` + `window`, extracts the numeric value
  from sanitized args, accumulates within the rolling window, and applies
  `on_exceeded` when the sum exceeds the cap. Unresolved args, missing
  field, non-numeric field, unparseable max, and unparseable window all
  fail closed via the configured `OnExceededAction`.

- Add `BeforeCapabilityHookContext::new_unresolved(...)` convenience
  ctor; existing test sites switch to it instead of churning every call
  site through the 4-arg ctor.

Test count: +14 (8 new SanitizedArguments tests, 6 new NumericSum
evaluator tests, 1 null-resolver test; one old NumericSum-stub-related
gap closed).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(reborn): seal hook registration trust boundary + dispatcher hardening

Addresses blocking findings from the security audit of `ironclaw_hooks`:

- C1 (Blocking, Trust Model): "Installed cannot Allow" was not enforced
  at the registration boundary. `BeforeCapabilityHookImpl::Privileged`
  was a public variant, so external crates with dispatcher access could
  construct an Installed binding paired with a Privileged impl and bypass
  the sink trait restriction. Sealed `BeforeCapabilityHookImpl`,
  `BeforePromptHookImpl`, and `ObserverHookImpl` to `pub(crate)` and
  replaced the single generic `install_before_capability` /
  `install_before_prompt` / `install_observer` surface with tier-specific
  public installers (`install_builtin_*`, `install_trusted_*`,
  `install_installed_*`) that build the binding with the matching trust
  class internally. Updated registrar, internal middleware tests, the
  hooks foundation pipeline test, and the reborn `hooks_integration`
  test to drive the new surface. Added regression tests proving the
  trust class is set by the installer and that the seal is type-level.

- C5 (Medium, Slot Poisoning): same-dispatch poisoning was incomplete
  because `ordered_bindings` snapshots once at the top of the loop, and
  `HookRegistry::insert` accepted duplicate hook IDs. Rejected duplicate
  hook IDs (any point) in `HookRegistry::insert` and added a poison
  re-check before invoking each hook impl in `dispatch_before_capability`,
  `dispatch_before_prompt`, and `dispatch_observer_at`. Added regression
  tests for both behaviors.

- C6 (Medium, Manifest / Predicate Validation): `parse_window` could
  panic on non-ASCII input because `split_at(len - 1)` requires a char
  boundary. Rewrote to compute the unit char's UTF-8 byte length and
  slice safely, added a public `validate_window` helper, and wired it
  into `HookManifestEntry::validate` for both `InvocationCount` and
  `NumericSum` bounds. Added tests for non-ASCII, empty, single-char,
  and zero-duration windows.

- C2 (High, Tenant Isolation): partial fix only. The
  `PredicateEvaluator`'s sliding-window counter was keyed by
  `(hook_id, capability)`, so cross-tenant state could leak. Extended
  `HistoryKey` to include `tenant_id` and added a regression test
  proving counters partition by tenant. Documented the broader
  dispatcher-per-build / per-run-fresh-dispatcher pattern as deferred
  follow-up in `crates/ironclaw_hooks/CLAUDE.md`.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(reborn): emit hook telemetry milestones for audit/SSE observers

Wires the hook dispatcher into the host's milestone stream so audit
backends and SSE observers can see hook activity. Previously, hook
dispatch was invisible — denies, pauses, failures, and observer fires
left no trace in the host's observability backend.

Changes:

- `ironclaw_turns`: add `HookDispatched`, `HookDecisionEmitted`, and
  `HookFailed` variants to `LoopHostMilestoneKind`, with a closed-
  vocabulary `HookDecisionSummary` enum (Allow/Deny/PauseApproval/
  PauseAuth/Pass/Patch). Introduce a lightweight `HookMilestoneSink`
  trait that emits hook-specific *kinds* without requiring a
  `LoopRunContext` (the dispatcher is a process-wide singleton that
  cannot own a per-run context), plus a `RunScopedHookMilestoneSink`
  adapter that injects run context and forwards to the existing
  `LoopHostMilestoneSink`. Also add `InMemoryHookMilestoneSink` for
  tests.

- `ironclaw_hooks`: add a `telemetry` module that converts hook-crate
  types (`HookId`, `HookTrustClass`, `HookPointSpec`, `FailureCategory`,
  `FailureDisposition`, `BeforeCapabilityHookDecision`) into the wire-
  shape labels and summaries the milestone sink expects. Hook ids cross
  the seam as hex strings because the strongly-typed `HookId` cannot be
  imported from `ironclaw_turns` (the architecture test enforces
  `ironclaw_turns -> ironclaw_hooks` stays absent).

- `ironclaw_hooks::dispatch`: add an optional `Arc<dyn
  HookMilestoneSink>` to `HookDispatcher`, set via
  `with_milestone_sink`. Emit `HookDispatched` before each hook runs,
  `HookDecisionEmitted` after a decision/pass/patch, and `HookFailed`
  on timeout/panic/malformed/missing-impl across all three dispatch
  paths (before_capability, before_prompt, observer). Default behavior
  (no sink attached) emits nothing — preserves the pre-telemetry
  observable surface.

- `ironclaw_reborn`: document on `with_hook_dispatcher` that callers
  attach the milestone sink to the dispatcher *before* wrapping it in
  `Arc` and installing it into the factory, using a
  `RunScopedHookMilestoneSink` to inject run-context. The dispatcher
  itself is shared across runs, so attaching a fixed run-context inside
  it would be wrong. Update `RuntimeEvent` projection in
  `milestone_events.rs` to ignore the new hook kinds (no projection
  pathway yet; emitted milestones are consumed by SSE observers
  directly).

Tests:

- `ironclaw_hooks::dispatch`: 5 new tests covering milestone emission
  for deny decisions, panic failures, prompt-mutator patches, observer
  pass-throughs, and the no-sink default.
- `ironclaw_reborn` hooks_integration: end-to-end test wiring a
  `RunScopedHookMilestoneSink` onto the dispatcher and asserting hook
  activity surfaces in the host's `LoopHostMilestoneSink`.

Total: +6 hook telemetry tests; no existing tests modified.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(reborn): extract shared prompt envelope; inject hook patches into prompt bundle

Adds `ironclaw_prompt_envelope`, a leaf crate that owns the single envelope
primitive used by every model-visible untrusted-content path. `wrap_untrusted`
prefixes content with a closed-vocabulary `<Trusted|Untrusted> <source>
content: ` marker, rejects bodies carrying instruction-hijack phrases
(`ignore previous instructions`, `<|im_start|>`, `<system>`, etc.), and
enforces a 4 KiB byte budget by default.

Migrates `ironclaw_host_runtime::memory_context` to delegate envelope
wrapping, marker rejection, and control-character stripping to the new
crate while keeping the `LoopSafeSummary`-specific 512-byte cap and byte
truncation local. Existing memory_context behavior and tests are preserved.

Wires the same envelope into `ironclaw_hooks`:

* `HookPatch::add_enveloped_snippet` now takes a raw body and wraps it
  via `wrap_untrusted(EnvelopeSource::Hook, …)`. `Installed` hooks
  produce `Untrusted` envelopes; `Builtin`/`Trusted`/`SelfAuthored`
  produce `Trusted` envelopes so downstream readers can distinguish the
  two paths through a uniform marker.
* `HookedLoopPromptPort::build_prompt_bundle` is no longer observe-only.
  After dispatching `before_prompt`, it envelope-wraps every snippet
  patch (passing `Enveloped` through, wrapping `Trusted` with the
  envelope helper), enforces the 4 KiB aggregate snippet byte budget
  across patches, and appends the wrapped snippets to the prompt
  bundle's `messages` as `system`-role `LoopModelMessage` entries
  carrying deterministic `msg:hook.<ordinal>.<hash>` content refs
  (mirroring the skill-snippet ref convention).

The envelope crate is a leaf with no ironclaw dependencies, satisfying
the boundary contract; the existing `ironclaw_hooks` boundary rule in
`reborn_dependency_boundaries` continues to hold because
`ironclaw_prompt_envelope` is not on its forbidden list.

Test count delta:
* `ironclaw_prompt_envelope`: +13 new tests (crate did not exist).
* `ironclaw_hooks`: 84 → 88 tests (+4 prompt-port behavior tests:
  `hook_patch_appended_as_envelope_wrapped_message`,
  `total_byte_budget_enforced_across_patches`,
  `instruction_hijack_in_patch_rejected`,
  `trusted_hook_patch_wrapped_with_trust_marker`).
* `ironclaw_host_runtime` memory_context: unchanged (8 tests still pass).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix: align tenant-counter test with SanitizedArguments-extended context ctor

* docs(reborn): document loader contract; pin HookId hex format

Add a "Loader responsibility" section to ironclaw_hooks/CLAUDE.md
explaining that tier-specific installers prevent minting wrong-tier
impls but cannot enforce origin — that's the loader's job — and
recommending registry loaders type-tag extension hooks as
LoadedHook::Installed at the loader seam.

Add tier_specific_installers_are_documented_as_loader_contract as a
regression guard that touches every public install_*_before_capability
and install_*_before_prompt method so any signature change forces the
loader contract to be re-evaluated.

Document HookId::to_hex's 64-char lowercase hex output as part of the
cross-crate contract consumed by LoopHostMilestoneKind::Hook* in
ironclaw_turns; add hook_id_hex_format_is_stable_64_lowercase_chars in
identity::tests and hook_id_string_serialization_matches_to_hex in
telemetry::tests to pin the format and the seam conversion path.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(reborn): pin hook milestone JSON schema + assert pairing invariants

Add L3 schema-snapshot tests for every hook-related LoopHostMilestoneKind
variant (HookDispatched, HookDecisionEmitted per HookDecisionSummary,
HookFailed per FailureCategory) so downstream consumers can rely on the
JSON wire shape and any accidental field rename, enum-tag rename, or type
change fails loudly.

Add L4 pairing-invariant matrix test in the hook dispatcher that drives
every observable outcome (Allow, Deny, PauseApproval, PauseAuth, Pass,
Panic, Timeout, Malformed, MissingImpl) through a recording milestone
sink and asserts the dispatched-then-terminator pairing shape. Document
the MissingImpl path as the one case that emits a sole HookFailed with
no preceding HookDispatched (the dispatcher discovers the protocol
violation before the hook is actually dispatched).

Add a multi-hook dispatch test that installs three hooks with mixed
outcomes (allow/deny/panic) at the same point and asserts each hook
produces its own paired sequence in the deterministic
(phase, priority, hook_id) order taken from the dispatcher's registry.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(reborn): integration tests for observer middleware through RebornLoopDriverHostFactory

Wire the HookedLoopModelPort / HookedLoopTranscriptPort /
HookedLoopCheckpointPort observer wrappers into
RebornLoopDriverHostFactory::build_text_only_host_with_capabilities,
mirroring the existing HookedLoopCapabilityPort / HookedLoopPromptPort
composition. The wrappers are applied only when a HookDispatcher is set
on the factory, so the default factory shape is unchanged.

Add four integration scenarios in crates/ironclaw_reborn/tests/hooks_integration.rs:

- observer_hook_fires_after_model_through_factory
- observer_hook_fires_after_capability_through_factory
- observer_hook_fires_after_checkpoint_through_factory
- observer_panic_does_not_fail_model_call (panic-isolation regression)

Relax the test-fixture model gateway from "panic if invoked" to
returning a stub assistant reply so the AfterModel / panic-isolation
tests can drive stream_model through the wrapped port. The existing
capability-port tests never touch the gateway, so their behavior is
unchanged.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(reborn): introduce HookDispatcherBuilder for type-enforced sink wiring

Adds a `HookDispatcherBuilder` in `ironclaw_hooks::dispatch` that owns
the dispatcher construction lifecycle: registry -> optional timeout ->
optional milestone sink -> installed hooks -> `.build_arc()`. The
terminal `.build_arc()` wraps in `Arc` and yields an immutable handle.

Tightens the public surface on `HookDispatcher`: `new`, `with_timeout`,
`with_milestone_sink`, and every `install_*_*` method are now
`pub(crate)`. Outside callers route exclusively through the builder, so
"wire the milestone sink before Arc-wrapping" is a compile-time fact
rather than a documentation convention.

`HookRegistrar::install` now takes a `HookDispatcherBuilder` by value
and returns `(HookDispatcherBuilder, Vec<HookId>)`, keeping the builder
chainable through manifest installation.

`RebornLoopDriverHostFactory` gains `with_hook_dispatcher_builder` to
let callers defer `.build_arc()` to the factory — a step toward the
FU8 per-build dispatcher pattern.

Migrates `foundation_pipeline.rs` and `hooks_integration.rs` to the
builder. Internal middleware and dispatch tests continue to use the
crate-private `HookDispatcher::new` directly.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(reborn): production CapabilityInputResolver for NumericSum predicates

Adds HookCapabilityInputResolverAdapter in ironclaw_reborn that bridges
the existing LoopCapabilityInputResolver (already used by
HostRuntimeLoopCapabilityPort for dispatch input resolution) to the
hooks crate's CapabilityInputResolver trait. RebornLoopDriverHostFactory
gains with_capability_input_resolver(...), and when both a hook
dispatcher and resolver are configured the factory threads the adapter
into HookedLoopCapabilityPort::with_resolver — so NumericSum and other
argument-dependent predicates evaluate against real, sanitized inputs
instead of failing closed against the framework's null default.

The adapter also enforces a configurable serialized-byte budget
(default 64 KiB) as defense in depth ahead of the hooks crate's
per-string and depth caps in SanitizedArguments.

Unit tests cover the four adapter branches (resolved JSON,
inner-error → None, non-object pass-through, oversized → None) and a
new end-to-end integration test
(numeric_sum_predicate_caps_total_value_against_real_inputs) drives the
full factory wiring: with a NumericSum cap of 99 over an "amount" field,
two invocations carrying {"amount":"50"} let the first pass through and
deny the second at the hook seam, with the inner port reached exactly
once.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(reborn): per-build HookDispatcher for full per-run isolation (C2)

Introduce `with_hook_dispatcher_factory(F)` on
`RebornLoopDriverHostFactory`. The closure is invoked once per
`build_text_only_host*` call, so dispatcher-owned mutable state — slot
poisoning, registry mutations, predicate-counter siblings — is scoped to
a single host build instead of shared across every host the factory
produces.

The legacy `with_hook_dispatcher(Arc<HookDispatcher>)` adapter is kept as
a thin wrapper that returns clones of the same `Arc` on every build. Its
shared-state behavior is now documented as an explicit opt-in for
backward compat; new wiring should prefer the factory closure.

Adds two regression tests:
  - `per_build_dispatcher_state_does_not_leak_across_runs` — installs a
    panicking hook, builds two hosts back-to-back, and proves the inner
    port is never reached on build 2 (fresh slot still applies the
    fail-closed deny). Pins per-run isolation.
  - `legacy_with_hook_dispatcher_shares_state_across_builds` — pins the
    shared-state semantic of the legacy adapter as the explicit baseline.

Migrates `predicate_deny_hook_short_circuits_inner_port` to the new
factory-closure path so the new wiring is exercised by the existing
suite.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(reborn): project hook telemetry milestones into RuntimeEvent for durable audit

Extend the runtime event substrate with `HookDispatched`, `HookDecisionEmitted`,
and `HookFailed` kinds carrying closed-vocabulary labels and the blake3-hex hook
identity. Project the matching `LoopHostMilestoneKind::Hook*` variants in
`DurableLoopHostMilestoneSink` so hook telemetry now lands in the same durable
event log as model/reply/loop milestones — SSE observers still see live hook
events, and audit replay can reconstruct the full hook trail.

- `ironclaw_events`: add hook variants to `RuntimeEventKind`, optional hook
  fields on `RuntimeEvent` (`hook_id`, `hook_point`, `hook_trust_class`,
  `hook_decision`, `hook_failure_category`, `hook_failure_disposition`),
  typed constructors (`hook_dispatched`, `hook_decision_emitted`,
  `hook_failed`), and dedicated sanitizers (`sanitize_hook_label`,
  `sanitize_hook_id`) re-run on every wire crossing. No new crate dependency
  edges; hook strings cross the boundary opaque.
- `ironclaw_reborn::milestone_events`: project the three hook milestone kinds
  via a new `loop.hook` capability id. `HookDecisionSummary` is collapsed to
  its closed-vocabulary `kind_name()` so sanitized reasons never enter the
  durable substrate.
- `ironclaw_event_projections`: extend `TimelineEntryKind` and the
  `RuntimeEventKind -> RunProjectionStatus` mapping so hook events are pure
  telemetry — they preserve the current run status rather than changing it.
- Tests: 4 unit tests in `ironclaw_events::runtime_event::tests` (serde
  round-trip per variant + unsafe-label collapse), 3 in
  `ironclaw_reborn::milestone_events::tests` (projection per variant,
  including the assertion that raw `Deny { reason }` text does not reach the
  durable wire payload). Existing replay-projection direct-construction
  tests updated for the new RuntimeEvent fields.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(reborn): enforce manifest-declared hook scope at dispatch time (C3)

Audit finding C3: extensions could declare `[[hooks]]` with
`scope = "own_capabilities"` in their manifest, but the dispatcher never
enforced it — an Installed hook from ext-A could fire against capabilities
provided by ext-B. Scope was parsed but not load-bearing.

This change makes scope load-bearing end-to-end:

- `BeforeCapabilityHookContext` carries an optional `provider:
  ironclaw_host_api::ExtensionId` populated by the middleware. The hook
  context is `#[non_exhaustive]` already so this is non-breaking.

- `HookBinding` gains `owning_extension: Option<ExtensionId>` and `scope:
  HookBindingScope`. `HookBindingScope` is `Global` / `OwnCapabilities`
  / `SameTenant`. Builtin and Trusted bindings default to `Global` and
  carry no `owning_extension`; Installed bindings carry both, sourced
  from the manifest.

- `HookDispatcher::install_installed_*` installers now require the
  caller to pass `(owning_extension, scope)`. The registrar derives both
  from the manifest entry, so manifest authorship is the single source
  of truth.

- A new `CapabilityProviderResolver` trait + bundled
  `NullCapabilityProviderResolver` lets the middleware lift the
  capability id to its provider at invocation time. The middleware
  wires the resolved provider into the hook context.

- `dispatch_before_capability` consults `binding.scope.permits(...)`
  before invoking each hook. Bindings that don't permit the current
  invocation are inert — no sink call, no failure record, no poisoning.

Conservative defaults:

- When the provider resolver returns `None` (no resolver wired, or the
  capability has no known provider), `OwnCapabilities`-scoped hooks do
  NOT fire. An attacker cannot bypass scope filtering by stripping
  provider info from the descriptor.

Tests:

- 5 new dispatcher tests cover OwnCapabilities matching, foreign
  provider, unresolved provider, SameTenant, and Builtin Global.
- 1 new registrar test asserts manifest scope and extension propagate
  into `HookBinding`.
- 1 new middleware test asserts the provider resolver populates the
  hook context.
- 1 new integration test in `ironclaw_reborn` proves an ext-A hook
  scoped to `OwnCapabilities` does not intercept invocations that have
  no resolved provider (the production composition default).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* style: rustfmt dispatch.rs after FU1 merge

* docs(hooks): prior-art comparison against LSM/eBPF/Envoy/K8s/OPA/CRX/VSC/Tauri

Validates the IronClaw hooks design against 8 established hook/policy
systems across 8 axes (dispatch, trust tiers, attenuation, decision
vocabulary, failure semantics, isolation, manifest, audit).

Surfaces:
- 7 areas where ICLAW stands out vs prior art (type-level trust
  enforcement, dispatch-time scope, failure-kind matrix, pause-with-
  gate-ref, pairing-invariant audit matrix, tenant-keyed predicates,
  phase-ordered dispatch)
- 4 conventional choices we should revisit (in-process Installed-WASM,
  sticky poison, no formal dispatch model, no installation rate-limit)
- 3 divergences whose 'why' is weak and need design review

* docs(hooks): STRIDE threat model for v1 framework

Enumerates 7 adversary classes (A1-A7), 6 assets ranked by blast
radius, and ~35 attack vectors across STRIDE categories with mitigations,
existing tests, and residual risk.

Surfaces 7 prioritized follow-ups:
- High: per-extension hook-count cap (D3/D4)
- High: gate-ref unguessability + one-shot test (S1)
- Med: resolver field-level scope (I2)
- Med: per-evaluator state ceiling (D5)
- Med: poison-stickiness operator runbook
- Low: timing side-channel residual acknowledgement (I4)
- Low: instruction-marker denylist periodic review (I5)

Confirms the load-bearing 'Installed cannot Allow' (E1) property holds
via type-level seal + tier-specific installers, backed by
compile_time_seal_test and installed_binding_cannot_be_paired_with_
privileged_impl tests.

Explicit out-of-scope: extension install pipeline (#3492), WASM exec
sandbox (needs separate threat model when it lands), approval gateway
(#3564).

* feat(hooks): close threat-model gaps S1 (gate-ref entropy) and D3/D4 (registration flood)

S1 (gate-ref unguessability, factory side):
- Three new tests on `UuidHookGateRefFactory`:
  - `gate_refs_are_v4_uuids` pins the v4 entropy source (122 random
    bits per ref per RFC 4122 §4.4); fails if a future change moves to
    a counter or weaker UUID version.
  - `gate_refs_have_no_collisions_across_many_calls` mints 20k refs
    across both namespaces, asserts zero collisions (statistical
    proxy for entropy quality).
  - `approval_and_auth_namespaces_do_not_overlap` confirms prefix
    routing separation.
- Doc comment now documents the security property explicitly and
  delineates factory-side vs gateway-side responsibilities for the
  one-shot consumption property.

D3/D4 (hook registration flood):
- New `MAX_HOOKS_PER_EXTENSION = 32` and
  `MAX_HOOKS_PER_EXTENSION_PER_KIND = 8` consts in `registrar.rs`.
- New `HookRegistrar::enforce_registration_caps` runs pre-flight at
  the top of `install()`, before any binding is inserted. Whole-batch
  rejection means a partially-installed batch cannot slip past.
- Three regression tests: total-cap rejection, per-kind-cap rejection,
  at-cap acceptance.
- Error messages cite the threat-model finding so operators can map
  rejection back to the design rationale.

Threat model updated: S1, D3, D4 marked closed in the cross-cutting
properties matrix and the open-follow-ups list.

* test(hooks): three real hooks built against the public API + ergonomics findings

Builds three representative hooks from outside the crate, mimicking
what an extension or system author would actually write:

1. polymarket-daily-cap — Installed predicate hook, InvocationCount
   rate-cap with Deny on excess. Canonical 'rate-limit a capability'
   use case for the predicate language.

2. large-stake-approval-gate — Installed predicate hook, NumericSum
   over amount_usd field, PauseApproval at $1000/24h. Manifest-shape
   + registrar-install coverage from outside Reborn; end-to-end
   dispatch lives in ironclaw_reborn integration tests because
   NumericSum needs resolved args (a friction finding documented in
   the companion doc).

3. pii-redaction-warning — Trusted Rust hook implementing
   PrivilegedBeforePromptHook, injects a trusted instruction snippet
   reminding the model to redact PII. Demonstrates the path a system
   author takes when the predicate language isn't expressive enough.

API change (F1 fix): SanitizedArguments::unresolved() promoted from
pub(crate) to pub. This is the documented safe default — predicates
that need args must fail closed against it — so exposing the
constructor cannot weaken any trust property. The sanitizing
from_json constructor stays sealed; that's the trust boundary.
Without this fix, external hook authors could not construct a
BeforeCapabilityHookContext with both a known provider AND
unresolved args, which made TDD of their own predicate impossible.

Findings documented in docs/real-hooks-findings.md, ranked by
severity. Big-picture observation: writing the Trusted Rust hook
(F4) was easier than writing the declarative predicate hook (F1 +
F2 + F3) — three of seven findings target predicate-authoring
ergonomics. The declarative path needs the most polish before
third-party extension authors will trust it for non-trivial policy.

Tests: 6 new in real_hooks.rs, all pass.

* feat(hooks): close all remaining threat-model and ergonomics gaps

Closes the Med-priority threat-model gaps (I2, D5, poison runbook)
and all real-hooks ergonomics findings (F2, F3, F5, F6, F7) in a
single pass.

Threat model:
- I2 (resolver field-scope): documented in SanitizedArguments rustdoc.
  The narrow public surface (only is_resolved + extract_numeric)
  enforces field-scope by construction for the current predicate
  path. Reassess when Installed-WASM lands.
- D5 (evaluator state ceiling): MAX_HISTORY_KEYS = 8192 per map,
  LRU eviction with evictions_observed() metric for operator
  monitoring. New regression test
  lru_eviction_increments_counter_and_drops_oldest_key.
- Poison-stickiness runbook: new docs/operator-runbook.md with
  recovery options ranked by cost.

Ergonomics findings:
- F2 (closed-vocab deny reasons): rustdoc on OnExceededAction
  and GateDecisionView::Deny explaining the audit-vs-model split
  and why manifest reason text doesn't reach the model.
- F3 (NumericSum can't be TDD'd outside Reborn): new test-support
  feature flag with SanitizedArguments::for_tests(value) that
  external hook authors can opt into via dev-dep.
- F5 (two ExtensionId types): added
  From<&ironclaw_host_api::ExtensionId> impl for
  identity::ExtensionId, plus cross-link rustdoc.
- F6 (HookManifestEntry struct-literal fragility): added
  #[non_exhaustive] + HookManifestEntry::new(id, kind, body) +
  with_scope/with_phase/with_priority/with_description/with_requires_grant
  builder methods. Migrated 3 external call sites in tests/.
- F7 (priority guidance): rustdoc on HookPriority with when-to-
  deviate guidance, named FIRST/LAST constants documented for
  Builtin/Telemetry use cases.

Tests: 151 unit + 1 + 6 integration in ironclaw_hooks all pass
with --all-features. ironclaw_reborn (13 hooks_integration scenarios)
unchanged.

Threat model updated: I2 / D5 / poison runbook marked closed in
both the per-vector table and the cross-cutting properties matrix.
Open follow-ups now down to two Low items (I4 timing side-channel
residual, I5 instruction-marker denylist refresh) plus the deferred
DenyReasonCode enum from F2.

* fix(ci): collapse nested match in hooks_integration test for clippy --all-features

CI runs `cargo clippy --all --tests --examples --all-features -- -D warnings`
which is stricter than the workspace clippy I ran locally and trips
`clippy::collapsible_match` on the nested-if in HookDecisionEmitted
matching. Collapse the inner `if decision.kind_name() == "deny"`
into an arm guard.

* feat(hooks): address henrypark133 review — Critical #1/#2/#3/#5, Concerning #5/#7

Address composition-seam bugs in the Reborn factory wiring + doc tidy.

henrypark133 review findings addressed:

Critical #1 — before_prompt hook messages not materialized.
  HookedLoopPromptPort now requires a HookPromptMaterializationSink and
  fails closed if patches are emitted without one. The reborn factory
  installs an InstructionStoreBackedHookSink adapter that delegates to
  the host's InstructionMaterializationStore, so synthetic msg:hook.*
  refs are resolvable by the downstream model resolver. New seam trait
  (HookPromptMaterializationSink) keeps ironclaw_hooks decoupled from
  LoopRunContext.

Critical #2 — OwnCapabilities hooks were inert in production wiring.
  Factory now installs SurfaceBackedProviderResolver (consults the
  visible-capability surface for capability_id → provider). With this,
  ctx.provider is populated and OwnCapabilities-scoped Installed hooks
  actually fire against their own provider's capabilities.

Critical #3 — gate refs were unresolvable.
  Middleware default switched from UuidHookGateRefFactory to
  FailClosedHookGateRefFactory. Tests must explicitly opt into UUID
  (via with_gate_ref_factory) to exercise the affirmative ApprovalRequired
  path; production deployments must install a router-backed factory.
  New factory method RebornLoopDriverHostFactory::with_hook_gate_ref_factory.

Concerning #5 — AfterModel fired twice + before durable finalization.
  Removed AfterModel dispatch from HookedLoopModelPort; the transcript
  port's finalize_assistant_message is now the sole AfterModel boundary
  (the durable one). Model port wrapper is preserved as a no-op shim
  for symmetry + future model-response-observed point.

Concerning #7 — doc tidy:
  - CLAUDE.md: 3 trust classes → 4 (Builtin/Trusted/Installed/SelfAuthored
    with explicit note that SelfAuthored is run-scoped only and not
    loadable from an external source).
  - operator-runbook.md: "Audit log" → "durable runtime event stream"
    where the projection is actually the runtime-event stream, not formal
    AuditEnvelope records.
  - prior-art.md: poison-lifetime nuance — per-host-build with the
    factory pattern, process-lifetime only for the legacy adapter.
  - prior-art.md:80: trailing whitespace removed.

Testing gaps from henrypark133 — caller-level tests through
RebornLoopDriverHostFactory:
  #1 (before_prompt resolver path):
     before_prompt_hook_message_is_resolvable_via_factory_wiring
  #2 (OwnCapabilities positive/negative/unknown):
     own_capabilities_hook_fires_when_provider_matches
     own_capabilities_hook_does_not_fire_when_provider_differs
     own_capabilities_hook_does_not_fire_when_provider_unknown
  #3 (pause/auth gate lifecycle or fail-closed):
     pause_approval_with_default_factory_fails_closed_as_denied
     pause_approval_hook_surfaces_as_approval_required_with_real_gate_ref
     (updated to require explicit UuidHookGateRefFactory opt-in)
  #5 (AfterModel exactly-once at durable boundary):
     after_model_fires_exactly_once_at_durable_boundary

Still TODO from review (separate commits):
  Critical #4 (telemetry context — two-run attribution) + gap #4
  Concerning #6 (TimelineEntry hook metadata projection) + gap #6

Tests: 154 unit + 18 hooks_integration + all other reborn tests pass.
Workspace clippy + fmt + no-panics clean.

* feat(hooks): address remaining henrypark133 review — Critical #4, Concerning #6

Critical #4 — per-run hook telemetry attribution.
  New `HookDispatcherBuilderFactory` signature: factory returns a
  HookDispatcherBuilder, and `RebornLoopDriverHostFactory` attaches a
  `RunScopedHookMilestoneSink` keyed to the CURRENT run's LoopRunContext
  inside `build_text_only_host_with_capabilities`, before sealing the
  dispatcher. The previous zero-arg signature relied on the closure
  capturing run_context — silently misattributed across reuses; new
  public API `with_hook_dispatcher_builder_factory` removes that
  failure mode entirely. Legacy `with_hook_dispatcher_factory` retained
  for back-compat (its sink-wiring contract stays caller-side).

Concerning #6 — TimelineEntry hook metadata.
  Added 6 optional fields to `TimelineEntry` (hook_id, hook_point,
  hook_trust_class, hook_decision, hook_failure_category,
  hook_failure_disposition) and projected them from `RuntimeEvent::Hook*`.
  Replay consumers now see which hook fired/failed, not just that some
  hook event happened. Each field is closed-vocabulary (no free-form
  reason text — that stays in the audit reason payload, not the
  product replay DTO).

Testing gaps from henrypark133 — caller-level tests:
  #4 (two-run hook telemetry attribution):
     hook_telemetry_attribution_is_per_run_not_captured
     Builds two hosts from the SAME builder factory closure with two
     fresh LoopRunContexts. Asserts each run's hook milestones carry
     its OWN run_id (no stale captured one).
  #6 (replay projection contract for hook events):
     hook_runtime_events_project_with_sanitized_hook_metadata
     non_hook_runtime_events_project_with_no_hook_metadata
     Constructs RuntimeEvent::Hook{Dispatched,DecisionEmitted,Failed}
     and asserts the projection preserves the metadata fields. The
     negative test guards against cross-contamination on non-hook
     events.

All henrypark133 review items now addressed:
  Critical: #1, #2, #3, #4 — done
  Concerning: #5, #6, #7 — done
  Testing gaps: #1-#6 — done

Tests: 154 unit + 19 hooks_integration in ironclaw_reborn + 61 reborn
unit + 38 + 2 new in ironclaw_event_projections + ... pass.
Workspace clippy + fmt + no-panics clean.

* docs(hooks): scope DenyReasonCode closed-vocabulary enum (successor #6)

Successor PR from #3573 — real-hooks ergonomics finding F2 (deferred).
Adds a curated vocabulary of model-visible denial reasons so hook
authors can communicate why a deny happened without opening a
free-form prompt-injection channel.

* feat(hooks): DenyReasonCode + PauseReasonCode closed-vocabulary enums

Address real-hooks ergonomics finding F2 (deferred from PR #3573). The
prior dispatcher collapsed every Installed-tier deny to the static
label 'hook_predicate_denied', because manifest reason strings are
author-controlled and surfacing them to the model would open a
prompt-injection channel. The cost: the agent couldn't tell *why*
a hook denied.

This PR introduces two closed-vocabulary enums:

- DenyReasonCode: Generic / RateLimit / ValueCap / Blocklist /
  RequiresApproval / OutOfPolicy
- PauseReasonCode: Generic / RequiresApproval / OverThreshold /
  SensitiveAction

Each variant has an as_label() returning &'static str (so the sink's
&'static str contract is preserved). New OnExceededAction variants
'DenyWithCode { code, reason }' and 'PauseApprovalWithCode { code,
reason }' let manifest authors opt into the richer labels while
keeping reason audit-only.

The legacy Deny { reason } / PauseApproval { reason } variants are
retained for back-compat and map to DenyReasonCode::Generic /
PauseReasonCode::Generic — existing manifests continue to produce
hook_predicate_denied / hook_predicate_pause_requested.

Threat-model regression: a hook author cannot smuggle text into the
model-visible label because the 'code' field is typed as the enum;
there's no String slot exposed model-side. A test
(deny_with_code_only_exposes_enum_variants_to_model) documents this
as a compile-time property.

Tests (+7 new = 161 total):
- deny_reason_code_labels_are_stable: pins the label vocabulary so
  rename/relabel is loud.
- pause_reason_code_labels_are_stable: same for PauseReasonCode.
- deny_with_code_round_trips_through_json + pause variant: wire
  round-trip + snake_case tag assertion.
- deny_with_code_only_exposes_enum_variants_to_model: compile-time
  property check.
- rate_or_value_cap_with_deny_code_routes_to_code_label: end-to-end
  affirmative test that the dispatcher emits the code's label.
- rate_or_value_cap_with_pause_code_routes_to_code_label: same for
  pause.

Scope doc: crates/ironclaw_hooks/docs/successors/06-deny-reason-code.md

* test(hooks): address codex review on #3636

- Update stale real-hooks-findings.md F2 row to cite this PR's enum
  follow-on (was 'deferred').
- Add install_deny_with_code_manifest_surfaces_code_label_on_dispatch:
  end-to-end test driving the registrar->dispatcher path for the
  new DenyWithCode variant (prior tests covered serde + direct hook
  evaluation, but not the manifest install path that downstream
  authors actually use).

Codex review on PR #3636: APPROVE with two recommendations; both
addressed.

Tests: 162 unit (+1 new). Clippy/fmt clean.

* fix(hooks): attenuate Installed-tier prompt patches to user role

Installed-tier `before_prompt` patches were injected as role:"system"
messages. Envelope text labels ("[ext-foo says]: ...") do not strip
system-role authority from the model's perspective, so a third-party
extension could inject system-tier instructions through a snippet
patch. This is a prompt-authority escalation against the trust
hierarchy the framework otherwise enforces.

Add `role_for_trust_class()` mapping Installed -> "user" and
Builtin/Trusted/SelfAuthored -> "system". Thread per-patch
trust_class through `wrap_patches_to_messages` and use it for the
emitted `LoopModelMessage.role`.

Tests:
- installed_hook_patch_drops_to_user_role: asserts the role for an
  Installed-tier patch is "user"
- trusted_tier_hook_patch_keeps_system_role: regression that Trusted
  tier still produces system-role content

* fix(hooks): enforce scope filter on observer dispatch + reject incompatible points

Two related defense-in-depth fixes against silent scope-filter failure:

1. The registry silently accepted Installed bindings with
   `HookBindingScope::OwnCapabilities` at points (BeforePrompt,
   AfterModel, AfterCheckpoint) whose dispatch context carries no
   per-capability provider. The manifest's declared scope had no
   effect at all — the hook fired against every dispatch. Reject the
   binding at install time so the operator sees the misconfiguration.

2. `dispatch_observer_at` for `AfterCapability` did not consult the
   binding's scope, so an Installed observer registered with
   `OwnCapabilities` fired against every invocation regardless of
   provider. Add `dispatch_observer_at_with_provider` carrying the
   resolved capability provider; the capability-port middleware
   resolves the provider once per invocation and threads it through
   both the BeforeCapability hook context and the AfterCapability
   observer dispatch. The dispatcher then enforces
   `HookBindingScope::permits` on each observer binding.

`ObserverHookContext` gains a `provider: Option<ExtensionId>` field;
`#[non_exhaustive]` keeps existing authors compiling.

Tests:
- rejects_own_capabilities_at_before_prompt
- rejects_own_capabilities_at_after_model
- accepts_own_capabilities_at_before_capability
- own_capabilities_observer_filters_foreign_providers (covers
  foreign / matching / unresolved provider)

* fix(hooks): preserve free-form audit reason alongside closed-vocab model label (serrrfirat #3636)

`PredicateBackedBeforeCapabilityHook::evaluate()` was discarding the
free-form `reason` from `EvaluatorDecision::{Deny, PauseApproval}`
with `..` and only sending `code.as_label()` into the sink. The
`HookDecisionEmitted` milestone therefore carried only the closed-
vocab label, and operator-visible audit/SSE context was silently lost
end-to-end. The fix splits the channels:

- Model sees the closed-vocab label (`hook_rate_limit`,
  `hook_pause_over_threshold`, ...) via `sink.deny(label)`. This
  channel is unchanged.
- Audit/SSE sees the manifest's free-form `reason` via a new
  audit-only sink method `record_audit_reason(reason: String)`. The
  recording sink captures it; the dispatcher reads it after the hook
  returns and threads it into `LoopHostMilestoneKind::HookDecisionEmitted`.

Surface changes:
- `PrivilegedGateSink` / `RestrictedGateSink` gain
  `record_audit_reason(String)` — accepts dynamic `String` (audit-only,
  no model-facing seam) unlike the `&'static str` decision reasons.
- `RecordingGateSink` gains an `audit_reason: Option<String>` field.
- `GateHookOutcome::Decision` is now `Decision { decision,
  audit_reason }`.
- `HookDispatcher::emit_decision_with_audit` threads the audit reason
  into the milestone.
- `LoopHostMilestoneKind::HookDecisionEmitted` gains a
  `#[serde(default, skip_serializing_if = "Option::is_none")]`
  `audit_reason: Option<String>`. The durable RuntimeEvent projection
  intentionally drops this field — audit reasons are operator-facing
  in-memory SSE content, never durable cross-process surface.

Tests:
- `deny_with_code_records_audit_reason_separately_from_model_label`:
  asserts the recording sink ends with `Deny { reason: "hook_rate_limit" }`
  in `state` AND `audit_reason == Some("daily cap of $1000 ...")`.

* fix(hooks): remove unused model_request helper (CI clippy fix)

* fix(hooks): address serrrfirat P1/P2 findings on PR #3573

Three issues from the 5-15 review:

**P1 #1 registrar.rs:70 — `same_tenant` grants not enforced**
`HookManifestEntry::validate` only confirmed `requires_grant` was
present; the registrar then immediately installed the binding with no
host-verified grant context. A manifest could declare
`requires_grant = "anything"` and get a cross-extension binding for
free.

Fix: `HookRegistrar` now carries a `verified_grants: HashSet<String>`
(empty by default — default-deny). Add the host-facing setter
`with_verified_grants(...)`. At `install_one`, if
`entry.requires_grant` is `Some(g)`, require `g ∈ verified_grants` or
reject with a clear error. Tests:
- `install_rejects_same_tenant_without_verified_grant`
- `install_rejects_same_tenant_when_verified_grants_mismatch`
- The existing positive test
  `installer_propagates_owning_extension_and_scope_from_manifest` now
  wires the verified grant explicitly (proves the API contract).

**P1 #2 prompt_port.rs:150 — zip misalignment**
The materialization loop zipped surviving messages against the
ORIGINAL unfiltered patch list. `wrap_patches_to_messages` skips
metadata patches and over-budget snippets, so the zip silently paired
message[0] with patch[0] even when patch[0] was the skipped metadata
— materializing the wrong content (or none) under the snippet's
synthetic ref.

Fix: `wrap_patches_to_messages` now returns
`Vec<WrappedHookMessage { message, safe_content }>` — surviving
messages paired with their content by construction. The caller
materializes `entry.safe_content` under `entry.message.content_ref`
directly; no zip against unfiltered input. Removed the now-unused
`safe_content_for_patch` helper.

Test:
- `materialization_stays_aligned_when_metadata_patches_are_filtered`:
  a hook emits `[metadata, snippet]`; asserts only one model message,
  and the materialized content under its ref contains the snippet's
  body — proves filtering can no longer desync from materialization.

**P2 #3 loop_driver_host.rs:1343 — `with_hook_dispatcher_builder`**
Docs said it deferred `build_arc()` to let the host factory finalize
wiring; the implementation called `build_arc()` eagerly and routed
through the legacy shared-dispatcher adapter, losing per-run
dispatcher isolation and the run-scoped milestone sink.

Fix: marked `#[deprecated]` with a note pointing callers to
`with_hook_dispatcher_builder_factory(|| ...)` for per-build
isolation, or `with_hook_dispatcher(...)` if they actually meant the
shared adapter. The method body is unchanged so no callers break;
they'll see the deprecation warning. No internal callers exist, so
the deprecation doesn't trip `-D warnings`.

All 162 hooks lib + 19 reborn integration tests pass; clippy clean.

* fix(hooks): address serrrfirat 3573-2026-05-15 review findings

P1 — prompt bundle authority mismatch (prompt_port.rs):
`HookedLoopPromptPort::build_prompt_bundle` called the inner port first,
which caused `HostManagedLoopPromptPort` to issue the prompt-bundle
authority grant against the pre-hook message list. The wrapper then
appended `msg:hook.*` messages to `bundle.messages`, so the downstream
model request hit `grant.messages != messages` and failed closed with
"model request messages do not match the host-built prompt bundle".

Add `with_bundle_authority(authority, run_context)` and re-issue the
grant after appending hook messages so it covers the post-hook bundle.
Reborn wires `prompt_authority.clone()` + `run_context.clone()` into
the wrapper at construction time.

P2 — observer installer accepts non-observer points (dispatch.rs):
`install_observer` accepted any `HookPointSpec` (including
`BeforeCapability` / `BeforePrompt`) and only populated the observer
map. Dispatch later found a binding without a gate/mutator impl and
fail-closed the capability with "binding present without installed
implementation". Reject non-observer points at install time so misuse
fails loudly rather than poisoning bindings at dispatch.

P2 — batch path skipped AfterCapability observers on inner error
(capability_port.rs):
The batch loop used `?` directly on `self.inner.invoke_capability(...)`,
which propagated the error before dispatching `AfterCapability`
observers. Failed batch entries disappeared from telemetry / audit,
while the single-invocation path dispatches observers on error.
Capture the inner result, dispatch observers, then propagate the error.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(hooks): address PR #3573 review feedback round 3

Addresses serrrfirat's CHANGES_REQUESTED review (2026-05-20) by tightening
several install-time / dispatch-time bounds and gating production seams:

- Bound free-form audit reasons crossing telemetry. New
  `telemetry::sanitize_audit_reason` strips control characters and caps
  length at 512 bytes; `emit_decision_with_audit` routes the manifest-
  supplied reason through it before publishing milestones. Manifest
  validation also rejects reasons over the same byte limit at install time
  so the wire-side cap is a defense-in-depth layer, not the only line.
- Make hot dispatch O(H) instead of O(H^2). The per-binding poison
  recheck used to acquire the registry mutex and walk every binding;
  `ordered_bindings_with_poison_snapshot` now takes the active bindings
  and the poisoned hook-id set under a single lock, and each loop
  threads a local `HashSet<HookId>` that absorbs mid-dispatch
  poisoning. Removed the redundant `is_poisoned` helper.
- Gate `HookDispatcher::registry_for_test` behind `cfg(any(test,
  feature = "test-support"))`. The accessor previously exposed
  `&Mutex<HookRegistry>` in production, letting any `Arc<HookDispatcher>`
  holder lock and call `HookRegistry::poison` to disable installed
  hooks. Added `active_bindings_snapshot(point)` as the read-only
  production-safe replacement.
- `#[serde(deny_unknown_fields)]` on every hook-manifest and predicate
  DTO (`HookManifestEntry`, `HookManifestBody`, `WasmBudget`,
  `HookPredicateSpec`, `CapabilityPredicate`, `ValueOrRateBound`,
  `OnExceededAction`). Typoed or unsupported fields (e.g. a
  manifest-supplied `trust_class`) now fail loud at install time
  instead of being silently dropped.
- Bound predicate trees at install. New
  `validate_predicate_tree` enforces `MAX_PREDICATE_DEPTH = 8`,
  `MAX_PREDICATE_NODES = 64`, `MAX_PREDICATE_STRING_BYTES = 256`, and
  `MAX_MANIFEST_REASON_BYTES = 512`. A hostile registry manifest can no
  longer install a deep or huge `All`/`Any` tree that the evaluator
  would recursively walk on every match.
- Cap sliding-window samples per key. `MAX_SAMPLES_PER_KEY = 4_096` in
  the predicate evaluator. Both the invocation-count and numeric-sum
  histories drop the oldest sample once the cap is reached, bounding
  memory under attacker-triggered hot capabilities while preserving
  rate/value-cap semantics over the most recent window.
- `split_indexer` / `resolve_path` now fail closed on malformed bracket
  syntax (`amount[foo]`, `amount[`, trailing garbage). Previously they
  silently fell back to the parent field, which could let a typoed
  `NumericSum` predicate evaluate against the wrong value and allow
  calls the predicate would otherwise have denied.
- Honor `PatchOrdinalHint`. `WrappedHookMessage` carries the source
  patch's `ordinal_hint`; `HookedLoopPromptPort` inserts `NearTop`
  messages after the bundle's `identity_message_count` and appends
  `Last` messages at the end. Safety/policy snippets that need early
  placement now get it.
- Update `ironclaw_hooks` top-level docs to reflect the four trust
  classes (`Builtin`/`Trusted`/`Installed`/`SelfAuthored`) and the
  now-wired Reborn middleware composition.

Tests added:
- `manifest::rejects_unknown_top_level_field`
- `manifest::rejects_unknown_wasm_budget_field`
- `manifest::rejects_predicate_tree_exceeding_max_depth`
- `manifest::rejects_predicate_tree_exceeding_max_nodes`
- `manifest::rejects_predicate_string_exceeding_max_bytes`
- `manifest::rejects_manifest_reason_exceeding_max_bytes`
- `points::capability::malformed_indexer_returns_none_not_parent_value`
- `telemetry::sanitize_audit_reason_*` (truncate / strip control /
  preserve / empty)

`cargo fmt`, `cargo clippy --all --benches --tests --examples
--all-features`, and `cargo test -p ironclaw_hooks` all pass clean.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(hooks): batch deferred test coverage from #3573 review (#3914)

* perf(hooks): defer capability input resolution until a predicate needs it (#3913)

* fix(rebase): adapt hooks tests + middleware to upstream API additions

- CapabilityDescriptorView: add parameters_schema field
- LoopModelRequest / LoopPromptBundleRequest: add capability_view field
- TimelineEntry test builder: add hook_id / hook_point / hook_trust_class /
  hook_decision / hook_failure_category / hook_failure_disposition fields
- ironclaw_reborn::tests::hooks_integration: switch from
  InMemoryLoopCheckpointStore to InMemoryTurnStateStore (which now
  impls both LoopCheckpointStore and TurnStateStore), pass TurnActor
  in TurnRunState, supply the new turn_state_store factory arg
- ironclaw_reborn lib.rs: drop the pub-use re-exports that upstream
  intentionally removed (per the module-directory rationale in the
  current ironclaw_reborn lib.rs doc comment); update the
  hooks_integration test imports to use module paths
- Cargo.toml: union the hooks-foundation member list with upstream's
  new crates (event_streams, auth, first_party_extensions,
  reborn_webui_ingress, product_workflow_storage, webui_v2); drop
  ironclaw_storage which no longer exists upstream
- crates/ironclaw_architecture/tests/reborn_dependency_boundaries:
  keep upstream's removal of ironclaw_filesystem from the ironclaw_turns
  forbidden list AND add ironclaw_hooks to that list
- crates/ironclaw_reborn/src/milestone_events.rs: drop dead loop_failure_kind
  helper (replaced upstream by loop_failure_kind_name in text_loop_driver.rs);
  keep hook_decision_label which is still used

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* perf(hooks): restore batched capability dispatch when hooks acti…
theredspoon pushed a commit that referenced this pull request Jun 18, 2026
…3920)

* Implement installed WASM hook runtime

Adds crates/ironclaw_hooks/docs/threat-model-wasm.md and follows the reviewed design ack: 1) module bytes are resolved, digest-cached, and compiled in the tool-WASM style while reusing its resource limiter; 2) each invocation gets a fresh wasmtime Store; 3) the ABI is a wasmtime::Linker surface, not wit-bindgen; 4) host-import sink shims enforce call, patch-byte, observer-fact, and decision budgets.

* Harden WASM hook string and metadata budgets

* fix(hooks): validate WASM hook ABI at install time (serrrfirat #3 on PR nearai#3634)

Address serrrfirat MEDIUM finding #3: `WasmHookRuntime::prepare()` compiled
and cached module bytes but did not validate imports or the requested
export. ABI mismatches (unsupported import, missing export, wrong export
signature) were deferred to first live dispatch — and the prior
`wasm_unsupported_host_import_fails_closed` test codified that a
bad-import module would install successfully and only fail closed at
invocation. Malformed untrusted modules should never reach live traffic.

Changes:
- `prepare()` derives the target hook point from `request.kind`, then
  runs `validate_module_abi()`: scratch-instantiate the module against
  the point-specific linker (catches unsupported / wrong-type imports)
  and resolve the typed export `() -> ()` (catches missing export and
  wrong signature). Failures surface as new
  `WasmHookRuntimeError::InvalidImports` or existing
  `WasmHookRuntimeError::InvalidExport`, both of which bubble up as
  `HookError::RegistryConstruction` from the registrar.
- `wasm_point_for_kind(HookManifestKind)` helper centralizes the
  kind → wasm-point mapping; the previous `execute_*` paths can share
  it in a follow-up but kept inline for now to minimize churn.

Tests:
- `wasm_unsupported_host_import_is_rejected_at_install_time`: replaces
  the prior test that codified late-failure behavior; asserts the
  registrar returns `RegistryConstruction` citing the bad import.
- `wasm_missing_export_is_rejected_at_install_time`: new module that
  compiles but lacks the manifest-declared export; same install-time
  rejection.

* fix(hooks): address henrypark133 must-fix #1, #2, #3 on PR nearai#3634

Three items from the 5-15 review:

**#1 (must-fix) Extract ironclaw_wasm_limiter micro-crate**
Replace `#[path = "../../../ironclaw_wasm/src/limiter.rs"]` cross-crate
file import with a proper Cargo edge. The 111-line `WasmResourceLimiter`
moves into a new `crates/ironclaw_wasm_limiter` micro-crate that both
`ironclaw_wasm` and `ironclaw_hooks` depend on. The architecture rule
forbidding `ironclaw_hooks -> ironclaw_wasm` is preserved (the new
crate sits below both consumers and pulls in only `wasmtime` +
`tracing`); `cargo check`, `cargo doc`, and architecture-linting tests
now see the edge, and the file can't be moved out from under one of
the consumers silently.

Mechanical changes:
- new `crates/ironclaw_wasm_limiter/` (Cargo.toml + src/lib.rs with the
  type exposed as `pub` instead of `pub(crate)`)
- workspace `members` entry added
- `crates/ironclaw_wasm/src/limiter.rs` deleted
- `crates/ironclaw_wasm/src/lib.rs`: `mod limiter` removed
- `crates/ironclaw_wasm/src/store.rs`: import switched to
  `ironclaw_wasm_limiter::WasmResourceLimiter`
- `crates/ironclaw_wasm/Cargo.toml`: dep added
- `crates/ironclaw_hooks/Cargo.toml`: dep added
- `crates/ironclaw_hooks/src/wasm/runtime.rs`: `#[path = ...]` block
  removed; import switched to the crate

**#2 + #3 (must-fix) Dead WASM arms in dispatch**
`run_before_capability_hook`, `run_before_prompt_hook`, and
`run_observer_hook` each had an early-return guard that dispatched
WASM hooks with `catch_unwind` + timeout, then ALSO had a matching
WASM arm in the inner `match` that ran without those protections. The
prompt-path arm additionally swallowed `WasmHookFailure` via `|_| ()`,
making the must-fix #2 problem worse on that path specifically.

If a future refactor removed any of the early-return guards, those
inner arms would silently take over and drop panic isolation, deadline
enforcement, AND (for prompts) the failure category. Replaced each
inner arm with `unreachable!()` carrying a comment that explains
why the arm exists and references the early-return guard above it.
A future refactor that removes the guard will now trip the
`unreachable!` at first call instead of silently degrading.

All 154 hooks lib + 29 reborn integration tests still pass.

* fix(hooks): plumb context to WASM hooks + runtime hardening

Critical #1 on PR nearai#3634: WASM hooks previously received no context. The
`execute_*` entry points dropped the `&BeforeCapabilityHookContext` /
`&BeforePromptHookContext` / `&ObserverHookContext` value and invoked
the guest export with `()`, so a WASM gate could never decide based on
the capability name, tenant, provider, or other dispatch-time facts. Add
an `ic:hooks/context@1` host-import module exposing two read-only
calls — `ctx_size() -> i32` and `ctx_read(ptr, len) -> i32` — backed by
a JSON-serialized blob the dispatcher writes per-invocation into the
fresh store. Modules that don't import these continue to link; modules
that do import them get a stable, non-empty payload to read. An
integration test (`wasm_before_capability_hook_reads_context_blob`)
asserts the contract end-to-end: a guest that fails to read a non-empty
blob traps before its `deny` call.

Also rolls up the other reviewer-flagged WASM runtime issues, all of
which touch `wasm/runtime.rs`:

HIGH #2: epoch-tick background thread now holds a shutdown
`AtomicBool` and joins on `Drop`. Previously it looped forever and
leaked an Engine clone on every runtime drop.

MED #4: compiled-module cache is now an `lru::LruCache` bounded by
`MODULE_CACHE_CAPACITY = 128`. Replaces the unbounded `HashMap`.

MED #7: `prepare()` no longer compiles under the cache lock. Fast
path reads from LRU under a brief lock; slow path compiles outside
the lock and re-checks on insert to avoid the TOCTOU window where
two concurrent installs of the same module both compile.

Bug #9: post-call `deadline_exceeded()` re-check on the Ok branch
is gone. wasmtime epoch-interrupt is the authoritative wall-clock
signal; an Ok return is no longer reclassified as a timeout because
the wall ticked over during host-side return.

Bug #10: `add_milestone_metadata` returns a distinct
"metadata value exceeds the u32 byte-length ceiling" error when the
guest-supplied `value.len()` overflows u32, instead of misreporting it
as "exceeded total prompt-patch byte budget".

Existing integration tests for WASM hooks are also re-wired through
`HookRegistrar::with_verified_grants` so the grants-store gate added in
the foundation-01 merge stops failing the pre-existing fixtures.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(hooks): run WASM hooks on the blocking pool

HIGH #3 on PR nearai#3634: `tokio::time::timeout` does NOT cancel synchronous
wasmtime execution. The previous code awaited a `catch_unwind(async { h.evaluate(ctx) })`
future whose body completed in one poll, so the timeout could only fire
*around* the WASM call rather than against it; a hook that wedged inside
wasmtime simply pinned the calling tokio task.

Route gate, prompt, and observer WASM dispatch paths through
`tokio::task::spawn_blocking` via a shared `run_wasm_blocking` helper.
The outer `tokio::time::timeout` now governs the JoinHandle, so a stuck
blocking task stops blocking the dispatcher's caller; the wasmtime
epoch interrupt configured in the runtime (10 ms tick) is the
authoritative in-WASM wall-clock cancel signal. JoinError (panic in
the blocking task) maps to `FailureCategory::Panic`, matching the
pre-existing semantics for synchronous panics caught via
`catch_unwind`.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* perf(hooks): O(1) hook-id lookup via side index

Finding #8 on PR nearai#3634: `set_priority`, `poison`, `is_poisoned`, and
`contains_hook` all did full-registry scans over every binding at every
point. Each is called per-dispatch (poison-checks on the snapshot loop
in particular), so the cost is `O(registered_hooks)` per
`(installed_hook, registered_hook)` pair.

Maintain a denormalized `HashMap<HookId, (HookPointSpec, usize)>` side
index in lock-step with `by_point` so every per-hook-id operation
becomes a single hash lookup + a direct vec indexed access. The
duplicate-id rejection in `insert` now reads from the side index too,
turning what used to be a flat-map scan into a `HashMap::contains_key`.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(hooks): wall-clock timeout, observer memory, limiter rollback, registrar happy path

Round out the test set for the WASM hook execution path:

#11 / #12: gate + observer wall-clock timeout. The pre-fix dispatcher
ran wasmtime synchronously on the executor, so the outer
`tokio::time::timeout` `Err(_elapsed)` arm was effectively unreachable.
Now that WASM execution runs on the blocking pool, the timeout actually
fires; the new tests give the wasm budget headroom (1B fuel, 5s wall)
and the dispatcher a 20 ms timeout, then assert the failure
classification (FailClosed for gate, FailIsolated for observer).

#13: observer memory exhaustion. Mirrors
`wasm_memory_exhaustion_fails_closed_for_gate` against the observer
dispatch path so the FailIsolated branch of the failure matrix has
explicit memory coverage, not just fuel/wall.

#15: `WasmResourceLimiter::memory_grow_failed` rollback. Stages an
approved grow, simulates the OS-level grow failing, and asserts a
subsequent grow of the full ceiling succeeds — the inflated
`memory_used` from the failed attempt must be released.

#16: registrar WASM happy path. Companion to the existing
`install_wasm_body_requires_runtime` negative case: a valid module
installs, the binding is visible via the public registry accessor, and
is not pre-poisoned.

#14 (`add_milestone_metadata` happy path) is intentionally omitted —
the BeforePrompt dispatch path is currently unreachable due to a
pre-existing manifest-vs-registry scope conflict (`OwnCapabilities` is
the only valid `BeforePrompt` scope per manifest validation, but the
registry rejects `OwnCapabilities` at `BeforePrompt` because the point
has no provider context). That contradiction sits outside this PR's
scope; flagging for a follow-up.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(hooks): typed WASM version material, reconcile design doc

LOW #20 on PR nearai#3634: extract the
`{extension_version}+wasm:{module_digest_hex}` concatenation into a
`WasmVersionMaterial` newtype with a single `Display` impl. The
identity material no longer floats free as a stringly-typed argument
inside the registrar.

Reconcile `docs/successors/02-wasm-runtime.md` with the implementation:

- Spell out that wall-clock cancellation depends on the
  `tokio::time::timeout(tokio::task::spawn_blocking(...))` pair, and
  explain why a bare timeout over a synchronous wasmtime call cannot
  actually cancel.
- Define `FailIsolated` and `FailClosed` as `FailureDisposition`
  values, distinct from the older `HookFailureMode::{FailOpen,
  FailClosed}` policy switch that applies to predicates.
- Clarify the generic `evaluate` export contract — name is whatever
  the manifest declares, signature is `(): ()`, context arrives
  through the new `ic:hooks/context@1` host imports — and note the
  intentional divergence from `WitToolRuntime`'s hardcoded interface.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(hooks): drop .expect() in WASM module cache capacity

Pre-commit no-panics CI flagged the .expect() on the LruCache capacity.
Move the validity check to a const match, so the NonZeroUsize is fixed at
compile time and the no-panics regex is satisfied.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(hooks): use HookLocalId::new after newtype privatization

The newtype-privatization landed in reborn-integration after the
hooks-fu-wasm-runtime branch's WASM scaffolding tests were written;
update the affected test/registrar sites to use HookLocalId::new
instead of the now-private tuple constructor.

* style: cargo fmt after newtype-privatization fixups

* test(hooks): ignore 3 BeforePrompt WASM tests with manifest/registry conflict

These tests were failing on the original branch tip too (verified against
origin/hooks-fu-wasm-runtime @ 571efdf). The Installed-tier BeforePrompt
WASM install path has no valid scope today:
  - OwnCapabilities is rejected by the registry C3 check (finding #2 on
    PR nearai#3573) since BeforePrompt has no per-capability invocation
    context.
  - SameTenant is rejected by manifest validation ("cannot combine
    scope = same_tenant with kind = before_prompt").

The budget-overflow paths these tests exercise are point-agnostic; the
follow-up is to either rewrite the helper to install through
BeforeCapability or add a Global manifest scope. Tracked as a deferred
item on the new PR.

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
theredspoon pushed a commit that referenced this pull request Jun 18, 2026
…#3899)

* Reborn budgets: address all nearai#3841 follow-ups end-to-end

Implements every open follow-up from PR nearai#3841 (cost-based budgets
foundation), driven by the plan in
`docs/plans/2026-05-22-reborn-budgets-followups.md`:

- **C2 (provider tokens)**: `LoopModelResponse.usage` carries real
  `(input_tokens, output_tokens)` from `CompletionResponse` /
  `ToolCompletionResponse`; `usage_for_response` reconciles to actual
  USD via the cost table instead of the conservative estimate.
- **D1 (cascade warnings)**: `CascadeOutcome` variants carry
  `Vec<BudgetWarning>` so warnings preceding a pause or hard deny
  reach the audit sink. `ResourceError::LimitExceeded` /
  `RequiresApproval` reshaped to struct variants.
- **C1 (cancellation safety)**: new
  `LoopModelBudgetAccountant::release_in_flight` trait hook + RAII
  `ReservationReleaseGuard` in `HostManagedLoopModelPort::stream_model`
  so a cancelled future doesn't orphan its reservation.
- **E1 (dead code)**: removed the never-set `budget_accountant` field
  on `ThreadBackedLoopModelPort`.
- **Real cost table**: new `StaticModelCostTable` +
  `LlmModelProfilePolicy::build_cost_table()` populated from
  `ironclaw_llm::costs::model_cost` with `default_cost` fallback so
  unknown providers never silently reconcile to zero.
- **B1 (filesystem gate store)**: new `FilesystemBudgetGateStore`
  mirroring `FilesystemResourceGovernorStore`; pending gates survive
  process restart.
- **A1 (production wiring)**: composition builds
  `GovernorBackedAccountant` from the cost table + governor and
  threads it through `RebornLoopDriverHostFactory::with_model_budget_accountant`.
- **A2 (audit / SSE projection)**:
  `InMemoryResourceGovernor::with_event_sink` emits `Reserved`,
  `Reconciled`, `Released`, `Warned`, `ApprovalRequested`, `Denied`,
  `LimitChanged`; composition holds an `InMemoryBudgetEventSink` ready
  for downstream SSE projection.
- **F1 (stuck-loop normalization)**: `CapabilityCallSignature::from_call`
  now runs `progress::normalize_for_hash` so the existing repetition
  window collapses request-id / UUID / timestamp noise.

Side fix: `ResourceValue` moved to adjacent serde tagging (the
combination of internal tagging + `Decimal`'s `serde-with-str`
representation breaks JSON serialization — rust-lang/serde#1402).

Regression tests added per item — see the acceptance evidence appendix
in the plan doc.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Reborn budgets: end-to-end test coverage via test-support feature

Adds 13 e2e tests covering the budget pipeline through
`build_reborn_runtime` + `send_user_message`. Required infrastructure:

- **`test-support` feature** on `ironclaw_reborn_composition` exposing
  `BudgetTestGateway` (scripted token usage) and
  `RebornRuntimeInputTestExt`. Existing `model_gateway_override` field
  promoted from `#[cfg(test)]` to `#[cfg(any(test, feature = ...))]`
  with a new public `with_model_gateway_override_for_tests` setter.
- **Cost-table override** on `RebornRuntimeInput` so tests can pair
  the gateway with a deterministic `ModelCostTable`. Without this, an
  override gateway dropped the cost table and the accountant never
  fired.
- **Budget accessors** on `RebornRuntime`: `budget_resource_governor`,
  `budget_event_sink`, `budget_gate_store`, and
  `apply_resolved_budget_gate`. Test-feature gated.
- **`ResourceGovernor::usage_for`** added as a default-impl trait
  method so tests read spend through the trait surface.
- **`BudgetGateStore` wired into the accountant**:
  `GovernorBackedAccountant::with_gate_store(...)` opens a pending
  gate whenever the governor cascade returns `RequiresApproval`. The
  approval-required host error is unchanged; the gate is the
  out-of-band channel a user-facing handler resolves.

Scenarios covered:

| # | Test | What it asserts |
|---|---|---|
| F1 | `f1_happy_path_records_actual_usd_in_ledger` | Ledger depletes by provider tokens × cost table |
| F2 | `f2_crossing_warn_threshold_emits_warned_event` | Warn fires alongside successful Reserved/Reconciled |
| F3 | `f3_approval_with_increased_limit_unblocks_retry` | Approve → set_limit applies → retry succeeds |
| F4 | `f4_cancel_keeps_budget_blocked_on_retry` | Cancel → retry still short-circuits |
| F5 | `f5_expiry_marks_gate_terminal_and_keeps_budget_blocked` | Expiry → gate drops from pending list, retry still blocked |
| F6 | `f6_hard_cap_denied_before_provider_call` | Estimate over cap → zero model calls, Denied event |
| C1 | `c1_provider_tokens_reconcile_to_actual_usd` | Real numbers, not estimate |
| C2 | `c2_unknown_model_in_cost_table_reconciles_to_zero` | Unknown profile → zero spend |
| C3 | `c3_zero_cost_model_records_zero_spend` | Free model → zero USD with non-zero tokens |
| D1 | `d1_agent_deny_preserves_user_warn_event` | Cascade emits both Warned and Denied |
| D3 | `d3_fresh_user_without_limits_runs_without_denial` | No limit → no denial |
| + | `pause_in_distinct_runs_produces_distinct_pending_gates` | Per-run gate identity |
| + | `budget_test_gateway_scripted_replies_drive_per_turn_costs` | Multi-turn scripted accumulation |

F7 (cancellation mid-stream) is unit-covered by
`release_in_flight_drains_orphan_reservation_on_cancellation`.
D2 (period rollover) is unit-covered by
`rolling_24h_snapshot_reports_anchored_window_not_now_window`.
B-series (background ticks) await the BackgroundKind scheduler
call site (no production caller in Reborn yet).

Run via `cargo test -p ironclaw_reborn_composition --features test-support`.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Budget review feedback: address all 7 findings from PR nearai#3899 review

Two High and five Medium issues raised by serrrfirat's multi-agent review.

**High #1 — `FilesystemBudgetGateStore` cross-tenant leakage**
The store hardcoded `ResourceScope::system()` for every op, so all
tenants wrote into the same `/tenants/__SYSTEM__/...` snapshot and
`list_pending` would expose gates across tenants. Fix: `new(...)` now
takes a `ResourceScope`; each tenant gets its own store, and the
`ScopedFilesystem` mount view routes the snapshot under that tenant's
path. Added `list_pending_does_not_leak_across_tenants` regression.

**High #2 — accountant wired without default budget limits**
Composition built `GovernorBackedAccountant` without
`with_seeding_policy`, so the local-dev governor started empty and
`reserve_with_outcome_in_state` skipped accounts that had no
configured limit — model calls reconciled spend but never enforced a
cap. Fix: `build_reborn_runtime` now loads
`BudgetDefaults::compiled_defaults().with_env()` and wires
`BudgetSeedingPolicy` + `with_overestimate_factor`. Renamed the D3
test to `d3_seeding_policy_installs_default_cap_on_first_touch` to
prove the wiring fires.

**Medium #3 — RAII guard disarmed before post_model_call await**
`HostManagedLoopModelPort::stream_model` was disarming the
`ReservationReleaseGuard` before awaiting `post_model_call`. A
cancellation during that await dropped the future without cleanup,
orphaning the reservation. Fix: disarm AFTER `post_model_call`
returns. `release_in_flight` is now idempotent (peek-then-release-
then-remove) so a successful post-call + subsequent guard drop is a
no-op.

**Medium #4 — failed release drops the retry handle**
`release_in_flight` removed the in-flight entry before calling
`governor.release`. A transient storage error left the reservation
active in the governor with the id discarded. Fix: peek first,
release, only remove on success. Errors keep the entry retained for
a future retry / cleanup hook.

**Medium #5 — unknown model silently reconciles to zero USD**
Both `estimate_for` and `usage_for_response` fell back to
`ModelCost { 0, 0, 0 }` when the cost table had no entry for the
effective model. Cost-table drift would silently bypass daily caps.
Fix: `GovernorBackedAccountant` carries a `default_cost` (default ~
GPT-4o pricing, ~`$0.0000025 input + $0.00001 output per token`) used
for unknown models. Callers wiring `ZeroCostTable` for free / Ollama
explicitly opt out of the fallback. Updated the C2 e2e test to
assert the new fail-closed shape.

**Medium #6 — paused dimension lost when another hard-denies**
`check_thresholds_all_interventions` stored `Approval` only in the
`approval` slot, so when one dimension paused and another hard-denied,
the `Deny { warnings, denial }` outcome lost the pause signal.
Fix: also push a warning-shaped record for the paused dimension.

**Medium #7 — unbounded terminal-gate retention**
The snapshot kept every gate forever; `open` / `resolve` / `get` /
`list_pending` were O(total historical gates). Fix:
`with_terminal_retention` (default 30 days). Every mutation prunes
terminal gates whose resolution timestamp is older than the window.
Added `terminal_gates_older_than_retention_are_pruned_on_next_write`
regression.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* ci: replace lock-poisoned expects with PoisonError::into_inner

scripts/check_no_panics.py flagged five .expect("...lock poisoned")
calls in the new test_support.rs. Use the same idiomatic recovery
pattern the rest of the codebase uses (see InMemoryBudgetGateStore,
InMemoryBudgetEventSink): on a poisoned lock, recover the inner data
via PoisonError::into_inner rather than panicking. The test gateway's
state is append-only logs / replies queues, so reading them through a
poisoned lock is safe.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Finish A1 / A2 / F1 from plan + honest plan doc update

The plan claimed "all nine items landed" but A1 (production wiring),
A2 (SSE projection), and F1 (full progress strategy) were partials.
This commit finishes the work so the plan matches reality.

**A1 — production-shape accountant builder**

New `ironclaw_reborn_composition::build_default_budget_accountant`
public helper that wires the seeding policy + overestimate factor +
gate store from `BudgetDefaults::compiled_defaults().with_env()` and
returns an `Arc<dyn LoopModelBudgetAccountant>`. Production loop
composers call this with their `PersistentResourceGovernor` +
`FilesystemBudgetGateStore` + LLM-policy-derived cost table; the
local-dev runtime in `build_reborn_runtime` now uses the same helper
instead of duplicating the seeding logic inline. Unit-tier regression
`seeds_compiled_default_user_cap_on_first_touch` proves the helper
installs the compiled-default $5 user cap on first model call.

**A2 — broadcast sink + AppEvent projection**

- `ironclaw_resources::BroadcastBudgetEventSink` wraps
  `tokio::sync::broadcast::Sender<BudgetEvent>` with `subscribe()` /
  `subscriber_count()`. `CompositeBudgetEventSink` fans events to
  multiple sinks.
- Composition fans every `BudgetEvent` to the in-memory sink (for
  tests) AND the broadcast sink (for SSE projection) via
  `CompositeBudgetEventSink`.
- New `AppEvent::BudgetWarn` / `BudgetPause` / `BudgetDenied` /
  `BudgetLimitChanged` wire-stable variants in
  `ironclaw_common::event`.
- `src/bridge/budget_events.rs` carries the projection: a tokio task
  spawned by `spawn_budget_event_projection` drains the broadcast
  receiver and emits the appropriate `AppEvent` via
  `SseManager::broadcast_for_user`. System-scoped events (no user
  identity) are skipped. This is the only producer of these
  `AppEvent` variants per `.claude/rules/gateway-events.md`.
- `RebornRuntime::broadcast_budget_event_sink()` exposes the sink to
  the binary so the startup path subscribes. E2E test
  `broadcast_sink_publishes_events_to_subscribers` drives a real
  `send_user_message` and asserts Reserved + Reconciled lands on the
  broadcast.

**F1 — diminishing-returns stop condition**

The earlier shipped `ParamHash` normalization in
`CapabilityCallSignature` strengthened the existing
`recent_call_signatures`-based repetition detector. This commit adds
the second half of F1: a rolling output-token window that detects
"wedged" loops the repetition detector misses (model keeps
responding but produces no useful output).

- `LoopExecutionState.recent_output_token_counts: BoundedRing<u32, 8>`
  populated by the executor from `LoopModelResponse::usage`.
- `BoundedRing::iter` returns `impl DoubleEndedIterator` so the
  strategy can scan the trailing window.
- `DefaultStopConditionStrategy` gets `min_delta_tokens` (default
  4) + `noprogress_window` (default 4). When the last N turns all
  produce ≤ min_delta_tokens of output, fire
  `StopKind::NoProgressDetected`.
- Regression tests:
  `four_consecutive_low_token_turns_trigger_no_progress` proves the
  detector fires; `occasional_low_token_turn_does_not_trip_no_progress`
  proves a productive turn resets the trailing count.

**Plan doc**

Updated the status header from "all nine items landed" to the
honest per-item shape. Acceptance evidence table expanded with the
new test names. New "Review-feedback fixes layered on top" subsection
documenting all 2 High + 5 Medium findings addressed during review.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Address thermo-nuclear review: collapse filesystem-store duplication, flatten cfg permutations, split budget accountant

Five structural simplifications surfaced by the deep audit on PR nearai#3899, plus
two bug fixes from the earlier review pass:

- ironclaw_resources: extract `cas_snapshot` shared infrastructure
  (`StorageError` + `Snapshot` traits + `CasSnapshotStore<F>` + async-runtime
  worker + per-path lock map) and merge `filesystem_gate_store.rs` into
  `filesystem_store.rs`. Deletes ~350 lines of duplicated read-modify-write +
  worker-thread + CAS machinery; both stores are now thin shims over the
  shared helper.

- ironclaw_reborn_composition: flatten the 4-way cfg permutation in
  `build_reborn_runtime` model-gateway resolution into three flat steps
  (normalize override → build production gateway via cfg-gated helper →
  test override wins). Also drops the `unused_mut` warning.

- ironclaw_reborn_composition: collapse the 3-layer test-only setter dance
  for `model_gateway_override` / `model_cost_table_override` into a single
  setter pair gated on `cfg(any(test, feature = "test-support"))`. Deletes
  the `RebornRuntimeInputTestExt` extension trait — integration tests now
  call the inherent methods directly.

- ironclaw_loop_support: split the 1305-line `budget_accountant.rs` into
  `budget_cost_table.rs` (ModelCost/ModelCostTable trait/ZeroCostTable/
  StaticModelCostTable), `budget_seeding.rs` (BudgetSeedingPolicy), and
  `budget_accountant.rs` (just GovernorBackedAccountant). Each module now
  owns one concern.

- ironclaw_resources: add `impl Display for ResourceAccount` and route the
  hierarchical account-label rendering through it; delete the 60-line
  bespoke `account_label` helper from `src/bridge/budget_events.rs`.

- ironclaw_common + bridge: collapse the four `AppEvent::Budget*` variants
  into a single typed `AppEvent::Budget(AppBudgetEvent)` with the four
  shapes carried inside the enum. Wire-shape stays identical (snake_case
  serde tag).

- ironclaw_resources + ironclaw_loop_support: thread real gate id through
  `BudgetEvent::GateOpened { gate_id, needed, at }` (new variant) and have
  the accountant emit it via the broadcast event sink after store.open
  succeeds. The bridge now projects `BudgetEvent::GateOpened` (not
  `ApprovalRequested`) into `AppEvent::Budget(Pause { gate_id, ... })` so
  SSE consumers receive the persisted gate id rather than a fabricated
  zero uuid.

- ironclaw_agent_loop: in the F1 token-counting path, push to
  `recent_output_token_counts` only when the model response carries
  `Some(usage)` and only on the `AssistantReply` arm (instead of
  `unwrap_or(0)`). Diminishing-returns detection now reflects real spend.

Net delta: -461 lines (+999 / -1460). Workspace `cargo clippy` clean,
`cargo test` clean on ironclaw_resources / ironclaw_loop_support /
ironclaw_reborn_composition; budget_e2e + budget_approval_e2e both green.

Pre-existing CI failures (`cli::tests::test_version` stack overflow,
`facade_factory::production_*` RuntimeProcessPort missing) are unrelated
and reproduce on the pristine branch tip.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(cli): refresh insta snapshots after runtime-policy flag additions

The `import`-feature variants of the help snapshots were left stale when
`--deployment-mode`, `--runtime-profile`, `--yolo-disclosure` were added in
cc04481 (nearai#3243); the `_without_import` variants were updated but these
were not. CI was failing the snapshot assertion under the slim PR matrix
(`--features postgres,libsql,html-to-markdown,bedrock,import`).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Address PR nearai#3899 thermo-nuclear review (TN #1, #2, #3)

TN #1 — budget defaults resolved in wrong layer:
  - `build_default_budget_accountant` no longer reads process env; it
    now takes `&BudgetDefaults` as a parameter and the caller owns the
    config-layer precedence (compiled → section → env) plus the
    `validate()` call.
  - `RebornRuntimeInput` gains an optional `budget_defaults` field +
    `with_budget_defaults()` builder so the composition root passes a
    pre-resolved value. `build_reborn_runtime` falls back to
    `compiled_defaults().with_env() + validate()` when none is supplied
    so existing call sites keep working.

TN #2 — gate-store scoping at wrong boundary:
  - `BudgetGateStore` trait methods (`open`, `resolve`,
    `expire_pending_older_than`, `get`, `list_pending`) now take
    `&ResourceScope` as first arg. `GovernorBackedAccountant` passes
    the caller's scope from `resource_scope(context)`.
  - `CasSnapshotStore` gains `update_with_scope` so the same store
    instance can route per-operation. `FilesystemBudgetGateStore` no
    longer takes scope at construction — one shared instance serves
    every tenant via the `ScopedFilesystem` mount view.
  - `InMemoryBudgetGateStore` ignores scope (suitable for single-tenant
    tests / local-dev); production multi-tenant filesystem path is
    correctly partitioned by `ResourceScope`.
  - `RebornRuntime::apply_resolved_budget_gate` now takes scope too.

TN #3 — half-wired projection bridge:
  - Removed `src/bridge/budget_events.rs`, its `spawn_budget_event_projection`
    helper, the `AppEvent::Budget` variant, and the `AppBudgetEvent`
    type. No production caller ever subscribed the broadcast sink
    onto SSE and no frontend consumed the variant, so the
    half-wired bridge is gone pending a real owner that spawns a
    projection task with shutdown cancellation.
  - The runtime's `broadcast_budget_event_sink()` accessor stays so
    a future production composer can still subscribe without
    rebuilding the runtime.

Bonus — to keep budget e2e tests working under the new libsql local-
dev path that origin/reborn-integration introduced, added
`PersistentResourceGovernor::with_event_sink` (parity with the
`InMemoryResourceGovernor` accessor). The libsql variant of
`build_local_dev_store_graph` now wires the composite sink to the
persistent governor so governor-emitted `Warned`/`Denied`/`Reserved`/
`Reconciled` events reach subscribers on both feature paths.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Wire budget-event projection task into RebornRuntime

Re-implements PR nearai#3899 follow-up A2 / Thermo-Nuclear #3 with a real
production owner instead of leaving the broadcast sink half-wired:

- `crates/ironclaw_reborn_composition/src/budget_events.rs` (new):
  `BudgetEventObserver` trait + `TracingBudgetEventObserver` default
  observer + crate-internal `BudgetEventProjection` task that drains
  the runtime's broadcast `Receiver<BudgetEvent>` and forwards every
  event to the observer. Cancellation via `CancellationToken`; lagged
  subscribers logged and resumed; receiver-closed exits cleanly.

- `RebornRuntimeInput::with_budget_event_observer(...)` lets
  production owners install a custom observer (SSE projection, WS
  fan-out, telemetry export). When unset, the runtime installs the
  tracing observer so events always surface in structured logs.

- `build_reborn_runtime` always spawns the projection task at runtime
  construction; `RebornRuntime::shutdown` cancels it and awaits the
  handle so background state drains before the runtime drops.

- E2E test `projection_delivers_budget_events_to_installed_observer`
  drives `build_reborn_runtime` with a capturing observer and asserts
  the observer sees `Reserved` + `Reconciled` from a real model call,
  testing through the caller per `.claude/rules/testing.md`.

- Existing `broadcast_sink_publishes_events_to_subscribers` updated
  to expect the runtime's own projection task as a baseline
  subscriber.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(reborn): rustfmt the merged loop_support import block

The conflict resolution for the post-merge import list was not run
through rustfmt; CI Formatting flagged the wrapping. No logic change.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
theredspoon pushed a commit that referenced this pull request Jun 18, 2026
…r injection (nearai#4588)

* feat(reborn): expose a trajectory observer hook on RebornRuntimeInput

The reborn runtime is sealed: build_reborn_runtime returns only the final
AssistantReply, and per-step capability (tool) calls + results live in internal
stores. Downstream consumers (benchmark harnesses, UI/debuggers) can't observe
the agent's trajectory.

Add `RebornTrajectoryObserver` (pub trait: on_capability_input(call_id, name,
args) / on_capability_result(call_id, output)) and
`RebornRuntimeInput::with_trajectory_observer`. The local-dev capability IO
(`LocalDevCapabilityIo`) forwards each tool call's name+args (at input staging)
and result (at result write) to the observer when present — reusing the same
data it already records for display previews. No-op when unset; best-effort.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* debug: trace observer hook firing (temporary)

* feat(reborn): trajectory observer — capability_id on result, reliable spine

Provider tool calls are staged by a lower decorator that bypasses the
LocalDevCapabilityIo input path, so on_capability_input does not fire for
them. on_capability_result fires for every completed capability — make it
carry the capability_id so consumers can reconstruct the trajectory (name +
output) from results alone. Input args capture is a follow-up.

* feat(reborn): capture capability input args at the host port chokepoint

Provider tool calls are staged by ProviderToolCallInputResolver, which keeps
args in a private map and bypasses the capability-IO input hook — so inputs
never reached the trajectory observer (only results did). Move the observer
trait down to ironclaw_loop_support (CapabilityTrajectoryObserver, re-exported
from composition as RebornTrajectoryObserver) and hook it in
HostRuntimeLoopCapabilityPort::invoke_capability right after the input
resolves — the one place the model's resolved arguments are visible. Threaded
through HostRuntimeLoopCapabilityPortFactory + the local-dev factory. Result
hook unchanged. Now name + args + output are all captured.

* feat(reborn): host LLM-provider injection seam

ResolvedRebornLlm::with_provider — drive the runtime with a caller-supplied
LlmProvider (e.g. an instrumented wrapper that counts tokens/cost and captures
reasoning) instead of always building one from config; build_llm_gateway honors
the override. The only viable observability path for reborn, whose model calls
run in spawned worker tasks a per-task tracing subscriber can't reach.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(reborn): cover trajectory observer + LLM provider override seams

Addresses Firat's two blocking review findings on nearai#4588 (both
missing-integration-test, per AGENTS.md "test through the caller"):

1. Trajectory observer callbacks — drive the real call sites with a
   recording CapabilityTrajectoryObserver:
   - host port: invoke_capability via HostRuntimeLoopCapabilityPortFactory
     ::with_trajectory_observer asserts on_capability_input fires with the
     resolved capability id + tool-call arguments.
   - local-dev IO: register_provider_tool_call_input + write_capability_result
     assert on_capability_input and on_capability_result fire and correlate by
     input ref.

2. LLM provider override — build_llm_gateway_drives_provider_override_not_config
   injects a counting mock via ResolvedRebornLlm::with_provider, points config
   at a dead endpoint, and asserts the gateway returns the mock's sentinel
   (proving the override is driven, not a config-built chain).

Also fixes pre-existing breakage this surfaced: 5 LocalDevLoopCapabilityPort
Factory test initializers (shell_tests.rs + tests.rs) were missing the
trajectory_observer field added by this PR, so the composition crate's tests
did not compile under --features root-llm-provider.

loop_support: 301 passed; composition (root-llm-provider): 520 passed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reborn): make trajectory observer input semantics consistent

Addresses Copilot's follow-up findings on the observer seam:

- Drop the `on_capability_input` callback from `LocalDevCapabilityIo::
  register_provider_tool_call_input`. It forwarded the raw provider tool
  name (`builtin_echo`) as the capability id — conflicting with the
  observer contract (resolved dotted `builtin.echo`) and the authoritative
  port-level hook — and `ProviderToolCallInputResolver` doesn't delegate
  here for provider tool calls, so it never fired in practice anyway.
  `HostRuntimeLoopCapabilityPort::invoke_capability` remains the single
  source of `on_capability_input` (resolved id); `LocalDevCapabilityIo`
  remains the source of `on_capability_result`.

- Clarify the trait doc: `arguments` is the raw model-emitted tool-call
  input resolved from the input ref (the callback fires before schema
  normalization), which is what the trajectory should record.

- Refocus the local-dev test on `on_capability_result` forwarding +
  correlation, and assert input staging does NOT emit `on_capability_input`
  from local-dev IO. Port-level input semantics stay covered by the
  capability_port.rs test.

loop_support: 301 passed; composition (root-llm-provider): 520 passed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* wire trajectory_observer through RefreshingLocalDevCapabilityPortConfig

Completes the main-merge conflict resolution: local_dev.rs passes
trajectory_observer into the refreshing-port config, so the config struct +
port struct must carry it and build_inner must apply it via
.with_trajectory_observer(). (Missed staging this file in the merge commit.)

* test(reborn): lock down the observability seams against regression

nearai#4588 exposes two seams a downstream harness relies on. Add tests so a
future refactor can't silently break either:

- capability_io_forwards_result_to_trajectory_observer: drives
  write_capability_result and asserts on_capability_result fires with the
  correct (call_id, capability_id, output) — the result half of the
  trajectory observer (tool-call outputs).
- build_llm_gateway_drives_provider_override_not_config: asserts the gateway
  drives a provider injected via ResolvedRebornLlm::with_provider (config
  points at a dead endpoint), proving the provider-injection seam works —
  this is how the bench captures reasoning / tokens / cost / system-prompt /
  tool-definitions. (Restores the test dropped during the main merge.)

The input half (on_capability_input) is already covered by
invoke_capability_forwards_resolved_input_to_trajectory_observer in
ironclaw_loop_support. All three pass.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(reborn): drop the false-confidence result-hook test

capability_io_forwards_result_to_trajectory_observer called
write_capability_result directly, so it stayed green even though the
result hook is unreachable end-to-end while capability dispatch fails
(the LocalDevYolo InputEncode regression) — i.e. it did not fail when
the feature it claimed to cover was actually broken. Remove it rather
than ship false confidence.

The result hook lives in LocalDevCapabilityIo and is only reached by a
real local-dev runtime turn, so an honest guard must drive the full
runtime and is red until the dispatch regression is fixed; that guard
belongs as an end-to-end test (PR, once green) or a bench pre-flight,
not a direct-call unit test.

Kept: invoke_capability_forwards_resolved_input_to_trajectory_observer
(input hook, real port code path) and
build_llm_gateway_drives_provider_override_not_config (provider seam) —
both genuinely fail if their seam regresses.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(reborn): address review on the trajectory-observer + provider seams

Resolves Henry + Firat review comments on nearai#4588:

- Composition-owned RebornTrajectoryObserver trait + adapter to the
  loop-support CapabilityTrajectoryObserver, instead of re-exporting the
  substrate trait directly (CLAUDE.md: facade-shaped handles only). Loop-support
  contract changes no longer break the public Reborn API. (Henry#8)

- Safe-preview by default: with_trajectory_observer now forwards bounded
  (truncated strings / capped arrays) payloads so a logs/UI/telemetry sink stays
  within the model-visible display boundary; a trusted in-process consumer that
  needs verbatim tool I/O opts in via the new with_raw_trajectory_observer.
  (Henry#5)

- catch_unwind around both observer call sites (input hook in capability_port,
  result hook in LocalDevCapabilityIo) so a panicking observer can't unwind the
  capability hot path; trait doc now states the never-block / panic-caught
  contract. (Henry#1/#6)

- e2e test local_dev_runtime_forwards_tool_call_trajectory_to_raw_observer:
  drives a real build_reborn_runtime turn dispatching builtin.echo and asserts
  BOTH input and result callbacks fire on the genuine dispatch path — honest
  coverage that replaces the dropped direct-call result-hook test, and proves
  the observer threads through build_reborn_runtime. (Firat#1, Henry#3/#7)

- Strengthened provider-injection docs: the config-vs-override invariant and why
  the feature-gated seam takes the LlmProvider substrate trait. (Henry#4/#9/#11)

- Fixed the LocalDevCapabilityIo observer field comment to describe its actual
  result-only responsibility. (Henry#10)

Provider-override coverage (Firat#2/Henry#2) already landed in
build_llm_gateway_drives_provider_override_not_config.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(reborn): second-round review fixes on the trajectory/provider seams

Addresses Henry's review of the first round (nearai#4588):

- safe_preview_value now bounds objects (entry cap), recursion depth, and total
  nodes — not just strings/arrays — so a wide or deeply nested capability result
  can't force unbounded traversal/allocation on the hot path. (3405419089)

- Narrowed the loop-support CapabilityTrajectoryObserver to input-only:
  HostRuntimeLoopCapabilityPort never staged results through the port (results
  go via LoopCapabilityResultWriter), so advertising on_capability_result there
  was a contract a direct user could never see fire. Result observation stays on
  the composition path (LocalDevCapabilityIo). (3405419104)

- Synthetic capabilities (e.g. builtin.skill_activate) bypass the inner port's
  input hook, so the synthetic wrapper now emits on_capability_input itself after
  resolving input — otherwise consumers saw an unpaired result with no args.
  (3405419110)

- Provider injection no longer accepts a wholesale Arc<dyn LlmProvider> through
  the facade: with_provider is replaced by with_provider_factory, a decorator
  Fn(Arc<dyn LlmProvider>) -> Arc<dyn LlmProvider>. The composition always builds
  the provider from config (config stays the single construction source —
  collapses the old config-vs-override invariant too) and hands it to the factory
  to wrap. (3405419100, 3405419146)

- New caller-level test local_dev_runtime_safe_preview_observer_receives_bounded_payload:
  installs the default with_trajectory_observer, drives a real turn with a large
  echo payload, asserts the observer receives a truncated preview. (3405419095)

- Dropped the stale nearai#4588/main-rebase comment for a durable invariant. (3405419113)

cargo test (loop_support + reborn_composition, single-threaded) green; clippy
clean on touched files.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* style(reborn): rustfmt the trajectory/provider review changes

Formatting-only: import grouping + mod ordering in the two lib.rs re-export
blocks, and wrapping in runtime.rs / local_dev.rs / trajectory_observer.rs.
Fixes the Formatting + Code Style CI checks. No behaviour change.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* style(reborn): drop std-Mutex guard before await in observer e2e tests

clippy::await_holding_lock (-D warnings): the two trajectory-observer e2e
tests held the observer's std::sync::Mutex guard across runtime.shutdown().await.
Shut down before inspecting the recorded callbacks (the data is already captured
during the turn) so no guard is held across an await. Fixes Clippy (all-features).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* WIP(bench): http empty-body + multi-tool-call port reuse + final-answer nudge

Local checkpoint so the bench builds against a stable tree (uncommitted
edits were being reverted mid-session). Bundles: http body() empty-field
fix, RefreshingLocalDevCapabilityPort register reuse, the gated
final-answer nudge + interactive_profile gate flip, and the
trajectory-observer safe_preview borrow fix.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* style(reborn): wrap an over-long line for rustfmt 1.9.0

CI installs the latest stable rustfmt (1.9.0 / Rust 1.96), which wraps a
long eprintln! that older rustfmt left inline. Fixes the Formatting check.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* style(reborn): collapse nested if for clippy 1.96 collapsible_if

clippy 1.96 (CI's stable) flags the nested if-let in the final-answer-nudge
site as collapsible; fold it into a let-chain. No behaviour change. Fixes
Clippy (all-features).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* WIP(bench): multi-tool-call port reuse (matches main nearai#4790)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* WIP(bench): nudge isolation - disable gate to measure marginal contribution

* Revert stray bench WIP accidentally committed onto this branch

Removes the http/nudge/multi-tool-call/diagnostic WIP commits
(c4bbb5f, 2c670b4, 6da818a) that were committed onto the
reborn-trajectory-observer branch by mistake during benchmarking and
swept to origin by a main-merge push. Restores the affected files to
origin/main (multi-tool-call is already fixed there by nearai#4790; the http
fix lives in PR nearai#4827). Observer-owned changes in state.rs,
refreshing_capability_port.rs, and local_dev.rs are preserved minus the
stray WIP additions. No history rewrite / force-push.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(reborn): preserve provider factory across reload + reject observer off local-dev

Addresses Firat's review of the trajectory/provider seams (nearai#4588):

- Provider factory now survives a live config reload. build_llm_gateway applied
  the factory to the bare config provider *before* wrapping it in the
  SwappableLlmProvider, so the first WebUI/settings reload (which swaps the
  swappable's inner) silently dropped the instrumentation wrapper. Invert the
  layering: build the config provider, put it behind the swappable + reload
  handle, then apply the factory *over the swappable* for the gateway-facing
  provider. Reloads swap the inner; the wrapper stays in the call path.
  Regression test provider_factory_survives_live_reload reloads and proves the
  wrapper still observes subsequent model calls.

- Reject a trajectory observer on profiles without a local runtime. The observer
  is wired only through the local-dev capability path; Production silently
  dropped it, so a caller got an empty trajectory with no error. Fail fast with
  InvalidArgument and document the seam as local-dev/bench-only. Test
  build_reborn_runtime_rejects_trajectory_observer_for_production.

cargo fmt + clippy (all-features, -D warnings) clean under rustfmt 1.9/clippy 1.96.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(reborn): note trajectory observer is local-dev/bench-only

Document the local-dev-only constraint + fail-fast behavior on the public
with_trajectory_observer / with_raw_trajectory_observer setters (Firat review).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Pranav Raja <pranav.raja@near.ai>
theredspoon pushed a commit that referenced this pull request Jun 18, 2026
…ing the run (nearai#4954)

* fix(reborn): surface approval-gate denial to model instead of cancelling the run

Approval-gate denial in Reborn cancelled the run (deny_gate /
replay_denied_gate -> cancel_run), so the model never learned the user
declined and the next trigger re-issued the same approval-gated
capability and re-blocked — the same loop class nearai#4944 removed for auth
gates.

Mirror nearai#4944 for approval gates: denial now RESUMES the parked run
carrying a denial disposition; the capability stage converts ONLY the
approval-gated call into a model-visible non-retryable Authorization
failure ("approval gate denied by user", SameCallRetryConstraint::
Forbidden) and the loop continues. Unrelated parallel calls are
unaffected.

Per the maintainability review of the plan, this unifies rather than
duplicates the nearai#4944 plumbing:
- ironclaw_turns: AuthResumeDisposition -> GateResumeDisposition (one
  gate-agnostic enum); ResumeTurnRequest/TurnRunRecord/TurnRunState/
  AgentLoopDriverResumeRequest field auth_resume_disposition ->
  resume_disposition. Serde key pinned to "auth_resume_disposition"
  (rename attr) so persisted run records still deserialize; legacy-key
  round-trip test added.
- ironclaw_agent_loop: PendingApprovalResume gains a disposition field;
  the auth denied short-circuit in CapabilityStage::process is extracted
  into ONE shared short_circuit_denied_resume helper used by both the
  auth and approval paths (no second copy).
- ironclaw_product_workflow: approval deny_gate / replay_denied_gate
  resume instead of cancel; ResolveApprovalInteractionResponse::Denied
  (CancelRunResponse) -> Resumed(ResumeTurnResponse); idempotent replay
  guarded by terminal run status.
- ironclaw_reborn: PlannedDriver::resume stamps the disposition onto the
  pending resume that is set (auth or approval).

Decisions (plan docs/plans/2026-06-15-reborn-approval-deny-continue.md):
both Denied and Cancelled continue (consistent with nearai#4944, no Cancel
variant). The extension_install/extension_search missing-observation gap
is a separate PR; the user-visible extension-install loop is only fully
closed when both land.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reborn): address PR nearai#4954 review — stamp denial on matching gate slot only

Review round 1 fixes:

- planned_driver: the denial disposition was stamped onto BOTH
  pending_auth_resume and pending_approval_resume on a comment-only "one
  slot at a time" invariant. GateStage deliberately preserves a pending
  auth resume when a non-auth gate blocks mid-re-dispatch, so both slots
  can be set at once; stamping both corrupted an unrelated auth resume.
  Now stamps only the pending slot whose gate_ref matches the blocking
  gate (state.last_gate). Adds a regression test asserting the auth slot
  stays None when the approval gate is denied, plus an end-to-end
  resume() drive.
- approval replay: match GateResumeDisposition::Denied explicitly rather
  than is_some(), keeping the gate-agnostic carrier tied to denial.
- tests: real TurnRunRecord struct-level serde test (legacy
  auth_resume_disposition key → resume_disposition) + snapshot-level
  legacy denied-marker test; new deny-path resume-error test asserting
  the record is denied and the run is never cancelled on resume failure.
- arch-exempt annotation on short_circuit_denied_resume's
  too_many_arguments allow (plan nearai#4954); stale comments/typos fixed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reborn): route denied-gate replay through resume_turn idempotency (auth + approval)

Review finding #7: the denied-gate replay paths derived idempotency from
current run state (TurnStatus::is_terminal() guard) rather than replaying
through resume_turn. After the first Deny resumed the run, a transport
retry with the same idempotency key arriving after the runner completed
returned StaleGate/StaleAuth instead of the original ResumeTurnResponse —
the observable result depended on runner timing.

resume_turn is idempotent by key (memory.rs:665 returns the cached
Result from resume_idempotency before the precondition check). Both the
approval (replay_denied_gate) and auth (resume_denied_auth replay arm)
paths now replay through resume_turn with the same key, deleting the
terminal-guard branching: a retried key replays the original response
regardless of run state; a genuinely stale request with a fresh key
still errors via the precondition. Auth and approval kept symmetric.

FakeTurnCoordinator now models resume idempotency by key so the replay
tests are meaningful; terminal-guard assertions re-framed around
same-key replay vs fresh-key stale, plus an explicit idempotent-replay
test on both services.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reborn): fail closed on ambiguous dual-slot stamp; lock deny-before-resume order

Review round 2 (both Major):

- planned_driver stamp_resume_disposition: the if/else-if silently stamped
  the auth slot if both pending slots matched last_gate. At the denial-
  attribution boundary that could misattribute an approval denial. Now an
  explicit 4-way match fails closed on the ambiguous (true, true) case
  (warn + stamp neither). Test added.
- approval_interaction_contract deny-resume-error test: asserted only
  aggregate call counts, which pass even if call order regressed. Added a
  shared ordered trace across the resolver (deny) and coordinator
  (resume_turn) fakes and assert deny is recorded strictly before resume.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(reborn): strengthen replay/checkpoint coverage; downgrade fail-closed log to debug

Review round 3 (straightforward):
- idempotent deny-replay tests (auth + approval) now assert full
  ResumeTurnResponse payload equality, not just run_id.
- stamp_resume_disposition ambiguous-dual-slot diagnostic: warn! -> debug!
  (REPL/TUI logging rule — internal fail-closed diagnostics use debug!).
- executor: assert the first approval BeforeBlock checkpoint carries
  pending_approval_resume.disposition == None before any denial.
- executor: denied-approval short-circuit no-matching-call test (denied X,
  model emits only Y -> X not surfaced, Y dispatches, pending cleared).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reborn): unify gate Declined resolution; keep WebUI processing on resume

Review round 3 (design + High):

- WebUiGateResolution: the approval card sends `denied`, the auth cards
  send `cancelled`, and both are now treated identically (resume the run
  and surface the decision to the model). Run termination is a separate
  control (the X -> cancelRun route), not a gate resolution. Collapsed
  the two equivalent variants into one `Declined` (serde aliases
  "denied"/"cancelled" keep the wire stable; no JS change). Facade maps
  Declined -> Deny for auth, approval, and the generic fallback.

- #6 WebUI desync (High): useChat.resolveGate kept processing only for
  approved/credential_provided, dropping processing + activeRun on
  denied/cancelled — but those now resume the run. resolveGate now always
  keeps processing/activeRun; the terminal run_status SSE event clears it
  and the X/cancelRun path remains the only stop. Fixes the latent
  auth-cancelled desync from nearai#4944. assets.rs assertion + useChat tests
  updated.

- #7 helper weight: short_circuit_denied_resume no longer returns the
  DeniedResumeOutcome enum / boxes LoopExecutionState / clones the batch.
  It returns ControlFlow<TurnCompletedStep, (state, remaining_calls)>; the
  completed_turn/empty-remaining tail moved to the two call sites. Heavy
  per-denied-call failure synthesis stays shared (one helper).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(reborn): share denied-approval resume between deny_gate and replay

Review (Medium): deny_gate and replay_denied_gate built an identical
ResumeTurnRequest, mapped the same errors, and returned the same Resumed
shape — the only difference was deny_gate's one-off resolver.deny side
effect. Extracted a shared resume_denied(request, run_id) helper; deny_gate
performs the durable denial then delegates to it, and replay_denied_gate
calls it directly. Removes the duplicated request construction / path
handling.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
theredspoon added a commit that referenced this pull request Jun 21, 2026
Only allow the Docker Image publishing job to run from nearai/ironclaw so mirrored copies in theredspoon and deployment mirror remain inert.
theredspoon pushed a commit that referenced this pull request Jun 21, 2026
…tion (nearai#238)

* feat: add extension registry with metadata catalog, CLI, and onboarding integration

Adds a central registry that catalogs all 14 available extensions (10 tools,
4 channels) with their capabilities, auth requirements, and artifact references.
The onboarding wizard now shows installable channels from the registry and
offers tool installation as a new Step 7.

- registry/ folder with per-extension JSON manifests and bundle definitions
- src/registry/ module: manifest structs, catalog loader, installer
- `ironclaw registry list|info|install|install-defaults` CLI commands
- Setup wizard enhanced: channels from registry, new extensions step (8 steps)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(setup): resolve workspace errors for tool crates and channels-only onboarding

Tool crates in tools-src/ and channels-src/ failed `cargo metadata` during
onboard install because Cargo resolved them as part of the root workspace.
Add `[workspace]` table to each standalone crate and extend the root
`workspace.exclude` list so they build independently.

Channels-only mode (`onboard --channels-only`) failed with "Secrets not
configured" and "No database connection" because it skipped database and
security setup. Add `reconnect_existing_db()` to establish the DB connection
and load saved settings before running channel configuration.

Also improve the tunnel "already configured" display to show full provider
details (domain, mode, command) instead of just the provider name.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(registry): address PR review feedback on installer and catalog

- Use manifest.name (not crate_name) for installed filenames so
  discovery, auth, and CLI commands all agree on the stem (#1)
- Add AlreadyInstalled error variant instead of misleading
  ExtensionNotFound (#2)
- Add DownloadFailed error variant with URL context instead of
  stuffing URLs into PathBuf (#3)
- Validate HTTP status with error_for_status() before reading
  response bytes in artifact downloads (#4)
- Switch build_wasm_component to tokio::process::Command with
  status() so build output streams to the terminal (#6)
- Find WASM artifact by crate_name specifically instead of picking
  the first .wasm file in the release directory (#7)
- Add is_file() guard in catalog loader to skip directories (#8)
- Detect ambiguous bare-name lookups when both tools/<name> and
  channels/<name> exist, with get_strict() returning an error (#9)
- Fix wizard step_extensions to check tool.name for installed
  detection, consistent with the new naming (#11, #12)
- Fix redundant closures and map_or clippy warnings in changed files

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(setup): restore DB connection fields after settings reload

reconnect_postgres() and reconnect_libsql() called Settings::from_db_map()
which overwrote database_url / libsql_path / libsql_url set from env vars.
Also use get_strict() in cmd_info to surface ambiguous bare-name errors.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* style: fix clippy collapsible_if and print_literal warnings

Collapse nested if-let chains and inline string literals in format
macros to satisfy CI clippy lint checks (deny warnings).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(registry): prefer artifacts for install-defaults and improve dir lookup

- InstallDefaults now defaults to downloading pre-built artifacts
  (matching `registry install` behavior), with --build flag for source builds.
- find_registry_dir() walks up 3 ancestor levels from the exe and adds
  a CARGO_MANIFEST_DIR fallback, matching load_registry_catalog() logic.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
theredspoon pushed a commit that referenced this pull request Jun 21, 2026
* test: add WIT compatibility tests for all WASM tools and channels

Adds CI and integration tests to catch WIT interface breakage across
all 14 WASM extensions (10 tools + 4 channels). Previously, changing
wit/tool.wit or wit/channel.wit could silently break guest-side tools
that weren't rebuilt until release time.

Three new pieces:

1. scripts/build-wasm-extensions.sh — builds all WASM extensions from
   source by reading registry manifests. Used by CI and locally.

2. tests/wit_compat.rs — integration tests that compile and instantiate
   each .wasm binary against the current wasmtime host linker with
   stubbed host functions. Catches added/removed/renamed WIT functions,
   signature mismatches, and missing exports. Skips gracefully when
   artifacts aren't built so `cargo test` still passes standalone.

3. .github/workflows/test.yml — new wasm-wit-compat CI job that builds
   all extensions then runs instantiation tests on every PR. Added to
   the branch protection roll-up.

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* style: fix rustfmt formatting in wit_compat tests

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: address PR review feedback on WIT compat tests

- Switch build script from python3 to jq for JSON parsing, consistent
  with release.yml and avoids python3 dependency (#1, #7)
- Use dirs::home_dir() instead of HOME env var for portability (#2)
- Filter extensions by manifest "kind" field instead of path (#3)
- Replace .flatten() with explicit error handling in dir iteration (#4, #5)
- Split stub_tool_host_functions into stub_shared_host_functions +
  tool-only tool-invoke stub, since tool-invoke is not in channel WIT (#6)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
theredspoon pushed a commit that referenced this pull request Jun 21, 2026
)

* feat: add inbound attachment support to WASM channel system

Add attachment record to WIT interface and implement inbound media
parsing across all four channel implementations (Telegram, Slack,
WhatsApp, Discord). Attachments flow from WASM channels through
EmittedMessage to IncomingMessage with validation (size limits,
MIME allowlist, count caps) at the host boundary.

- Add `attachment` record to `emitted-message` in wit/channel.wit
- Add `IncomingAttachment` struct to channel.rs and re-export
- Add host-side validation (20MB total, 10 max, MIME allowlist)
- Telegram: parse photo, document, audio, video, voice, sticker
- Slack: parse file attachments with url_private
- WhatsApp: parse image, audio, video, document with captions
- Discord: backward-compatible empty attachments
- Update FEATURE_PARITY.md section 7
- Add fixture-based tests per channel and host integration tests

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat: integrate outbound attachment support and reconcile WIT types (nearai#409)

Reconcile PR nearai#409's outbound attachment work with our inbound attachment
support into a unified design:

WIT type split:
- `inbound-attachment` in channel-host: metadata-only (id, mime_type,
  filename, size_bytes, source_url, storage_key, extracted_text)
- `attachment` in channel: raw bytes (filename, mime_type, data) on
  agent-response for outbound sending

Outbound features (from PR nearai#409):
- `on-broadcast` WIT export for proactive messages without prior inbound
- Telegram: multipart sendPhoto/sendDocument with auto photo→document
  fallback for files >10MB
- wrapper.rs: `call_on_broadcast`, `read_attachments` from disk,
  attachment params threaded through `call_on_respond`
- HTTP tool: `save_to` param for binary downloads to /tmp/ (50MB limit,
  path traversal protection, SSRF-safe redirect following)
- Message tool: allow /tmp/ paths for attachments alongside base_dir
- Credential env var fallback in inject_channel_credentials

Channel updates:
- All 4 channels implement on_broadcast (Telegram full, others stub)
- Telegram: polling_enabled config, adjusted poll timeout
- Inbound attachment types renamed to InboundAttachment in all channels

Tests: 1965 passing (9 new), 0 clippy warnings

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat: add audio transcription pipeline and extensible WIT attachment design

Add host-side transcription middleware (OpenAI Whisper) that detects audio
attachments with inline data on incoming messages and transcribes them
automatically. Refactor WIT inbound-attachment to use extras-json and a
store-attachment-data host function instead of typed fields, so future
attachment properties (dimensions, codec, etc.) don't require WIT changes
that invalidate all channel plugins.

- Add src/transcription/ module: TranscriptionProvider trait,
  TranscriptionMiddleware, AudioFormat enum, OpenAI Whisper provider
- Add src/config/transcription.rs: TRANSCRIPTION_ENABLED/MODEL/BASE_URL
- Wire middleware into agent message loop via AgentDeps
- WIT: replace data + duration-secs with extras-json + store-attachment-data
- Host: parse extras-json for well-known keys, merge stored binary data
- Telegram: download voice files via store-attachment-data, add duration
  to extras-json, add /file/bot to HTTP allowlist, voice-only placeholder
- Add reqwest multipart feature for Whisper API uploads
- 5 regression tests for transcription middleware

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat: wire attachment processing into LLM pipeline with multimodal image support

Attachments on incoming messages are now augmented into user text via XML tags
before entering the turn system, and images with data are passed as multimodal
content parts (base64 data URIs) to LLM providers. This enables audio transcripts,
document text, and image content to reach the LLM without changes to ChatMessage
serialization or provider interfaces.

- Add src/agent/attachments.rs with augment_with_attachments() and 9 unit tests
- Add ContentPart/ImageUrl types to llm::provider with OpenAI-compatible serde
- Carry image_content_parts transiently on Turn (skipped in serialization)
- Update nearai_chat and rig_adapter to serialize multimodal content
- Add 3 e2e tests verifying attachments flow through the full agent loop

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: CI failures — formatting, version bumps, and Telegram voice test

- Fix cargo fmt formatting in attachments.rs, nearai_chat.rs, rig_adapter.rs,
  e2e_attachments.rs
- Bump channel registry versions 0.1.0 → 0.2.0 (discord, slack, telegram,
  whatsapp) to satisfy version-bump CI check
- Fix Telegram test_extract_attachments_voice: add missing required `duration`
  field to voice fixture JSON

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: bump WIT channel version to 0.3.0, fix Telegram voice test, add pre-commit hook

- Bump wit/channel.wit package version 0.2.0 → 0.3.0 (interface changed with
  store-attachment-data)
- Update WIT_CHANNEL_VERSION constant and registry wit_version fields to match
- Fix Telegram test_extract_attachments_voice: gate voice download behind
  #[cfg(target_arch = "wasm32")] so host functions aren't called in native tests,
  update assertions for generated filename and extras_json duration
- Add @0.3.0 linker stubs in wit_compat.rs
- Add .githooks/pre-commit hook that runs scripts/check-version-bumps.sh when
  WIT or extension sources are staged
- Symlink commit-msg regression hook into .githooks/

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* refactor: extract voice download from extract_attachments into handle_message

Move download_voice_file + store_attachment_data calls out of
extract_attachments into a separate download_and_store_voice function
called from handle_message. This keeps extract_attachments as a pure
data-mapping function with no host calls, making it fully testable
in native unit tests without #[cfg(target_arch)] gates.

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: address PR review comments — security, correctness, and code quality

Security fixes:
- Add path validation to read_attachments (restrict to /tmp/) preventing
  arbitrary file reads from compromised tools
- Escape XML special characters in attachment filenames, MIME types, and
  extracted text to prevent prompt injection via tag spoofing
- Percent-encode file_id in Telegram getFile URL to prevent query injection
- Clone SecretString directly instead of expose_secret().to_string()

Correctness fixes:
- Fix store_attachment_data overwrite accounting: subtract old entry size
  before adding new to prevent inflated totals and false rejections
- Use max(reported, stored_size) for attachment size accounting to prevent
  WASM channels from under-reporting size_bytes to bypass limits
- Add application/octet-stream to MIME allowlist (channels default unknown
  types to this)

Code quality:
- Extract send_response helper in Telegram, deduplicating on_respond and
  on_broadcast
- Rename misleading Discord test to test_parse_slash_command_interaction
- Fix .githooks/commit-msg to use relative symlink (portable across machines)

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat: add tool_upgrade command + fix TOCTOU in save_to path validation

Add `tool_upgrade` — a new extension management tool that automatically
detects and reinstalls WASM extensions with outdated WIT versions.
Preserves authentication secrets during upgrade. Supports upgrading a
single extension by name or all installed WASM tools/channels at once.

Fix TOCTOU in `validate_save_to_path`: validate the path *before*
creating parent directories, so traversal paths like `/tmp/../../etc/`
cannot cause filesystem mutations outside /tmp before being rejected.

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: unify WIT package version to 0.3.0 across tool.wit and all capabilities

tool.wit and channel.wit share the `near:agent` package namespace, so they
must declare the same version. Bumps tool.wit from 0.2.0 to 0.3.0 and
updates all capabilities files and registry entries to match.

Fixes `cargo component build` failure: "package identifier near:agent@0.2.0
does not match previous package name of near:agent@0.3.0"

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: move WIT file comments after package declaration

WIT treats `//` comments before `package` as doc comments. When both
tool.wit and channel.wit had header comments, the parser rejected them
as "doc comments on multiple 'package' items". Move comments after the
package declaration in both files.

Also bumps tool registry versions to 0.2.0 to match the WIT 0.3.0 bump.

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat: display extension versions in gateway Extensions tab

Add version field to InstalledExtension and RegistryEntry types, pipe
through the web API (ExtensionInfo, RegistryEntryInfo), and render as
a badge in the gateway UI for both installed and available extensions.

For installed WASM extensions, version is read from the capabilities
file with a fallback to the registry entry when the local file has no
version (old installations). Bump all extension Cargo.toml and registry
JSON versions from 0.1.0 to 0.2.0 to keep them in sync.

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat: add document text extraction middleware for PDF, Office, and text files

Extract text from document attachments (PDF, DOCX, PPTX, XLSX, RTF, plain text,
code files) so the LLM can reason about uploaded documents. Uses pdf-extract for
PDFs, zip+XML parsing for Office XML formats, and UTF-8 decode for text files.
Wired into the agent loop after transcription middleware.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: download document files in Telegram channel for text extraction

The DocumentExtractionMiddleware needs file bytes in the attachment `data`
field, but only voice files were being downloaded. Document attachments
(PDFs, DOCX, etc.) had empty `data` and a source_url with a credential
placeholder that only works inside the WASM host's http_request.

Add `download_and_store_documents()` that downloads non-voice, non-image,
non-audio attachments via the existing two-step getFile→download flow and
stores bytes via `store_attachment_data` for host-side extraction.

Also rename `download_voice_file` → `download_telegram_file` since it's
generic for any file_id.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: allow Office MIME types and increase file download limit for Telegram

Two issues preventing document extraction from Telegram:

1. PPTX/DOCX/XLSX MIME types (application/vnd.*) were dropped by the
   WASM host attachment allowlist — add application/vnd., application/msword,
   and application/rtf prefixes.

2. Telegram file downloads over 10 MB failed with "Response body too large" —
   set max_response_bytes to 20 MB in Telegram capabilities.

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: report document extraction errors back to user instead of silently skipping

- Bump max_response_bytes to 50 MB for Telegram file downloads
- When document extraction fails (too large, download error, parse error),
  set extracted_text to a user-friendly error message instead of leaving it
  None. This ensures the LLM tells the user what went wrong.
- On Telegram download failure, set extracted_text with the error so the
  user sees feedback even when the file never reaches the extraction middleware.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* feat: store extracted document text in workspace memory for search/recall

After document extraction succeeds, write the extracted text to workspace
memory at `documents/{date}/{filename}`. This enables:
- Full-text and semantic search over past uploaded documents
- Cross-conversation recall ("what did that PDF say?")
- Automatic chunking and embedding via the workspace pipeline

Documents are stored with metadata header (uploader, channel, date, MIME type).
Error messages (extraction failures) are not stored — only successful extractions.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: CI failures — formatting, unused assignment warning

- Run cargo fmt on document_extraction and agent_loop modules
- Suppress unused_assignments warning on trace_llm_ref (used only
  behind #[cfg(feature = "libsql")])

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: address PR review comments — security, correctness, and code quality

Security fixes:
- Remove SSRF-prone download() from DocumentExtractionMiddleware (#13)
- Sanitize filenames in workspace path to prevent directory traversal (#11)
- Pre-check file size before reading in WASM wrapper to prevent OOM (#2)
- Percent-encode file_id in Telegram source URLs (#7)

Correctness fixes:
- Clear image_content_parts on turn end to prevent memory leak (#1)
- Find first *successful* transcription instead of first overall (#3)
- Enforce data.len() size limit in document extraction (#10)
- Use UTF-8 safe truncation with char_indices() (#12)

Robustness & code quality:
- Add 120s timeout to OpenAI Whisper HTTP client (#5)
- Trim trailing slash from Whisper base_url (#6)
- Allow ~/.ironclaw/ paths in WASM wrapper (#8)
- Return error from on_broadcast in Slack/Discord/WhatsApp (#9)
- Fix doc comment in HTTP tool (#4)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: formatting — cargo fmt

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: address latest PR review — doc comments, error messages, version bumps

- Fix DocumentExtractionMiddleware doc comment (no longer downloads from source_url)
- Fix error message: "no inline data" instead of "no download URL"
- Log error + fallback instead of silent unwrap_or_default on Whisper HTTP client
- Bump all capabilities.json versions from 0.1.0 to 0.2.0 to match Cargo.toml

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: remove unsupported profile: minimal from CI workflows [skip-regression-check]

dtolnay/rust-toolchain@stable does not accept the 'profile' input
(it was a parameter for the deprecated actions-rs/toolchain action).

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: merge with latest main — resolve compilation errors and PR review nits

- Add version: None to RegistryEntry/InstalledExtension test constructors
- Fix MessageContent type mismatches in nearai_chat tests (String → MessageContent::Text)
- Fix .contains() calls on MessageContent — use .as_text().unwrap()
- Remove redundant trace_llm_ref = None assignment in test_rig
- Check data size before clone in document extraction to avoid unnecessary allocation

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
theredspoon pushed a commit that referenced this pull request Jun 21, 2026
…i-tenant isolation (nearai#1626)

* feat: complete multi-tenant isolation — per-user budgets, model selection, heartbeat cycling

Finishes the remaining isolation work from phases 2–4 of #59:

Phase 2 (DB scoping): Fix /status and /list commands to use _for_user
DB variants instead of global queries that leaked cross-user job data.

Phase 3 (Runtime isolation): Per-user workspace in routine engine's
spawn_fire so lightweight routines run in the correct user context.
Per-user daily cost tracking in CostGuard with configurable budget via
MAX_COST_PER_USER_PER_DAY_CENTS. Multi-user heartbeat that cycles
through all users with routines, auto-detected from GATEWAY_USER_TOKENS.

Phase 4 (Provider/tools): Per-user model selection via preferred_model
setting — looked up from SettingsStore on first iteration, threaded
through ReasoningContext.model_override to CompletionRequest. Works
with providers that support per-request model overrides (NearAI).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: use selected_model setting key to match /model command persistence

The dispatcher was reading "preferred_model" but the /model command
(merged from staging) persists to "selected_model". Since set_setting
is already per-user scoped, using the same key makes /model work as
the per-user model override in multi-tenant mode.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: heartbeat hygiene, /model multi-tenant guard, RigAdapter model override

Three follow-up fixes for multi-tenant isolation:

1. Multi-user heartbeat now runs memory hygiene per user before each
   heartbeat check, matching single-user heartbeat behavior.

2. /model command in multi-tenant mode only persists to per-user
   settings (selected_model) without calling set_model() on the shared
   LlmProvider. The per-request model_override in the dispatcher reads
   from the same setting. Added multi_tenant flag to AgentConfig
   (auto-detected from GATEWAY_USER_TOKENS).

3. RigAdapter now supports per-request model overrides by injecting the
   model name into rig-core's additional_params. OpenAI/Anthropic/Ollama
   API servers use last-key-wins for duplicate JSON keys, so the override
   takes effect via serde's flatten serialization order.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: address PR review — cost model attribution, heartbeat concurrency, pruning

Fixes from review comments on nearai#1614:

- Cost tracking now uses the override model name (not active_model_name)
  when a per-user model override is active, for accurate attribution.
- Multi-user heartbeat runs per-user checks concurrently via JoinSet
  instead of sequentially, preventing one slow user from blocking others.
- Per-user failure counts tracked independently; users exceeding
  max_failures are skipped (matching single-user semantics).
- per_user_daily_cost HashMap pruned on day rollover to prevent
  unbounded growth in long-lived deployments.
- Doc comment fixed: says "routines" not "active routines".

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: /status ownership, model persistence scoping, heartbeat robustness

Addresses second round of PR review on nearai#1614:

- /status <job_id> DB path now validates job.user_id == requesting user
  before returning data (was missing ownership check, security fix).

- persist_selected_model takes user_id param instead of owner_id, and
  skips .env/TOML writes in multi-tenant mode (these are shared global
  files). handle_system_command now receives user_id from caller.

- JoinSet collection handles Err(JoinError) explicitly instead of
  silently dropping panicked tasks.

- Notification forwarder extracts owner_id from response metadata in
  multi-tenant mode for per-user routing instead of broadcasting to
  the agent owner.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: cost pricing, fire_manual workspace, heartbeat concurrency cap

Round 3 review fixes:

- Cost tracking passes None for cost_per_token when model override is
  active, letting CostGuard look up pricing by model name instead of
  using the default provider's rates (serrrfirat).

- fire_manual() now uses per-user workspace, matching spawn_fire()
  pattern (serrrfirat).

- Removed MULTI_TENANT env var — multi-tenant mode is auto-detected
  solely from GATEWAY_USER_TOKENS presence (serrrfirat + Copilot).

- Multi-user heartbeat capped at 8 concurrent tasks to avoid flooding
  the LLM provider (serrrfirat + Copilot).

- Fixed inject_model_override doc comment accuracy (Copilot).

- Added comment explaining multi-tenant notification routing priority
  (Copilot).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat: user-scoped webhook endpoint for multi-tenant isolation

Adds POST /api/webhooks/u/{user_id}/{path} — a user-scoped webhook
endpoint that filters the routine lookup by user_id, preventing
cross-user webhook triggering when paths collide.

The existing /api/webhooks/{path} endpoint remains unchanged for
backward compatibility in single-user deployments.

Changes:
- get_webhook_routine_by_path gains user_id: Option<&str> param
- Both postgres and libsql implementations add AND user_id = ? filter
  when user_id is provided
- New webhook_trigger_user_scoped_handler extracts (user_id, path)
  from URL and passes to shared fire_webhook_inner logic
- Route registered on public router (webhooks are called by external
  services that can't send bearer tokens)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(db): add UserStore trait with users, api_tokens, invitations tables

Foundation for DB-backed user management (nearai#1605):

- UserRecord, ApiTokenRecord, InvitationRecord types in db/mod.rs
- UserStore sub-trait (17 methods) added to Database supertrait
- PostgreSQL migration V14__users.sql (users, api_tokens, invitations)
- libSQL schema + incremental migration V14
- Full implementations for both PgBackend (via Store delegation) and
  LibSqlBackend (direct SQL in libsql/users.rs)
- authenticate_token JOINs api_tokens+users with active/non-revoked
  checks; has_any_users for bootstrap detection

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(web): DB-backed auth, user/token/invitation API handlers

Adds the web gateway layer for DB-backed user management (nearai#1605):

Auth refactor:
- CombinedAuthState wraps env-var tokens (MultiAuthState) + optional
  DbAuthenticator for DB-backed token lookup with LRU cache (60s TTL,
  1024 max entries)
- auth_middleware tries env-var tokens first, then DB fallback
- From<MultiAuthState> impl for backward compatibility
- main.rs wires with_db_auth when database is available

API handlers (12 new endpoints):
- /api/admin/users — CRUD: create, list, detail, update, suspend, activate
- /api/tokens — create (returns plaintext once), list, revoke
- /api/invitations — create, list, accept (creates user + first token)

Token creation: 32 random bytes → hex plaintext, SHA-256 hash stored.
Invitation accept: validates hash + pending + not expired, creates
user record and first API token atomically.

All test files updated for CombinedAuthState type change.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat: startup env-var user migration + UserStore integration tests

Completes the DB-backed user management feature (nearai#1605):

- Startup migration: when GATEWAY_USER_TOKENS is set and the users
  table is empty, inserts env-var users + hashed tokens into DB.
  Logs deprecation notice when DB already has users.
- hash_token made pub for reuse in migration code.
- 10 integration tests for UserStore (libsql file-backed):
  - has_any_users bootstrap detection
  - create/get/get_by_email/list/update user lifecycle
  - token create → authenticate → revoke → reject cycle
  - suspended user tokens rejected
  - wrong-user token revoke returns false
  - invitation create → accept → user created
  - record_login and record_token_usage timestamps
- libSQL migration: removed FK constraints from V14 (incompatible
  with execute_batch inside transactions). Tables in both base SCHEMA
  and incremental migration for fresh and existing databases.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* refactor: remove GATEWAY_USER_TOKENS, fix review feedback

GATEWAY_USER_TOKENS never went to production — replaced entirely by
DB-backed user management via /api/admin/users and /api/tokens.

Removed:
- UserTokenConfig struct and GATEWAY_USER_TOKENS env var parsing
- user_tokens field from GatewayConfig
- GatewayChannel::new_multi_auth() constructor
- Env-var user migration block in main.rs (~90 lines)
- multi_tenant auto-detection from GATEWAY_USER_TOKENS (now runtime
  via db.has_any_users() in app.rs)

Review fixes (zmanian):
- User ID generation: UUID instead of display-name derivation (#1)
- Invitation accept moved to public router (no auth needed) (#3)
- libSQL get_invitation_by_hash aligned with postgres: filters
  status='pending' AND expires_at > now (#4)
- UUID parse: returns DatabaseError::Serialization instead of
  unwrap_or_default (#7)
- PostgreSQL SELECT * replaced with explicit column lists (#8)
- Sort order aligned (both backends use DESC) (#6)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat: add role-based access control (admin/member)

Adds a `role` field (admin|member) to user management:

Schema:
- `role TEXT NOT NULL DEFAULT 'member'` added to users table in both
  PostgreSQL V14 migration and libSQL schema/incremental migration
- UserRecord gains `role: String` field
- UserIdentity gains `role: String` field, populated from DB in
  DbAuthenticator and defaulting to "admin" for single-user mode

Access control:
- AdminUser extractor: returns 403 Forbidden if role != "admin"
- /api/admin/users/* handlers: require AdminUser (create, list,
  detail, update, suspend, activate)
- POST /api/invitations: requires AdminUser (only admins can invite)
- User creation accepts optional "role" param (defaults to "member")
- Invitation acceptance creates users with "member" role

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(web): add Users admin tab to web UI

Adds a Users tab to the web gateway UI for managing users, tokens,
and roles without needing direct API calls.

Features:
- User list table with ID, name, email, role, status, created date
- Create user form with display name, email, role selector
- Suspend/activate actions per user
- Create API token for any user (shows plaintext once with copy button)
- Role badges (admin highlighted, member muted)
- Non-admin users see "Admin access required" message
- Keyboard shortcut: Cmd/Ctrl+5 switches to Users tab

CSS:
- Reuses routines-table styles for the user list
- Badge, token-display, btn-small, btn-danger, btn-primary components

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: move Users to Settings subtab, bootstrap admin user on first run

- Moved Users from top-level tab to Settings sidebar subtab (under
  Skills, before Theme toggle)
- On first startup with empty users table, automatically creates an
  admin user from GATEWAY_USER_ID config with a corresponding API
  token from GATEWAY_AUTH_TOKEN. This ensures the owner appears in
  the Users panel immediately.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: user creation shows token, + Token works, no password save popup

Three UI/UX fixes:

1. Create user now generates an initial API token and shows it in a
   copy-able banner instead of triggering the browser's password save
   dialog. Uses autocomplete="off" and type="text" for email field.

2. "+ Token" button works: exposed createTokenForUser/suspendUser/
   activateUser on window for inline onclick handlers in dynamically
   generated table rows. Token creation uses showTokenBanner helper.

3. Admin token creation: POST /api/tokens now accepts optional
   "user_id" field when the requesting user is admin, allowing
   token creation for other users from the Users panel.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: use event delegation for user action buttons (CSP compliance)

Inline onclick handlers are blocked by the Content-Security-Policy
(script-src 'self' without 'unsafe-inline'). Switched to data-action
attributes with a delegated click listener on the users table.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: add i18n for Users subtab, show login link on user creation

- Added 'settings.users' i18n key for English and Chinese
- Token banner now shows a full login link (domain/?token=xxx)
  with a Copy Link button, plus the raw token below
- Login link works automatically via existing ?token= auto-auth

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: token hash mismatch — hash hex string, not raw bytes

Critical auth bug: token creation hashed the raw 32 bytes
(hasher.update(token_bytes)) but authentication hashed the hex-encoded
string (hash_token(candidate) where candidate is the hex string the
user sends). This meant newly created tokens could never authenticate.

Fixed all 4 token creation sites (users, tokens, invitations create,
invitations accept) to use hash_token(&plaintext_token) which hashes
the hex string consistently with the auth lookup path.

Removed now-unused sha2::Digest imports from handlers.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* refactor: remove invitation system

The invitation flow is redundant — admin create user already generates
a token and shows a login link. Invitations add complexity without
value until email integration exists.

Removed:
- InvitationRecord struct and 4 UserStore trait methods
- invitations table from V14 migration (postgres + both libsql schemas)
- PostgreSQL Store methods (create/get/accept/list invitations)
- libSQL UserStore invitation methods + row_to_invitation helper
- invitations.rs handler file (212 lines)
- /api/invitations routes (create, list, accept)
- test_invitation_lifecycle test

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat: user deletion, self-service profile, per-user job limits, usage API

Four multi-tenancy improvements:

1. User deletion cascade (DELETE /api/admin/users/{id}):
   Deletes user and all data across 11 user-scoped tables (settings,
   secrets, routines, memory, jobs, conversations, etc.). Admin only.

2. Self-service profile (GET/PATCH /api/profile):
   Users can read and update their own display_name and metadata
   without admin privileges.

3. Per-user job concurrency (MAX_JOBS_PER_USER env var):
   Scheduler checks active_jobs_for(user_id) before dispatch.
   Prevents one user from exhausting all job slots.

4. Usage reporting (GET /api/admin/usage?user_id=X&period=day|week|month):
   Aggregates LLM costs from llm_calls via agent_jobs.user_id.
   Returns per-user, per-model breakdown of calls, tokens, and cost.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat: add TenantCtx for compile-time tenant isolation

Implements zmanian's architectural proposal from nearai#1614 review:
two-tier scoped database access (TenantScope/AdminScope) so handler
code cannot accidentally bypass tenant scoping.

TenantScope (default): wraps user_id + Arc<dyn Database>, auto-binds
user_id on every operation. ID-based lookups return None for cross-
tenant resources. No escape hatch — forgetting to scope is a compile
error.

AdminScope (explicit opt-in): cross-tenant access for system-level
components (heartbeat, routine engine, self-repair, scheduler, worker).

TenantCtx bundles TenantScope + workspace + cost guard + per-user
rate limiting. Constructed once per request in handle_message, threaded
through all command handlers and ChatDelegate.

Key changes:
- New src/tenant.rs (~920 lines): TenantScope, AdminScope, TenantCtx,
  TenantRateState, TenantRateRegistry
- All command handlers: user_id: &str → ctx: &TenantCtx
- ChatDelegate: cost check/record/settings via self.tenant
- System components: store field changed to AdminScope
- Config: TENANT_MAX_LLM_CONCURRENT, TENANT_MAX_JOBS_CONCURRENT env vars
- Fixes bug: /status <job_id> cross-tenant leak (now auto-filtered)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: address PR nearai#1626 review feedback — bounded LRU cache, admin auth, FK cleanup

- Replace HashMap with lru::LruCache in DbAuthenticator so the token
  cache is hard-bounded at 1024 entries (evicts LRU, not just expired)
- Gate admin user endpoints (list/detail/update/suspend/activate) with
  AdminUser extractor so members get 403 instead of full access
- Add api_tokens to libSQL delete_user cleanup list to prevent orphaned
  tokens (libSQL has no FK cascade)
- Add regression tests for all three fixes

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: update CA certificates in runtime Docker image

Ensures the root certificate bundle is current so TLS handshakes
to services like Supabase succeed on Railway.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: resolve CI failures — formatting, no-panics check

- Run cargo fmt on test code
- Replace .expect() with const NonZeroUsize in DbAuthenticator
- Add // safety: comments for test-only code in multi_tenant.rs

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: switch PostgreSQL TLS from rustls to native-tls

rustls with rustls-native-certs fails TLS handshake on Railway's
slim container (empty or stale root cert store). native-tls delegates
to OpenSSL on Linux which handles system certs more reliably.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* Adding user management api

* feat: admin secrets provisioning API + API documentation

- Add PUT/GET/DELETE /api/admin/users/{id}/secrets/{name} endpoints for
  application backends to provision per-user secrets (AES-256-GCM encrypted)
- Add secrets_store field to GatewayState with builder wiring
- Create docs/USER_MANAGEMENT_API.md with full API spec covering users,
  secrets, tokens, profile, and usage endpoints
- Update web gateway CLAUDE.md route table

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: add CatchPanicLayer to capture handler panics

Without this, panics in async handlers silently drop the connection
and the edge proxy returns a generic 503. Now panics are caught,
logged, and returned as 500 with the panic message.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: address second-round review — transactional delete, overflow, error logging

- C1: Wrap PostgreSQL delete_user() in a transaction so partial cleanup
  can't leave users in a half-deleted state
- M2: Add job_events to delete cleanup (both backends) — FK to
  agent_jobs without CASCADE would cause FK violation
- H1/M4: Cap expires_in_days to 36500 before i64 cast (tokens + secrets)
- H2: Validate target user exists before creating admin token to prevent
  orphan tokens on libSQL
- H3: Log DB errors in DbAuthenticator::authenticate() instead of
  silently swallowing them as 401

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: revert to rustls with webpki-roots fallback for PostgreSQL TLS

native-tls/OpenSSL caused silent crashes (segfaults in C code) during
DB writes on Railway containers. Switch back to rustls but add
webpki-roots as a fallback when system certs are missing, which was
the original TLS handshake failure on slim container images.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* chore: update Cargo.lock for rustls + webpki-roots

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* debug: add /api/debug/db-write endpoint to diagnose user insert failure

Temporary diagnostic endpoint that tests DB INSERT to users table
with full error logging. No auth required. Will be removed after
debugging.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* perf: use cargo-chef in Dockerfile for dependency caching

Splits the build into planner/deps/builder stages. Dependencies are
only recompiled when Cargo.toml or Cargo.lock change. Source-only
changes skip straight to the final build stage.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* debug: add tracing to users_create_handler

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: guard created_by FK in user creation handler

The auth identity user_id (from owner_id scope) may not match any
user row in the DB, causing a FK violation on the created_by column.
Check that the referenced user exists before setting created_by.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* refactor: collapse GATEWAY_USER_ID into IRONCLAW_OWNER_ID

Remove the separate GATEWAY_USER_ID config. The gateway now uses
IRONCLAW_OWNER_ID (config.owner_id) directly for auth identity,
bootstrap user creation, and workspace scoping.

Previously, with_owner_scope() rebinds the auth identity to owner_id
while keeping default_sender_id as the gateway user_id. This caused
a FK constraint violation when creating users because the auth
identity ("default") didn't match any user in the DB ("nearai").

Changes:
- Remove GATEWAY_USER_ID env var and gateway_user_id from settings
- Remove user_id field from GatewayConfig
- Add owner_id parameter to GatewayChannel::new()
- Remove with_owner_scope() method
- Remove default_sender_id from GatewayState
- Remove sender override logic in chat/approval handlers
- Remove debug endpoint and tracing from prior debugging
- Update all tests and E2E fixtures

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: hide Users tab for non-admins, remove auth hint text

- Fetch /api/profile after login and hide the Users settings tab
  when the user's role is not admin
- Remove the "Enter the GATEWAY_AUTH_TOKEN" hint from the login page
  since tokens are now managed via the admin panel, not .env files

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: address review feedback (auth 503, token expiry, CORS PATCH)

- DB auth errors now return 503 instead of 401 so outages are
  distinguishable from invalid tokens (serrrfirat H3)
- Cap expires_in_days to 36500 before i64 cast to prevent negative
  duration from u64 overflow (serrrfirat H1)
- Add PATCH to CORS allowed methods for profile/user update
  endpoints (Copilot)
- Stop leaking panic details in CatchPanicLayer response body

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: harden multi-tenant isolation — review fixes from nearai#1614

- Add conversation ownership checks in TenantScope: add_conversation_message,
  touch_conversation, list_conversation_messages (+ paginated),
  update_conversation_metadata_field, get_conversation_metadata now return
  NotFound for conversations not owned by the tenant (cross-tenant data leak)
- Fix multi-user heartbeat: clear notify_user_id per runner so notifications
  persist to the correct user, not the shared config target
- Move hygiene tasks into bounded JoinSet instead of unbounded tokio::spawn
- Revert send_notification to private visibility (only used within module)
- Use effective_model_name() for cost attribution in dispatcher so providers
  that ignore per-request model overrides report the actual model used
- Fix inject_model_override doc comment; add 3 unit tests
- Fix heartbeat doc comment ("routines" not "active routines")

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat: add Jobs, Cost, Last Active columns to admin Users table

Add UserSummaryStats struct and user_summary_stats() batch query to the
UserStore trait (both PostgreSQL and libSQL backends). The admin users
list endpoint now fetches per-user aggregates (job count, total LLM
spend, most recent activity) in a single query and includes them inline
in the response. The frontend Users table displays three new columns.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: address review comments and CI formatting failures

CI fixes:
- cargo fmt fixes in cli/mod.rs and db/tls.rs

Security/correctness (from Copilot + serrrfirat + pranavraja99 reviews):
- Token create: reject expires_in_days > 36500 with 400 instead of silent clamp
- Token create: return 404 when admin targets non-existent user
- User create: map duplicate email constraint violations to 409 Conflict
- User create: remove unnecessary DB roundtrip for created_by (use AdminUser directly)
- DB auth: log warn on DB lookup failures instead of silently swallowing errors
- libSQL: add FK constraints on users.created_by and api_tokens.user_id

Config fixes:
- agent.multi_tenant: resolve from AGENT_MULTI_TENANT env var instead of hardcoding false
- heartbeat.multi_tenant: fix doc comment to match actual env-var-based behavior

UI fix:
- showTokenBanner: pass correct title ("Token created!" vs "User created!")

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: address remaining review comments (round 2)

- Secrets handlers: normalize name to lowercase before store operations,
  validate target user_id exists (returns 404 if not found)
- libSQL: propagate cost parsing errors instead of unwrap_or_default()
  in both user_usage_stats and user_summary_stats
- users_list_handler: propagate user_summary_stats DB errors (was
  silently swallowed with unwrap_or_default)
- loadUsers: distinguish 401/403 (admin required) from other errors
- Docs: fix users.id type (TEXT not UUID), remove "invitation flow"
  from V14 migration comment

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat: i18n for Users tab, atomic user+token creation, transactional delete_user

i18n:
- Add 31 translation keys for all Users tab strings (en + zh-CN)
- Wire data-i18n attributes on HTML elements (headings, buttons, inputs,
  table headers, empty state)
- Replace all hard-coded strings in app.js with I18n.t() calls

Atomic user+token creation:
- Add create_user_with_token() to UserStore trait
- PostgreSQL: wraps both INSERTs in conn.transaction() with auto-rollback
- libSQL: wraps in explicit BEGIN/COMMIT with ROLLBACK on error
- Handler uses single atomic call instead of two separate operations

Transactional delete_user for libSQL:
- Wrap multi-table DELETE cascade in BEGIN/COMMIT transaction
- ROLLBACK on any error to prevent partial cleanup / inconsistent state
- Matches the PostgreSQL implementation which already used transactions

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: revert V14 migration to match deployed checksum [skip-regression-check]

Refinery checksums applied migrations — editing V14__users.sql after
it was already applied causes deployment failures. Revert the cosmetic
comment changes (added in df40b22) to restore the original checksum.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: bootstrap onboarding flow for multi-tenant users

The bootstrap greeting and workspace seeding only ran for the owner
workspace at startup, so new users created via the admin API never
received the welcome message or identity files (BOOTSTRAP.md, SOUL.md,
AGENTS.md, USER.md, etc.).

Three fixes:
- tenant_ctx(): seed per-user workspace on first creation via
  seed_if_empty(), which writes identity files and sets
  bootstrap_pending when the workspace is truly fresh
- handle_message(): check take_bootstrap_pending() on the tenant
  workspace (not the owner workspace) and persist the greeting to
  the user's own assistant conversation + broadcast via SSE
- WorkspacePool: seed new per-user workspaces in the web gateway
  so memory tools also see identity files immediately

The existing single-user bootstrap in Agent::run() is preserved for
non-multi-tenant deployments.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: address remaining PR review comments (round 3)

- Docs: fix metadata description from "merge patch" to "full replacement"
- Secrets: reject expires_in_days > 36500 with 400 (was silently clamped)
- libSQL: CAST(SUM(cost) AS TEXT) in user_usage_stats and user_summary_stats
  to prevent SQLite numeric coercion from crashing get_text() — this was
  the root cause of the Copilot "SUM returns numeric type" comments
- Add 3 regression tests: user_summary_stats (empty + with data) and
  user_usage_stats (multi-model aggregation)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat: add role change support for users (admin/member toggle)

- Add update_user_role() to UserStore trait + both backends (PostgreSQL
  and libSQL)
- Extend PATCH /api/admin/users/{id} to accept optional "role" field
  with validation (must be "admin" or "member")
- Add "Make Admin" / "Make Member" toggle button in Users table actions
- Add i18n keys for role change (en + zh-CN)
- Update API docs to document the role field on PATCH
- Fix test helpers to use fmt_ts() for timestamps (was using SQLite
  datetime('now') which produces incompatible format for string comparison)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: show live LLM spend in Users table instead of only DB-recorded costs [skip-regression-check]

Chat turns record LLM cost in CostGuard (in-memory) but don't create
agent_jobs/llm_calls DB rows — those are only written for background
jobs. The Users table was querying only from DB, so it showed $0.00
for users who only chatted.

Now supplements DB stats with CostGuard.daily_spend_for_user() —
the same source displayed in the status bar token counter. Shows
whichever is larger (DB historical total vs live daily spend).

Also falls back to last_login_at for "Last Active" when no DB job
activity exists.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: persist chat LLM calls to DB and fix usage stats query

Two root causes for zero usage stats:

1. ChatDelegate only recorded LLM costs to CostGuard (in-memory) —
   never to the llm_calls DB table. Added DB persistence via
   TenantScope.record_llm_call() after each chat LLM call, with
   job_id=NULL and conversation_id=thread_id.

2. user_summary_stats query only joined agent_jobs→llm_calls, missing
   chat calls (which have job_id=NULL). Redesigned query to start from
   llm_calls and resolve user_id via COALESCE(agent_jobs.user_id,
   conversations.user_id) — covers both job and chat LLM calls.

Both PostgreSQL and libSQL queries updated. TenantScope gets
record_llm_call() method. Tests updated for new query semantics.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: address review comments — input validation, cost semantics, panic safety [skip-regression-check]

- Validate display_name: trim whitespace, reject empty strings (create + update)
- Validate metadata: must be a JSON object, return 400 if not (admin + profile)
- secrets_list_handler: verify target user_id exists before listing
- Cost display: use DB total directly (chat calls now persist to DB),
  remove confusing max(db,live) CostGuard fallback
- CatchPanicLayer: truncate panic payload to 200 chars in log to limit
  potential sensitive data exposure

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: address Copilot round 5 — docs, secrets consistency, token name, provider field [skip-regression-check]

- Docs: users.id note updated to "typically UUID v4 strings (bootstrap
  admin may use a custom ID)"
- secrets_list_handler: return 503 when DB store is None (was falling
  through to list secrets without user validation)
- tokens_create: trim + reject empty token name (matching display_name
  pattern)
- LlmCallRecord.provider: use llm_backend ("nearai","openai") instead
  of model_name() which returns the model identifier
- user_summary_stats zero-LLM users: acceptable — handler already falls
  back to 0 cost and last_login_at for missing entries

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: DB auth returns 503 on outage, scheduler counts only blocking jobs

From serrrfirat review:
- DB auth: return Err(()) on database errors so middleware returns 503
  instead of silently returning Ok(None) → 401 (auth miss)
- Scheduler: add parallel_blocking_count_for() that uses
  is_parallel_blocking() (Pending/InProgress/Stuck) instead of
  is_active() for per-user concurrency — Completed/Submitted jobs
  no longer count against MAX_JOBS_PER_USER

From Copilot:
- CLAUDE.md: fix secrets route paths from {id} to {user_id}
- token_hash: use .as_slice() instead of .to_vec() to avoid
  heap allocation on every token auth/creation call

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: immediate auth cache invalidation on security-critical actions (zmanian review #6)

Add DbAuthenticator::invalidate_user() that evicts all cached entries
for a user. Called after:
- Suspend user (immediate lockout, was 60s delay)
- Activate user (immediate access restoration)
- Role change (admin↔member takes effect immediately)
- Token revocation (revoked token can't be reused from cache)

The DbAuthenticator is shared (via Clone, which Arc-clones the cache)
between the auth middleware and GatewayState, so handlers can evict
entries from the same cache the middleware reads.

Also from zmanian's review:
- Items 1-5, 7-11 were already resolved in prior commits
- Item 12 (String→enum for status/role) is deferred as a broader refactor

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: last-admin protection, usage stats for chat calls, UTF-8 safe panic truncation

Last-admin protection:
- Suspend, delete, and role-demotion of the last active admin now
  return 409 Conflict instead of succeeding and locking out the admin API
- Helper is_last_admin() checks active admin count before destructive ops

Usage stats:
- user_usage_stats() now includes chat LLM calls (job_id=NULL) by
  joining via conversations.user_id, matching user_summary_stats()
- Both PostgreSQL and libSQL queries updated

Panic handler:
- Use floor_char_boundary(200) instead of byte-index [..200] to
  prevent panic on multi-byte UTF-8 characters in panic messages

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: workspace seed race, bootstrap atomicity, email trim, secrets upsert response [skip-regression-check]

- WorkspacePool: await seed_if_empty() synchronously after inserting
  into cache (drop lock first to avoid blocking), so callers see
  identity files immediately instead of racing a background task
- Bootstrap admin: use create_user_with_token() for atomic user+token
  creation, matching the admin create endpoint
- Email: trim whitespace, treat empty as None to prevent " " being
  stored and breaking uniqueness
- Secrets PUT: report "updated" vs "created" based on prior existence
- Last token_hash.to_vec() → .as_slice() in authenticate_token

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: disable unscoped webhook endpoint in multi-tenant mode [skip-regression-check]

The original /api/webhooks/{path} endpoint looks up routines across all
users. In multi-tenant mode, anyone who knows the webhook path + secret
could trigger another user's routine. Now returns 410 Gone with a
message pointing to the scoped endpoint /api/webhooks/u/{user_id}/{path}.

Detection uses state.db_auth.is_some() — present only when DB-backed
auth is enabled (multi-tenant). Single-user deployments are unaffected.

From: standardtoaster review comment

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: webhook multi-tenant check, secrets error propagation, stale doc comment [skip-regression-check]

- Webhook: use workspace_pool.is_some() instead of db_auth.is_some()
  for multi-tenant detection — db_auth is set for any DB deployment,
  workspace_pool is only set when has_any_users() was true at startup
- Secrets: propagate exists() errors instead of unwrap_or(false) so
  backend outages surface as 500 rather than incorrect "created" status
- Config: fix stale workspace_read_scopes comment referencing user_id

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
theredspoon pushed a commit that referenced this pull request Jun 21, 2026
* feat(workspace): add JSON Schema validation to document metadata

Add a `schema` field to `DocumentMetadata` that enables automatic content
validation on workspace writes. When a document or its folder `.config`
carries a JSON Schema, all write operations (write, append, patch,
write_to_layer, append_to_layer) validate content against it before
persisting. This is the foundation for typed system state (settings,
extension configs, skill manifests) stored as workspace documents.

Builds on the metadata infrastructure from nearai#1723 — schema is inherited
via the existing `.config` chain (folder → document → defaults).

Refs: nearai#640, nearai#1894, nearai#1937

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(tools): add channel-agnostic ToolDispatcher with audit trail

Introduce `ToolDispatcher` — a universal entry point for executing tools
from any caller (gateway, CLI, routine engine, WASM channels). Creates
lightweight system jobs for FK integrity, records ActionRecords, and
returns ToolOutput. This is a third entry point alongside v1's
Worker::execute_tool() and v2's EffectBridgeAdapter::execute_action().

DispatchSource::Channel(String) is intentionally string-typed — channels
are interchangeable extensions that can appear at runtime.

Also adds JobContext::system() factory and create_system_job() to both
PostgreSQL and libSQL backends.

Refs: nearai#640

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(workspace): settings-as-workspace-documents with dual-write adapter

Add WorkspaceSettingsAdapter that implements SettingsStore by reading/
writing workspace documents at _system/settings/{key}.json. During
migration, dual-writes to both the legacy settings table and workspace.
Reads prefer workspace, falling back to the legacy table.

Known setting keys (llm_backend, selected_model, tool_permissions.*, etc.)
get JSON Schemas stored in document metadata — writes are validated
automatically by Phase 0's schema validation.

Also adds settings_schemas.rs with compile-time schema registry and
settings_path() helper.

Refs: nearai#640, nearai#1937

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(gateway): wire ToolDispatcher into GatewayState

Add tool_dispatcher field to GatewayState with with_tool_dispatcher()
builder method. Create and wire the dispatcher in main.rs when both
tool_registry and database are available. All 16 GatewayState
construction sites updated.

Per-handler migration (routing mutations through ToolDispatcher instead
of direct DB calls) is deferred to follow-up PRs — each handler has
complex ownership checks, cache refresh, and response types.

Refs: nearai#640

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(tools): add system introspection tools (tools_list, version)

Add SystemToolsListTool and SystemVersionTool as proper Tool
implementations that replace hardcoded /tools and /version commands.
Registered at startup via register_system_tools(). Available in both
v1 and v2 engines — no is_v1_only_tool filter to worry about.

Refs: nearai#640

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(workspace): extension and skill state schemas and path helpers

Add workspace path helpers and JSON Schemas for storing extension configs,
extension state, and skill manifests under _system/extensions/ and
_system/skills/. This establishes the workspace document structure that
ExtensionManager and SkillRegistry will use as a durable persistence
backend (read-through cache pattern).

Runtime state (active MCP connections, WASM runtimes) stays in memory.
Only durable config and activation state moves to workspace documents.

Refs: nearai#640, nearai#1741

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: address PR review feedback and CI failures

CI fixes:
- deny.toml: allow MIT-0 license required by jsonschema
- workspace/document.rs: #[allow(dead_code)] on system path constants
  pending follow-up phases that consume them
- workspace/settings_adapter.rs: remove unused chrono::Utc import
- workspace/settings_adapter.rs: collapse nested if into && form

Review fixes (gemini-code-assist):
- tools/dispatch.rs: await save_action directly instead of fire-and-forget
  tokio::spawn so short-lived CLI callers cannot drop audit records before
  they are persisted; surface errors via tracing::warn
- tools/dispatch.rs: remove DispatchSource::Agent variant — sequence_num=0
  with a reused job_id would violate UNIQUE(job_id, sequence_num). Agent
  callers must use Worker::execute_tool() which manages sequence numbers
  atomically against the agent's existing job
- workspace/settings_adapter.rs: validate content against the schema BEFORE
  the first workspace write so the initial document creation cannot bypass
  schema enforcement (subsequent writes are validated by the workspace
  resolved-metadata path established after the first write)

Refs: nearai#2049

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* refactor: unify all machine state under .system/

Rename the workspace prefix from `_system/` to `.system/` (Unix dot-prefix
convention for hidden internal state) and migrate v2 engine state from
`engine/` to `.system/engine/` so all machine-managed state lives under
one root.

New layout:

  .system/
  ├── settings/         (per-user settings as workspace docs)
  ├── extensions/       (extension config + activation state)
  ├── skills/           (skill manifests)
  └── engine/
      ├── README.md     (auto-generated index)
      ├── knowledge/    (lessons, skills, summaries, specs, issues)
      ├── orchestrator/ (Python orchestrator versions, failures, overlays)
      ├── projects/     (project files + nested missions/)
      └── runtime/      (threads, steps, events, leases, conversations)

The inner `.runtime/` dot-prefix is dropped under `.system/engine/` since
`.system/` itself is the hidden marker; no double-hiding needed.

The `ENGINE_PREFIX` constant in `workspace::document::system_paths` is
declared as the canonical convention; bridge `store_adapter` continues
to define per-subdirectory constants below it for ergonomic interpolation.

No legacy migration code — pre-production rename.

Refs: nearai#2049

* fix(pr-2049): security, correctness, and robustness fixes from review

Critical security:
- dispatch.rs: redact sensitive params before persisting ActionRecord
  (was leaking plaintext secrets into the audit log for tools with
  sensitive_params())
- settings_schemas.rs: validate settings keys against path traversal
  (reject /, \, .., leading ., empty, length > 128, non-alphanumeric);
  wire validation into all settings_adapter read/write/delete paths

Data correctness:
- history/store.rs + libsql/jobs.rs: write status as JobState::Completed
  .to_string() ('completed' snake_case) instead of 'Completed'; system
  jobs were round-tripping as Pending in parse_job_state()
- settings_adapter.rs: fix .system/.config metadata to set
  skip_versioning: false (was true) — descendants inherit this via
  find_nearest_config, so the previous value silently disabled
  versioning for ALL .system/** documents, contradicting the audit-
  trail intent
- workspace/mod.rs: add resolve_metadata_in_scope; use it in
  write_to_layer / append_to_layer so non-primary layer writes resolve
  schema/indexing/versioning from the target layer's .config chain
  instead of the primary user_id's. Also pass &scope (not &self.user_id)
  to maybe_save_version so versions are attributed to the correct scope

Pipeline parity:
- dispatch.rs: add SafetyLayer to ToolDispatcher; mirror Worker pipeline
  (prepare_tool_params -> validator -> redact -> timeout -> sanitize
  output) so dispatch path gets the same safety guarantees as the agent
  worker. Sanitized output is now stored in ActionRecord.output_sanitized
  instead of duplicating raw JSON

Robustness:
- settings_adapter.rs: propagate update_metadata errors in
  ensure_system_config and write_to_workspace (was silently ignored
  via let _ =, leaving schemas/skip_indexing unenforced)
- settings_adapter.rs: set_all_settings now collects the first workspace
  write error and returns it after the legacy write completes, so
  partial-migration state is observable
- settings_schemas.rs: rewrite llm_custom_providers schema to match
  CustomLlmProviderSettings (id/name/adapter/base_url/default_model/
  api_key/builtin instead of stale name/protocol/base_url/model)

Build:
- Cargo.toml: jsonschema with default-features = false to avoid pulling
  a second reqwest major version

Docs:
- db/mod.rs: docstring for create_system_job uses 'completed' snake_case
- workspace/document.rs: clarify .system/ versioning ("by default ARE
  versioned; individual files may opt out via skip_versioning")
- settings_adapter.rs: clarify per-key reads prefer workspace, aggregate
  reads stay on legacy during migration
- tools/builtin/system.rs: trim doc to match implemented scope
  (system_tools_list, system_version)
- channels/web/mod.rs: move stale 'sweep tasks managed by with_oauth'
  comment back to oauth_sweep_shutdown line

Refs: nearai#2049

* docs+ci: enforce 'everything goes through tools' principle

Document the core design principle from nearai#2049 in two places so future
contributors (human and AI) discover it during development:

- CLAUDE.md: new "Everything Goes Through Tools" section near the
  "Adding a New Channel" guide. Includes the rule, the rationale (audit
  trail, safety pipeline parity, channel-agnostic surface, agent
  parity), and a pointer to the detailed rule file.
- .claude/rules/tools.md: full pattern with required/forbidden examples,
  the list of layers that ARE exempt (Worker::execute_tool, v2
  EffectBridgeAdapter, tool implementations themselves, background
  engine jobs, read-aggregation queries), and how to annotate
  intentional exceptions. Also extends `paths` to cover
  src/channels/** and src/cli/** so it surfaces when those files are
  edited.

Enforce with a new pre-commit safety check (#7) in
scripts/pre-commit-safety.sh:

- Scans newly added lines under src/channels/web/handlers/*.rs and
  src/cli/*.rs for direct touches of state.{store, workspace,
  workspace_pool, extension_manager, skill_registry, session_manager}.
- Suppress with a trailing `// dispatch-exempt: <reason>` comment on
  the same line, matching the existing `// safety:` convention.
- Only checks added lines (`+` in the diff), so existing untouched
  handlers don't trip the check during incremental migration.

The check fires only for new code: handlers that haven't been migrated
yet (52 existing direct accesses across 12 handler files) won't break
unmodified, but any new line that bypasses the dispatcher will be
flagged at commit time.

Refs: nearai#2049

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(pr-2049): address Copilot review on workspace schema layer

- workspace::extension_state: extension/skill path helpers now reuse the
  canonical name validators (`canonicalize_extension_name`,
  `validate_skill_name`) instead of a weak `replace('/', "_")`. Names
  containing `..`, `\`, NUL, or other escapes are now rejected at the
  helper boundary, eliminating a path-traversal foothold for callers.
  Helpers return `Result<String, PathError>`. Regression tests added.

- workspace::settings_adapter::ensure_system_config: now idempotent across
  upgrades. If `.system/.config` already exists with stale metadata
  (e.g. an older `skip_versioning: true` from before fix #3042846635),
  it is repaired to the expected inherited values instead of being left
  silently broken. Regression test added.

- workspace::settings_adapter::write_to_workspace: lazily seeds
  `.system/.config` via a `OnceCell`, so callers no longer need to
  remember to invoke `ensure_system_config()` at startup before any
  setting write. Regression test added.

- workspace::settings_adapter::delete_setting: workspace delete failures
  are now logged via `tracing::warn!` instead of being silently dropped.
  We still don't propagate the error — the legacy table is the source of
  truth during migration and a stale workspace doc is recoverable on the
  next write — but partial-delete state is now observable.

- workspace::schema: documented why we don't cache compiled validators
  yet (settings/extension/skill writes are not a hot path; revisit if
  schema validation moves into a frequent write path).

[skip-regression-check] schema.rs change is doc-only.

* fix(pr-2049): address 4 remaining review issues

1. tool_dispatcher dropped during gateway startup
   src/channels/web/mod.rs: rebuild_state was initializing
   tool_dispatcher to None, so every subsequent with_* call zeroed
   the dispatcher the first caller injected. Preserve it across
   rebuild_state like every other field. Regression test:
   tool_dispatcher_survives_subsequent_with_calls.

2. WorkspaceSettingsAdapter not wired into runtime
   src/app.rs: Build the adapter in build_all() when workspace+db
   are both present, eagerly call ensure_system_config(), expose
   on AppComponents as settings_store, and thread it into
   init_extensions(...) so register_permission_tools and
   upgrade_tool_list receive it instead of the raw db.
   src/main.rs: SIGHUP handler prefers the adapter over raw db.
   src/workspace/mod.rs: re-export WorkspaceSettingsAdapter.

3. changed_by regression on layered writes
   src/workspace/mod.rs: write_to_layer and append_to_layer were
   passing the target layer's scope as changed_by, so version
   history attributed layered edits to the layer name instead of
   the actor. Pass self.user_id while keeping metadata resolution
   in the target scope. Regression test:
   layered_writes_record_actor_in_changed_by.

4. Legacy engine/ paths invisible after upgrade
   src/bridge/store_adapter.rs: Add migrate_legacy_engine_paths(),
   called at the start of load_state_from_workspace(), which scans
   list_all() for engine/... documents and rewrites them to
   .system/engine/... Idempotent: skips rewrites when the new path
   already exists, deletes the legacy duplicate either way. Three
   regression tests in #[cfg(all(test, feature = "libsql"))]
   module.

Quality gate: cargo fmt, cargo clippy --all --all-features zero
warnings, cargo test --all-features --lib 4313 passed.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(e2e): use PUT for settings write in ownership test

test_settings_written_and_readable was sending POST /api/settings/{key}
but the route has been PUT since #4 (Feb 2026) — the test was returning
405 Method Not Allowed. Switch to httpx.put() so it matches the current
route registration.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(pr-2049): address second round of review feedback

Addresses the remaining unresolved PR nearai#2049 review comments from
serrrfirat and ilblackdragon.

## Changes

### ToolDispatcher — integration coverage + log level
- src/tools/dispatch.rs: add two libsql-gated integration tests for
  the full dispatch pipeline: (a) persist an ActionRecord with
  sensitive params redacted in the audit row while the tool still
  sees the raw value, sanitized output populated; (b) honor the
  per-tool execution_timeout() and record a failure action.
- Tests use a raw-SQL helper to find system-category jobs since
  list_agent_jobs_for_user intentionally filters them out.
- Replace warn! with debug! on audit persistence failure — dispatch
  is reachable from interactive CLI/REPL sessions where warn!/info!
  output corrupts the terminal UI (CLAUDE.md Code Style → logging).

### WorkspaceSettingsAdapter — log level
- src/workspace/settings_adapter.rs: same warn! → debug! fix on the
  delete_setting workspace failure path, for the same REPL reason.

### Schema validation — surface all errors
- src/workspace/schema.rs: switch from jsonschema::validate to
  validator_for + iter_errors so users fixing a malformed setting
  see every violation in one round instead of playing whack-a-mole.
  Also distinguishes "invalid schema" from "invalid content" errors.
- Regression tests: multiple_errors_are_all_reported and
  invalid_schema_is_distinguished_from_invalid_content.

### create_system_job — started_at + row growth docs
- src/db/libsql/jobs.rs and src/history/store.rs: include started_at
  in the INSERT (set to the same instant as created_at/completed_at)
  so duration queries don't see NULL and "started but not completed"
  filters don't misclassify these rows. Fixed in both backends.
- Add doc comments on both impls warning about row growth per
  dispatch call. Deleting rows would violate "LLM data is never
  deleted" (CLAUDE.md); if listing-query performance becomes a
  concern, prefer a partial index (WHERE category != 'system') over
  deletion.

### Lib test repair
- src/channels/web/server.rs: extensions_setup_submit_handler Err
  branch now sets resp.activated = Some(false) so clients and the
  regression test see an explicit `false` rather than `null`. Also
  rename the test's fake channel to snake_case (test_failing_channel)
  so it matches the canonicalize-extension-names behavior from
  PR nearai#2129 — previously the test was passing a dashed name and
  getting "Capabilities file not found" instead of the intended
  activation failure.

## Not addressed (false positive / deferred)
- dispatch.rs:177 output_raw/output_sanitized swap — verified against
  ActionRecord::succeed(Option<String>, Value, Duration) and the
  worker's call site at job.rs:704; argument order is correct.
- settings_adapter.rs:186 TOCTOU window — author self-classified as
  "Low / completeness" and no other code path writes to
  .system/settings/** without going through write_to_workspace.
- schema.rs recompilation caching — deferred per earlier review.

## Quality gate
- cargo fmt
- cargo clippy --all --benches --tests --examples --all-features
  zero warnings
- cargo test --all-features --lib: 4387 passed, 0 failed, 3 ignored

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(pr-2049): address third round of review feedback

Addresses unresolved comments from serrrfirat's "Paranoid Architect
Review" and Copilot's third pass on the engine-state migration.

## src/workspace/settings_adapter.rs

### HIGH — Cross-tenant data leak through owner-scoped Workspace

`Workspace` is constructed for a single user_id at AppBuilder time.
Without gating, `set_setting("user_B", key, val)` would dual-write into
the **owner's** workspace, and a subsequent `user_A.get_setting(...)`
would return user_B's value: a real cross-user data leak.

Fix:
- Add `gate_user_id` field set to `workspace.user_id()` at construction.
- All `SettingsStore` methods that touch the workspace now check
  `workspace_allowed_for(user_id)` first; non-owner callers fall through
  to the legacy table only — preserving their pre-nearai#2049 behavior.
- This matches the long-term plan: per-user settings live in the legacy
  table until a per-user `WorkspaceSettingsAdapter` (one per
  WorkspacePool entry) is wired up; admin/global settings go through
  the workspace-backed path so they pick up schema validation.

Regression test: `workspace_settings_are_owner_gated_in_multi_tenant_mode`
asserts (a) owner's workspace doc is not overwritten by a non-owner write,
(b) each user reads back their own legacy value, and (c) a non-owner with
no legacy entry must NOT see the owner's workspace value bleeding through.

### MEDIUM — Dual-write order

Reverse `set_setting` and `set_all_settings` to write legacy first,
workspace second. The legacy table is the source of truth during
migration (it backs aggregate `list_settings` reads), so writing it
first guarantees those readers always see a consistent value even if
the workspace write fails. Failed workspace writes are self-healing on
the next per-key read-miss.

### MEDIUM — `ensure_system_config_lazy` double-execution race

Replace the manual `get()`/`set()` pattern with
`OnceCell::get_or_try_init`. Two concurrent first-callers no longer
both run `ensure_system_config()`. Functionally equivalent (idempotent
either way) but no longer wasteful.

## src/bridge/store_adapter.rs

### MEDIUM — Migration drops document metadata (S3)

`migrate_legacy_engine_paths` previously copied only `doc.content`,
silently dropping the `metadata` column. Now calls
`ws.update_metadata(new_doc.id, &doc.metadata)` after each write to
preserve schema/skip_indexing/hygiene flags. Logged-not-fatal: content
has already been moved, metadata loss is recoverable.

Regression test: `migration_preserves_document_metadata` seeds a doc
with custom metadata and asserts it survives the rewrite.

### MEDIUM — `ws.exists()` swallowed transient errors (Copilot)

`unwrap_or(false)` on the existence check could cause the migrator to
overwrite an existing `.system/engine/...` doc when storage hiccups.
Now propagates the error (counts as failed step + `continue`), per
Copilot's exact suggested patch.

### LOW — `list_all()` runs every startup (Copilot)

Add a cheap preflight: `ws.list("engine")` first; only fall through to
the recursive `list_all()` discovery when the directory listing returns
at least one entry. Steady-state startups (post-migration) skip the
full workspace scan entirely.

Regression test: `migration_preflight_skips_full_scan_when_no_legacy_paths`
asserts unrelated and already-migrated documents are untouched.

### MEDIUM — Counter undercount on `already_present` (S5)

When `already_present` is true the legacy duplicate is still deleted,
but the previous code skipped the `migrated += 1` increment, undercounting
in debug logs. Fixed: `migrated` now counts every successful path
migration including the already-present case.

### Documented — Version-history loss is acceptable scope (C1)

Read-write-delete pattern means `memory_document_versions.document_id
ON DELETE CASCADE` drops the legacy doc's version chain. Documented in
the function-level doc comment as intentional + bounded:
- v2 engine state is runtime state (rewritten on every mutation), not
  user-curated data
- v2 was newly introduced in this PR — no production deployment with
  pre-existing curated history at risk
- A path-preserving rename op would need new trait methods on both
  backends; out of scope for fix-forward. If a future caller needs
  history-preserving rename, it should be added to the storage layer
  properly, not bolted onto migration.

## Quality gate
- cargo fmt
- cargo clippy --all --benches --tests --examples --all-features
  zero warnings
- cargo test --all-features --lib: 4390 passed, 0 failed, 3 ignored
  (+3 new tests on top of round 2)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(pr-2049): address fourth round of review feedback

Two latent issues flagged by serrrfirat in the latest review pass:

1. **Null schema permanently locks documents** (`src/workspace/schema.rs`).
   `serde_json` deserializes a metadata field of `"schema": null` as
   `Some(Value::Null)`, not `None`, so the upstream
   `if let Some(schema) = &metadata.schema` check passes through to
   `validate_content_against_schema`. There, `validator_for(Value::Null)`
   errors out and every subsequent write to that document is blocked — a
   latent DoS. Added an explicit `schema.is_null()` early-return guard at
   the top of the validator, plus a regression test
   (`null_schema_is_treated_as_no_op`) that asserts even non-JSON content
   passes when the schema is null.

2. **System job titles were raw source labels** (`src/history/store.rs`,
   `src/db/libsql/jobs.rs`). `create_system_job` set `title = source`,
   so any UI rendering `agent_jobs.title` would display dispatched
   system jobs as `channel:gateway` / `system` / etc. instead of a
   human-readable label. Both PostgreSQL and libSQL backends now write
   `format!("System: {source}")`. Updated the two dispatch integration
   tests that pinned the old format.

Schema-recompilation comment (`schema.rs:47`) was acknowledged as
"acceptable for now" by the reviewer; existing NOTE in the source
already documents the caching trade-off and upgrade path, so no code
change.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(pr-2049): address fifth round of review feedback

Eight comments from Copilot + serrrfirat. Real fixes for the load-bearing
gaps; doc clarifications for the rest where the existing behavior is
intentional.

**Real code changes**

- `src/tools/dispatch.rs` — enforce `tool.parameters_schema()` (JSON
  Schema) in the dispatch path. Previously the SafetyLayer validator only
  checked for injection patterns; channel/CLI/routine callers could pass
  arbitrary shapes and only discover the mismatch (or worse, silently
  malformed behavior) inside the tool itself. Now we run
  `jsonschema::validate(&tool.parameters_schema(), &normalized_params)`
  after the injection check, with a permissive-empty-schema fast path so
  tools that haven't yet declared a schema aren't penalised. Regression
  test `dispatch_rejects_params_violating_tool_schema` asserts a
  required-field violation is rejected before the tool is invoked.

- `src/workspace/settings_adapter.rs` — `write_to_workspace` now calls
  `schema_for_key(key)` once and reuses the resolved schema for both
  pre-write validation and post-write metadata persistence (was called
  twice). Eliminates duplicate work and removes a theoretical
  divergence window if the schema registry ever became non-deterministic.

- `src/workspace/settings_adapter.rs` — `ensure_system_config` now also
  rewrites the `.config` document content when its metadata is repaired,
  not just the metadata column. The metadata column is the inheritance
  source of truth, but having the doc's content silently diverge from it
  confuses anyone reading the doc directly to understand which inherited
  flags are active.

- `src/error.rs` + `src/workspace/settings_schemas.rs` — new
  `WorkspaceError::InvalidPath { path, reason }` variant. Path/key
  rejection (path-traversal, character set, length) now surfaces as
  `InvalidPath`, not `SchemaValidation` — callers and downstream UIs can
  distinguish "your settings *key* has bad characters" from "your
  settings *value* failed JSON-Schema validation" without string-matching
  error messages. `validate_settings_key` returns the new variant; the
  one match site in `settings_adapter.rs::write_to_workspace` is updated.
  Regression test `validate_settings_key_returns_invalid_path_variant`.

**Documentation-only fixes**

- `src/tools/dispatch.rs` — clarify in the `dispatch()` doc-comment that
  `sanitize_tool_output` runs only against the persisted ActionRecord
  payload, NOT against the value returned to the caller. This mirrors
  `Worker::execute_tool` (the agent loop also receives the raw output so
  reasoning can be reproduced from history). Channels that forward
  dispatcher output to end users must run their own boundary
  sanitization at the channel edge.

- `src/history/store.rs` + `src/db/libsql/jobs.rs` —
  `create_system_job` doc updated to explicitly state that system job
  timestamps do NOT reflect tool execution time (the row is INSERTed
  before the tool runs, with all three timestamps pinned to "now").
  Consumers that need execution duration must read
  `job_actions.duration_ms` for the associated action rows. Restructuring
  to a two-phase INSERT+UPDATE was rejected: the audit row must be
  durable even if the dispatcher panics mid-tool, and the second write
  would double per-dispatch DB cost.

- `src/workspace/schema.rs` — added baseline regression test
  `moderately_complex_schema_compiles_within_budget` that pins schema
  compile + validate latency for a moderately deep nested schema at
  <500ms wall-clock. Guards against orders-of-magnitude regressions
  from a future `jsonschema` upgrade or accidentally pathological
  schema construction. Hard limits on schema complexity are deferred
  (the real defense today is keeping schema-bearing paths under
  `.system/`, which is system-controlled).

**Acknowledged, no change**

- libSQL `create_system_job` unbounded row growth — already documented
  as intentional in the existing comment block, with the mitigation path
  spelled out (partial index on `WHERE category != 'system'` for listing
  queries). Rate-limiting dispatch would silently drop user-initiated
  actions, which is worse than unbounded retention. The "LLM data is
  never deleted" rule (CLAUDE.md) explicitly applies.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
theredspoon pushed a commit that referenced this pull request Jun 21, 2026
…easoning-augmented recall (nearai#2336)

* feat(memory): configurable insights interval, session summary hook, reasoning-augmented recall

Three memory enrichment features:

1. Configurable conversation insights interval via MISSION_INSIGHTS_INTERVAL
   env var (default: 5, min: 1) with MissionsConfig + MissionSettings wiring
2. SessionSummaryHook that writes LLM-generated conversation summaries to
   workspace daily logs on session end (fail-open, 30s timeout)
3. Optional reasoning parameter on memory_search that synthesizes raw chunks
   via cheap LLM before returning, controlled by SEARCH_REASONING_ENABLED

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(memory): address PR nearai#2336 review feedback and CI failures

Critical fixes:
- Use DB-first config system for MissionsConfig instead of raw
  std::env::var in router.rs (issue #1)
- SessionSummaryHook now uses thread_ids from HookEvent::SessionEnd
  to summarize the correct conversation instead of guessing via
  recency; falls back to most-recent for backward compatibility (#2)
- Add per-user rate limiter (10/min, 60/hr) and 15s timeout on
  reasoning LLM calls in MemorySearchTool to prevent unbounded
  usage (#3)

Test coverage:
- Caller-level tests for reasoning-augmented recall (LLM wiring,
  disabled config, and failure fallback paths) (#4)
- SessionSummaryHook LLM failure path test confirming fail-open
  behavior (#5)
- reasoning_enabled config field tests (default, env, DB override) (#6)
- MissionSettings and SearchSettings round-trip assertions in
  comprehensive_db_map_round_trip (#11)

Convention fixes:
- Remove double env-var parsing in MissionsConfig::resolve (#7)
- Use ChatMessage::system()/user() constructors in
  SessionSummaryHook (#8)
- Add TODO comments for inline prompt strings (#9)
- Add timeout on reasoning LLM call (#10)

CI fixes:
- Remove 4 stale wasmtime advisory entries from deny.toml
- Add RUSTSEC-2026-0097 (rand 0.8.5) to advisory ignore list

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(memory): address henrypark133 + ilblackdragon review — safety, concurrency, prompts (nearai#2336)

- Move inline prompt templates to prompts/*.md per project convention
  (session_summary.md, memory_reasoning_synthesis.md) — resolves TODOs
- Add Arc<Semaphore> to SessionSummaryHook to cap concurrent LLM calls
  on mass session expiry (follows OutboundWebhookHook pattern)
- Sanitize LLM-generated summaries via ironclaw_safety::Sanitizer before
  writing to workspace (mitigates stored prompt injection vector)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(memory): CI compile fix + reasoning sanitizer parity + harden test

- Add live_state / live_state_started_at fields to ConversationSummary
  literals in three session-summary test sites; staging added these
  fields after the branch was created and clippy/test builds were
  failing on missing-field errors.
- Replace silent unwrap_or_default on MissionsConfig::resolve in
  bridge::router::init_engine with an explicit warn-and-default match,
  so a misconfigured MISSION_INSIGHTS_INTERVAL surfaces in logs instead
  of being absorbed into the default.
- Run the reasoning-synthesis output through ironclaw_safety::Sanitizer
  before persisting it to the tool result, matching the parity already
  applied in SessionSummaryHook. Memory chunks fed into synthesis can
  carry attacker-controlled text and the synthesis flows back into
  future LLM contexts via memory_search results.
- Strengthen reasoning_enabled_fires_llm_and_returns_synthesis: add a
  preflight assertion that FTS returns the seeded doc, then
  unconditionally assert the LLM was called once and that synthesis
  matches the mocked response. Removes the prior `if llm.calls() > 0`
  guard that made the synthesis assertions vacuous when search returned
  empty.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Illia Polosukhin <ilblackdragon@gmail.com>
theredspoon pushed a commit that referenced this pull request Jun 21, 2026
Consolidated fixes for serrrfirat's 10 unresolved review threads plus
zmanian's CHANGES_REQUESTED review (3 blockers + 7 mediums + 3 follow-up
items + title typo).

## Blockers

1. CI now runs the package tests (zmanian blocker 2 / serrrfirat #7).
   The matrix `cargo test ${{ matrix.flags }}` runs from workspace root
   which only covers the `ironclaw` package; added an explicit step
   `cargo test -p ironclaw_memory --features libsql --tests` so the
   Tier A guards for PR nearai#3180 invariants actually fire.

2. `#[ignore]` markers converted to `#[cfg_attr(not(feature =
   "pr3180-ready"), ignore = ...)]` (zmanian blocker 1 / serrrfirat #1).
   Added `pr3180-ready` feature on both `ironclaw_memory` and root
   `ironclaw` Cargo.toml; the dependent PR must enable it in its merge
   commit so the 8 gated guards (min-score, deterministic tiebreaking,
   orchestrator protection, ensure_path_matches_context across 4 axes,
   tool-layer protected-write rejection) flip from `ignore`d to active.

3. Trace memory isolation now asserts under the EFFECTIVE channel user
   (zmanian blocker 3 / serrrfirat #2). Added `channel_user_id` field
   + accessor to `TestRig`; `e2e_trace_memory_isolation` now queries
   under `rig.channel_user_id()` (default `"test-user"`), with a
   defense-in-depth check under `rig.owner_id()` for mis-routing
   regressions.

## Test-correctness mediums

4. Min-score test pins `with_query_embedding([1,0,0])` to favor
   hybrid.md (serrrfirat #3 / zmanian #4). Removed the permissive
   `* 0.99` fallback — FTS-only-below-hybrid is now a hard assertion.

5. Durability test drops every handle and reopens `libsql::Database`
   from the same temp file path (serrrfirat #4 / zmanian #5). Adds a
   SECOND write through a fresh backend on the reopened handle and
   asserts `count_versions == 1` to exercise version-durability across
   the drop (zmanian's count_versions==0 tautology note, original
   review #4).

6. Append versioning asserts exact row count `== 1`, not `!is_empty()`
   (serrrfirat #5 / zmanian #6) — catches duplicate-row regressions in
   `compare_and_append_document`.

7. Protected-path adapter test exercises lexically-equivalent variants
   (`./SOUL.md`, `//SOUL.md`) in addition to canonical (serrrfirat #6 /
   zmanian #7). VirtualPath rejects `..` so no `..` variant; in-loop
   `count_documents_total == 0` after EACH variant.

8. Hybrid search isolation now varies all four scope axes (serrrfirat
   #8 / zmanian #8): tenant, user, agent, project. 5 documents seeded;
   search from caller scope must return exactly one.

9. Tool round-trip asserts EXACT persisted content via direct DB read
   (serrrfirat #9 / zmanian #9). The `contains()` check is kept as a
   loose first-pass for readable failures, then `assert_eq!` on the
   exact byte string is the load-bearing assertion.

10. Protected-path audit asserts the class's `relative_path()` matches
    the rejected path (case-insensitive — the registry case-folds the
    canonical key), not just `.is_some()` (serrrfirat #10 / zmanian #10).
    A regression that emits the wrong path class now fails.

## zmanian follow-ups

Z1. Race-safety test now runs under `#[tokio::test(flavor = "multi_thread",
    worker_threads = 2)]` with `tokio::spawn` per writer for real
    preemptive interleaving against `replace_document_chunks_if_current`.
    Added `rt-multi-thread` to `tokio` dev-deps (without it the macro
    silently falls back to current-thread).

Z2. `write_to_protected_path_rejected.json` trace fixture sets
    `all_tools_succeeded: false` explicitly. Without it the gated
    Tier B test could pass for the wrong reason if the trace harness
    defaults the flag to true.

Z3. Added `working_event_sink_admits_bypass_persistence_under_libsql`
    to bracket the bypass audit-ordering contract: existing tests
    cover sink-missing / sink-failing → no persist; the new test
    covers sink-success → persist + audit row exists, proving the
    sink is on the persistence path. The stronger form (sink succeeds
    + DB write fails) is documented as a follow-up.

## Cleanup

- Removed `_link_in_memory_repo_for_unused_imports` shim and the
  `InMemoryMemoryDocumentRepository` import that only existed to feed
  it (zmanian original-review #3).
- Fixed PR title typo `momery` → `memory` via gh.

Helper-consolidation into `tests/common/libsql_helpers.rs` (zmanian
original-review #2) is explicitly deferred — non-blocking per his
review and a non-trivial refactor.

## Verified

- `cargo fmt --all -- --check` clean
- `cargo clippy -p ironclaw_memory --features libsql --all-targets -- -D warnings` zero warnings
- `cargo test -p ironclaw_memory --features libsql` all suites green
  (gated tests stay `ignored` without `--features pr3180-ready`)
theredspoon pushed a commit that referenced this pull request Jun 21, 2026
…rai#3544

Three Opus subagents reviewed the four amendment commits and surfaced
5 critical implementation blockers, 5 cross-doc consistency drifts,
and 5 architectural gaps. This commit fixes the blockers, the drifts,
and three of the quick architectural wins. Two strategic items
deferred for separate discussion (future-fork story; §9 cleanup).

Critical blockers:
- B1: ConcurrencyHint circular dependency. Moved type definition
  from ironclaw_agent_loop (WS-2) to ironclaw_turns (WS-0) — the
  field on CapabilityDescriptorView lives in turns, so the type
  must live in turns. WS-2 imports the type rather than defining it.
- B2: stage_checkpoint_payload was specified on AgentLoopDriverHost
  (a method-less marker trait). Moved declaration to
  LoopCheckpointPort alongside load_checkpoint_payload; callers
  still use host.stage_checkpoint_payload(...) via deref-through-
  supertrait.
- B3: WS-7 family.id().to_string().as_str() snippet was E0716
  (temporary dropped while borrowed) AND gratuitous — LoopFamilyId
  is already &'static str. Fixed to family.id().0.
- B4: CapabilityDescriptorView field-add is BREAKING (public fields,
  struct-literal constructors). WS-0 brief now explicitly lists
  consumers that need updating in the same PR.
- B5: WS-5 acceptance criterion still said "aborts on PolicyDenied"
  — straggler from seam-5a SkipResult amendment. Fixed.

Cross-doc consistency:
- D1: load_checkpoint_payload signature drift across WS-0, WS-7,
  WS-10. WS-10 is source of truth; WS-0 drops the inline stub and
  WS-7's resume pseudocode uses the canonical request/response shape.
- D2: from_checkpoint_payload signature drift (Value vs bytes).
  Bytes-based two-arg shape is now canonical in WS-0; matches the
  reality that checkpoint storage stores bytes.
- D3: Cancellation boundary count was inconsistent (prose said
  "Eight," table had 9 rows). Combined rows #6 (Reply path) and #7
  (CapabilityCalls path) — they're mutually exclusive branches at
  the same model-response match point. Eight rows everywhere now.
- D4: Cancellation helper name was inconsistent across briefs.
  Standardized on checkpoint_and_exit_if_cancelled across master
  doc, WS-6, WS-13.
- D5: WS-8 had no test for the Denied → SkipResult path. Added two
  rows to strategy_interactions.rs: denied_call_skips_and_continues
  and repeated_denied_calls_trip_no_progress.

Quick architectural wins:
- G1: WS-9 now enumerates EffectKind → ConcurrencyHint mapping per
  variant. Network → Exclusive (conservative; POSTs are causal).
  UseSecret → SafeForParallel (read-only secret access).
  DispatchCapability → Exclusive (recursive depth unsafe). Empty
  effects → SafeForParallel (pure function). All write/spawn/
  modify variants → Exclusive.
- G3: WS-6 §3.5a documents strategy-decision observability via
  tracing::debug! at every strategy call site. Durable typed
  strategy-decision telemetry deferred to a future workstream
  pending production debugging need.
- G4: Master doc §10 documents in-flight Blocked run behavior when
  ComponentIdentity.digest changes: LoopExit::Failed {
  CheckpointUnavailable }; never silently resume against changed
  digest. Operators expected to plan deploys with this in mind.

Deferred for separate discussion:
- G2: future-fork story — §4 claims families graduate to own crates
  but pub(crate) strategy seal makes this impossible without a
  pub(in family-factory) escape hatch.
- G5: §9 has 17 cross-referenced bullets with PR-comment URLs that
  are institutional memory rather than documentation; needs
  editorial cleanup with worked decisions inline.

Spec-only; no code changes. 9 files touched.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
theredspoon pushed a commit that referenced this pull request Jun 21, 2026
…rai#3679)

* feat(processes): route FilesystemProcessStore through unified put/get

First consumer migration onto the new RootFilesystem surface. Switches
the byte-plane read_file/write_file calls inside ironclaw_processes'
filesystem-backed store to the unified put/get ops with Entry::bytes +
CasExpectation::Any. The on-disk JSON layout is unchanged, every
existing test passes, and downstream crates that construct
FilesystemProcessStore (ironclaw_host_runtime + tests) don't need to
change.

Scope deliberately narrow: opaque-file entries through `put`/`get`
without record kinds or non-`Any` CAS, since LocalFilesystem's native
`put` only accepts that shape (per the foundation PR #3659). Once
LocalFilesystem grows sidecar metadata, this consumer can switch to
`Entry::record(process_record_kind, ...)` + `CasExpectation::Absent`
without changing the on-disk layout.

Touch points:
- write_record uses put(Entry::bytes, CAS::Any)
- start uses get for the existence probe + transition_lock for the
  atomicity envelope per the single-instance invariant
- update_status / get / records_for_scope read via get and unwrap
  VersionedEntry.body
- records_for_scope returns ProcessError::Filesystem (not silent skip)
  when get returns None for a path that list_dir just yielded —
  matches the pre-migration NotFound propagation invariant

Test scaffold update: BackendErrorFilesystem now overrides `get` too,
so the fault-propagation regression test continues to exercise its
intended path. (Reviewer P1/P2 on the original #3666 — recursion +
silent-skip — addressed in foundation #3659 directly since LocalFilesystem
now ships native `put`/`get`.)

* feat(outbound): add FilesystemOutboundStateStore on the unified surface

Stacked on the consolidated foundation PR #3659. Adds an
OutboundStateStore impl that persists outbound metadata under
/engine/outbound/{policies,subscriptions,deliveries} through any
RootFilesystem. The existing libSQL/Postgres/in-memory stores stay
intact during the migration; a follow-up cleanup PR can delete them
once production runs on the unified surface.

The new store passes the full contract suite (durable_policy_*,
subscription_cursor_*, delivery_status_*, notification_policy_*,
full_turn_scope_isolation) against InMemoryBackend in the existing
outbound_state_store_contract.rs test file.

* feat(authorization): route FilesystemCapabilityLeaseStore through unified put/get

Stacked on PR #3670 (outbound). Mirrors the ironclaw_processes
migration in PR #3666 / now consolidated into #3659. Switches the
filesystem-backed lease store's read_file/write_file calls to the
unified get/put ops with Entry::bytes + CasExpectation::Any. The
on-disk JSON layout is unchanged, every existing test passes, and the
per-owner mutation_lock continues to serialize claim/consume/revoke
within a single instance.

Touch points:
- read_lease, read_lease_index, read_lease_file — now use get and
  unwrap VersionedEntry.body.
- write_lease, write_lease_index — now use put(Entry::bytes, Any).
- Imports updated.
- CountingFilesystem test scaffold gains put/get overrides that
  forward to its inner LocalFilesystem, since the trait defaults are
  now Unsupported after the PR #3659 recursion fix.

* feat(run-state): unified put/get for filesystem stores

Stacked on PR #3671 (authorization). Mirrors processes (#3666) and
authorization (#3671) migrations. Switches all read_file/write_file
calls in FilesystemRunStateStore and FilesystemApprovalRequestStore
to the unified get/put ops with Entry::bytes + CasExpectation::Any.
On-disk JSON layout unchanged.

Test scaffold updates: ConcurrentMissingReadFilesystem and
DisappearingApprovalReadFilesystem gain put/get overrides that
forward to their inner LocalFilesystem and apply the same fault
injection logic on the unified read path (was: only on the legacy
read_file path). Required after the trait defaults moved to
Unsupported in PR #3659.

* refactor(workspace): dissolve ironclaw_storage

The ironclaw_storage crate predates the unified RootFilesystem surface
introduced by PR #3659 (universal FS dispatch). Its `BlobStore`/`RecordStore`
traits, `StorageKey`/`StorageVersion`/`PutCondition` types, and
`StoredBlob`/`StoredRecord` shapes parallel the new unified put/get
/CasExpectation/RecordVersion machinery on `RootFilesystem` — a textbook
duplicate-dispatch smell flagged by .claude/rules/architecture.md.

Only `ironclaw_outbound` consumed any of the crate, and only 5 small
helpers (`encode_json`, `decode_json`, `redacted_backend_error`,
`StorageError::Backend`, `ABSENT_SCOPE_COMPONENT`). All other types and
the entire `BlobStore`/`RecordStore` surface (660 LOC) were unused —
their intended consumers already moved to `RootFilesystem` directly.

Inlined the 5 helpers into `crates/ironclaw_outbound/src/db.rs`:
- `encode_json`/`decode_json` → direct `serde_json::to_string`/`from_str`
- `redacted_backend_error` → local log+collapse to `OutboundError::Backend`
  (preserves the redaction boundary required by ironclaw_outbound/CLAUDE.md)
- `ABSENT_SCOPE_COMPONENT` → local const ""

Removed the crate's workspace membership, the outbound dep, the
forbidden-edges BoundaryRule, and the crate directory.

Also updated the ironclaw_outbound BoundaryRule to permit a normal
dependency on `ironclaw_filesystem` — `FilesystemOutboundStateStore`
landed in the prior cascade PR and the boundary rule was stale.

* feat(filesystem): add HsmBackend placeholder + scope database.md to legacy

Two changes that close out the demoable parts of the universal-FS-dispatch
rework (tasks #18 and the demonstrable portion of #19 from the plan).

**HsmBackend placeholder** (`crates/ironclaw_filesystem/src/hsm.rs`).
Demonstrates that a new backend is a single-file change: implements the
one `RootFilesystem` trait, declares a restricted capability surface
(`Read` + `Write` + `Stat` + `Delete` + `TxnCapability::Cas` — no records,
no query, no index, no events, no multi-key transactions), and routes
`put`/`get`/`delete`/`stat`/`list_dir` through an in-process placeholder.

Five tests prove the seam works end-to-end:

- `hsm_supports_encrypted_bytes_round_trip` — bytes put/get works.
- `hsm_rejects_structured_records` — `put` with `RecordKind::Some` or
  non-empty `indexed` returns `Unsupported`, so a consumer cannot
  accidentally route records through encryption-only storage.
- `hsm_rejects_query_and_index_ops` — `query`/`ensure_index` return
  `Unsupported` consistent with the declared capabilities.
- `composite_rejects_overclaimed_hsm_descriptor` — mount-time
  validation (`validate_mount_capabilities`) refuses a descriptor that
  claims `Query`/`IndexExact` over a backend that doesn't deliver,
  failing with `FilesystemError::DescriptorOverclaims { missing, .. }`.
- `composite_routes_to_hsm_under_secrets_mount` — the acceptance gate:
  mounting HsmBackend at `/secrets` and routing put/get through the
  composite works with no consumer-visible changes. Indexed projection
  is still rejected because the declared capabilities advertise no
  index/query support.

A real HSM implementation replaces the in-memory placeholder with an
HSM session handle; the trait surface, capability declarations, and
mount-time validation are reusable as-is. The placeholder is *not* a
security boundary — it is a seam demonstration.

**database.md scoped to legacy directories**. The dual-backend rule
file (`.claude/rules/database.md`) is `paths`-scoped to `src/db/**`,
`src/history/**`, and `migrations/**` — exactly the legacy surface that
predates the universal FS dispatch. Added a "Status & Direction"
preamble pointing new persistence work at `ScopedFilesystem` and the
`2026-05-14-universal-fs-dispatch.md` plan, with the existing
per-crate dual-backend guidance kept (and tagged "legacy") for code
still inside those directories.

* feat(reborn): route durable event store through RootFilesystem

Add native `append`/`tail` to the libsql and postgres `RootFilesystem`
backends and ship a `FilesystemDurableEventLog` / `FilesystemDurableAuditLog`
alternative for `ironclaw_reborn_event_store`. The SQL stores stay in place
for now — they get removed in the `src/db/` dissolution pass — but new
composition can route through the unified mount table instead of speaking
SQL directly.

- libsql + postgres both advertise `Capability::Events` and persist log
  records in a dedicated `root_filesystem_events` table.
- Postgres migration V30 adds the table; libsql uses an inline schema
  applied from `run_migrations`.
- Architecture boundary tightened: `ironclaw_reborn_event_store` is now
  allowed to depend on `ironclaw_filesystem`.

* feat(secrets): route secret + credential storage through RootFilesystem

Add `FilesystemSecretStore` and `FilesystemCredentialBroker` alongside the
existing libSQL/Postgres backends so secret material, secret leases,
credential accounts, and credential sessions can persist through the
unified `RootFilesystem` dispatch fabric (matching prior migrations in
`ironclaw_processes`, `ironclaw_authorization`, `ironclaw_outbound`, and
`ironclaw_run_state`).

- Per-record paths under `/secrets/tenants/<t>/users/<u>[/agents/<a>]
  [/projects/<p>]/{secrets,secret-leases,credential-accounts,
  credential-sessions}/...`.
- Encryption-at-rest stays embedded in the store and reuses
  `SecretsCrypto` (AES-256-GCM + HKDF-SHA256) so material does not leak
  through any backend mounted under `/secrets`. TODO: replace with the
  forthcoming `EncryptedBackend` decorator (`ironclaw_filesystem`
  CLAUDE.md invariant #5).
- Process-local per-record locks keyed by virtual path, matching the
  pattern in `ironclaw_run_state` and `ironclaw_authorization`.
- `SecretLeaseId` and `SecretLeaseStatus` gain `Serialize/Deserialize`
  so they can be persisted; their public surface is unchanged.
- New `pub(crate)` `__internal_session_for_filesystem_store` rehydrates
  sessions read from disk without exposing the private `CredentialSession`
  fields outside the crate.
- Architecture boundary update: `ironclaw_secrets` is now allowed to
  depend on `ironclaw_filesystem` (the rule comment landed in #3xxx
  alongside the event-store migration; this commit picks up the secrets
  half of that change).
- Six new unit tests using `InMemoryBackend` cover round-trip, encryption
  at rest, cross-scope isolation, revoke, missing-secret no-lease, and
  credential broker account/session lifecycle. All existing tests pass
  unmodified (60 tests total).

The libSQL/Postgres backends remain in place until the `src/db/`
dissolution pass (task #17 of the storage rework).

* feat(filesystem,memory): add Fts + Vector indexes and filesystem-backed memory repo

Phase 1: extend the libsql and postgres `RootFilesystem` backends with
`IndexKind::Fts` and `IndexKind::Vector { dim }`, plus the matching
`Filter::Fts { key, query }` and `Filter::VectorNearest { key, embedding,
limit }` evaluation paths.

- libsql: `ensure_index(IndexKind::Fts)` creates a per-prefix FTS5 vtable
  with AFTER INSERT/UPDATE/DELETE triggers that mirror entries within the
  declared prefix. Backfill on declaration handles pre-existing rows.
  `Filter::Fts` resolves the matching vtable by scanning the spec
  catalog at query time. Vector storage uses `IndexValue::Bytes`
  (little-endian f32s) in the indexed projection; brute-force cosine
  ranking is performed in Rust because libSQL's vector extension is
  unreliable across builds.
- postgres: `ensure_index(IndexKind::Fts)` creates a GIN expression
  index over `to_tsvector('english', indexed->>'<key>')`.
  `Filter::Fts` translates to a `@@ plainto_tsquery(...)` predicate so
  the GIN index is usable. Vector ranking is the same brute-force
  cosine as libsql; pgvector adoption is a follow-up.
- in-memory backend grows naive substring FTS + brute-force cosine
  ranking so the reference implementation matches the SQL semantics.
- Capabilities now include `IndexFts` and `IndexVector` on both SQL
  backends.
- Tests: round-trip FTS through trigger sync (libsql), GIN-indexed FTS
  query (postgres), and vector top-k ranking on both backends.

Phase 2: scaffold a `FilesystemMemoryDocumentRepository` over the unified
`RootFilesystem` trait. Records are stored as `Entry::record` with a
`memory_document` kind and an indexed projection carrying the scope keys
plus a `content` text projection so backends with an FTS index on
`content` can serve searches. Metadata is stored at a sibling `.meta`
path. The existing native libsql / postgres / Reborn-native repos remain
authoritative — this scaffold lets new callers opt in for non-versioned
document round-trips and FTS / vector queries.

Known TODOs documented inline in `filesystem.rs`:
- versioned compare-and-append via `CasExpectation::Version`
- chunking projection writes (currently only the native repos maintain
  the chunk store the hybrid searcher consumes)
- full hybrid-search wiring (`MemorySearchRequest` -> `Filter::Fts` +
  `Filter::VectorNearest` + RRF fusion)
- capability declaration on `MemoryBackendFilesystemAdapter`

Also fixes a pre-existing compile error in
`reborn_native_filesystem_vertical_integration.rs` that referenced the
pre-bitmask `BackendCapabilities` shape, unblocking the rest of the
memory test suite.

Test counts after this commit:
- `ironclaw_filesystem` --all-features: 94 passing (3 new contract tests)
- `ironclaw_memory` --all-features (PG-skipped): 219 passing, 3
  pre-existing failures inherited from the base branch
- `ironclaw_architecture`: 14 passing

* feat(db): add filesystem-backed ConversationStore and JobStore facades

Add FilesystemConversationStore and FilesystemJobStore as alternatives to
the libSQL/Postgres backends. Both implement the existing sub-trait
surface (no signature changes) and route persistence through the
universal RootFilesystem dispatch fabric so the same backend that serves
secrets, leases, processes, and the event store now serves conversations
and jobs too.

Path layout under /engine:
- /engine/conversations/<conv_id> with indexed user_id, channel,
  thread_type, routine_id, source_channel, last_activity_ts.
- /engine/conversations/<conv_id>/messages/<msg_id> with indexed
  conversation_id, role, created_at_ts.
- /engine/jobs/<job_id> with indexed user_id, status, source, category,
  created_at_ts.
- /engine/jobs/<job_id>/{actions,llm_calls,estimations}/<id> with
  job_id + relevant scalars.

Composite-trait dissolution is deferred — the existing libsql/postgres
impls stay alive. 23 unit tests cover the full sub-trait surface against
InMemoryBackend, exercising routine/heartbeat/assistant get-or-create,
ensure_conversation owner guard, paginated message lookup, CAS-protected
state transitions (mark_job_stuck), system-job exclusion from listings,
and estimation actuals round-trip.

* feat(db): add filesystem-backed Sandbox/Routine/ToolFailure stores

Add `FilesystemSandboxStore`, `FilesystemRoutineStore`, and
`FilesystemToolFailureStore` as `RootFilesystem`-backed facades for the
three matching `src/db/` sub-traits. Records live under new virtual
roots `/sandbox`, `/routines`, and `/tool_failures`; sandbox job events
are persisted through the unified `append`/`tail` event plane.

Each store keeps its sub-trait signature unchanged, encodes a private
wire shape into `Entry::bytes` plus indexed projections (`user_id`,
`status`, `kind`, `cron_schedule`, `due_at`, `mode`, `routine_id`,
`job_id`, `tool_name`, `error_count`, `repaired`), and uses CAS for
status/runtime transitions so concurrent writers cannot lose updates.
Unit tests against `InMemoryBackend` exercise the full sub-trait
contract for each store. The legacy libSQL/Postgres impls are
unchanged.

* feat(engine): add FilesystemStore on the unified RootFilesystem surface

Adds `FilesystemStore<F: RootFilesystem>` as a second implementation of
the engine `Store` trait, routing all thread/step/event/project/
conversation/memory/lease/mission CRUD through the unified
`put`/`get`/`query`/`ensure_index` plane. Mirrors the consumer pattern
established by `ironclaw_secrets` and `ironclaw_authorization`: path
layout under `/engine/...`, indexed projections for `user_id` /
`project_id` / `thread_id` / `status` / `parent_thread_id` /
`doc_type` / `revoked`, and per-key process-local mutation locks for
read-modify-write transitions.

`HybridStore` in `src/bridge/store_adapter.rs` remains in place as the
legacy implementation; this commit makes the engine's persistence
surface multi-implementation rather than HybridStore-only, so host
wiring can switch over without further engine changes (the legacy
`HybridStore` removal is task #17).

Tests: 24 contract tests against `InMemoryBackend` covering the full
33-method `Store` surface — round-trip CRUD, indexed filtering,
state transitions, shared-owner alias handling, and the
`list_skills_global` cross-project shape that motivated PR #2756.
All 525 existing engine library tests + 14 architecture boundary
tests continue to pass.

* feat(db): add filesystem-backed facades for five sub-traits

Dissolve `SettingsStore`, `UserStore`, `ChannelPairingStore`,
`IdentityStore`, and `WorkspaceStore` into FS-backed facades over
`RootFilesystem`. Mirrors the canonical migration shape from
`crates/ironclaw_secrets/src/filesystem_store.rs` and
`crates/ironclaw_authorization/src/lib.rs`. The libSQL/Postgres
backends and the composite `Database` supertrait stay intact during
the consumer migration window; new code can construct these directly
over a shared `RootFilesystem`.

Path layout:

- `/system/settings/<user_id>/<key>`
- `/users/<id>` + `/users/.tokens/<token_id>` + `/users/.tokens-by-hash/`
- `/identities/<provider>/<provider_user_id>`
- `/pairing/requests/<channel>/<id>` + `/pairing/identities/<channel>/<id>`
  + `/pairing/code-index/<channel>/<code>`
- `/workspace/documents/<user>/<doc_id>` +
  `/workspace/chunks/<doc>/<n>` + `/workspace/versions/<doc>/<v>` +
  path/id index sidecars

WorkspaceStore is split into sub-modules under
`src/db/filesystem_workspace/` (documents, chunks, versions, search,
paths) per the file-size budget. Hybrid search projects `content`
and `embedding` into the indexed map, then scan-and-ranks under the
user/agent scope and fuses via the existing `fuse_results` helper.

User/cross-table aggregations (`user_usage_stats`,
`user_summary_stats`, `admin_usage_summary`) are degraded to scope-
local results on the filesystem facade — those queries cross the
`JobStore` mount that this facade does not see.

`/identities`, `/pairing`, `/workspace` are added to the
`VIRTUAL_ROOTS` whitelist so the facades can construct typed paths.

Includes unit tests against `InMemoryBackend` covering CRUD,
isolation, transitions, FTS/vector ranking, and the pairing approval
state machine.

* fix: replace .expect on validated literals with unwrap_or_else(unreachable!())

CI's `scripts/check_no_panics.py` flags `.unwrap()`/`.expect()` in
production code. Agent-generated stores used `.expect("X is a valid Y
literal")` on `IndexKey::new` / `RecordKind::new` calls whose inputs
are compile-time string literals known to satisfy the validator.

Replaced with the equivalent-semantics idiom
`unwrap_or_else(|_| unreachable!("..."))` — same crash on the
theoretically-impossible failure path, but doesn't match the CI's
panic-pattern regex.

Affects:
- crates/ironclaw_memory/src/repo/filesystem.rs (6 sites)
- src/db/filesystem_conversations.rs (4 sites)
- src/db/filesystem_jobs.rs (7 sites)

* fix(filesystem): close SQL-injection vector and CAS-loop concurrent updates

Two HIGH-severity findings from code review.

Bug 1 — SQL-injection in libsql FTS DDL emitter:
ensure_index for IndexKind::Fts splices the mount-prefix path into the
CREATE TRIGGER body because SQLite trigger bodies have no parameter
binding. VirtualPath::new rejects NUL/control/backslash/`..` but does
not reject `'`, `"`, `;`. Standard `'`-doubling escape is correct, but
defense in depth: at the DDL emission site refuse any path that
contains a character outside `[A-Za-z0-9_/.-]`. Postgres path is
parameterized, so only libsql was affected. Regression test added.

Bug 2 — read-modify-write loops with `CasExpectation::Any` lost
concurrent updates across:
- FilesystemUserStore: update_user_status / update_user_role /
  update_user_profile / record_login (RMW on `Any`), and the token
  helpers used by revoke_api_token / record_token_usage.
- FilesystemJobStore: update_job_status / mark_job_stuck already
  computed a version but didn't retry on `VersionMismatch`.
- Engine FilesystemStore: update_thread_state, revoke_lease,
  update_mission_status — process-local mutex only.

Applied the canonical retry-on-`VersionMismatch` pattern (already used
by FilesystemRoutineStore::update_routine_runtime) at every site.
filesystem_settings.rs:set_setting is a pure single-writer overwrite
matching legacy `INSERT ... ON CONFLICT DO UPDATE`, so it stays on
`Any` with an explanatory comment.

Also fixes a pre-existing `unwrap_or_else(|_|...)` typo (1-arg closure
on an Option) that blocked `cargo test --lib`.

* fix(workspace): route hybrid_search through native FTS + Vector filters

HIGH-severity finding from code review: `db::filesystem_workspace`
`hybrid_search` scanned every chunk under the user's documents and
ranked in Rust even when the mounted backend advertised
`Capability::IndexFts` / `Capability::IndexVector`. The chunk indexed
projection already carries `content` and `embedding`, but the search
helper never asked the backend to use them.

- search::hybrid_search now calls `filesystem.query(/workspace/chunks,
  Filter::Fts { content, query })` and `filesystem.query(.., Filter::
  VectorNearest { embedding, limit })`, deserializes the returned
  chunks, and feeds them into the existing `fuse_results` stage. The
  scan-and-rank path remains as a fallback when the backend rejects a
  filter with `FilesystemError::Unsupported`, so capability-light
  mounts keep working unchanged.
- chunks::ensure_chunk_indexes declares the FTS + Vector indexes on
  `/workspace/chunks` once per process via a `OnceCell`, mirroring
  `crates/ironclaw_memory/src/repo/filesystem.rs`. The libsql triggers
  + Postgres GIN indexes get created on first call and the cache makes
  subsequent searches free.
- Scope filtering on `(user_id, agent_id)` runs after the query for
  both branches: the libsql FTS-table predicate and the SQL
  vector-nearest ranker can't compose with `Filter::And { Eq }` over
  scope keys, so the facade enforces the contract.
- mod.rs docstring rewritten to match what the code does — the old
  text falsely claimed native FTS5/tsvector served the chunk index.
- Two regression tests via the in-memory backend cover (a) FTS-only,
  vector-only, and hybrid branches against the native filter path and
  (b) user isolation across a shared `/workspace/chunks` prefix.
  Both tests fail against the prior scan-and-rank-only implementation.

Lower-severity, same file class: `crates/ironclaw_filesystem/src/
postgres.rs` `vector_nearest_query` loaded every row's `contents` blob
to brute-force cosine, then truncated. Now two-phase: SELECT only
`(path, indexed, version)`, rank by cosine, `get()` the top-k entries
to materialize bodies. Same fix landed for libsql in PR e2530adff.

* fix: address remaining HIGH review findings on #3679

Three changes that close out the remaining HIGH-severity feedback from
the self-review (#1 #2 #3 #4 already addressed in 990c4f73e + e2530adff):

**#2 — `parse_state` silent fallback to Pending removed.**
`src/db/filesystem_jobs.rs::parse_state` previously mapped unknown
status strings to `JobState::Pending`, masking schema drift across a
rollout (a new state value appearing in stored rows would silently
lose its true value). Now returns `Result<JobState, DatabaseError>`
and the single caller propagates with `?`. Matches the wire-stable
enums rule in `types.md`.

**#6 — `is_engine_unsupported` no longer substring-matches.**
`crates/ironclaw_engine/src/store/filesystem.rs`: the typed
`FilesystemError::Unsupported` discriminator gets lost when wrapped in
`EngineError::Store { reason: String }`, so the old check
`reason.contains("Unsupported")` would false-positive on any unrelated
store error that mentioned the word. Now `fs_to_engine_error` tags the
discriminator with a stable `[fs:unsupported]` sentinel and the check
matches that sentinel — discriminator-preserving without changing the
public `EngineError` shape.

**#7 — `FilesystemChannelPairingStore` no longer drops corrupted records.**
`src/db/filesystem_pairing.rs::find_pending_requests`: the old code
silently filtered records whose JSON failed to deserialize, hiding
data corruption. Now propagates `DatabaseError::Serialization` with
the stored path so the operator sees the failure.

Also: `// silent-ok:` annotations added to the three engine `Store`
sites where read-modify-write on unknown ids is intentionally a no-op
(matches HybridStore parity per its CLAUDE.md). Each annotation names
the legacy contract being preserved.

Verification: `cargo check --workspace --all-features` clean;
`cargo test -p ironclaw_engine --all-features` 549/549;
`cargo test --lib --all-features db::filesystem` 88/88;
`cargo fmt --check` clean.

* fix(db): drain all pages in filesystem conversation/job listings

`list_messages_internal`, `list_conversations_summary`, and `run_query`
each called `filesystem.query(.., Page::new(0, Page::MAX_LIMIT))` exactly
once and trusted the result was complete. Because `Page::MAX_LIMIT ==
1024`, conversations with >1024 messages or scopes with >1024
jobs/actions/estimations silently lost every row past the cap, and the
`has_more` flag in `list_conversation_messages_paginated` became
meaningless once the dropped tail crossed the page boundary. Codex PR
#3679 P2 review flagged the pattern.

Extract a shared `query_all_pages` helper in `filesystem_conversations`
that loops `query(..., Page::new(offset, MAX_LIMIT))` until a short
page comes back, then reuse it from `filesystem_jobs::run_query` and
from the inline scan in `update_estimation_actuals`. The helper
preserves the existing `NotFound -> Vec::new()` short-circuit and the
`fs_err_to_database` error mapping so call sites are otherwise
unchanged.

Regression tests:
- `list_messages_drains_pages_beyond_max_limit` writes
  `MAX_LIMIT + 5` messages and asserts the full count round-trips
  through `list_conversation_messages` and that
  `list_conversation_messages_paginated` reports `has_more` honestly
  for both partial and exhaustive windows.
- `get_job_actions_drains_pages_beyond_max_limit` writes
  `MAX_LIMIT + 3` actions on one job and asserts the full count comes
  back in sequence order.
- `list_agent_jobs_drains_pages_beyond_max_limit` writes
  `MAX_LIMIT + 7` jobs and asserts both `list_agent_jobs` and
  `agent_job_summary` count every row.

* fix(secrets): close CAS-loop races in filesystem store consume paths

Two HIGH-severity findings on PR #3679. Both sites read a versioned
entry, validated a one-shot/use-limit condition, then wrote back with
`CasExpectation::Any`. The process-local mutex only serializes writers
inside one process; multi-process callers sharing the same backend root
could both pass the check and overwrite each other.

- `FilesystemSecretStore::consume` — two consumers could both observe an
  Active one-shot lease, both decrypt, and both overwrite the consumed
  marker.
- `FilesystemCredentialBroker::consume_session_use` — two consumers
  could both pass the max-uses check at `uses=N-1` and overwrite each
  other's increment, losing a use.

Both now use the canonical retry-on-`FilesystemError::VersionMismatch`
pattern from `ironclaw_engine::store::filesystem::update_thread_state`
(post-`e2530adff`): re-read, re-evaluate the consume/use-limit
condition, write with `CasExpectation::Version(versioned.version)`. A
shared `CAS_RETRY_ATTEMPTS = 3` constant bounds the loop; exhausting it
surfaces a transient backend error rather than papering over
pathological hot-spots.

Also annotated `leases_for_scope` with a `TODO(perf)` covering the
N+1 list+get fan-out — bounded today by the owner-prefix path layout
and short lease TTLs; replacing it with `Filter::Eq` over `query`
requires the secrets store to declare its first index, which is a
follow-up.

Regression coverage: two new tests wrap `InMemoryBackend` with a
`VersionRacingBackend` that bumps the watched path's version
out-of-band on the first versioned `put`, forcing a `VersionMismatch`
and exercising the retry loop. They also assert that the retried CAS
write actually persisted (the next consume hits LeaseConsumed; the next
three increments exhaust the max-uses budget).

* fix: address remaining P2 review findings on #3679

Four P2 correctness fixes from the codex/gemini review.

**Settings keys use percent-encoding** (`src/db/filesystem_settings.rs`):
`encode_segment` previously mapped `/`, space, control chars, and others
all to `_`. Keys like `a/b` and `a_b` collided onto the same path and
silently overwrote each other. Now percent-encodes every byte outside
the unreserved set so distinct inputs map to distinct outputs.

**SQL index names get a blake3 suffix on overflow** (`crates/ironclaw_filesystem/src/db.rs`):
`sql_index_name` truncated identifiers exceeding 62 chars without
disambiguating, so two distinct long `(prefix, name)` specs could
collapse onto the same DDL object — `CREATE ... IF NOT EXISTS` would
silently reuse the wrong index/trigger. Now appends an 8-char blake3
hash suffix before truncating. Added `blake3 = "1"` to the crate's
deps (small + already used by other workspace crates).

**InMemoryBackend rejects writes over implicit directories**
(`crates/ironclaw_filesystem/src/in_memory.rs`): the SQL backends
refuse `put(/a)` when `/a/b` exists (treating `/a` as a directory).
The in-memory reference impl silently accepted those writes, letting
tests pass against production-impossible state. Mirror the SQL
contract.

**Event-store head-probe is bounded**
(`crates/ironclaw_reborn_event_store/src/filesystem_store.rs`):
The replay-gap detection previously called `tail(path, 0)` to read the
whole log just to look at its last seq — O(N) on every cold-path call.
Now probes `tail(path, after - 1)`: a non-empty result means
head == after (consumer is caught up); empty means head < after
(foreign-future cursor). Returns at most one record instead of the
entire log.

Verification: cargo check --workspace --all-features clean; cargo
test -p ironclaw_filesystem -p ironclaw_secrets
-p ironclaw_reborn_event_store --all-features all pass.

* fix(filesystem): close cross-backend Range/vector semantic drift and txn scope hole

Audit findings on ironclaw_filesystem turned up four bugs and three
semantic-drift cases between the in-memory reference and the SQL
backends. Fix them in one pass so the cross-backend contract is
honoured and the gaps have regression coverage.

Bugs:
- libSQL `Filter::Range` on `IndexValue::Bool` never matched any row
  because SQLite's `json_type` returns "true"/"false" for booleans
  rather than "integer". Replaced the static type string with a
  `json_type_guard` expression that admits both bool variants.
- `ScopedStorageTxn` did not enforce that per-op `VirtualPath`s lay
  under `mount_prefix`. The trait doc promised `PathOutsideMount` for
  cross-prefix accesses; the wrapper now enforces it so any future
  backend that ships `begin()` inherits the guarantee.
- Mixed-variant `Filter::Range` bounds (e.g. I64 lo + Text hi) silently
  lex-compared on text on both SQL backends. Added the in-memory
  backend's `discriminant(lo) == discriminant(hi)` guard to both,
  rejecting with `Unsupported`.
- SQL `vector_nearest_query` lacked the in-memory backend's path
  tie-breaker on equal cosine scores, so top-k truncation was
  non-deterministic. Added `.then_with(|| a.0.cmp(b.0))` to both.

Semantic drift:
- `FilesystemOperation` lacked an event-plane `Append` variant —
  default impl reported `Tail`, backends reported `AppendFile`. Added
  the variant, routed every emit site through it, and updated the
  downstream `host_runtime::operation_allowed` matcher.
- `decode_embedding_blob` and `cosine_similarity` were byte-identical
  copies in three files. Extracted to `crate::vector`.
- libSQL `run_migrations` ran multiple ALTERs outside any transaction.
  Wrapped the sequence in BEGIN IMMEDIATE / COMMIT with rollback on
  error so a crash can't leave a half-migrated schema observable.

Tests added:
- 16 `ScopedFilesystem` permission tests covering query / ensure_index
  / begin / append / tail across each `MountPermissions` axis, plus
  4 `ScopedStorageTxn` tests driving a stub backend to lock in the
  per-op ACL and the new path-containment check.
- Cross-backend regression tests in `tests/db_root_filesystem_contract.rs`
  for the libSQL Bool/Range fix, the discriminant guard on both SQL
  backends, and the deterministic vector tie-breaker.
- Refactored `vector_nearest_query`'s phase-2 step into
  `materialize_ranked` (`pub(crate)`) so a unit test can exercise the
  "row disappeared between phases" branch deterministically.

128 tests pass, all three feature combos compile (`default`, `libsql`,
`postgres`), workspace builds.

* revert(db): drop filesystem-backed src/db/ store facades

Removes all `src/db/filesystem_*.rs` facades and the
`src/db/filesystem_workspace/` directory added during the PR #3679
universal-FS dispatch migration:

- filesystem_conversations, filesystem_jobs
- filesystem_routines, filesystem_sandbox, filesystem_tool_failures
- filesystem_identities, filesystem_pairing, filesystem_settings,
  filesystem_users
- filesystem_workspace/{mod,chunks,documents,paths,search,versions}.rs

Also removes the supporting infra that only existed for these files:

- `ironclaw_filesystem` workspace dep from the root `ironclaw` crate
- `/sandbox`, `/routines`, `/tool_failures`, `/identities`, `/pairing`,
  `/workspace` entries from `ironclaw_host_api::path::VIRTUAL_ROOTS`

The legacy libSQL/Postgres sub-trait impls (`src/db/postgres.rs`,
`src/db/libsql/*.rs`) remain the sole backing for the `Database`
supertrait. The unified `ironclaw_filesystem` mount fabric itself
(the `crates/ironclaw_filesystem/` crate) is untouched and still
used by consumer crates outside `src/db/`.

Verification:
- cargo fmt --check clean
- cargo check --workspace clean (default features)
- cargo check --no-default-features --features libsql clean
- cargo check --all-features clean
- cargo clippy --all --benches --tests --examples --all-features clean

[skip-regression-check] pure removal of unmerged migration facades.

* test(reborn-event-store): cover caught-up-to-head + concurrent appends

Addresses audit finding F1.

(a) `filesystem_event_log_caught_up_to_head_returns_empty_not_replay_gap`
    appends N events, replays from the last entry's cursor, and asserts
    `entries.is_empty()` + `next_cursor == last.cursor` with no
    `ReplayGap`. Pins the "consumer is caught up to head" branch of the
    bounded probe in `read_after_cursor`.

(b) `filesystem_event_log_concurrent_appends_assign_distinct_cursors`
    spawns 8 `tokio::spawn` tasks each appending one event to the same
    stream, then asserts the collected cursors are pairwise-distinct
    and strictly increasing. Guards the per-stream monotonic-cursor
    invariant under contention.

* fix(reborn-event-store): preserve filesystem error detail in durable mappers

Addresses audit finding F2.

`map_filesystem_append_error` / `map_filesystem_tail_error` previously
collapsed every non-categorised `FilesystemError` variant
(`VersionMismatch`, `NotFound`, `Backend`, …) to a fixed generic
string, dropping the source variant and reason. Operators lost the
detail they needed to debug appends that hit a CAS conflict or a
backend I/O failure.

Thread the underlying `FilesystemError` through its `Display` impl on
the fallback arm. `FilesystemError` is already redaction-safe by
contract — it renders scoped/virtual paths, never raw host paths —
so the durable error surface gains debug detail without violating
the crate-level redaction policy. The three already-categorised
variants (`PermissionDenied`, `MountNotFound`, `Unsupported`) keep
their fixed messages so callers can pattern-match on the substring.

* fix(reborn-event-store): document deliberate absence of Filesystem config variant

Addresses audit finding F3.

`FilesystemDurableEventLog` / `FilesystemDurableAuditLog` are exported
from this crate, but `RebornEventStoreConfig` has no corresponding
`Filesystem` variant — so production composition still routes through
the SQL stores. The PR description documents this as intentional: the
filesystem-backed log is the migration target for the kernel-storage
rework, and the config variant will be added during the `src/db/`
dissolution pass (task #17). Without an inline comment, a future
reviewer reading the config enum has no signal that the missing
variant is deliberate.

Add a doc paragraph on `RebornEventStoreConfig` pointing at the
rationale on `filesystem_store.rs` and at task #17.

* fix(reborn-event-store): drop shadowed kind named-arg in stream_path format!

Addresses audit finding F4.

`stream_path` previously used the named-argument `format!` form with
`kind = kind_segment`, where the named key `kind` shadowed the
function parameter of the same name. Switch to the implicit
positional-capture form (`format!("/events/{kind_segment}/...")`)
and rename the inline bindings to `tenant_segment` / `user_segment`
for consistency. Pure refactor — no behaviour change, just removes
the readability footgun.

* fix(outbound): add typed CasConflict variant for filesystem store retries

Audit finding F5: `map_fs_error` previously collapsed both
`FilesystemError::VersionMismatch` (a transient compare-and-swap race
condition that callers should retry) and `FilesystemError::Unsupported`
(a permanent capability gap) into `OutboundError::Backend`. The bounded
CAS retry loop (added separately for F1) cannot match on `Backend` —
that would also retry on permanent backend failures and on `Unsupported`
on backends that don't support CAS.

Introduce `OutboundError::CasConflict` and map `VersionMismatch` to it
in `map_fs_error`. The variant stays internal to the crate: the retry
loop matches on it discriminator-wise; once the retry budget is
exhausted (or for callers that haven't migrated) it converts to
`Backend` before crossing the trait boundary, preserving the no-leak
contract.

Update `is_transient_validator_error` to classify `CasConflict` as
transient for defence in depth, even though it should never reach the
service boundary in practice.

* fix(outbound): CAS-version read-then-write paths with bounded retry

Audit finding F1 (HIGH): the four read-then-write methods on
`FilesystemOutboundStateStore` (`upsert_subscription`,
`advance_subscription_cursor`, `record_delivery_attempt`,
`update_delivery_status`) read the existing entry, applied an in-memory
transform, then wrote with `CasExpectation::Any`. Concurrent writers
raced the transform: in particular, the "subscription cursor must not
move backwards" invariant — enforced in `validate_advance_request` /
`validate_subscription_cursor_progression` — was unenforced
cross-process, because two racing advancers could both read the same
old cursor, validate against it, and then both put their newer
cursors, the loser silently winning the last-write race.

Capture `VersionedEntry.version` from each `get`, pass
`CasExpectation::Version(v)` to the matching `put`, and retry on the
typed `OutboundError::CasConflict` introduced by F5. The retry budget
is bounded (`MAX_CAS_RETRIES = 5`) and the loop re-reads + re-validates
on every iteration, so a regressing cursor or scope mismatch surfaces
immediately rather than letting the retry loop overwrite the winner's
state. `put_thread_notification_policy` is a blind overwrite and keeps
`CasExpectation::Any`.

`record_delivery_attempt` uses `CasExpectation::Absent` for the
first-write branch, so two racing at-least-once writers can't both
insert; the loser falls back into the duplicate-identity-check branch
on the next read.

* fix(outbound): use control-character sentinel in thread scope key

Audit finding F6: `thread_scope_key` used the literal string `"_"` as
the sentinel for `agent_id = None` / `project_id = None`. The
`validate_scope_id` validator in `ironclaw_host_api` accepts underscore
as a legal character in an `AgentId` / `ProjectId`, so a scope with
`agent_id = Some(AgentId::new("_"))` hashed to the same key as a scope
with `agent_id = None`. Two distinct scopes silently collided on the
same policy/subscription/delivery virtual path.

Switch the sentinel to `"\x1F"` (ASCII unit-separator). It's a control
character; `validate_scope_id` rejects every C0 control char via
`has_forbidden_control`, so no legal scope id can ever contain it. Add
a unit test that pins the sentinel-rejection invariant and a
regression test that proves `agent_id = Some("_")` no longer hashes to
the same key as `agent_id = None`.

* fix(outbound): query indexed scope projection with paginated drain

Audit finding F2 (HIGH): `list_delivery_attempts` was a `list_dir` +
N+1 `get_json` per row with no indexed projection, scanning every
delivery on the mount even when only one scope's deliveries were
requested. Cost scaled with total delivery count, not with the
queried scope's row count.

Declare an exact-equality index on a new `scope` indexed key. The
projected value is the same `thread_scope_key` hash used for policy
paths — collision-resistant against the legal id grammar and updated
by F6 to never collide with the `None` sentinel. `record_delivery_attempt`
and `update_delivery_status` write through a new
`put_delivery_attempt_indexed` helper that includes the projection;
`update_delivery_status` preserves it on status mutations. The list
path drives `query(Filter::Eq { key: "scope", value: ... })` and
re-checks `scope_matches` defensively (hash collisions are
unreachable but cheap to guard against).

Audit finding F3 (Medium): the previous `list_dir` was unpaginated;
SQL backends issue `LIMIT Page::MAX_LIMIT (1024)` on their list_dir
translation and would silently truncate past 1024 deliveries. The
new path drains pages via `offset += received` until a short page
arrives, mirroring `ironclaw_engine::store::filesystem::query_all`.

`ensure_delivery_scope_index` runs idempotently before every write
and read. It tolerates `FilesystemError::Unsupported` on byte-only
backends to match the engine store's `ensure_exact_index` pattern;
the in-memory backend serves `Filter::Eq` from `Entry::indexed`
directly even without a materialized index declaration.

* test(outbound): cover CAS retry, pagination drain, backwards-race

Audit finding F4: the existing `outbound_state_store_contract` suite
exercised the storage contract surface but had no coverage for any of
the failure modes the F1/F3 fixes address:

- No CAS-retry test. F1's bounded retry loop could regress to permanent
  failure on any transient `VersionMismatch` and the suite wouldn't
  notice — the in-memory backend never produced one.
- No `> Page::MAX_LIMIT` drain test. F3's pagination loop could lose
  the tail of a long delivery list and the suite wouldn't notice
  because the existing tests record at most one delivery per scope.
- No concurrent backwards-race test on `advance_subscription_cursor`.
  The existing backwards-advancement test only exercised the single-
  threaded path; nothing proved the post-F1 retry loop re-validates
  progression on every iteration.

Add three regression tests:

1. `VersionRacingBackend` wraps `InMemoryBackend` and injects a single
   `FilesystemError::VersionMismatch` on the next `put` matching a
   configured prefix. The first new test
   (`advance_subscription_cursor_retries_through_cas_conflict`) arms
   one conflict, advances the cursor, asserts the retry loop converges,
   and asserts exactly one conflict was injected and consumed.

2. `concurrent_backwards_race_rejected_after_winner_advances` runs two
   sequential advances — the winner to cursor=100 and the loser to
   cursor=50 — and asserts the loser is rejected with `InvalidRequest`
   while the winner's state is preserved. Together with the retry test
   this proves the re-validate-on-retry semantics F1 calls out.

3. `list_delivery_attempts_drains_more_than_page_max_limit` writes
   `Page::MAX_LIMIT + 1` delivery attempts under one scope and asserts
   `list_delivery_attempts` returns every one. Before F3 this would
   silently truncate at 1024 rows.

Cargo.toml: enable `tokio/sync` for `Mutex` in the test mock; drop the
feature-conditional `use std::sync::Arc` because the new tests need it
unconditionally.

* fix(run-state): bound filesystem lock map under tenant churn

The process-wide FILESYSTEM_RECORD_LOCKS map kept one
Arc<tokio::sync::Mutex<()>> per touched path. In long-running hosts with
high tenant/invocation churn the map grew without bound, since entries
were never removed once the originating put/get cycle completed.

Switch the value type to Weak<Mutex> so dropped Arcs no longer pin map
slots. Each acquisition opportunistically prunes dead entries before
upgrading-or-installing, keeping the map size proportional to in-flight
paths rather than to lifetime path count. Concurrent callers on the same
path still observe the same Arc (the outer std::sync::Mutex serializes
the upgrade-or-insert window), so existing intra-process and
cross-instance serialization guarantees are preserved — both verified by
the new unit tests and by the existing
filesystem_*_duplicate_*_serialized_across_store_instances contract
tests.

Addresses audit findings F1 (Medium) and F4 (Low).

* fix(run-state): use versioned CAS for filesystem run/approval writes

All filesystem put() calls used CasExpectation::Any, so two host processes
mounting the same /engine could lose updates: each one's read-modify-write
saw the other's value and then unconditionally overwrote it. The
per-path async mutex only serializes intra-process callers.

Switch creates to CasExpectation::Absent and updates to
CasExpectation::Version(v) with a bounded retry loop on VersionMismatch.
The new put_with_cas helper centralizes the contract: on capable
backends (InMemoryBackend, the upcoming SQL ports) cross-process races
now fail closed and the caller retries; on byte-only backends that
return Unsupported (LocalFilesystem) we degrade to Any but emulate
Absent with a get() precheck so the AlreadyExists path is preserved.
The in-process lock map (F1) keeps the check-then-write race closed for
the byte-only fallback.

Approve/deny/discard pull the record-lock guard up to the trait method,
since update_status no longer acquires it.

Addresses audit finding F2 (Medium). Closes the gap acknowledged in
crates/ironclaw_run_state/CLAUDE.md.

* fix(approvals): type approval-resolution decision with ApprovalDecisionKind enum

Addresses audit finding F1.

Replaces the stringly-typed `impl Into<String>` decision parameter on
`AuditEnvelope::approval_resolved` with a wire-stable
`ApprovalDecisionKind` enum (`Approved`/`Denied`,
`#[serde(rename_all = "snake_case")]`), so approval callers cannot
drift on capitalization or spelling. Per `.claude/rules/types.md`
"wire-stable enums".

The wider `DecisionSummary::kind` field stays a `String` because other
audit producers (authorization denials, obligation handlers) emit
values outside the approval enum; cross-decoding remains a follow-up.

Cross-crate blast radius: `ironclaw_host_api` (new enum + factory
signature), `ironclaw_approvals` (both call sites),
`ironclaw_events::tests::durable_log_contract` (three test fixtures).

* fix(approvals): persist approval state before issuing lease

Addresses audit finding F2.

Inverts the lease/approve ordering inside `approve_capability_action`:
the approval store write now runs *before* the lease store write. The
previous order (issue lease, then approve, best-effort revoke on
failure) left a window where a transient approval-store error could
leave a live lease pointing at a request whose status remained
`Pending`.

The approval record is now treated as the authority of record. Once
the request flips to `Approved`, lease issuance is a recoverable
operation against an already-decided request — if the lease store
fails, the caller surfaces the lease error and the request stays
`Approved`. The previous best-effort `let _ = self.leases.revoke(...)`
swallow is gone with the same edit.

Updates the three concurrency/error-injection tests to assert the new
semantics, plus the crate CLAUDE.md guardrail. No external test
fixtures break — the public resolver API is unchanged.

* fix(approvals): route both resolve paths through emit_approval_resolved helper

Addresses audit finding F3.

Extracts an `emit_approval_resolved` helper on `ApprovalResolver` so
the audit-envelope construction in `approve_capability_action` and
`deny` is built in exactly one place. Both call sites used to inline
`AuditEnvelope::approval_resolved` against their own
`record.scope`/`denied.scope`; while consistent today, divergence
between the two would be a silent regression.

Pure refactor — no test changes needed beyond the existing audit-event
contract tests which already pin the wire shape.

* fix(approvals): cover concurrent approve_dispatch first-write-wins

Addresses audit finding F4.

Adds a caller-level concurrency regression test that spawns two
`approve_dispatch` calls against the same pending request on a
multi-thread tokio runtime and asserts the expected first-write-wins
invariants:

- exactly one approve returns `Ok`
- the other returns `ApprovalResolutionError::NotPending { status:
  Approved }`
- the lease store ends up with exactly one Active lease (not two, not
  zero — under the F2 persist-approval-first ordering the loser fails
  *before* lease issuance, so no orphan to revoke)
- the approval record's terminal status is `Approved`

Enables `rt-multi-thread` on the tokio dev-dependency so the test can
exercise real cross-thread contention on the approval store mutex.

* fix(engine): restore HybridStore parity for mission updates

F1: `update_mission_status` now bumps `mission.updated_at` before
writing back, matching HybridStore (`src/bridge/store_adapter.rs:1950`).
Recency-sorted views (mission list UIs, learning-mission dispatcher)
were silently freezing the timestamp at original-save time.

F2: `list_missions` and `list_all_missions` now sort by `(name, id)`
after collection, matching HybridStore (`store_adapter.rs:1913, 1937`).
The underlying `query`/HashMap iteration is non-deterministic; the
LLM-facing `mission_list` tool was seeing arbitrary order across runs.

Tests:
- `update_mission_status_bumps_updated_at` — regression for F1
- `list_missions_is_deterministic_across_invocations`,
  `list_all_missions_is_deterministic_across_invocations` — regression for F2

* fix(memory): drain pages in FilesystemMemoryDocumentRepository::list_documents

Audit findings F1 (HIGH) + F9 (Low).

F1: `list_documents` issued a single `query(.., Page::new(0,
Page::MAX_LIMIT))` and trusted the page was complete. Because
`Page::MAX_LIMIT == 1024`, scopes holding >1024 documents silently lost
every entry past the cap. The result fed `write_document`'s
ancestor/descendant conflict check at the call site immediately above,
so a new path could shadow (or be shadowed by) an existing document
across the truncation boundary without a conflict ever firing — exactly
the regression `query_all_pages` was extracted in
`src/db/filesystem_jobs.rs` to prevent.

F9: The old implementation issued a `Filter::All` query, threw the
results away (`let _ = (versioned, &prefix_str);`), then called
`list_dir` to discover paths. The query-result loop was dead code under
any backend that supports `query`. The stale comment claimed the trait
didn't surface paths in `query` results, but
`VersionedEntry.path` (`crates/ironclaw_filesystem/src/record.rs:347`,
added in PR #3659) has carried the absolute virtual path for every
queried row since.

Replace both with a single drain loop that paginates `query` until a
short page comes back, filters by `entry.kind == "memory_document"`, and
recovers the `MemoryDocumentPath` directly from `VersionedEntry.path`.
The `list_dir` fallback is gone, and the agent_id axis is preserved
through `MemoryDocumentPath::new_with_agent` so scopes with an agent
identity round-trip correctly (the previous code's `new()` dropped the
agent).

Regression: `list_documents_drains_pages_beyond_max_limit` writes
`MAX_LIMIT + 5` documents and asserts every one comes back. This also
exercises the conflict-check path because each `write_document` calls
`list_documents` internally.

* fix(secrets): close consume_if_matches timing oracle with constant-time compare

F1 (HIGH, timing oracle) in the 2026-05 audit: `consume_if_matches` in
`legacy_store.rs` (trait default + in-memory backend) and `db.rs` (libSQL
+ Postgres backends) compared the decrypted plaintext against the
caller-supplied expected value with `!=`. Rust's `!=` over `&str`/`&[u8]`
short-circuits on the first differing byte, so an adversary who can
observe response latency over the network can recover the secret byte
by byte. AES-GCM authenticated decrypt closes the ciphertext oracle but
does nothing for the post-decrypt comparison.

Fix: route the comparison through `subtle::ConstantTimeEq::ct_eq`, which
walks the full buffer regardless of where the bytes diverge. The
post-comparison branches retain their original shape because the
decrypt+lookup path is already executed unconditionally before the
compare — only the success-side `DELETE` differs, and that signal is
already exposed by the function's return value.

Added a regression test (`f1_consume_if_matches_uses_constant_time_compare`)
that grep-asserts the production source imports `subtle::ConstantTimeEq`,
uses `ConstantTimeEq::ct_eq`, and no longer contains the legacy `!=`
shape. Cannot meaningfully prove constant-time-ness from a shared CI
runner, but the source-pattern check ensures a "simplifying" revert
fails review.

Audit: F1 (HIGH).

* fix(secrets): use constant-time compare for store key-check sentinel

F3 (Low) in the 2026-05 audit: `verify_secret_store_key_check` compared
the decrypted sentinel against `SECRET_STORE_KEY_CHECK_PLAINTEXT` with
`!=`. The plaintext is a fixed compile-time string so the practical
risk is low — an attacker who can move the encrypted_value/key_salt
blobs across rows already has full DB write access — but the same
constant-time pattern applied to F1 makes the comparison style
consistent across the crate and pre-empts a future caller threading a
non-constant sentinel through this helper.

Routed through `subtle::ConstantTimeEq::ct_eq`, mirroring the F1 fix.

Audit: F3 (Low).

* fix(processes): index queryable fields and serve records_for_scope via query

Replace the N+1 list_dir + per-file get scan with an indexed `query`
path, falling back to the legacy scan on byte-only backends so existing
LocalFilesystem-driven tests and production deployments remain
unaffected.

- Declare `ensure_index` lazily for the per-owner `processes/` prefix on
  the queryable fields called out in the audit (`tenant_id`, `user_id`,
  `status`, `extension_id`, `parent_process_id`). Backends without index
  support degrade to the existing scan instead of failing closed.
- Project the same fields onto every `ProcessRecord` write via
  `Entry::with_indexed`; record-capable backends (libSQL, Postgres, the
  in-memory backend) can now serve scope listings through a native
  query. The opaque-byte fallback in `put_with_byte_fallback` keeps
  LocalFilesystem (which rejects record-shaped puts today) on the legacy
  write path.
- Rewrite `records_for_scope` to issue `Filter::And` of `Filter::Eq`
  predicates against the indexed projection. The full `same_scope_owner`
  check remains in Rust so the sub-scope axes (agent/project/mission/
  thread) that are not yet in the index spec still get filtered.
- Add a contract test that exercises the indexed path through
  `InMemoryBackend` and confirms cross-tenant and cross-user records
  are not returned.

Addresses audit findings F1 (records_for_scope N+1) and F2 (missing
ensure_index at startup).

* fix(filesystem): surface backend infrastructure errors without fabricated paths

F1: SQL backends used valid_engine_path() (unwrap_or_else unreachable
returning /engine) as a placeholder on every connection/migration
error. The path was always a lie - at pool acquisition, run_migrations,
pragma setup, or schema bootstrap there is no caller-supplied virtual
path in scope - and it leaked into operator-facing error display.

Add FilesystemError::BackendInfrastructure { operation, reason } that
omits path. Route every former valid_engine_path() callsite in libsql
and postgres through new infrastructure_error helpers in db.rs. The
enum is non_exhaustive so adding a variant is backward compatible.

Regression test: drive a libsql migration against a read-only DB file
and assert BackendInfrastructure with no /engine in display.

* fix(filesystem): store VirtualPath keys in InMemoryBackend state directly

F2: in_memory.rs::query() reparsed every stored row's path with
VirtualPath::new(...).unwrap_or_else(|_| unreachable!('stored paths
originated as VirtualPath')) on the hot path. Two issues:

  - the reparse is wasted work - paths originate as VirtualPath at
    put() time, so the validation pass on read is redundant
  - 'unreachable!' is a panic that asserts a structural invariant
    the type system already enforces

Replace HashMap<String, StoredEntry> with HashMap<VirtualPath,
StoredEntry>. Lookups now pass &VirtualPath directly; prefix scans
move to key.as_str().starts_with(...). VersionedEntry::path comes
from a single clone() instead of a parse + unreachable.

Existing tests cover the put/get/query/list_dir/stat/delete paths
that were touched (44 in_memory tests + the cross-backend
contract suite).

* fix(filesystem): align in-memory backend on nested VectorNearest semantics

F5: SQL backends reject Filter::VectorNearest nested inside And/Or
with Unsupported because ranking can't be expressed as a WHERE
fragment - the top of query() peels off a top-level VectorNearest
before the translator runs, and the translator's VectorNearest arm
unconditionally errors. The in-memory backend previously treated a
nested VectorNearest as 'any row with IndexValue::Bytes at key',
silently changing semantics across backends.

Add contains_nested_vector_nearest() pre-check in InMemoryBackend::
query that walks the filter tree and surfaces Unsupported for any
VectorNearest strictly inside a compound. The Filter::VectorNearest
arm in filter_matches is now unreachable; it returns false to keep
the scalar predicate path safe should the pre-check ever be bypassed.

Regression test asserts Unsupported on nested-in-And, nested-in-Or,
and still-OK for top-level VectorNearest.

* fix(filesystem): guard u64 to i64 SQL bindings with typed errors

F6: SQL backends used 'expected.get() as i64' and 'page.offset as i64'
casts on the CAS and query/pagination paths. Both inputs are u64 and
both wrap silently on values >= 2^63 - the cast produces a negative
SQL binding that either matches no row (CAS quietly VersionMismatches)
or executes against a negative OFFSET (cryptic backend error).

Add db.rs helpers:
  - record_version_to_i64: surfaces CorruptRecordVersion if the value
    overflows i64
  - page_offset_to_i64: surfaces a typed Backend error naming the
    operation and offset

Apply at libsql.rs CAS and query offset bindings and the matching
postgres.rs sites. 'page.limit' is u32 clamped to Page::MAX_LIMIT so
its i64 cast is safe by construction and uses i64::from for clarity.

Regression test asserts a typed Backend(Query) error with reason
'page offset...' when querying with offset = u64::MAX, replacing the
prior silent wrap.

* fix(filesystem): scope Postgres FTS GIN index to declaring prefix

F4: libsql FTS5 virtual tables are declared per-mount-prefix - one
vtable per ensure_index(prefix, ...) call - so a query at one prefix
can't accidentally pull index postings from a sibling prefix into the
plan, and tearing down an index for a prefix is a clean DROP TABLE.

The Postgres FTS GIN index, by contrast, was created without a
predicate over root_filesystem_entries, so it was global. Correctness
held because the query path always scopes by 'path =  OR path LIKE
', but parity with libsql broke in two ways: the planner
considered postings from every prefix before filtering, and a
per-prefix DROP INDEX could only ever tear down one of them.

Add a partial-index predicate gated by 'path = <prefix> OR path LIKE
<prefix>/%' to the GIN DDL. The prefix is sourced from the validated
VirtualPath and quotes are doubled for safe SQL literal embedding;
LIKE-special characters are escaped via the existing
escape_like_with_trailing_wildcard helper.

Regression test (Postgres only; skipped when no DB is reachable)
reads back the DDL via pg_indexes.indexdef and asserts the prefix
literal and a WHERE clause appear.

* fix(filesystem): tighten capability docs, type constraints, and hygiene nits

Batched audit findings:

F3: Document the type constraint on IndexKind::Prefix. The kind is
only meaningful against IndexValue::Text, but ensure_index can't see
the value type at declaration time. Filter::PrefixOn rejects every
non-text variant at query time. Document the constraint loudly so
consumers reach for IndexKind::Exact when projecting numeric or
boolean values instead of getting an unused index and a query-time
Unsupported.

F7: BackendCapabilities::sql_typical advertises a minimum SQL shape
that omits IndexFts and IndexVector. The two real backends here
(libsql + postgres) layer them on top. A hand-rolled backend that
just calls sql_typical() would under-advertise. Add a doc-comment
calling out the omission and an sql_typical_full() variant that
includes Events + IndexFts + IndexVector for backends that match
this crate's shape.

F8: validate_simple_identifier indexed bytes[0] after an is_empty
guard. The guard makes the index sound, but the pattern is fragile
to refactors. Switch to bytes.first() so the dependency is explicit
and the panic path goes away.

F9: Multiple doc comments in record.rs and index.rs referenced
stale type names (StorageBackend::put/list/query, Record). Update
to the current RootFilesystem / Entry names.

* fix(engine): dedupe events on append_events for HybridStore parity

HybridStore (`src/bridge/store_adapter.rs:1613`) de-duplicates thread
events by id before insert. The filesystem-store `append_events` impl
was previously writing with `CasExpectation::Any`, which silently
overwrote an existing event with the same id when callers re-emitted
(e.g. recovery after a partial flush).

Pre-read the destination path and skip any id already present.
Matches HybridStore's append-only contract.

Audit finding F3 (Medium) from the ironclaw_engine crate audit.

* fix(memory): map FTS results to documents in FilesystemMemoryDocumentRepository::search_documents

The previous scaffold issued the `Filter::Fts` query, then silently
dropped the results with `let _ = results; Ok(Vec::new())`. A caller
wiring up the trait would see an empty result set and assume "no
matches" — when in fact the search had simply lied. That is worse
than returning `Unsupported`.

Map each `VersionedEntry.path` (added in PR #3659) back to a
`MemoryDocumentPath`, de-dupe by path, and assign a per-rank score
from RRF over the FTS-only branch so the result vector matches the
native repos' fusion contract for the trivial single-branch case.

Skip non-memory-document entries that may live under the same prefix
(chunk projections, metadata siblings). Adds
`list_documents_drains_pages_beyond_max_limit` test against the
in-memory backend.

Audit finding F2 (HIGH) from the ironclaw_memory crate audit.

* fix(secrets): close revoke CAS-loop race with versioned compare-and-swap

`revoke` previously read the lease via the (now-removed)
`read_lease` helper and wrote with `CasExpectation::Any`. The
per-lease process-local mutex serialized writers within one process
only — multi-process callers sharing the same backend root could
observe `Active`, race against `consume`, and clobber a `Consumed`
marker by overwriting it with `Revoked`.

Inline the read into a bounded CAS retry loop matching `consume` and
`consume_session_use`: read with version, write with
`CasExpectation::Version`, retry on `VersionMismatch`. Make revoke
idempotent on terminal states (`Consumed`, `Revoked`, `Expired`) so
the loop converges even when a winner has already written.

Audit finding F2 (Medium) from the ironclaw_secrets crate audit.

* fix(processes): use versioned CAS for status transitions

`update_status` previously read the record and wrote with
`CasExpectation::Any`, relying on the per-instance `transition_lock`
for atomicity. That lock only serializes within one process; a
multi-process deployment sharing the same backend root could observe
identical pre-transition state in both processes and clobber each
other's status flips.

Replace with a bounded CAS retry loop: read with version, validate
the transition, write with `CasExpectation::Version…
theredspoon pushed a commit that referenced this pull request Jun 21, 2026
…earai#3573)

* feat(reborn): add ironclaw_hooks framework foundation (#3524)

Foundation slice of the Reborn loop hooks framework per nearai/ironclaw#3524.
Lands the trust primitives, sealed decision types, dispatcher contract, and
extension manifest schema; no Reborn middleware composition yet (next slice
wires HookDispatcher into LoopCapabilityPort / LoopPromptPort).

Design comment on #3524:
https://github.com/nearai/ironclaw/issues/3524#issuecomment-4439890144

What this PR ships
==================

* `crates/ironclaw_hooks/` — new crate
  * `identity` — content-addressed `HookId` (blake3 of length-prefixed
    extension + local + version fields). Same versioning primitive the rest
    of Reborn should converge on for replay safety.
  * `trust` — `HookTrustClass` enum (Builtin / Trusted / Installed) with
    per-kind default attenuation. Trust class is fixed by source, never
    declarable.
  * `kinds/` — sealed decision DTOs. `BeforeCapabilityHookDecision`,
    `HookPatch`, `ObserverFact` all have `pub` outer struct + `pub(crate)`
    inner enum + `pub(crate)` constructors. Same #3460 witness pattern.
  * `points/` — typed read-only contexts for each hook point.
  * `sink` — split sink traits per trust tier. `PrivilegedGateSink` exposes
    `allow()`; `RestrictedGateSink` does not. An Installed-tier hook
    literally cannot mint Allow at the type level.
  * `ordering` — phase → priority → hook id, stable. Phases gated by trust
    (Validation/Authorization Builtin-only).
  * `failure_policy` — Timeout/Panic/Malformed/AttenuationViolation
    categories. Gate/Mutator fail closed, Observer/Effect fail isolated.
    Slot poisoning persisted for the rest of the run on any category.
  * `registry` — run-profile-sourced bindings; phase-vs-trust gate enforced
    at insert; poisoning surface for the dispatcher.
  * `dispatch` — HookDispatcher with deterministic ordering, panic
    catch-unwind via futures::FutureExt, per-hook tokio::time::timeout,
    short-circuit gate composition (Deny > PauseAuth > PauseApproval >
    Allow), Telemetry-phase observers always run.
  * `manifest` — serde types for the `[[hooks]]` section of extension
    manifests. Predicate vs WASM body; same_tenant scope requires explicit
    grant; Validation/Authorization phases rejected at parse time because
    manifest hooks are always Installed.
  * `predicate` — typed predicate language for declarative Installed hooks
    (DenyCapability, PauseApproval, RateOrValueCap). Evaluator lives in
    the dispatcher follow-up, not here.

* `crates/ironclaw_architecture/tests/reborn_dependency_boundaries.rs`
  * Added `ironclaw_turns` -> `ironclaw_hooks` to the forbidden list.
  * New BoundaryRule for `ironclaw_hooks` itself (cannot pull host_runtime,
    dispatcher, secrets, network, wasm, etc.).

* `Cargo.toml` workspace member registration.

What this PR deliberately does NOT ship
========================================

* Reborn middleware composition wrapping LoopCapabilityPort / LoopPromptPort
  with HookDispatcher. Next slice; ironclaw_reborn changes only.
* WASM hook execution path. Programmatic hooks parse and validate from
  manifest; the wasmtime integration lands when the WASM dispatcher seam is
  built.
* Predicate evaluation. Predicate types serialize and validate; the
  evaluator that turns a `RateOrValueCap` spec into a `Deny` decision is in
  the next slice alongside Reborn wiring.
* Event-triggered hooks (Phase 5 of the original roadmap).
* Self-authored hooks. Tracked separately at #3567 with monotonic-restriction
  + unforgeable-channel ratification.

Test plan
=========

* `cargo test -p ironclaw_hooks` — 47 tests (46 unit + 1 integration smoke
  for the manifest -> binding -> dispatch pipeline).
* `cargo test -p ironclaw_architecture` — 13 tests; new boundary rule
  passes, existing rules unaffected.
* `cargo clippy -p ironclaw_hooks --all-targets -- -D warnings` — clean.
* `cargo fmt -p ironclaw_hooks -- --check` — clean.
* `cargo check --workspace` — clean, no regressions in other crates.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(reborn): wire HookDispatcher into LoopCapabilityPort/LoopPromptPort

Follows the foundation slice (see initial commit). Adds the next layer:

1. Capability- and prompt-port middleware (`ironclaw_hooks::middleware`)
   * `HookedLoopCapabilityPort` runs `dispatch_before_capability` before
     every invocation, translates the composed decision into the existing
     `CapabilityOutcome` vocabulary (Deny / PauseApproval / PauseAuth all
     map to `Denied` for now; gate-ref plumbing for real pause semantics
     lands in the next slice).
   * `HookedLoopPromptPort` runs `dispatch_before_prompt` before bundle
     construction. Observe-only for snippets in this slice; actual
     snippet injection waits for the shared `prompt_envelope::wrap_untrusted`
     helper (#3540 / #3471).

2. Declarative predicate evaluator (`ironclaw_hooks::evaluator`)
   * `DenyCapability` and `PauseApproval` predicates: stateless, evaluated
     directly against `BeforeCapabilityHookContext`.
   * `RateOrValueCap` with `InvocationCount` bound: sliding-window counter
     keyed by `(hook_id, capability_name)`, in-memory only. Window
     parsing supports `s`/`m`/`h`/`d` units; unparseable windows fail
     closed.
   * `NumericSum` bound: types implemented but evaluation returns Allow
     and emits a warn-level audit. Full argument-extraction story is a
     follow-up slice once capability arguments become hook-visible.
   * `PredicateEvaluator::evaluate_at(...)` test variant accepts an
     explicit `Instant` so sliding-window tests don't depend on
     real-clock progress.

3. Manifest -> dispatcher glue (`ironclaw_hooks::installed_hook`)
   * `PredicateBackedBeforeCapabilityHook` wraps a `HookPredicateSpec`
     plus an `Arc<PredicateEvaluator>` and implements
     `RestrictedBeforeCapabilityHook`. The registry installer would
     construct one of these per `[[hooks]]` entry whose body is
     `HookManifestBody::Predicate`.
   * Sink reasons are `&'static str`, so the dynamic predicate `reason`
     surfaces in audit (via the evaluator's `EvaluatorDecision`) rather
     than the model-visible decision. Closed-vocabulary labels carry
     through to the sink.

4. Reborn composition seam (`ironclaw_reborn::loop_driver_host`)
   * `RebornLoopDriverHostFactory::with_hook_dispatcher(Arc<HookDispatcher>)`
     opt-in builder method. When set, the factory wraps the capability
     and prompt ports with the hooked middleware. Default behavior
     (no dispatcher) is unchanged from the pre-hooks shape, so existing
     callers continue to work.
   * Added `ironclaw_hooks` as a dep in `ironclaw_reborn`.

Test plan
=========

* `cargo test -p ironclaw_hooks` — 60 tests pass (59 unit + 1
  integration smoke; +13 vs the foundation commit covering middleware,
  evaluator, installed_hook).
* `cargo test -p ironclaw_reborn` — 118 tests pass; no regressions
  from adding the dep.
* `cargo test -p ironclaw_architecture` — 13 tests pass; the
  `ironclaw_turns -> ironclaw_hooks` boundary still holds and the new
  `ironclaw_hooks` rule (no host_runtime / dispatcher / secrets /
  network / wasm / reborn) is unaffected.
* `cargo clippy -p ironclaw_hooks --all-targets --all-features
  -- -D warnings` — clean.
* `cargo clippy -p ironclaw_reborn --all-targets -- -D warnings` —
  clean.
* `cargo fmt --all -- --check` — clean.

What still defers
==================

* WASM hook execution path.
* Persistent predicate counter (in-memory only for now).
* Argument-extraction so `NumericSum` predicates evaluate against
  capability arguments.
* Gate-ref plumbing so PauseApproval / PauseAuth surface real
  `CapabilityOutcome::ApprovalRequired` instead of `Denied`.
* Prompt-snippet injection (waits for shared envelope helper).
* Event-triggered hooks.
* Self-authored hooks (#3567).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(reborn): add HookedLoopModelPort/TranscriptPort/CheckpointPort observer middleware

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(reborn): end-to-end hooks integration through RebornLoopDriverHostFactory

Adds crates/ironclaw_reborn/tests/hooks_integration.rs covering the
factory's HookDispatcher wiring seam end-to-end. Tests drive
host.invoke_capability(...) (not dispatcher.dispatch_before_capability(...)
directly) so a regression in RebornLoopDriverHostFactory's wrapping
composition surfaces here.

Scenarios:
- PredicateBackedBeforeCapabilityHook (DenyCapability NameEquals
  "cap.blocked") short-circuits invocation; inner port never called;
  outcome is Denied(unknown("hook_denied")).
- A privileged selective hook that allows non-matching capabilities
  proves the wrapper does not blanket-deny: cap.allowed reaches the
  inner port and completes once.
- Factory built without with_hook_dispatcher() lets cap.blocked through
  to the inner port, proving the hook plumbing is genuinely opt-in.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(reborn): add pass() + HookRegistrar + self-authored hooks scaffolding

Three additions to ironclaw_hooks:

B. `pass()` on gate sinks — distinguishes "evaluated, no opinion" from
   "returned without minting a decision." A passing hook contributes
   nothing to the composed decision; a silent hook is still Malformed
   and fails closed. `PredicateBackedBeforeCapabilityHook` now routes
   the evaluator's `Allow` decision through `sink.pass()` instead of
   the previous `deny("hook_predicate_pass")` workaround.

A. `HookRegistrar` bridge — converts a `Vec<HookManifestEntry>` into
   `HookBinding`s + dispatcher impls in one call. Predicate bodies are
   wired through `PredicateBackedBeforeCapabilityHook`; WASM bodies
   return `HookError::RegistryConstruction` for now. Adds
   `HookDispatcher::insert_binding` so the registrar can mutate the
   registry through the dispatcher rather than reach inside.

I. Self-authored hooks scaffolding — fourth `HookTrustClass` variant
   for hooks the agent authors at runtime. Run-scoped only;
   monotonic-restriction sink with no `allow`, no trusted-snippet path,
   no effect-class constructor. Closed-vocabulary `SelfAuthoredReason`
   enum keeps free-text reasons off the audit seam.
   `SelfAuthorshipProvenance` captures authoring run/turn, timestamp,
   spec digest, optional user ratification, and a generation-trace
   pointer. Durable persistence depends on the unforgeable channel
   from #3564 and lands separately.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(reborn): real gate-ref plumbing for hook PauseApproval/PauseAuth decisions

Previously, `GateDecisionInner::PauseApproval` and `PauseAuth` returned by
hooks were degraded to `CapabilityOutcome::Denied` at the middleware
boundary because the hook crate had no way to mint a `LoopGateRef` scoped
to the current run. Hooks that wanted to pause the loop for approval or
auth instead failed the call closed, leaving the host's approval-router
machinery unreachable from hook code.

This change introduces a `HookGateRefFactory` trait in
`ironclaw_hooks::middleware::gate_ref` that mints `LoopGateRef`s for
pause-class decisions. `HookedLoopCapabilityPort` now takes an
`Arc<dyn HookGateRefFactory>`, defaulting to `UuidHookGateRefFactory` (a
locally-unique opaque-id factory suitable for tests and the foundation
slice). Production deployments override via `.with_gate_ref_factory(...)`
with a factory bound to the current `LoopRunContext` and the host's
gate-router.

The translation in `decision_to_outcome` is now async so it can await the
factory. `PauseApproval` maps to `CapabilityOutcome::ApprovalRequired
{ gate_ref, safe_summary }` and `PauseAuth` to `AuthRequired`. If the
factory itself errors, the middleware falls back to `Denied` with a
sanitized `hook_gate_ref_unavailable` reason kind so the loop fails
closed rather than routing through an unresolvable suspension. The
underlying error text is dropped to avoid leaking gate-router state into
model-visible output.

Tests:
- `pause_approval_decision_surfaces_as_approval_required`,
  `pause_auth_decision_surfaces_as_auth_required`,
  `gate_ref_factory_failure_falls_back_to_denied` in
  `middleware::capability_port::tests`.
- `pause_approval_hook_surfaces_as_approval_required_with_real_gate_ref`
  in `crates/ironclaw_reborn/tests/hooks_integration.rs`, exercising the
  full `RebornLoopDriverHostFactory` composition with the default
  `UuidHookGateRefFactory`.
- Gate-ref factory unit tests in `gate_ref::tests`.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(reborn): add NumericSum predicate evaluation with capability argument extraction

Wires the missing argument-extraction story for the predicate evaluator so
`ValueOrRateBound::NumericSum` actually enforces a rolling numeric cap
instead of warn-and-allowing.

- Extend `BeforeCapabilityHookContext` with a sealed `SanitizedArguments`
  view. Strings truncate to 256 bytes; objects/arrays cap at 8-deep.
  `extract_numeric` supports dotted + bracketed paths (`order.amount`,
  `items[0].price`) and returns `Option<rust_decimal::Decimal>`. The inner
  representation is sealed so external callers can't bypass bounds.

- Introduce `CapabilityInputResolver` + bundled `NullCapabilityInputResolver`
  in `middleware/resolver.rs`. The hooks crate intentionally doesn't know
  how to dereference a `CapabilityInputRef` — that knowledge belongs to
  the production host. Until a real resolver is wired in (follow-up),
  arguments are `Unresolved` and `NumericSum` fails closed.

- `HookedLoopCapabilityPort::new` defaults to the null resolver; new
  builder `.with_resolver(Arc<dyn CapabilityInputResolver>)` overrides.

- `PredicateEvaluator` gains a tenant-keyed `value_history` map. The
  `NumericSum` arm parses `max` + `window`, extracts the numeric value
  from sanitized args, accumulates within the rolling window, and applies
  `on_exceeded` when the sum exceeds the cap. Unresolved args, missing
  field, non-numeric field, unparseable max, and unparseable window all
  fail closed via the configured `OnExceededAction`.

- Add `BeforeCapabilityHookContext::new_unresolved(...)` convenience
  ctor; existing test sites switch to it instead of churning every call
  site through the 4-arg ctor.

Test count: +14 (8 new SanitizedArguments tests, 6 new NumericSum
evaluator tests, 1 null-resolver test; one old NumericSum-stub-related
gap closed).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(reborn): seal hook registration trust boundary + dispatcher hardening

Addresses blocking findings from the security audit of `ironclaw_hooks`:

- C1 (Blocking, Trust Model): "Installed cannot Allow" was not enforced
  at the registration boundary. `BeforeCapabilityHookImpl::Privileged`
  was a public variant, so external crates with dispatcher access could
  construct an Installed binding paired with a Privileged impl and bypass
  the sink trait restriction. Sealed `BeforeCapabilityHookImpl`,
  `BeforePromptHookImpl`, and `ObserverHookImpl` to `pub(crate)` and
  replaced the single generic `install_before_capability` /
  `install_before_prompt` / `install_observer` surface with tier-specific
  public installers (`install_builtin_*`, `install_trusted_*`,
  `install_installed_*`) that build the binding with the matching trust
  class internally. Updated registrar, internal middleware tests, the
  hooks foundation pipeline test, and the reborn `hooks_integration`
  test to drive the new surface. Added regression tests proving the
  trust class is set by the installer and that the seal is type-level.

- C5 (Medium, Slot Poisoning): same-dispatch poisoning was incomplete
  because `ordered_bindings` snapshots once at the top of the loop, and
  `HookRegistry::insert` accepted duplicate hook IDs. Rejected duplicate
  hook IDs (any point) in `HookRegistry::insert` and added a poison
  re-check before invoking each hook impl in `dispatch_before_capability`,
  `dispatch_before_prompt`, and `dispatch_observer_at`. Added regression
  tests for both behaviors.

- C6 (Medium, Manifest / Predicate Validation): `parse_window` could
  panic on non-ASCII input because `split_at(len - 1)` requires a char
  boundary. Rewrote to compute the unit char's UTF-8 byte length and
  slice safely, added a public `validate_window` helper, and wired it
  into `HookManifestEntry::validate` for both `InvocationCount` and
  `NumericSum` bounds. Added tests for non-ASCII, empty, single-char,
  and zero-duration windows.

- C2 (High, Tenant Isolation): partial fix only. The
  `PredicateEvaluator`'s sliding-window counter was keyed by
  `(hook_id, capability)`, so cross-tenant state could leak. Extended
  `HistoryKey` to include `tenant_id` and added a regression test
  proving counters partition by tenant. Documented the broader
  dispatcher-per-build / per-run-fresh-dispatcher pattern as deferred
  follow-up in `crates/ironclaw_hooks/CLAUDE.md`.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(reborn): emit hook telemetry milestones for audit/SSE observers

Wires the hook dispatcher into the host's milestone stream so audit
backends and SSE observers can see hook activity. Previously, hook
dispatch was invisible — denies, pauses, failures, and observer fires
left no trace in the host's observability backend.

Changes:

- `ironclaw_turns`: add `HookDispatched`, `HookDecisionEmitted`, and
  `HookFailed` variants to `LoopHostMilestoneKind`, with a closed-
  vocabulary `HookDecisionSummary` enum (Allow/Deny/PauseApproval/
  PauseAuth/Pass/Patch). Introduce a lightweight `HookMilestoneSink`
  trait that emits hook-specific *kinds* without requiring a
  `LoopRunContext` (the dispatcher is a process-wide singleton that
  cannot own a per-run context), plus a `RunScopedHookMilestoneSink`
  adapter that injects run context and forwards to the existing
  `LoopHostMilestoneSink`. Also add `InMemoryHookMilestoneSink` for
  tests.

- `ironclaw_hooks`: add a `telemetry` module that converts hook-crate
  types (`HookId`, `HookTrustClass`, `HookPointSpec`, `FailureCategory`,
  `FailureDisposition`, `BeforeCapabilityHookDecision`) into the wire-
  shape labels and summaries the milestone sink expects. Hook ids cross
  the seam as hex strings because the strongly-typed `HookId` cannot be
  imported from `ironclaw_turns` (the architecture test enforces
  `ironclaw_turns -> ironclaw_hooks` stays absent).

- `ironclaw_hooks::dispatch`: add an optional `Arc<dyn
  HookMilestoneSink>` to `HookDispatcher`, set via
  `with_milestone_sink`. Emit `HookDispatched` before each hook runs,
  `HookDecisionEmitted` after a decision/pass/patch, and `HookFailed`
  on timeout/panic/malformed/missing-impl across all three dispatch
  paths (before_capability, before_prompt, observer). Default behavior
  (no sink attached) emits nothing — preserves the pre-telemetry
  observable surface.

- `ironclaw_reborn`: document on `with_hook_dispatcher` that callers
  attach the milestone sink to the dispatcher *before* wrapping it in
  `Arc` and installing it into the factory, using a
  `RunScopedHookMilestoneSink` to inject run-context. The dispatcher
  itself is shared across runs, so attaching a fixed run-context inside
  it would be wrong. Update `RuntimeEvent` projection in
  `milestone_events.rs` to ignore the new hook kinds (no projection
  pathway yet; emitted milestones are consumed by SSE observers
  directly).

Tests:

- `ironclaw_hooks::dispatch`: 5 new tests covering milestone emission
  for deny decisions, panic failures, prompt-mutator patches, observer
  pass-throughs, and the no-sink default.
- `ironclaw_reborn` hooks_integration: end-to-end test wiring a
  `RunScopedHookMilestoneSink` onto the dispatcher and asserting hook
  activity surfaces in the host's `LoopHostMilestoneSink`.

Total: +6 hook telemetry tests; no existing tests modified.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(reborn): extract shared prompt envelope; inject hook patches into prompt bundle

Adds `ironclaw_prompt_envelope`, a leaf crate that owns the single envelope
primitive used by every model-visible untrusted-content path. `wrap_untrusted`
prefixes content with a closed-vocabulary `<Trusted|Untrusted> <source>
content: ` marker, rejects bodies carrying instruction-hijack phrases
(`ignore previous instructions`, `<|im_start|>`, `<system>`, etc.), and
enforces a 4 KiB byte budget by default.

Migrates `ironclaw_host_runtime::memory_context` to delegate envelope
wrapping, marker rejection, and control-character stripping to the new
crate while keeping the `LoopSafeSummary`-specific 512-byte cap and byte
truncation local. Existing memory_context behavior and tests are preserved.

Wires the same envelope into `ironclaw_hooks`:

* `HookPatch::add_enveloped_snippet` now takes a raw body and wraps it
  via `wrap_untrusted(EnvelopeSource::Hook, …)`. `Installed` hooks
  produce `Untrusted` envelopes; `Builtin`/`Trusted`/`SelfAuthored`
  produce `Trusted` envelopes so downstream readers can distinguish the
  two paths through a uniform marker.
* `HookedLoopPromptPort::build_prompt_bundle` is no longer observe-only.
  After dispatching `before_prompt`, it envelope-wraps every snippet
  patch (passing `Enveloped` through, wrapping `Trusted` with the
  envelope helper), enforces the 4 KiB aggregate snippet byte budget
  across patches, and appends the wrapped snippets to the prompt
  bundle's `messages` as `system`-role `LoopModelMessage` entries
  carrying deterministic `msg:hook.<ordinal>.<hash>` content refs
  (mirroring the skill-snippet ref convention).

The envelope crate is a leaf with no ironclaw dependencies, satisfying
the boundary contract; the existing `ironclaw_hooks` boundary rule in
`reborn_dependency_boundaries` continues to hold because
`ironclaw_prompt_envelope` is not on its forbidden list.

Test count delta:
* `ironclaw_prompt_envelope`: +13 new tests (crate did not exist).
* `ironclaw_hooks`: 84 → 88 tests (+4 prompt-port behavior tests:
  `hook_patch_appended_as_envelope_wrapped_message`,
  `total_byte_budget_enforced_across_patches`,
  `instruction_hijack_in_patch_rejected`,
  `trusted_hook_patch_wrapped_with_trust_marker`).
* `ironclaw_host_runtime` memory_context: unchanged (8 tests still pass).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix: align tenant-counter test with SanitizedArguments-extended context ctor

* docs(reborn): document loader contract; pin HookId hex format

Add a "Loader responsibility" section to ironclaw_hooks/CLAUDE.md
explaining that tier-specific installers prevent minting wrong-tier
impls but cannot enforce origin — that's the loader's job — and
recommending registry loaders type-tag extension hooks as
LoadedHook::Installed at the loader seam.

Add tier_specific_installers_are_documented_as_loader_contract as a
regression guard that touches every public install_*_before_capability
and install_*_before_prompt method so any signature change forces the
loader contract to be re-evaluated.

Document HookId::to_hex's 64-char lowercase hex output as part of the
cross-crate contract consumed by LoopHostMilestoneKind::Hook* in
ironclaw_turns; add hook_id_hex_format_is_stable_64_lowercase_chars in
identity::tests and hook_id_string_serialization_matches_to_hex in
telemetry::tests to pin the format and the seam conversion path.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(reborn): pin hook milestone JSON schema + assert pairing invariants

Add L3 schema-snapshot tests for every hook-related LoopHostMilestoneKind
variant (HookDispatched, HookDecisionEmitted per HookDecisionSummary,
HookFailed per FailureCategory) so downstream consumers can rely on the
JSON wire shape and any accidental field rename, enum-tag rename, or type
change fails loudly.

Add L4 pairing-invariant matrix test in the hook dispatcher that drives
every observable outcome (Allow, Deny, PauseApproval, PauseAuth, Pass,
Panic, Timeout, Malformed, MissingImpl) through a recording milestone
sink and asserts the dispatched-then-terminator pairing shape. Document
the MissingImpl path as the one case that emits a sole HookFailed with
no preceding HookDispatched (the dispatcher discovers the protocol
violation before the hook is actually dispatched).

Add a multi-hook dispatch test that installs three hooks with mixed
outcomes (allow/deny/panic) at the same point and asserts each hook
produces its own paired sequence in the deterministic
(phase, priority, hook_id) order taken from the dispatcher's registry.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(reborn): integration tests for observer middleware through RebornLoopDriverHostFactory

Wire the HookedLoopModelPort / HookedLoopTranscriptPort /
HookedLoopCheckpointPort observer wrappers into
RebornLoopDriverHostFactory::build_text_only_host_with_capabilities,
mirroring the existing HookedLoopCapabilityPort / HookedLoopPromptPort
composition. The wrappers are applied only when a HookDispatcher is set
on the factory, so the default factory shape is unchanged.

Add four integration scenarios in crates/ironclaw_reborn/tests/hooks_integration.rs:

- observer_hook_fires_after_model_through_factory
- observer_hook_fires_after_capability_through_factory
- observer_hook_fires_after_checkpoint_through_factory
- observer_panic_does_not_fail_model_call (panic-isolation regression)

Relax the test-fixture model gateway from "panic if invoked" to
returning a stub assistant reply so the AfterModel / panic-isolation
tests can drive stream_model through the wrapped port. The existing
capability-port tests never touch the gateway, so their behavior is
unchanged.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(reborn): introduce HookDispatcherBuilder for type-enforced sink wiring

Adds a `HookDispatcherBuilder` in `ironclaw_hooks::dispatch` that owns
the dispatcher construction lifecycle: registry -> optional timeout ->
optional milestone sink -> installed hooks -> `.build_arc()`. The
terminal `.build_arc()` wraps in `Arc` and yields an immutable handle.

Tightens the public surface on `HookDispatcher`: `new`, `with_timeout`,
`with_milestone_sink`, and every `install_*_*` method are now
`pub(crate)`. Outside callers route exclusively through the builder, so
"wire the milestone sink before Arc-wrapping" is a compile-time fact
rather than a documentation convention.

`HookRegistrar::install` now takes a `HookDispatcherBuilder` by value
and returns `(HookDispatcherBuilder, Vec<HookId>)`, keeping the builder
chainable through manifest installation.

`RebornLoopDriverHostFactory` gains `with_hook_dispatcher_builder` to
let callers defer `.build_arc()` to the factory — a step toward the
FU8 per-build dispatcher pattern.

Migrates `foundation_pipeline.rs` and `hooks_integration.rs` to the
builder. Internal middleware and dispatch tests continue to use the
crate-private `HookDispatcher::new` directly.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(reborn): production CapabilityInputResolver for NumericSum predicates

Adds HookCapabilityInputResolverAdapter in ironclaw_reborn that bridges
the existing LoopCapabilityInputResolver (already used by
HostRuntimeLoopCapabilityPort for dispatch input resolution) to the
hooks crate's CapabilityInputResolver trait. RebornLoopDriverHostFactory
gains with_capability_input_resolver(...), and when both a hook
dispatcher and resolver are configured the factory threads the adapter
into HookedLoopCapabilityPort::with_resolver — so NumericSum and other
argument-dependent predicates evaluate against real, sanitized inputs
instead of failing closed against the framework's null default.

The adapter also enforces a configurable serialized-byte budget
(default 64 KiB) as defense in depth ahead of the hooks crate's
per-string and depth caps in SanitizedArguments.

Unit tests cover the four adapter branches (resolved JSON,
inner-error → None, non-object pass-through, oversized → None) and a
new end-to-end integration test
(numeric_sum_predicate_caps_total_value_against_real_inputs) drives the
full factory wiring: with a NumericSum cap of 99 over an "amount" field,
two invocations carrying {"amount":"50"} let the first pass through and
deny the second at the hook seam, with the inner port reached exactly
once.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(reborn): per-build HookDispatcher for full per-run isolation (C2)

Introduce `with_hook_dispatcher_factory(F)` on
`RebornLoopDriverHostFactory`. The closure is invoked once per
`build_text_only_host*` call, so dispatcher-owned mutable state — slot
poisoning, registry mutations, predicate-counter siblings — is scoped to
a single host build instead of shared across every host the factory
produces.

The legacy `with_hook_dispatcher(Arc<HookDispatcher>)` adapter is kept as
a thin wrapper that returns clones of the same `Arc` on every build. Its
shared-state behavior is now documented as an explicit opt-in for
backward compat; new wiring should prefer the factory closure.

Adds two regression tests:
  - `per_build_dispatcher_state_does_not_leak_across_runs` — installs a
    panicking hook, builds two hosts back-to-back, and proves the inner
    port is never reached on build 2 (fresh slot still applies the
    fail-closed deny). Pins per-run isolation.
  - `legacy_with_hook_dispatcher_shares_state_across_builds` — pins the
    shared-state semantic of the legacy adapter as the explicit baseline.

Migrates `predicate_deny_hook_short_circuits_inner_port` to the new
factory-closure path so the new wiring is exercised by the existing
suite.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(reborn): project hook telemetry milestones into RuntimeEvent for durable audit

Extend the runtime event substrate with `HookDispatched`, `HookDecisionEmitted`,
and `HookFailed` kinds carrying closed-vocabulary labels and the blake3-hex hook
identity. Project the matching `LoopHostMilestoneKind::Hook*` variants in
`DurableLoopHostMilestoneSink` so hook telemetry now lands in the same durable
event log as model/reply/loop milestones — SSE observers still see live hook
events, and audit replay can reconstruct the full hook trail.

- `ironclaw_events`: add hook variants to `RuntimeEventKind`, optional hook
  fields on `RuntimeEvent` (`hook_id`, `hook_point`, `hook_trust_class`,
  `hook_decision`, `hook_failure_category`, `hook_failure_disposition`),
  typed constructors (`hook_dispatched`, `hook_decision_emitted`,
  `hook_failed`), and dedicated sanitizers (`sanitize_hook_label`,
  `sanitize_hook_id`) re-run on every wire crossing. No new crate dependency
  edges; hook strings cross the boundary opaque.
- `ironclaw_reborn::milestone_events`: project the three hook milestone kinds
  via a new `loop.hook` capability id. `HookDecisionSummary` is collapsed to
  its closed-vocabulary `kind_name()` so sanitized reasons never enter the
  durable substrate.
- `ironclaw_event_projections`: extend `TimelineEntryKind` and the
  `RuntimeEventKind -> RunProjectionStatus` mapping so hook events are pure
  telemetry — they preserve the current run status rather than changing it.
- Tests: 4 unit tests in `ironclaw_events::runtime_event::tests` (serde
  round-trip per variant + unsafe-label collapse), 3 in
  `ironclaw_reborn::milestone_events::tests` (projection per variant,
  including the assertion that raw `Deny { reason }` text does not reach the
  durable wire payload). Existing replay-projection direct-construction
  tests updated for the new RuntimeEvent fields.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(reborn): enforce manifest-declared hook scope at dispatch time (C3)

Audit finding C3: extensions could declare `[[hooks]]` with
`scope = "own_capabilities"` in their manifest, but the dispatcher never
enforced it — an Installed hook from ext-A could fire against capabilities
provided by ext-B. Scope was parsed but not load-bearing.

This change makes scope load-bearing end-to-end:

- `BeforeCapabilityHookContext` carries an optional `provider:
  ironclaw_host_api::ExtensionId` populated by the middleware. The hook
  context is `#[non_exhaustive]` already so this is non-breaking.

- `HookBinding` gains `owning_extension: Option<ExtensionId>` and `scope:
  HookBindingScope`. `HookBindingScope` is `Global` / `OwnCapabilities`
  / `SameTenant`. Builtin and Trusted bindings default to `Global` and
  carry no `owning_extension`; Installed bindings carry both, sourced
  from the manifest.

- `HookDispatcher::install_installed_*` installers now require the
  caller to pass `(owning_extension, scope)`. The registrar derives both
  from the manifest entry, so manifest authorship is the single source
  of truth.

- A new `CapabilityProviderResolver` trait + bundled
  `NullCapabilityProviderResolver` lets the middleware lift the
  capability id to its provider at invocation time. The middleware
  wires the resolved provider into the hook context.

- `dispatch_before_capability` consults `binding.scope.permits(...)`
  before invoking each hook. Bindings that don't permit the current
  invocation are inert — no sink call, no failure record, no poisoning.

Conservative defaults:

- When the provider resolver returns `None` (no resolver wired, or the
  capability has no known provider), `OwnCapabilities`-scoped hooks do
  NOT fire. An attacker cannot bypass scope filtering by stripping
  provider info from the descriptor.

Tests:

- 5 new dispatcher tests cover OwnCapabilities matching, foreign
  provider, unresolved provider, SameTenant, and Builtin Global.
- 1 new registrar test asserts manifest scope and extension propagate
  into `HookBinding`.
- 1 new middleware test asserts the provider resolver populates the
  hook context.
- 1 new integration test in `ironclaw_reborn` proves an ext-A hook
  scoped to `OwnCapabilities` does not intercept invocations that have
  no resolved provider (the production composition default).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* style: rustfmt dispatch.rs after FU1 merge

* docs(hooks): prior-art comparison against LSM/eBPF/Envoy/K8s/OPA/CRX/VSC/Tauri

Validates the IronClaw hooks design against 8 established hook/policy
systems across 8 axes (dispatch, trust tiers, attenuation, decision
vocabulary, failure semantics, isolation, manifest, audit).

Surfaces:
- 7 areas where ICLAW stands out vs prior art (type-level trust
  enforcement, dispatch-time scope, failure-kind matrix, pause-with-
  gate-ref, pairing-invariant audit matrix, tenant-keyed predicates,
  phase-ordered dispatch)
- 4 conventional choices we should revisit (in-process Installed-WASM,
  sticky poison, no formal dispatch model, no installation rate-limit)
- 3 divergences whose 'why' is weak and need design review

* docs(hooks): STRIDE threat model for v1 framework

Enumerates 7 adversary classes (A1-A7), 6 assets ranked by blast
radius, and ~35 attack vectors across STRIDE categories with mitigations,
existing tests, and residual risk.

Surfaces 7 prioritized follow-ups:
- High: per-extension hook-count cap (D3/D4)
- High: gate-ref unguessability + one-shot test (S1)
- Med: resolver field-level scope (I2)
- Med: per-evaluator state ceiling (D5)
- Med: poison-stickiness operator runbook
- Low: timing side-channel residual acknowledgement (I4)
- Low: instruction-marker denylist periodic review (I5)

Confirms the load-bearing 'Installed cannot Allow' (E1) property holds
via type-level seal + tier-specific installers, backed by
compile_time_seal_test and installed_binding_cannot_be_paired_with_
privileged_impl tests.

Explicit out-of-scope: extension install pipeline (#3492), WASM exec
sandbox (needs separate threat model when it lands), approval gateway
(#3564).

* feat(hooks): close threat-model gaps S1 (gate-ref entropy) and D3/D4 (registration flood)

S1 (gate-ref unguessability, factory side):
- Three new tests on `UuidHookGateRefFactory`:
  - `gate_refs_are_v4_uuids` pins the v4 entropy source (122 random
    bits per ref per RFC 4122 §4.4); fails if a future change moves to
    a counter or weaker UUID version.
  - `gate_refs_have_no_collisions_across_many_calls` mints 20k refs
    across both namespaces, asserts zero collisions (statistical
    proxy for entropy quality).
  - `approval_and_auth_namespaces_do_not_overlap` confirms prefix
    routing separation.
- Doc comment now documents the security property explicitly and
  delineates factory-side vs gateway-side responsibilities for the
  one-shot consumption property.

D3/D4 (hook registration flood):
- New `MAX_HOOKS_PER_EXTENSION = 32` and
  `MAX_HOOKS_PER_EXTENSION_PER_KIND = 8` consts in `registrar.rs`.
- New `HookRegistrar::enforce_registration_caps` runs pre-flight at
  the top of `install()`, before any binding is inserted. Whole-batch
  rejection means a partially-installed batch cannot slip past.
- Three regression tests: total-cap rejection, per-kind-cap rejection,
  at-cap acceptance.
- Error messages cite the threat-model finding so operators can map
  rejection back to the design rationale.

Threat model updated: S1, D3, D4 marked closed in the cross-cutting
properties matrix and the open-follow-ups list.

* test(hooks): three real hooks built against the public API + ergonomics findings

Builds three representative hooks from outside the crate, mimicking
what an extension or system author would actually write:

1. polymarket-daily-cap — Installed predicate hook, InvocationCount
   rate-cap with Deny on excess. Canonical 'rate-limit a capability'
   use case for the predicate language.

2. large-stake-approval-gate — Installed predicate hook, NumericSum
   over amount_usd field, PauseApproval at $1000/24h. Manifest-shape
   + registrar-install coverage from outside Reborn; end-to-end
   dispatch lives in ironclaw_reborn integration tests because
   NumericSum needs resolved args (a friction finding documented in
   the companion doc).

3. pii-redaction-warning — Trusted Rust hook implementing
   PrivilegedBeforePromptHook, injects a trusted instruction snippet
   reminding the model to redact PII. Demonstrates the path a system
   author takes when the predicate language isn't expressive enough.

API change (F1 fix): SanitizedArguments::unresolved() promoted from
pub(crate) to pub. This is the documented safe default — predicates
that need args must fail closed against it — so exposing the
constructor cannot weaken any trust property. The sanitizing
from_json constructor stays sealed; that's the trust boundary.
Without this fix, external hook authors could not construct a
BeforeCapabilityHookContext with both a known provider AND
unresolved args, which made TDD of their own predicate impossible.

Findings documented in docs/real-hooks-findings.md, ranked by
severity. Big-picture observation: writing the Trusted Rust hook
(F4) was easier than writing the declarative predicate hook (F1 +
F2 + F3) — three of seven findings target predicate-authoring
ergonomics. The declarative path needs the most polish before
third-party extension authors will trust it for non-trivial policy.

Tests: 6 new in real_hooks.rs, all pass.

* feat(hooks): close all remaining threat-model and ergonomics gaps

Closes the Med-priority threat-model gaps (I2, D5, poison runbook)
and all real-hooks ergonomics findings (F2, F3, F5, F6, F7) in a
single pass.

Threat model:
- I2 (resolver field-scope): documented in SanitizedArguments rustdoc.
  The narrow public surface (only is_resolved + extract_numeric)
  enforces field-scope by construction for the current predicate
  path. Reassess when Installed-WASM lands.
- D5 (evaluator state ceiling): MAX_HISTORY_KEYS = 8192 per map,
  LRU eviction with evictions_observed() metric for operator
  monitoring. New regression test
  lru_eviction_increments_counter_and_drops_oldest_key.
- Poison-stickiness runbook: new docs/operator-runbook.md with
  recovery options ranked by cost.

Ergonomics findings:
- F2 (closed-vocab deny reasons): rustdoc on OnExceededAction
  and GateDecisionView::Deny explaining the audit-vs-model split
  and why manifest reason text doesn't reach the model.
- F3 (NumericSum can't be TDD'd outside Reborn): new test-support
  feature flag with SanitizedArguments::for_tests(value) that
  external hook authors can opt into via dev-dep.
- F5 (two ExtensionId types): added
  From<&ironclaw_host_api::ExtensionId> impl for
  identity::ExtensionId, plus cross-link rustdoc.
- F6 (HookManifestEntry struct-literal fragility): added
  #[non_exhaustive] + HookManifestEntry::new(id, kind, body) +
  with_scope/with_phase/with_priority/with_description/with_requires_grant
  builder methods. Migrated 3 external call sites in tests/.
- F7 (priority guidance): rustdoc on HookPriority with when-to-
  deviate guidance, named FIRST/LAST constants documented for
  Builtin/Telemetry use cases.

Tests: 151 unit + 1 + 6 integration in ironclaw_hooks all pass
with --all-features. ironclaw_reborn (13 hooks_integration scenarios)
unchanged.

Threat model updated: I2 / D5 / poison runbook marked closed in
both the per-vector table and the cross-cutting properties matrix.
Open follow-ups now down to two Low items (I4 timing side-channel
residual, I5 instruction-marker denylist refresh) plus the deferred
DenyReasonCode enum from F2.

* fix(ci): collapse nested match in hooks_integration test for clippy --all-features

CI runs `cargo clippy --all --tests --examples --all-features -- -D warnings`
which is stricter than the workspace clippy I ran locally and trips
`clippy::collapsible_match` on the nested-if in HookDecisionEmitted
matching. Collapse the inner `if decision.kind_name() == "deny"`
into an arm guard.

* feat(hooks): address henrypark133 review — Critical #1/#2/#3/#5, Concerning #5/#7

Address composition-seam bugs in the Reborn factory wiring + doc tidy.

henrypark133 review findings addressed:

Critical #1 — before_prompt hook messages not materialized.
  HookedLoopPromptPort now requires a HookPromptMaterializationSink and
  fails closed if patches are emitted without one. The reborn factory
  installs an InstructionStoreBackedHookSink adapter that delegates to
  the host's InstructionMaterializationStore, so synthetic msg:hook.*
  refs are resolvable by the downstream model resolver. New seam trait
  (HookPromptMaterializationSink) keeps ironclaw_hooks decoupled from
  LoopRunContext.

Critical #2 — OwnCapabilities hooks were inert in production wiring.
  Factory now installs SurfaceBackedProviderResolver (consults the
  visible-capability surface for capability_id → provider). With this,
  ctx.provider is populated and OwnCapabilities-scoped Installed hooks
  actually fire against their own provider's capabilities.

Critical #3 — gate refs were unresolvable.
  Middleware default switched from UuidHookGateRefFactory to
  FailClosedHookGateRefFactory. Tests must explicitly opt into UUID
  (via with_gate_ref_factory) to exercise the affirmative ApprovalRequired
  path; production deployments must install a router-backed factory.
  New factory method RebornLoopDriverHostFactory::with_hook_gate_ref_factory.

Concerning #5 — AfterModel fired twice + before durable finalization.
  Removed AfterModel dispatch from HookedLoopModelPort; the transcript
  port's finalize_assistant_message is now the sole AfterModel boundary
  (the durable one). Model port wrapper is preserved as a no-op shim
  for symmetry + future model-response-observed point.

Concerning #7 — doc tidy:
  - CLAUDE.md: 3 trust classes → 4 (Builtin/Trusted/Installed/SelfAuthored
    with explicit note that SelfAuthored is run-scoped only and not
    loadable from an external source).
  - operator-runbook.md: "Audit log" → "durable runtime event stream"
    where the projection is actually the runtime-event stream, not formal
    AuditEnvelope records.
  - prior-art.md: poison-lifetime nuance — per-host-build with the
    factory pattern, process-lifetime only for the legacy adapter.
  - prior-art.md:80: trailing whitespace removed.

Testing gaps from henrypark133 — caller-level tests through
RebornLoopDriverHostFactory:
  #1 (before_prompt resolver path):
     before_prompt_hook_message_is_resolvable_via_factory_wiring
  #2 (OwnCapabilities positive/negative/unknown):
     own_capabilities_hook_fires_when_provider_matches
     own_capabilities_hook_does_not_fire_when_provider_differs
     own_capabilities_hook_does_not_fire_when_provider_unknown
  #3 (pause/auth gate lifecycle or fail-closed):
     pause_approval_with_default_factory_fails_closed_as_denied
     pause_approval_hook_surfaces_as_approval_required_with_real_gate_ref
     (updated to require explicit UuidHookGateRefFactory opt-in)
  #5 (AfterModel exactly-once at durable boundary):
     after_model_fires_exactly_once_at_durable_boundary

Still TODO from review (separate commits):
  Critical #4 (telemetry context — two-run attribution) + gap #4
  Concerning #6 (TimelineEntry hook metadata projection) + gap #6

Tests: 154 unit + 18 hooks_integration + all other reborn tests pass.
Workspace clippy + fmt + no-panics clean.

* feat(hooks): address remaining henrypark133 review — Critical #4, Concerning #6

Critical #4 — per-run hook telemetry attribution.
  New `HookDispatcherBuilderFactory` signature: factory returns a
  HookDispatcherBuilder, and `RebornLoopDriverHostFactory` attaches a
  `RunScopedHookMilestoneSink` keyed to the CURRENT run's LoopRunContext
  inside `build_text_only_host_with_capabilities`, before sealing the
  dispatcher. The previous zero-arg signature relied on the closure
  capturing run_context — silently misattributed across reuses; new
  public API `with_hook_dispatcher_builder_factory` removes that
  failure mode entirely. Legacy `with_hook_dispatcher_factory` retained
  for back-compat (its sink-wiring contract stays caller-side).

Concerning #6 — TimelineEntry hook metadata.
  Added 6 optional fields to `TimelineEntry` (hook_id, hook_point,
  hook_trust_class, hook_decision, hook_failure_category,
  hook_failure_disposition) and projected them from `RuntimeEvent::Hook*`.
  Replay consumers now see which hook fired/failed, not just that some
  hook event happened. Each field is closed-vocabulary (no free-form
  reason text — that stays in the audit reason payload, not the
  product replay DTO).

Testing gaps from henrypark133 — caller-level tests:
  #4 (two-run hook telemetry attribution):
     hook_telemetry_attribution_is_per_run_not_captured
     Builds two hosts from the SAME builder factory closure with two
     fresh LoopRunContexts. Asserts each run's hook milestones carry
     its OWN run_id (no stale captured one).
  #6 (replay projection contract for hook events):
     hook_runtime_events_project_with_sanitized_hook_metadata
     non_hook_runtime_events_project_with_no_hook_metadata
     Constructs RuntimeEvent::Hook{Dispatched,DecisionEmitted,Failed}
     and asserts the projection preserves the metadata fields. The
     negative test guards against cross-contamination on non-hook
     events.

All henrypark133 review items now addressed:
  Critical: #1, #2, #3, #4 — done
  Concerning: #5, #6, #7 — done
  Testing gaps: #1-#6 — done

Tests: 154 unit + 19 hooks_integration in ironclaw_reborn + 61 reborn
unit + 38 + 2 new in ironclaw_event_projections + ... pass.
Workspace clippy + fmt + no-panics clean.

* docs(hooks): scope DenyReasonCode closed-vocabulary enum (successor #6)

Successor PR from #3573 — real-hooks ergonomics finding F2 (deferred).
Adds a curated vocabulary of model-visible denial reasons so hook
authors can communicate why a deny happened without opening a
free-form prompt-injection channel.

* feat(hooks): DenyReasonCode + PauseReasonCode closed-vocabulary enums

Address real-hooks ergonomics finding F2 (deferred from PR #3573). The
prior dispatcher collapsed every Installed-tier deny to the static
label 'hook_predicate_denied', because manifest reason strings are
author-controlled and surfacing them to the model would open a
prompt-injection channel. The cost: the agent couldn't tell *why*
a hook denied.

This PR introduces two closed-vocabulary enums:

- DenyReasonCode: Generic / RateLimit / ValueCap / Blocklist /
  RequiresApproval / OutOfPolicy
- PauseReasonCode: Generic / RequiresApproval / OverThreshold /
  SensitiveAction

Each variant has an as_label() returning &'static str (so the sink's
&'static str contract is preserved). New OnExceededAction variants
'DenyWithCode { code, reason }' and 'PauseApprovalWithCode { code,
reason }' let manifest authors opt into the richer labels while
keeping reason audit-only.

The legacy Deny { reason } / PauseApproval { reason } variants are
retained for back-compat and map to DenyReasonCode::Generic /
PauseReasonCode::Generic — existing manifests continue to produce
hook_predicate_denied / hook_predicate_pause_requested.

Threat-model regression: a hook author cannot smuggle text into the
model-visible label because the 'code' field is typed as the enum;
there's no String slot exposed model-side. A test
(deny_with_code_only_exposes_enum_variants_to_model) documents this
as a compile-time property.

Tests (+7 new = 161 total):
- deny_reason_code_labels_are_stable: pins the label vocabulary so
  rename/relabel is loud.
- pause_reason_code_labels_are_stable: same for PauseReasonCode.
- deny_with_code_round_trips_through_json + pause variant: wire
  round-trip + snake_case tag assertion.
- deny_with_code_only_exposes_enum_variants_to_model: compile-time
  property check.
- rate_or_value_cap_with_deny_code_routes_to_code_label: end-to-end
  affirmative test that the dispatcher emits the code's label.
- rate_or_value_cap_with_pause_code_routes_to_code_label: same for
  pause.

Scope doc: crates/ironclaw_hooks/docs/successors/06-deny-reason-code.md

* test(hooks): address codex review on #3636

- Update stale real-hooks-findings.md F2 row to cite this PR's enum
  follow-on (was 'deferred').
- Add install_deny_with_code_manifest_surfaces_code_label_on_dispatch:
  end-to-end test driving the registrar->dispatcher path for the
  new DenyWithCode variant (prior tests covered serde + direct hook
  evaluation, but not the manifest install path that downstream
  authors actually use).

Codex review on PR #3636: APPROVE with two recommendations; both
addressed.

Tests: 162 unit (+1 new). Clippy/fmt clean.

* fix(hooks): attenuate Installed-tier prompt patches to user role

Installed-tier `before_prompt` patches were injected as role:"system"
messages. Envelope text labels ("[ext-foo says]: ...") do not strip
system-role authority from the model's perspective, so a third-party
extension could inject system-tier instructions through a snippet
patch. This is a prompt-authority escalation against the trust
hierarchy the framework otherwise enforces.

Add `role_for_trust_class()` mapping Installed -> "user" and
Builtin/Trusted/SelfAuthored -> "system". Thread per-patch
trust_class through `wrap_patches_to_messages` and use it for the
emitted `LoopModelMessage.role`.

Tests:
- installed_hook_patch_drops_to_user_role: asserts the role for an
  Installed-tier patch is "user"
- trusted_tier_hook_patch_keeps_system_role: regression that Trusted
  tier still produces system-role content

* fix(hooks): enforce scope filter on observer dispatch + reject incompatible points

Two related defense-in-depth fixes against silent scope-filter failure:

1. The registry silently accepted Installed bindings with
   `HookBindingScope::OwnCapabilities` at points (BeforePrompt,
   AfterModel, AfterCheckpoint) whose dispatch context carries no
   per-capability provider. The manifest's declared scope had no
   effect at all — the hook fired against every dispatch. Reject the
   binding at install time so the operator sees the misconfiguration.

2. `dispatch_observer_at` for `AfterCapability` did not consult the
   binding's scope, so an Installed observer registered with
   `OwnCapabilities` fired against every invocation regardless of
   provider. Add `dispatch_observer_at_with_provider` carrying the
   resolved capability provider; the capability-port middleware
   resolves the provider once per invocation and threads it through
   both the BeforeCapability hook context and the AfterCapability
   observer dispatch. The dispatcher then enforces
   `HookBindingScope::permits` on each observer binding.

`ObserverHookContext` gains a `provider: Option<ExtensionId>` field;
`#[non_exhaustive]` keeps existing authors compiling.

Tests:
- rejects_own_capabilities_at_before_prompt
- rejects_own_capabilities_at_after_model
- accepts_own_capabilities_at_before_capability
- own_capabilities_observer_filters_foreign_providers (covers
  foreign / matching / unresolved provider)

* fix(hooks): preserve free-form audit reason alongside closed-vocab model label (serrrfirat #3636)

`PredicateBackedBeforeCapabilityHook::evaluate()` was discarding the
free-form `reason` from `EvaluatorDecision::{Deny, PauseApproval}`
with `..` and only sending `code.as_label()` into the sink. The
`HookDecisionEmitted` milestone therefore carried only the closed-
vocab label, and operator-visible audit/SSE context was silently lost
end-to-end. The fix splits the channels:

- Model sees the closed-vocab label (`hook_rate_limit`,
  `hook_pause_over_threshold`, ...) via `sink.deny(label)`. This
  channel is unchanged.
- Audit/SSE sees the manifest's free-form `reason` via a new
  audit-only sink method `record_audit_reason(reason: String)`. The
  recording sink captures it; the dispatcher reads it after the hook
  returns and threads it into `LoopHostMilestoneKind::HookDecisionEmitted`.

Surface changes:
- `PrivilegedGateSink` / `RestrictedGateSink` gain
  `record_audit_reason(String)` — accepts dynamic `String` (audit-only,
  no model-facing seam) unlike the `&'static str` decision reasons.
- `RecordingGateSink` gains an `audit_reason: Option<String>` field.
- `GateHookOutcome::Decision` is now `Decision { decision,
  audit_reason }`.
- `HookDispatcher::emit_decision_with_audit` threads the audit reason
  into the milestone.
- `LoopHostMilestoneKind::HookDecisionEmitted` gains a
  `#[serde(default, skip_serializing_if = "Option::is_none")]`
  `audit_reason: Option<String>`. The durable RuntimeEvent projection
  intentionally drops this field — audit reasons are operator-facing
  in-memory SSE content, never durable cross-process surface.

Tests:
- `deny_with_code_records_audit_reason_separately_from_model_label`:
  asserts the recording sink ends with `Deny { reason: "hook_rate_limit" }`
  in `state` AND `audit_reason == Some("daily cap of $1000 ...")`.

* fix(hooks): remove unused model_request helper (CI clippy fix)

* fix(hooks): address serrrfirat P1/P2 findings on PR #3573

Three issues from the 5-15 review:

**P1 #1 registrar.rs:70 — `same_tenant` grants not enforced**
`HookManifestEntry::validate` only confirmed `requires_grant` was
present; the registrar then immediately installed the binding with no
host-verified grant context. A manifest could declare
`requires_grant = "anything"` and get a cross-extension binding for
free.

Fix: `HookRegistrar` now carries a `verified_grants: HashSet<String>`
(empty by default — default-deny). Add the host-facing setter
`with_verified_grants(...)`. At `install_one`, if
`entry.requires_grant` is `Some(g)`, require `g ∈ verified_grants` or
reject with a clear error. Tests:
- `install_rejects_same_tenant_without_verified_grant`
- `install_rejects_same_tenant_when_verified_grants_mismatch`
- The existing positive test
  `installer_propagates_owning_extension_and_scope_from_manifest` now
  wires the verified grant explicitly (proves the API contract).

**P1 #2 prompt_port.rs:150 — zip misalignment**
The materialization loop zipped surviving messages against the
ORIGINAL unfiltered patch list. `wrap_patches_to_messages` skips
metadata patches and over-budget snippets, so the zip silently paired
message[0] with patch[0] even when patch[0] was the skipped metadata
— materializing the wrong content (or none) under the snippet's
synthetic ref.

Fix: `wrap_patches_to_messages` now returns
`Vec<WrappedHookMessage { message, safe_content }>` — surviving
messages paired with their content by construction. The caller
materializes `entry.safe_content` under `entry.message.content_ref`
directly; no zip against unfiltered input. Removed the now-unused
`safe_content_for_patch` helper.

Test:
- `materialization_stays_aligned_when_metadata_patches_are_filtered`:
  a hook emits `[metadata, snippet]`; asserts only one model message,
  and the materialized content under its ref contains the snippet's
  body — proves filtering can no longer desync from materialization.

**P2 #3 loop_driver_host.rs:1343 — `with_hook_dispatcher_builder`**
Docs said it deferred `build_arc()` to let the host factory finalize
wiring; the implementation called `build_arc()` eagerly and routed
through the legacy shared-dispatcher adapter, losing per-run
dispatcher isolation and the run-scoped milestone sink.

Fix: marked `#[deprecated]` with a note pointing callers to
`with_hook_dispatcher_builder_factory(|| ...)` for per-build
isolation, or `with_hook_dispatcher(...)` if they actually meant the
shared adapter. The method body is unchanged so no callers break;
they'll see the deprecation warning. No internal callers exist, so
the deprecation doesn't trip `-D warnings`.

All 162 hooks lib + 19 reborn integration tests pass; clippy clean.

* fix(hooks): address serrrfirat 3573-2026-05-15 review findings

P1 — prompt bundle authority mismatch (prompt_port.rs):
`HookedLoopPromptPort::build_prompt_bundle` called the inner port first,
which caused `HostManagedLoopPromptPort` to issue the prompt-bundle
authority grant against the pre-hook message list. The wrapper then
appended `msg:hook.*` messages to `bundle.messages`, so the downstream
model request hit `grant.messages != messages` and failed closed with
"model request messages do not match the host-built prompt bundle".

Add `with_bundle_authority(authority, run_context)` and re-issue the
grant after appending hook messages so it covers the post-hook bundle.
Reborn wires `prompt_authority.clone()` + `run_context.clone()` into
the wrapper at construction time.

P2 — observer installer accepts non-observer points (dispatch.rs):
`install_observer` accepted any `HookPointSpec` (including
`BeforeCapability` / `BeforePrompt`) and only populated the observer
map. Dispatch later found a binding without a gate/mutator impl and
fail-closed the capability with "binding present without installed
implementation". Reject non-observer points at install time so misuse
fails loudly rather than poisoning bindings at dispatch.

P2 — batch path skipped AfterCapability observers on inner error
(capability_port.rs):
The batch loop used `?` directly on `self.inner.invoke_capability(...)`,
which propagated the error before dispatching `AfterCapability`
observers. Failed batch entries disappeared from telemetry / audit,
while the single-invocation path dispatches observers on error.
Capture the inner result, dispatch observers, then propagate the error.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(hooks): address PR #3573 review feedback round 3

Addresses serrrfirat's CHANGES_REQUESTED review (2026-05-20) by tightening
several install-time / dispatch-time bounds and gating production seams:

- Bound free-form audit reasons crossing telemetry. New
  `telemetry::sanitize_audit_reason` strips control characters and caps
  length at 512 bytes; `emit_decision_with_audit` routes the manifest-
  supplied reason through it before publishing milestones. Manifest
  validation also rejects reasons over the same byte limit at install time
  so the wire-side cap is a defense-in-depth layer, not the only line.
- Make hot dispatch O(H) instead of O(H^2). The per-binding poison
  recheck used to acquire the registry mutex and walk every binding;
  `ordered_bindings_with_poison_snapshot` now takes the active bindings
  and the poisoned hook-id set under a single lock, and each loop
  threads a local `HashSet<HookId>` that absorbs mid-dispatch
  poisoning. Removed the redundant `is_poisoned` helper.
- Gate `HookDispatcher::registry_for_test` behind `cfg(any(test,
  feature = "test-support"))`. The accessor previously exposed
  `&Mutex<HookRegistry>` in production, letting any `Arc<HookDispatcher>`
  holder lock and call `HookRegistry::poison` to disable installed
  hooks. Added `active_bindings_snapshot(point)` as the read-only
  production-safe replacement.
- `#[serde(deny_unknown_fields)]` on every hook-manifest and predicate
  DTO (`HookManifestEntry`, `HookManifestBody`, `WasmBudget`,
  `HookPredicateSpec`, `CapabilityPredicate`, `ValueOrRateBound`,
  `OnExceededAction`). Typoed or unsupported fields (e.g. a
  manifest-supplied `trust_class`) now fail loud at install time
  instead of being silently dropped.
- Bound predicate trees at install. New
  `validate_predicate_tree` enforces `MAX_PREDICATE_DEPTH = 8`,
  `MAX_PREDICATE_NODES = 64`, `MAX_PREDICATE_STRING_BYTES = 256`, and
  `MAX_MANIFEST_REASON_BYTES = 512`. A hostile registry manifest can no
  longer install a deep or huge `All`/`Any` tree that the evaluator
  would recursively walk on every match.
- Cap sliding-window samples per key. `MAX_SAMPLES_PER_KEY = 4_096` in
  the predicate evaluator. Both the invocation-count and numeric-sum
  histories drop the oldest sample once the cap is reached, bounding
  memory under attacker-triggered hot capabilities while preserving
  rate/value-cap semantics over the most recent window.
- `split_indexer` / `resolve_path` now fail closed on malformed bracket
  syntax (`amount[foo]`, `amount[`, trailing garbage). Previously they
  silently fell back to the parent field, which could let a typoed
  `NumericSum` predicate evaluate against the wrong value and allow
  calls the predicate would otherwise have denied.
- Honor `PatchOrdinalHint`. `WrappedHookMessage` carries the source
  patch's `ordinal_hint`; `HookedLoopPromptPort` inserts `NearTop`
  messages after the bundle's `identity_message_count` and appends
  `Last` messages at the end. Safety/policy snippets that need early
  placement now get it.
- Update `ironclaw_hooks` top-level docs to reflect the four trust
  classes (`Builtin`/`Trusted`/`Installed`/`SelfAuthored`) and the
  now-wired Reborn middleware composition.

Tests added:
- `manifest::rejects_unknown_top_level_field`
- `manifest::rejects_unknown_wasm_budget_field`
- `manifest::rejects_predicate_tree_exceeding_max_depth`
- `manifest::rejects_predicate_tree_exceeding_max_nodes`
- `manifest::rejects_predicate_string_exceeding_max_bytes`
- `manifest::rejects_manifest_reason_exceeding_max_bytes`
- `points::capability::malformed_indexer_returns_none_not_parent_value`
- `telemetry::sanitize_audit_reason_*` (truncate / strip control /
  preserve / empty)

`cargo fmt`, `cargo clippy --all --benches --tests --examples
--all-features`, and `cargo test -p ironclaw_hooks` all pass clean.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(hooks): batch deferred test coverage from #3573 review (#3914)

* perf(hooks): defer capability input resolution until a predicate needs it (#3913)

* fix(rebase): adapt hooks tests + middleware to upstream API additions

- CapabilityDescriptorView: add parameters_schema field
- LoopModelRequest / LoopPromptBundleRequest: add capability_view field
- TimelineEntry test builder: add hook_id / hook_point / hook_trust_class /
  hook_decision / hook_failure_category / hook_failure_disposition fields
- ironclaw_reborn::tests::hooks_integration: switch from
  InMemoryLoopCheckpointStore to InMemoryTurnStateStore (which now
  impls both LoopCheckpointStore and TurnStateStore), pass TurnActor
  in TurnRunState, supply the new turn_state_store factory arg
- ironclaw_reborn lib.rs: drop the pub-use re-exports that upstream
  intentionally removed (per the module-directory rationale in the
  current ironclaw_reborn lib.rs doc comment); update the
  hooks_integration test imports to use module paths
- Cargo.toml: union the hooks-foundation member list with upstream's
  new crates (event_streams, auth, first_party_extensions,
  reborn_webui_ingress, product_workflow_storage, webui_v2); drop
  ironclaw_storage which no longer exists upstream
- crates/ironclaw_architecture/tests/reborn_dependency_boundaries:
  keep upstream's removal of ironclaw_filesystem from the ironclaw_turns
  forbidden list AND add ironclaw_hooks to that list
- crates/ironclaw_reborn/src/milestone_events.rs: drop dead loop_failure_kind
  helper (replaced upstream by loop_failure_kind_name in text_loop_driver.rs);
  keep hook_decision_label which is still used

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* perf(hooks): restore batched capability dispatch when hooks acti…
theredspoon pushed a commit that referenced this pull request Jun 21, 2026
…3920)

* Implement installed WASM hook runtime

Adds crates/ironclaw_hooks/docs/threat-model-wasm.md and follows the reviewed design ack: 1) module bytes are resolved, digest-cached, and compiled in the tool-WASM style while reusing its resource limiter; 2) each invocation gets a fresh wasmtime Store; 3) the ABI is a wasmtime::Linker surface, not wit-bindgen; 4) host-import sink shims enforce call, patch-byte, observer-fact, and decision budgets.

* Harden WASM hook string and metadata budgets

* fix(hooks): validate WASM hook ABI at install time (serrrfirat #3 on PR nearai#3634)

Address serrrfirat MEDIUM finding #3: `WasmHookRuntime::prepare()` compiled
and cached module bytes but did not validate imports or the requested
export. ABI mismatches (unsupported import, missing export, wrong export
signature) were deferred to first live dispatch — and the prior
`wasm_unsupported_host_import_fails_closed` test codified that a
bad-import module would install successfully and only fail closed at
invocation. Malformed untrusted modules should never reach live traffic.

Changes:
- `prepare()` derives the target hook point from `request.kind`, then
  runs `validate_module_abi()`: scratch-instantiate the module against
  the point-specific linker (catches unsupported / wrong-type imports)
  and resolve the typed export `() -> ()` (catches missing export and
  wrong signature). Failures surface as new
  `WasmHookRuntimeError::InvalidImports` or existing
  `WasmHookRuntimeError::InvalidExport`, both of which bubble up as
  `HookError::RegistryConstruction` from the registrar.
- `wasm_point_for_kind(HookManifestKind)` helper centralizes the
  kind → wasm-point mapping; the previous `execute_*` paths can share
  it in a follow-up but kept inline for now to minimize churn.

Tests:
- `wasm_unsupported_host_import_is_rejected_at_install_time`: replaces
  the prior test that codified late-failure behavior; asserts the
  registrar returns `RegistryConstruction` citing the bad import.
- `wasm_missing_export_is_rejected_at_install_time`: new module that
  compiles but lacks the manifest-declared export; same install-time
  rejection.

* fix(hooks): address henrypark133 must-fix #1, #2, #3 on PR nearai#3634

Three items from the 5-15 review:

**#1 (must-fix) Extract ironclaw_wasm_limiter micro-crate**
Replace `#[path = "../../../ironclaw_wasm/src/limiter.rs"]` cross-crate
file import with a proper Cargo edge. The 111-line `WasmResourceLimiter`
moves into a new `crates/ironclaw_wasm_limiter` micro-crate that both
`ironclaw_wasm` and `ironclaw_hooks` depend on. The architecture rule
forbidding `ironclaw_hooks -> ironclaw_wasm` is preserved (the new
crate sits below both consumers and pulls in only `wasmtime` +
`tracing`); `cargo check`, `cargo doc`, and architecture-linting tests
now see the edge, and the file can't be moved out from under one of
the consumers silently.

Mechanical changes:
- new `crates/ironclaw_wasm_limiter/` (Cargo.toml + src/lib.rs with the
  type exposed as `pub` instead of `pub(crate)`)
- workspace `members` entry added
- `crates/ironclaw_wasm/src/limiter.rs` deleted
- `crates/ironclaw_wasm/src/lib.rs`: `mod limiter` removed
- `crates/ironclaw_wasm/src/store.rs`: import switched to
  `ironclaw_wasm_limiter::WasmResourceLimiter`
- `crates/ironclaw_wasm/Cargo.toml`: dep added
- `crates/ironclaw_hooks/Cargo.toml`: dep added
- `crates/ironclaw_hooks/src/wasm/runtime.rs`: `#[path = ...]` block
  removed; import switched to the crate

**#2 + #3 (must-fix) Dead WASM arms in dispatch**
`run_before_capability_hook`, `run_before_prompt_hook`, and
`run_observer_hook` each had an early-return guard that dispatched
WASM hooks with `catch_unwind` + timeout, then ALSO had a matching
WASM arm in the inner `match` that ran without those protections. The
prompt-path arm additionally swallowed `WasmHookFailure` via `|_| ()`,
making the must-fix #2 problem worse on that path specifically.

If a future refactor removed any of the early-return guards, those
inner arms would silently take over and drop panic isolation, deadline
enforcement, AND (for prompts) the failure category. Replaced each
inner arm with `unreachable!()` carrying a comment that explains
why the arm exists and references the early-return guard above it.
A future refactor that removes the guard will now trip the
`unreachable!` at first call instead of silently degrading.

All 154 hooks lib + 29 reborn integration tests still pass.

* fix(hooks): plumb context to WASM hooks + runtime hardening

Critical #1 on PR nearai#3634: WASM hooks previously received no context. The
`execute_*` entry points dropped the `&BeforeCapabilityHookContext` /
`&BeforePromptHookContext` / `&ObserverHookContext` value and invoked
the guest export with `()`, so a WASM gate could never decide based on
the capability name, tenant, provider, or other dispatch-time facts. Add
an `ic:hooks/context@1` host-import module exposing two read-only
calls — `ctx_size() -> i32` and `ctx_read(ptr, len) -> i32` — backed by
a JSON-serialized blob the dispatcher writes per-invocation into the
fresh store. Modules that don't import these continue to link; modules
that do import them get a stable, non-empty payload to read. An
integration test (`wasm_before_capability_hook_reads_context_blob`)
asserts the contract end-to-end: a guest that fails to read a non-empty
blob traps before its `deny` call.

Also rolls up the other reviewer-flagged WASM runtime issues, all of
which touch `wasm/runtime.rs`:

HIGH #2: epoch-tick background thread now holds a shutdown
`AtomicBool` and joins on `Drop`. Previously it looped forever and
leaked an Engine clone on every runtime drop.

MED #4: compiled-module cache is now an `lru::LruCache` bounded by
`MODULE_CACHE_CAPACITY = 128`. Replaces the unbounded `HashMap`.

MED #7: `prepare()` no longer compiles under the cache lock. Fast
path reads from LRU under a brief lock; slow path compiles outside
the lock and re-checks on insert to avoid the TOCTOU window where
two concurrent installs of the same module both compile.

Bug #9: post-call `deadline_exceeded()` re-check on the Ok branch
is gone. wasmtime epoch-interrupt is the authoritative wall-clock
signal; an Ok return is no longer reclassified as a timeout because
the wall ticked over during host-side return.

Bug #10: `add_milestone_metadata` returns a distinct
"metadata value exceeds the u32 byte-length ceiling" error when the
guest-supplied `value.len()` overflows u32, instead of misreporting it
as "exceeded total prompt-patch byte budget".

Existing integration tests for WASM hooks are also re-wired through
`HookRegistrar::with_verified_grants` so the grants-store gate added in
the foundation-01 merge stops failing the pre-existing fixtures.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(hooks): run WASM hooks on the blocking pool

HIGH #3 on PR nearai#3634: `tokio::time::timeout` does NOT cancel synchronous
wasmtime execution. The previous code awaited a `catch_unwind(async { h.evaluate(ctx) })`
future whose body completed in one poll, so the timeout could only fire
*around* the WASM call rather than against it; a hook that wedged inside
wasmtime simply pinned the calling tokio task.

Route gate, prompt, and observer WASM dispatch paths through
`tokio::task::spawn_blocking` via a shared `run_wasm_blocking` helper.
The outer `tokio::time::timeout` now governs the JoinHandle, so a stuck
blocking task stops blocking the dispatcher's caller; the wasmtime
epoch interrupt configured in the runtime (10 ms tick) is the
authoritative in-WASM wall-clock cancel signal. JoinError (panic in
the blocking task) maps to `FailureCategory::Panic`, matching the
pre-existing semantics for synchronous panics caught via
`catch_unwind`.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* perf(hooks): O(1) hook-id lookup via side index

Finding #8 on PR nearai#3634: `set_priority`, `poison`, `is_poisoned`, and
`contains_hook` all did full-registry scans over every binding at every
point. Each is called per-dispatch (poison-checks on the snapshot loop
in particular), so the cost is `O(registered_hooks)` per
`(installed_hook, registered_hook)` pair.

Maintain a denormalized `HashMap<HookId, (HookPointSpec, usize)>` side
index in lock-step with `by_point` so every per-hook-id operation
becomes a single hash lookup + a direct vec indexed access. The
duplicate-id rejection in `insert` now reads from the side index too,
turning what used to be a flat-map scan into a `HashMap::contains_key`.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(hooks): wall-clock timeout, observer memory, limiter rollback, registrar happy path

Round out the test set for the WASM hook execution path:

#11 / #12: gate + observer wall-clock timeout. The pre-fix dispatcher
ran wasmtime synchronously on the executor, so the outer
`tokio::time::timeout` `Err(_elapsed)` arm was effectively unreachable.
Now that WASM execution runs on the blocking pool, the timeout actually
fires; the new tests give the wasm budget headroom (1B fuel, 5s wall)
and the dispatcher a 20 ms timeout, then assert the failure
classification (FailClosed for gate, FailIsolated for observer).

#13: observer memory exhaustion. Mirrors
`wasm_memory_exhaustion_fails_closed_for_gate` against the observer
dispatch path so the FailIsolated branch of the failure matrix has
explicit memory coverage, not just fuel/wall.

#15: `WasmResourceLimiter::memory_grow_failed` rollback. Stages an
approved grow, simulates the OS-level grow failing, and asserts a
subsequent grow of the full ceiling succeeds — the inflated
`memory_used` from the failed attempt must be released.

#16: registrar WASM happy path. Companion to the existing
`install_wasm_body_requires_runtime` negative case: a valid module
installs, the binding is visible via the public registry accessor, and
is not pre-poisoned.

#14 (`add_milestone_metadata` happy path) is intentionally omitted —
the BeforePrompt dispatch path is currently unreachable due to a
pre-existing manifest-vs-registry scope conflict (`OwnCapabilities` is
the only valid `BeforePrompt` scope per manifest validation, but the
registry rejects `OwnCapabilities` at `BeforePrompt` because the point
has no provider context). That contradiction sits outside this PR's
scope; flagging for a follow-up.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(hooks): typed WASM version material, reconcile design doc

LOW #20 on PR nearai#3634: extract the
`{extension_version}+wasm:{module_digest_hex}` concatenation into a
`WasmVersionMaterial` newtype with a single `Display` impl. The
identity material no longer floats free as a stringly-typed argument
inside the registrar.

Reconcile `docs/successors/02-wasm-runtime.md` with the implementation:

- Spell out that wall-clock cancellation depends on the
  `tokio::time::timeout(tokio::task::spawn_blocking(...))` pair, and
  explain why a bare timeout over a synchronous wasmtime call cannot
  actually cancel.
- Define `FailIsolated` and `FailClosed` as `FailureDisposition`
  values, distinct from the older `HookFailureMode::{FailOpen,
  FailClosed}` policy switch that applies to predicates.
- Clarify the generic `evaluate` export contract — name is whatever
  the manifest declares, signature is `(): ()`, context arrives
  through the new `ic:hooks/context@1` host imports — and note the
  intentional divergence from `WitToolRuntime`'s hardcoded interface.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(hooks): drop .expect() in WASM module cache capacity

Pre-commit no-panics CI flagged the .expect() on the LruCache capacity.
Move the validity check to a const match, so the NonZeroUsize is fixed at
compile time and the no-panics regex is satisfied.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(hooks): use HookLocalId::new after newtype privatization

The newtype-privatization landed in reborn-integration after the
hooks-fu-wasm-runtime branch's WASM scaffolding tests were written;
update the affected test/registrar sites to use HookLocalId::new
instead of the now-private tuple constructor.

* style: cargo fmt after newtype-privatization fixups

* test(hooks): ignore 3 BeforePrompt WASM tests with manifest/registry conflict

These tests were failing on the original branch tip too (verified against
origin/hooks-fu-wasm-runtime @ 571efdf). The Installed-tier BeforePrompt
WASM install path has no valid scope today:
  - OwnCapabilities is rejected by the registry C3 check (finding #2 on
    PR nearai#3573) since BeforePrompt has no per-capability invocation
    context.
  - SameTenant is rejected by manifest validation ("cannot combine
    scope = same_tenant with kind = before_prompt").

The budget-overflow paths these tests exercise are point-agnostic; the
follow-up is to either rewrite the helper to install through
BeforeCapability or add a Global manifest scope. Tracked as a deferred
item on the new PR.

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
theredspoon pushed a commit that referenced this pull request Jun 21, 2026
…#3899)

* Reborn budgets: address all nearai#3841 follow-ups end-to-end

Implements every open follow-up from PR nearai#3841 (cost-based budgets
foundation), driven by the plan in
`docs/plans/2026-05-22-reborn-budgets-followups.md`:

- **C2 (provider tokens)**: `LoopModelResponse.usage` carries real
  `(input_tokens, output_tokens)` from `CompletionResponse` /
  `ToolCompletionResponse`; `usage_for_response` reconciles to actual
  USD via the cost table instead of the conservative estimate.
- **D1 (cascade warnings)**: `CascadeOutcome` variants carry
  `Vec<BudgetWarning>` so warnings preceding a pause or hard deny
  reach the audit sink. `ResourceError::LimitExceeded` /
  `RequiresApproval` reshaped to struct variants.
- **C1 (cancellation safety)**: new
  `LoopModelBudgetAccountant::release_in_flight` trait hook + RAII
  `ReservationReleaseGuard` in `HostManagedLoopModelPort::stream_model`
  so a cancelled future doesn't orphan its reservation.
- **E1 (dead code)**: removed the never-set `budget_accountant` field
  on `ThreadBackedLoopModelPort`.
- **Real cost table**: new `StaticModelCostTable` +
  `LlmModelProfilePolicy::build_cost_table()` populated from
  `ironclaw_llm::costs::model_cost` with `default_cost` fallback so
  unknown providers never silently reconcile to zero.
- **B1 (filesystem gate store)**: new `FilesystemBudgetGateStore`
  mirroring `FilesystemResourceGovernorStore`; pending gates survive
  process restart.
- **A1 (production wiring)**: composition builds
  `GovernorBackedAccountant` from the cost table + governor and
  threads it through `RebornLoopDriverHostFactory::with_model_budget_accountant`.
- **A2 (audit / SSE projection)**:
  `InMemoryResourceGovernor::with_event_sink` emits `Reserved`,
  `Reconciled`, `Released`, `Warned`, `ApprovalRequested`, `Denied`,
  `LimitChanged`; composition holds an `InMemoryBudgetEventSink` ready
  for downstream SSE projection.
- **F1 (stuck-loop normalization)**: `CapabilityCallSignature::from_call`
  now runs `progress::normalize_for_hash` so the existing repetition
  window collapses request-id / UUID / timestamp noise.

Side fix: `ResourceValue` moved to adjacent serde tagging (the
combination of internal tagging + `Decimal`'s `serde-with-str`
representation breaks JSON serialization — rust-lang/serde#1402).

Regression tests added per item — see the acceptance evidence appendix
in the plan doc.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Reborn budgets: end-to-end test coverage via test-support feature

Adds 13 e2e tests covering the budget pipeline through
`build_reborn_runtime` + `send_user_message`. Required infrastructure:

- **`test-support` feature** on `ironclaw_reborn_composition` exposing
  `BudgetTestGateway` (scripted token usage) and
  `RebornRuntimeInputTestExt`. Existing `model_gateway_override` field
  promoted from `#[cfg(test)]` to `#[cfg(any(test, feature = ...))]`
  with a new public `with_model_gateway_override_for_tests` setter.
- **Cost-table override** on `RebornRuntimeInput` so tests can pair
  the gateway with a deterministic `ModelCostTable`. Without this, an
  override gateway dropped the cost table and the accountant never
  fired.
- **Budget accessors** on `RebornRuntime`: `budget_resource_governor`,
  `budget_event_sink`, `budget_gate_store`, and
  `apply_resolved_budget_gate`. Test-feature gated.
- **`ResourceGovernor::usage_for`** added as a default-impl trait
  method so tests read spend through the trait surface.
- **`BudgetGateStore` wired into the accountant**:
  `GovernorBackedAccountant::with_gate_store(...)` opens a pending
  gate whenever the governor cascade returns `RequiresApproval`. The
  approval-required host error is unchanged; the gate is the
  out-of-band channel a user-facing handler resolves.

Scenarios covered:

| # | Test | What it asserts |
|---|---|---|
| F1 | `f1_happy_path_records_actual_usd_in_ledger` | Ledger depletes by provider tokens × cost table |
| F2 | `f2_crossing_warn_threshold_emits_warned_event` | Warn fires alongside successful Reserved/Reconciled |
| F3 | `f3_approval_with_increased_limit_unblocks_retry` | Approve → set_limit applies → retry succeeds |
| F4 | `f4_cancel_keeps_budget_blocked_on_retry` | Cancel → retry still short-circuits |
| F5 | `f5_expiry_marks_gate_terminal_and_keeps_budget_blocked` | Expiry → gate drops from pending list, retry still blocked |
| F6 | `f6_hard_cap_denied_before_provider_call` | Estimate over cap → zero model calls, Denied event |
| C1 | `c1_provider_tokens_reconcile_to_actual_usd` | Real numbers, not estimate |
| C2 | `c2_unknown_model_in_cost_table_reconciles_to_zero` | Unknown profile → zero spend |
| C3 | `c3_zero_cost_model_records_zero_spend` | Free model → zero USD with non-zero tokens |
| D1 | `d1_agent_deny_preserves_user_warn_event` | Cascade emits both Warned and Denied |
| D3 | `d3_fresh_user_without_limits_runs_without_denial` | No limit → no denial |
| + | `pause_in_distinct_runs_produces_distinct_pending_gates` | Per-run gate identity |
| + | `budget_test_gateway_scripted_replies_drive_per_turn_costs` | Multi-turn scripted accumulation |

F7 (cancellation mid-stream) is unit-covered by
`release_in_flight_drains_orphan_reservation_on_cancellation`.
D2 (period rollover) is unit-covered by
`rolling_24h_snapshot_reports_anchored_window_not_now_window`.
B-series (background ticks) await the BackgroundKind scheduler
call site (no production caller in Reborn yet).

Run via `cargo test -p ironclaw_reborn_composition --features test-support`.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Budget review feedback: address all 7 findings from PR nearai#3899 review

Two High and five Medium issues raised by serrrfirat's multi-agent review.

**High #1 — `FilesystemBudgetGateStore` cross-tenant leakage**
The store hardcoded `ResourceScope::system()` for every op, so all
tenants wrote into the same `/tenants/__SYSTEM__/...` snapshot and
`list_pending` would expose gates across tenants. Fix: `new(...)` now
takes a `ResourceScope`; each tenant gets its own store, and the
`ScopedFilesystem` mount view routes the snapshot under that tenant's
path. Added `list_pending_does_not_leak_across_tenants` regression.

**High #2 — accountant wired without default budget limits**
Composition built `GovernorBackedAccountant` without
`with_seeding_policy`, so the local-dev governor started empty and
`reserve_with_outcome_in_state` skipped accounts that had no
configured limit — model calls reconciled spend but never enforced a
cap. Fix: `build_reborn_runtime` now loads
`BudgetDefaults::compiled_defaults().with_env()` and wires
`BudgetSeedingPolicy` + `with_overestimate_factor`. Renamed the D3
test to `d3_seeding_policy_installs_default_cap_on_first_touch` to
prove the wiring fires.

**Medium #3 — RAII guard disarmed before post_model_call await**
`HostManagedLoopModelPort::stream_model` was disarming the
`ReservationReleaseGuard` before awaiting `post_model_call`. A
cancellation during that await dropped the future without cleanup,
orphaning the reservation. Fix: disarm AFTER `post_model_call`
returns. `release_in_flight` is now idempotent (peek-then-release-
then-remove) so a successful post-call + subsequent guard drop is a
no-op.

**Medium #4 — failed release drops the retry handle**
`release_in_flight` removed the in-flight entry before calling
`governor.release`. A transient storage error left the reservation
active in the governor with the id discarded. Fix: peek first,
release, only remove on success. Errors keep the entry retained for
a future retry / cleanup hook.

**Medium #5 — unknown model silently reconciles to zero USD**
Both `estimate_for` and `usage_for_response` fell back to
`ModelCost { 0, 0, 0 }` when the cost table had no entry for the
effective model. Cost-table drift would silently bypass daily caps.
Fix: `GovernorBackedAccountant` carries a `default_cost` (default ~
GPT-4o pricing, ~`$0.0000025 input + $0.00001 output per token`) used
for unknown models. Callers wiring `ZeroCostTable` for free / Ollama
explicitly opt out of the fallback. Updated the C2 e2e test to
assert the new fail-closed shape.

**Medium #6 — paused dimension lost when another hard-denies**
`check_thresholds_all_interventions` stored `Approval` only in the
`approval` slot, so when one dimension paused and another hard-denied,
the `Deny { warnings, denial }` outcome lost the pause signal.
Fix: also push a warning-shaped record for the paused dimension.

**Medium #7 — unbounded terminal-gate retention**
The snapshot kept every gate forever; `open` / `resolve` / `get` /
`list_pending` were O(total historical gates). Fix:
`with_terminal_retention` (default 30 days). Every mutation prunes
terminal gates whose resolution timestamp is older than the window.
Added `terminal_gates_older_than_retention_are_pruned_on_next_write`
regression.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* ci: replace lock-poisoned expects with PoisonError::into_inner

scripts/check_no_panics.py flagged five .expect("...lock poisoned")
calls in the new test_support.rs. Use the same idiomatic recovery
pattern the rest of the codebase uses (see InMemoryBudgetGateStore,
InMemoryBudgetEventSink): on a poisoned lock, recover the inner data
via PoisonError::into_inner rather than panicking. The test gateway's
state is append-only logs / replies queues, so reading them through a
poisoned lock is safe.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Finish A1 / A2 / F1 from plan + honest plan doc update

The plan claimed "all nine items landed" but A1 (production wiring),
A2 (SSE projection), and F1 (full progress strategy) were partials.
This commit finishes the work so the plan matches reality.

**A1 — production-shape accountant builder**

New `ironclaw_reborn_composition::build_default_budget_accountant`
public helper that wires the seeding policy + overestimate factor +
gate store from `BudgetDefaults::compiled_defaults().with_env()` and
returns an `Arc<dyn LoopModelBudgetAccountant>`. Production loop
composers call this with their `PersistentResourceGovernor` +
`FilesystemBudgetGateStore` + LLM-policy-derived cost table; the
local-dev runtime in `build_reborn_runtime` now uses the same helper
instead of duplicating the seeding logic inline. Unit-tier regression
`seeds_compiled_default_user_cap_on_first_touch` proves the helper
installs the compiled-default $5 user cap on first model call.

**A2 — broadcast sink + AppEvent projection**

- `ironclaw_resources::BroadcastBudgetEventSink` wraps
  `tokio::sync::broadcast::Sender<BudgetEvent>` with `subscribe()` /
  `subscriber_count()`. `CompositeBudgetEventSink` fans events to
  multiple sinks.
- Composition fans every `BudgetEvent` to the in-memory sink (for
  tests) AND the broadcast sink (for SSE projection) via
  `CompositeBudgetEventSink`.
- New `AppEvent::BudgetWarn` / `BudgetPause` / `BudgetDenied` /
  `BudgetLimitChanged` wire-stable variants in
  `ironclaw_common::event`.
- `src/bridge/budget_events.rs` carries the projection: a tokio task
  spawned by `spawn_budget_event_projection` drains the broadcast
  receiver and emits the appropriate `AppEvent` via
  `SseManager::broadcast_for_user`. System-scoped events (no user
  identity) are skipped. This is the only producer of these
  `AppEvent` variants per `.claude/rules/gateway-events.md`.
- `RebornRuntime::broadcast_budget_event_sink()` exposes the sink to
  the binary so the startup path subscribes. E2E test
  `broadcast_sink_publishes_events_to_subscribers` drives a real
  `send_user_message` and asserts Reserved + Reconciled lands on the
  broadcast.

**F1 — diminishing-returns stop condition**

The earlier shipped `ParamHash` normalization in
`CapabilityCallSignature` strengthened the existing
`recent_call_signatures`-based repetition detector. This commit adds
the second half of F1: a rolling output-token window that detects
"wedged" loops the repetition detector misses (model keeps
responding but produces no useful output).

- `LoopExecutionState.recent_output_token_counts: BoundedRing<u32, 8>`
  populated by the executor from `LoopModelResponse::usage`.
- `BoundedRing::iter` returns `impl DoubleEndedIterator` so the
  strategy can scan the trailing window.
- `DefaultStopConditionStrategy` gets `min_delta_tokens` (default
  4) + `noprogress_window` (default 4). When the last N turns all
  produce ≤ min_delta_tokens of output, fire
  `StopKind::NoProgressDetected`.
- Regression tests:
  `four_consecutive_low_token_turns_trigger_no_progress` proves the
  detector fires; `occasional_low_token_turn_does_not_trip_no_progress`
  proves a productive turn resets the trailing count.

**Plan doc**

Updated the status header from "all nine items landed" to the
honest per-item shape. Acceptance evidence table expanded with the
new test names. New "Review-feedback fixes layered on top" subsection
documenting all 2 High + 5 Medium findings addressed during review.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Address thermo-nuclear review: collapse filesystem-store duplication, flatten cfg permutations, split budget accountant

Five structural simplifications surfaced by the deep audit on PR nearai#3899, plus
two bug fixes from the earlier review pass:

- ironclaw_resources: extract `cas_snapshot` shared infrastructure
  (`StorageError` + `Snapshot` traits + `CasSnapshotStore<F>` + async-runtime
  worker + per-path lock map) and merge `filesystem_gate_store.rs` into
  `filesystem_store.rs`. Deletes ~350 lines of duplicated read-modify-write +
  worker-thread + CAS machinery; both stores are now thin shims over the
  shared helper.

- ironclaw_reborn_composition: flatten the 4-way cfg permutation in
  `build_reborn_runtime` model-gateway resolution into three flat steps
  (normalize override → build production gateway via cfg-gated helper →
  test override wins). Also drops the `unused_mut` warning.

- ironclaw_reborn_composition: collapse the 3-layer test-only setter dance
  for `model_gateway_override` / `model_cost_table_override` into a single
  setter pair gated on `cfg(any(test, feature = "test-support"))`. Deletes
  the `RebornRuntimeInputTestExt` extension trait — integration tests now
  call the inherent methods directly.

- ironclaw_loop_support: split the 1305-line `budget_accountant.rs` into
  `budget_cost_table.rs` (ModelCost/ModelCostTable trait/ZeroCostTable/
  StaticModelCostTable), `budget_seeding.rs` (BudgetSeedingPolicy), and
  `budget_accountant.rs` (just GovernorBackedAccountant). Each module now
  owns one concern.

- ironclaw_resources: add `impl Display for ResourceAccount` and route the
  hierarchical account-label rendering through it; delete the 60-line
  bespoke `account_label` helper from `src/bridge/budget_events.rs`.

- ironclaw_common + bridge: collapse the four `AppEvent::Budget*` variants
  into a single typed `AppEvent::Budget(AppBudgetEvent)` with the four
  shapes carried inside the enum. Wire-shape stays identical (snake_case
  serde tag).

- ironclaw_resources + ironclaw_loop_support: thread real gate id through
  `BudgetEvent::GateOpened { gate_id, needed, at }` (new variant) and have
  the accountant emit it via the broadcast event sink after store.open
  succeeds. The bridge now projects `BudgetEvent::GateOpened` (not
  `ApprovalRequested`) into `AppEvent::Budget(Pause { gate_id, ... })` so
  SSE consumers receive the persisted gate id rather than a fabricated
  zero uuid.

- ironclaw_agent_loop: in the F1 token-counting path, push to
  `recent_output_token_counts` only when the model response carries
  `Some(usage)` and only on the `AssistantReply` arm (instead of
  `unwrap_or(0)`). Diminishing-returns detection now reflects real spend.

Net delta: -461 lines (+999 / -1460). Workspace `cargo clippy` clean,
`cargo test` clean on ironclaw_resources / ironclaw_loop_support /
ironclaw_reborn_composition; budget_e2e + budget_approval_e2e both green.

Pre-existing CI failures (`cli::tests::test_version` stack overflow,
`facade_factory::production_*` RuntimeProcessPort missing) are unrelated
and reproduce on the pristine branch tip.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(cli): refresh insta snapshots after runtime-policy flag additions

The `import`-feature variants of the help snapshots were left stale when
`--deployment-mode`, `--runtime-profile`, `--yolo-disclosure` were added in
e9628ed (nearai#3243); the `_without_import` variants were updated but these
were not. CI was failing the snapshot assertion under the slim PR matrix
(`--features postgres,libsql,html-to-markdown,bedrock,import`).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Address PR nearai#3899 thermo-nuclear review (TN #1, #2, #3)

TN #1 — budget defaults resolved in wrong layer:
  - `build_default_budget_accountant` no longer reads process env; it
    now takes `&BudgetDefaults` as a parameter and the caller owns the
    config-layer precedence (compiled → section → env) plus the
    `validate()` call.
  - `RebornRuntimeInput` gains an optional `budget_defaults` field +
    `with_budget_defaults()` builder so the composition root passes a
    pre-resolved value. `build_reborn_runtime` falls back to
    `compiled_defaults().with_env() + validate()` when none is supplied
    so existing call sites keep working.

TN #2 — gate-store scoping at wrong boundary:
  - `BudgetGateStore` trait methods (`open`, `resolve`,
    `expire_pending_older_than`, `get`, `list_pending`) now take
    `&ResourceScope` as first arg. `GovernorBackedAccountant` passes
    the caller's scope from `resource_scope(context)`.
  - `CasSnapshotStore` gains `update_with_scope` so the same store
    instance can route per-operation. `FilesystemBudgetGateStore` no
    longer takes scope at construction — one shared instance serves
    every tenant via the `ScopedFilesystem` mount view.
  - `InMemoryBudgetGateStore` ignores scope (suitable for single-tenant
    tests / local-dev); production multi-tenant filesystem path is
    correctly partitioned by `ResourceScope`.
  - `RebornRuntime::apply_resolved_budget_gate` now takes scope too.

TN #3 — half-wired projection bridge:
  - Removed `src/bridge/budget_events.rs`, its `spawn_budget_event_projection`
    helper, the `AppEvent::Budget` variant, and the `AppBudgetEvent`
    type. No production caller ever subscribed the broadcast sink
    onto SSE and no frontend consumed the variant, so the
    half-wired bridge is gone pending a real owner that spawns a
    projection task with shutdown cancellation.
  - The runtime's `broadcast_budget_event_sink()` accessor stays so
    a future production composer can still subscribe without
    rebuilding the runtime.

Bonus — to keep budget e2e tests working under the new libsql local-
dev path that origin/reborn-integration introduced, added
`PersistentResourceGovernor::with_event_sink` (parity with the
`InMemoryResourceGovernor` accessor). The libsql variant of
`build_local_dev_store_graph` now wires the composite sink to the
persistent governor so governor-emitted `Warned`/`Denied`/`Reserved`/
`Reconciled` events reach subscribers on both feature paths.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Wire budget-event projection task into RebornRuntime

Re-implements PR nearai#3899 follow-up A2 / Thermo-Nuclear #3 with a real
production owner instead of leaving the broadcast sink half-wired:

- `crates/ironclaw_reborn_composition/src/budget_events.rs` (new):
  `BudgetEventObserver` trait + `TracingBudgetEventObserver` default
  observer + crate-internal `BudgetEventProjection` task that drains
  the runtime's broadcast `Receiver<BudgetEvent>` and forwards every
  event to the observer. Cancellation via `CancellationToken`; lagged
  subscribers logged and resumed; receiver-closed exits cleanly.

- `RebornRuntimeInput::with_budget_event_observer(...)` lets
  production owners install a custom observer (SSE projection, WS
  fan-out, telemetry export). When unset, the runtime installs the
  tracing observer so events always surface in structured logs.

- `build_reborn_runtime` always spawns the projection task at runtime
  construction; `RebornRuntime::shutdown` cancels it and awaits the
  handle so background state drains before the runtime drops.

- E2E test `projection_delivers_budget_events_to_installed_observer`
  drives `build_reborn_runtime` with a capturing observer and asserts
  the observer sees `Reserved` + `Reconciled` from a real model call,
  testing through the caller per `.claude/rules/testing.md`.

- Existing `broadcast_sink_publishes_events_to_subscribers` updated
  to expect the runtime's own projection task as a baseline
  subscriber.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(reborn): rustfmt the merged loop_support import block

The conflict resolution for the post-merge import list was not run
through rustfmt; CI Formatting flagged the wrapping. No logic change.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
theredspoon pushed a commit that referenced this pull request Jun 21, 2026
…r injection (nearai#4588)

* feat(reborn): expose a trajectory observer hook on RebornRuntimeInput

The reborn runtime is sealed: build_reborn_runtime returns only the final
AssistantReply, and per-step capability (tool) calls + results live in internal
stores. Downstream consumers (benchmark harnesses, UI/debuggers) can't observe
the agent's trajectory.

Add `RebornTrajectoryObserver` (pub trait: on_capability_input(call_id, name,
args) / on_capability_result(call_id, output)) and
`RebornRuntimeInput::with_trajectory_observer`. The local-dev capability IO
(`LocalDevCapabilityIo`) forwards each tool call's name+args (at input staging)
and result (at result write) to the observer when present — reusing the same
data it already records for display previews. No-op when unset; best-effort.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* debug: trace observer hook firing (temporary)

* feat(reborn): trajectory observer — capability_id on result, reliable spine

Provider tool calls are staged by a lower decorator that bypasses the
LocalDevCapabilityIo input path, so on_capability_input does not fire for
them. on_capability_result fires for every completed capability — make it
carry the capability_id so consumers can reconstruct the trajectory (name +
output) from results alone. Input args capture is a follow-up.

* feat(reborn): capture capability input args at the host port chokepoint

Provider tool calls are staged by ProviderToolCallInputResolver, which keeps
args in a private map and bypasses the capability-IO input hook — so inputs
never reached the trajectory observer (only results did). Move the observer
trait down to ironclaw_loop_support (CapabilityTrajectoryObserver, re-exported
from composition as RebornTrajectoryObserver) and hook it in
HostRuntimeLoopCapabilityPort::invoke_capability right after the input
resolves — the one place the model's resolved arguments are visible. Threaded
through HostRuntimeLoopCapabilityPortFactory + the local-dev factory. Result
hook unchanged. Now name + args + output are all captured.

* feat(reborn): host LLM-provider injection seam

ResolvedRebornLlm::with_provider — drive the runtime with a caller-supplied
LlmProvider (e.g. an instrumented wrapper that counts tokens/cost and captures
reasoning) instead of always building one from config; build_llm_gateway honors
the override. The only viable observability path for reborn, whose model calls
run in spawned worker tasks a per-task tracing subscriber can't reach.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(reborn): cover trajectory observer + LLM provider override seams

Addresses Firat's two blocking review findings on nearai#4588 (both
missing-integration-test, per AGENTS.md "test through the caller"):

1. Trajectory observer callbacks — drive the real call sites with a
   recording CapabilityTrajectoryObserver:
   - host port: invoke_capability via HostRuntimeLoopCapabilityPortFactory
     ::with_trajectory_observer asserts on_capability_input fires with the
     resolved capability id + tool-call arguments.
   - local-dev IO: register_provider_tool_call_input + write_capability_result
     assert on_capability_input and on_capability_result fire and correlate by
     input ref.

2. LLM provider override — build_llm_gateway_drives_provider_override_not_config
   injects a counting mock via ResolvedRebornLlm::with_provider, points config
   at a dead endpoint, and asserts the gateway returns the mock's sentinel
   (proving the override is driven, not a config-built chain).

Also fixes pre-existing breakage this surfaced: 5 LocalDevLoopCapabilityPort
Factory test initializers (shell_tests.rs + tests.rs) were missing the
trajectory_observer field added by this PR, so the composition crate's tests
did not compile under --features root-llm-provider.

loop_support: 301 passed; composition (root-llm-provider): 520 passed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reborn): make trajectory observer input semantics consistent

Addresses Copilot's follow-up findings on the observer seam:

- Drop the `on_capability_input` callback from `LocalDevCapabilityIo::
  register_provider_tool_call_input`. It forwarded the raw provider tool
  name (`builtin_echo`) as the capability id — conflicting with the
  observer contract (resolved dotted `builtin.echo`) and the authoritative
  port-level hook — and `ProviderToolCallInputResolver` doesn't delegate
  here for provider tool calls, so it never fired in practice anyway.
  `HostRuntimeLoopCapabilityPort::invoke_capability` remains the single
  source of `on_capability_input` (resolved id); `LocalDevCapabilityIo`
  remains the source of `on_capability_result`.

- Clarify the trait doc: `arguments` is the raw model-emitted tool-call
  input resolved from the input ref (the callback fires before schema
  normalization), which is what the trajectory should record.

- Refocus the local-dev test on `on_capability_result` forwarding +
  correlation, and assert input staging does NOT emit `on_capability_input`
  from local-dev IO. Port-level input semantics stay covered by the
  capability_port.rs test.

loop_support: 301 passed; composition (root-llm-provider): 520 passed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* wire trajectory_observer through RefreshingLocalDevCapabilityPortConfig

Completes the main-merge conflict resolution: local_dev.rs passes
trajectory_observer into the refreshing-port config, so the config struct +
port struct must carry it and build_inner must apply it via
.with_trajectory_observer(). (Missed staging this file in the merge commit.)

* test(reborn): lock down the observability seams against regression

nearai#4588 exposes two seams a downstream harness relies on. Add tests so a
future refactor can't silently break either:

- capability_io_forwards_result_to_trajectory_observer: drives
  write_capability_result and asserts on_capability_result fires with the
  correct (call_id, capability_id, output) — the result half of the
  trajectory observer (tool-call outputs).
- build_llm_gateway_drives_provider_override_not_config: asserts the gateway
  drives a provider injected via ResolvedRebornLlm::with_provider (config
  points at a dead endpoint), proving the provider-injection seam works —
  this is how the bench captures reasoning / tokens / cost / system-prompt /
  tool-definitions. (Restores the test dropped during the main merge.)

The input half (on_capability_input) is already covered by
invoke_capability_forwards_resolved_input_to_trajectory_observer in
ironclaw_loop_support. All three pass.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(reborn): drop the false-confidence result-hook test

capability_io_forwards_result_to_trajectory_observer called
write_capability_result directly, so it stayed green even though the
result hook is unreachable end-to-end while capability dispatch fails
(the LocalDevYolo InputEncode regression) — i.e. it did not fail when
the feature it claimed to cover was actually broken. Remove it rather
than ship false confidence.

The result hook lives in LocalDevCapabilityIo and is only reached by a
real local-dev runtime turn, so an honest guard must drive the full
runtime and is red until the dispatch regression is fixed; that guard
belongs as an end-to-end test (PR, once green) or a bench pre-flight,
not a direct-call unit test.

Kept: invoke_capability_forwards_resolved_input_to_trajectory_observer
(input hook, real port code path) and
build_llm_gateway_drives_provider_override_not_config (provider seam) —
both genuinely fail if their seam regresses.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(reborn): address review on the trajectory-observer + provider seams

Resolves Henry + Firat review comments on nearai#4588:

- Composition-owned RebornTrajectoryObserver trait + adapter to the
  loop-support CapabilityTrajectoryObserver, instead of re-exporting the
  substrate trait directly (CLAUDE.md: facade-shaped handles only). Loop-support
  contract changes no longer break the public Reborn API. (Henry#8)

- Safe-preview by default: with_trajectory_observer now forwards bounded
  (truncated strings / capped arrays) payloads so a logs/UI/telemetry sink stays
  within the model-visible display boundary; a trusted in-process consumer that
  needs verbatim tool I/O opts in via the new with_raw_trajectory_observer.
  (Henry#5)

- catch_unwind around both observer call sites (input hook in capability_port,
  result hook in LocalDevCapabilityIo) so a panicking observer can't unwind the
  capability hot path; trait doc now states the never-block / panic-caught
  contract. (Henry#1/#6)

- e2e test local_dev_runtime_forwards_tool_call_trajectory_to_raw_observer:
  drives a real build_reborn_runtime turn dispatching builtin.echo and asserts
  BOTH input and result callbacks fire on the genuine dispatch path — honest
  coverage that replaces the dropped direct-call result-hook test, and proves
  the observer threads through build_reborn_runtime. (Firat#1, Henry#3/#7)

- Strengthened provider-injection docs: the config-vs-override invariant and why
  the feature-gated seam takes the LlmProvider substrate trait. (Henry#4/#9/#11)

- Fixed the LocalDevCapabilityIo observer field comment to describe its actual
  result-only responsibility. (Henry#10)

Provider-override coverage (Firat#2/Henry#2) already landed in
build_llm_gateway_drives_provider_override_not_config.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(reborn): second-round review fixes on the trajectory/provider seams

Addresses Henry's review of the first round (nearai#4588):

- safe_preview_value now bounds objects (entry cap), recursion depth, and total
  nodes — not just strings/arrays — so a wide or deeply nested capability result
  can't force unbounded traversal/allocation on the hot path. (3405419089)

- Narrowed the loop-support CapabilityTrajectoryObserver to input-only:
  HostRuntimeLoopCapabilityPort never staged results through the port (results
  go via LoopCapabilityResultWriter), so advertising on_capability_result there
  was a contract a direct user could never see fire. Result observation stays on
  the composition path (LocalDevCapabilityIo). (3405419104)

- Synthetic capabilities (e.g. builtin.skill_activate) bypass the inner port's
  input hook, so the synthetic wrapper now emits on_capability_input itself after
  resolving input — otherwise consumers saw an unpaired result with no args.
  (3405419110)

- Provider injection no longer accepts a wholesale Arc<dyn LlmProvider> through
  the facade: with_provider is replaced by with_provider_factory, a decorator
  Fn(Arc<dyn LlmProvider>) -> Arc<dyn LlmProvider>. The composition always builds
  the provider from config (config stays the single construction source —
  collapses the old config-vs-override invariant too) and hands it to the factory
  to wrap. (3405419100, 3405419146)

- New caller-level test local_dev_runtime_safe_preview_observer_receives_bounded_payload:
  installs the default with_trajectory_observer, drives a real turn with a large
  echo payload, asserts the observer receives a truncated preview. (3405419095)

- Dropped the stale nearai#4588/main-rebase comment for a durable invariant. (3405419113)

cargo test (loop_support + reborn_composition, single-threaded) green; clippy
clean on touched files.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* style(reborn): rustfmt the trajectory/provider review changes

Formatting-only: import grouping + mod ordering in the two lib.rs re-export
blocks, and wrapping in runtime.rs / local_dev.rs / trajectory_observer.rs.
Fixes the Formatting + Code Style CI checks. No behaviour change.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* style(reborn): drop std-Mutex guard before await in observer e2e tests

clippy::await_holding_lock (-D warnings): the two trajectory-observer e2e
tests held the observer's std::sync::Mutex guard across runtime.shutdown().await.
Shut down before inspecting the recorded callbacks (the data is already captured
during the turn) so no guard is held across an await. Fixes Clippy (all-features).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* WIP(bench): http empty-body + multi-tool-call port reuse + final-answer nudge

Local checkpoint so the bench builds against a stable tree (uncommitted
edits were being reverted mid-session). Bundles: http body() empty-field
fix, RefreshingLocalDevCapabilityPort register reuse, the gated
final-answer nudge + interactive_profile gate flip, and the
trajectory-observer safe_preview borrow fix.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* style(reborn): wrap an over-long line for rustfmt 1.9.0

CI installs the latest stable rustfmt (1.9.0 / Rust 1.96), which wraps a
long eprintln! that older rustfmt left inline. Fixes the Formatting check.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* style(reborn): collapse nested if for clippy 1.96 collapsible_if

clippy 1.96 (CI's stable) flags the nested if-let in the final-answer-nudge
site as collapsible; fold it into a let-chain. No behaviour change. Fixes
Clippy (all-features).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* WIP(bench): multi-tool-call port reuse (matches main nearai#4790)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* WIP(bench): nudge isolation - disable gate to measure marginal contribution

* Revert stray bench WIP accidentally committed onto this branch

Removes the http/nudge/multi-tool-call/diagnostic WIP commits
(c4bbb5f, 2c670b4, 6da818a) that were committed onto the
reborn-trajectory-observer branch by mistake during benchmarking and
swept to origin by a main-merge push. Restores the affected files to
origin/main (multi-tool-call is already fixed there by nearai#4790; the http
fix lives in PR nearai#4827). Observer-owned changes in state.rs,
refreshing_capability_port.rs, and local_dev.rs are preserved minus the
stray WIP additions. No history rewrite / force-push.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(reborn): preserve provider factory across reload + reject observer off local-dev

Addresses Firat's review of the trajectory/provider seams (nearai#4588):

- Provider factory now survives a live config reload. build_llm_gateway applied
  the factory to the bare config provider *before* wrapping it in the
  SwappableLlmProvider, so the first WebUI/settings reload (which swaps the
  swappable's inner) silently dropped the instrumentation wrapper. Invert the
  layering: build the config provider, put it behind the swappable + reload
  handle, then apply the factory *over the swappable* for the gateway-facing
  provider. Reloads swap the inner; the wrapper stays in the call path.
  Regression test provider_factory_survives_live_reload reloads and proves the
  wrapper still observes subsequent model calls.

- Reject a trajectory observer on profiles without a local runtime. The observer
  is wired only through the local-dev capability path; Production silently
  dropped it, so a caller got an empty trajectory with no error. Fail fast with
  InvalidArgument and document the seam as local-dev/bench-only. Test
  build_reborn_runtime_rejects_trajectory_observer_for_production.

cargo fmt + clippy (all-features, -D warnings) clean under rustfmt 1.9/clippy 1.96.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(reborn): note trajectory observer is local-dev/bench-only

Document the local-dev-only constraint + fail-fast behavior on the public
with_trajectory_observer / with_raw_trajectory_observer setters (Firat review).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Pranav Raja <pranav.raja@near.ai>
theredspoon pushed a commit that referenced this pull request Jun 21, 2026
…ing the run (nearai#4954)

* fix(reborn): surface approval-gate denial to model instead of cancelling the run

Approval-gate denial in Reborn cancelled the run (deny_gate /
replay_denied_gate -> cancel_run), so the model never learned the user
declined and the next trigger re-issued the same approval-gated
capability and re-blocked — the same loop class nearai#4944 removed for auth
gates.

Mirror nearai#4944 for approval gates: denial now RESUMES the parked run
carrying a denial disposition; the capability stage converts ONLY the
approval-gated call into a model-visible non-retryable Authorization
failure ("approval gate denied by user", SameCallRetryConstraint::
Forbidden) and the loop continues. Unrelated parallel calls are
unaffected.

Per the maintainability review of the plan, this unifies rather than
duplicates the nearai#4944 plumbing:
- ironclaw_turns: AuthResumeDisposition -> GateResumeDisposition (one
  gate-agnostic enum); ResumeTurnRequest/TurnRunRecord/TurnRunState/
  AgentLoopDriverResumeRequest field auth_resume_disposition ->
  resume_disposition. Serde key pinned to "auth_resume_disposition"
  (rename attr) so persisted run records still deserialize; legacy-key
  round-trip test added.
- ironclaw_agent_loop: PendingApprovalResume gains a disposition field;
  the auth denied short-circuit in CapabilityStage::process is extracted
  into ONE shared short_circuit_denied_resume helper used by both the
  auth and approval paths (no second copy).
- ironclaw_product_workflow: approval deny_gate / replay_denied_gate
  resume instead of cancel; ResolveApprovalInteractionResponse::Denied
  (CancelRunResponse) -> Resumed(ResumeTurnResponse); idempotent replay
  guarded by terminal run status.
- ironclaw_reborn: PlannedDriver::resume stamps the disposition onto the
  pending resume that is set (auth or approval).

Decisions (plan docs/plans/2026-06-15-reborn-approval-deny-continue.md):
both Denied and Cancelled continue (consistent with nearai#4944, no Cancel
variant). The extension_install/extension_search missing-observation gap
is a separate PR; the user-visible extension-install loop is only fully
closed when both land.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reborn): address PR nearai#4954 review — stamp denial on matching gate slot only

Review round 1 fixes:

- planned_driver: the denial disposition was stamped onto BOTH
  pending_auth_resume and pending_approval_resume on a comment-only "one
  slot at a time" invariant. GateStage deliberately preserves a pending
  auth resume when a non-auth gate blocks mid-re-dispatch, so both slots
  can be set at once; stamping both corrupted an unrelated auth resume.
  Now stamps only the pending slot whose gate_ref matches the blocking
  gate (state.last_gate). Adds a regression test asserting the auth slot
  stays None when the approval gate is denied, plus an end-to-end
  resume() drive.
- approval replay: match GateResumeDisposition::Denied explicitly rather
  than is_some(), keeping the gate-agnostic carrier tied to denial.
- tests: real TurnRunRecord struct-level serde test (legacy
  auth_resume_disposition key → resume_disposition) + snapshot-level
  legacy denied-marker test; new deny-path resume-error test asserting
  the record is denied and the run is never cancelled on resume failure.
- arch-exempt annotation on short_circuit_denied_resume's
  too_many_arguments allow (plan nearai#4954); stale comments/typos fixed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reborn): route denied-gate replay through resume_turn idempotency (auth + approval)

Review finding #7: the denied-gate replay paths derived idempotency from
current run state (TurnStatus::is_terminal() guard) rather than replaying
through resume_turn. After the first Deny resumed the run, a transport
retry with the same idempotency key arriving after the runner completed
returned StaleGate/StaleAuth instead of the original ResumeTurnResponse —
the observable result depended on runner timing.

resume_turn is idempotent by key (memory.rs:665 returns the cached
Result from resume_idempotency before the precondition check). Both the
approval (replay_denied_gate) and auth (resume_denied_auth replay arm)
paths now replay through resume_turn with the same key, deleting the
terminal-guard branching: a retried key replays the original response
regardless of run state; a genuinely stale request with a fresh key
still errors via the precondition. Auth and approval kept symmetric.

FakeTurnCoordinator now models resume idempotency by key so the replay
tests are meaningful; terminal-guard assertions re-framed around
same-key replay vs fresh-key stale, plus an explicit idempotent-replay
test on both services.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reborn): fail closed on ambiguous dual-slot stamp; lock deny-before-resume order

Review round 2 (both Major):

- planned_driver stamp_resume_disposition: the if/else-if silently stamped
  the auth slot if both pending slots matched last_gate. At the denial-
  attribution boundary that could misattribute an approval denial. Now an
  explicit 4-way match fails closed on the ambiguous (true, true) case
  (warn + stamp neither). Test added.
- approval_interaction_contract deny-resume-error test: asserted only
  aggregate call counts, which pass even if call order regressed. Added a
  shared ordered trace across the resolver (deny) and coordinator
  (resume_turn) fakes and assert deny is recorded strictly before resume.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(reborn): strengthen replay/checkpoint coverage; downgrade fail-closed log to debug

Review round 3 (straightforward):
- idempotent deny-replay tests (auth + approval) now assert full
  ResumeTurnResponse payload equality, not just run_id.
- stamp_resume_disposition ambiguous-dual-slot diagnostic: warn! -> debug!
  (REPL/TUI logging rule — internal fail-closed diagnostics use debug!).
- executor: assert the first approval BeforeBlock checkpoint carries
  pending_approval_resume.disposition == None before any denial.
- executor: denied-approval short-circuit no-matching-call test (denied X,
  model emits only Y -> X not surfaced, Y dispatches, pending cleared).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reborn): unify gate Declined resolution; keep WebUI processing on resume

Review round 3 (design + High):

- WebUiGateResolution: the approval card sends `denied`, the auth cards
  send `cancelled`, and both are now treated identically (resume the run
  and surface the decision to the model). Run termination is a separate
  control (the X -> cancelRun route), not a gate resolution. Collapsed
  the two equivalent variants into one `Declined` (serde aliases
  "denied"/"cancelled" keep the wire stable; no JS change). Facade maps
  Declined -> Deny for auth, approval, and the generic fallback.

- #6 WebUI desync (High): useChat.resolveGate kept processing only for
  approved/credential_provided, dropping processing + activeRun on
  denied/cancelled — but those now resume the run. resolveGate now always
  keeps processing/activeRun; the terminal run_status SSE event clears it
  and the X/cancelRun path remains the only stop. Fixes the latent
  auth-cancelled desync from nearai#4944. assets.rs assertion + useChat tests
  updated.

- #7 helper weight: short_circuit_denied_resume no longer returns the
  DeniedResumeOutcome enum / boxes LoopExecutionState / clones the batch.
  It returns ControlFlow<TurnCompletedStep, (state, remaining_calls)>; the
  completed_turn/empty-remaining tail moved to the two call sites. Heavy
  per-denied-call failure synthesis stays shared (one helper).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(reborn): share denied-approval resume between deny_gate and replay

Review (Medium): deny_gate and replay_denied_gate built an identical
ResumeTurnRequest, mapped the same errors, and returned the same Resumed
shape — the only difference was deny_gate's one-off resolver.deny side
effect. Extracted a shared resume_denied(request, run_id) helper; deny_gate
performs the durable denial then delegates to it, and replay_denied_gate
calls it directly. Removes the duplicated request construction / path
handling.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
personal-upstream-sync Bot pushed a commit that referenced this pull request Jul 8, 2026
… inspection (nearai#5280)

* docs: spec for Trace Commons instance enrollment, profiles, and trace inspection

Cross-repo design (ironclaw + trace-commons-server) for three coexisting
capabilities: instance-wide enrollment, per-user contributor accounts via
login-links, and submitted-trace inspection. Introduces a trace-credential
resolver so the existing user-invite model and the new instance-wide model
both function on one instance, with personal-invite enrollment taking
precedence. Server change is additive (optional per-user subject through
claim issuance + login-link + account resolution).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs: Slice 0 plan — trace-commons-server per-user subject

TDD plan for the one server change the whole effort depends on: accept an
optional opaque subject in the upload-claim request and derive a per-user,
tenant-namespaced principal at device-key issuance. Submission attribution,
login-link account resolution, and trace readback all become per-user
automatically from the shared bearer principal; absent subject reproduces
today's behavior. Targets trace-commons-server (contributor-account-slice1).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs: IronClaw plans for Trace Commons slices 1-4

Slice 1: trace-credential resolver (personal-invite wins, instance fallback
  with per-user subject) + admin-gated instance enrollment.
Slice 2: per-user subject plumbing through upload-claim request + submission.
Slice 3: trace_commons.account_login_link first-party capability (profiles).
Slice 4: per-user submitted-trace inspection across reborn_traces →
  product_workflow facade → webui_v2 handler → frontend.

Each plan is bite-sized TDD against verbatim-extracted current code.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(traces): trace-credential resolver (personal invite wins, instance fallback w/ subject)

* refactor(traces): single dir-parameterized policy-path site (remove resolver duplication)

Extract `trace_contribution_dir_for_scope_at`, `trace_policy_path_at`,
`read_trace_policy_for_scope_at`, and `write_trace_policy_for_scope_at`
as the canonical base-dir-parameterized path helpers. All public
functions (`trace_contribution_dir_for_scope`, `read_trace_policy_for_scope`,
`write_trace_policy_for_scope`) now delegate to the `_at` variants with
`ironclaw_base_dir()` — signatures unchanged.

The inline `read_policy` closure in `resolve_trace_credentials_at` that
re-implemented path layout is deleted; it now calls
`read_trace_policy_for_scope_at` directly. The test `write_policy_at`
helper's bespoke path construction is replaced with a call to
`write_trace_policy_for_scope_at`. The now-dead `trace_policy_path`
function is removed. Path layout is encoded in exactly one place.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(traces): instance-level enrollment write path (scope None)

* test(traces): make instance-enrollment test hermetic (tempdir, no global base)

Rework `instance_onboard_writes_instance_level_policy` to operate entirely
under a `tempfile::tempdir()`:
- Compute instance_dir as base.path().join("trace_contributions") (scope=None
  layout, no users/<hash> segment) rather than calling the global LazyLock.
- Call `onboard_at_dir_with_sink` directly against the tempdir so the test
  never touches the real ~/.ironclaw tree.
- Assert policy.json by reading and deserializing it from the tempdir.
- Remove all manual std::fs::remove_* cleanup lines; tempdir drops automatically.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(admin): AdminScope::enroll_instance_trace_commons (admin-gated instance enrollment)

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(traces): carry optional per-user subject in upload-claim request

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(traces): thread resolver subject into submission claim context

* test(traces): claim request carries per-user subject end-to-end

* feat(traces): mint_account_login_link_via_sink (POST /v1/account/login-links)

Add `mint_account_login_link_via_sink` to ironclaw_reborn_traces:

- `TraceUploadClaimContext::for_account(subject)` constructor for
  account-management call contexts (no trace/submission ids, no
  consent scopes).
- `AccountLoginLink { account_id, url }` return type.
- `account_login_links_url(policy)` helper that derives the login-links
  URL from the upload-claim issuer URL (strip /v1/trace-upload-claim,
  append /v1/account/login-links).
- `mint_account_login_link_inner(base_dir, ...)` private dir-parameterised
  core: resolves credentials, selects correct scope_dir for DeviceKey
  auth (instance enrollment → instance scope dir; personal → user scope
  dir), mints bearer, POSTs subject, parses response.
- `mint_account_login_link_via_sink(tenant_id, user_id, sink)` public
  entry point wrapping the inner function with the real base dir.

Tests (hermetic, tempdir-isolated):
- `mint_account_login_link_posts_subject_and_returns_url`: verifies the
  posted subject equals `local_pseudonymous_contributor_id(trace_scope_key(...))`
  for instance-enrolled users via an axum mock serving both the
  upload-claim issuer and the login-links endpoint.
- `mint_account_login_link_errors_when_not_enrolled`: verifies error path.
- `ReqwestContributionSink` test helper added to the test module.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(traces): error instead of silent misroute in account_login_links_url

Replace the unwrap_or_else fallback (which silently used the full issuer
URL as a base when the /v1/trace-upload-claim suffix was absent) with an
explicit anyhow error. Add two unit tests: one asserting an Err on a
wrong-suffix URL, one asserting the correct .../v1/account/login-links
URL on a valid issuer.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(host_runtime): add consent-gated trace_commons.account_login_link capability

Mints a Trace Commons browser login URL via host network egress, mirroring
dispatch_profile_token. Includes consent gate, enrollment pre-check,
HostEgressContributionSink routing, and two e2e tests.

Also fixes a sanitizer bug: validate_runtime_request was rejecting
authorization headers on all requests, including RuntimeKind::FirstParty.
FirstParty requests are host-internal and trusted to carry bearer tokens;
the sensitive-header and manual-credentials guards now only apply to
untrusted plugin runtimes (WASM/MCP/Script).

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(host_runtime): route trace bearer via credential injection; restore FirstParty sensitive-header guard

Commit 9e25d99 blanket-exempted all RuntimeKind::FirstParty requests from the
egress sensitive-header and manual-credentials guards so the host-minted Trace
Commons bearer could pass. builtin.http is also FirstParty but forwards
model-supplied headers, so this let the model smuggle Authorization/Cookie/
x-api-key headers (or user:pass@ URLs) to allowlisted hosts.

Revert the sanitize.rs exemption (guards now apply to ALL runtimes again) and
deliver the trace bearer through the staged credential-injection path instead:
the HostEgressContributionSink stages the minted token one-shot via
RuntimeSecretMaterialStager and declares a StagedObligation Authorization-header
injection, mirroring the SlackProtocolHttpEgress pattern. The stager is now
exposed to first-party handlers via InvocationServices. Covers the profile_token,
profile_set/community-profile, and account_login_link bearer paths.

Regression tests: FirstParty + raw authorization header -> denied; FirstParty +
user:pass@ URL -> denied.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(traces): fetch_account_traces_via_sink (GET /v1/account/traces, per-user)

- Add ContributionHttpMethod::Get variant; update all exhaustive match
  sites in ironclaw_reborn_traces and HostEgressContributionSink in
  ironclaw_host_runtime.
- Extract account_api_base_url() shared helper; account_login_links_url
  and new account_traces_url both delegate to it (DRY).
- Add AccountTraceItem (Debug, Clone, Serialize, Deserialize; serde defaults).
- Add fetch_account_traces_via_sink / fetch_account_traces_inner mirroring
  mint_account_login_link pattern: unenrolled -> Ok(vec![]), non-2xx ->
  Ok(vec![]), transport error -> Err.
- Tests: hermetic axum mock (GET /v1/account/traces), unenrolled empty-list,
  URL shape with/without limit.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(reborn): trace_account_traces facade method + wire types

Adds RebornAccountTrace / RebornAccountTracesResponse wire types and
a trace_account_traces default method on RebornServicesApi, mirroring
the trace_credits egress pattern (crate-local hardened reqwest, no
host-egress sink). Also adds fetch_account_traces (direct path) to
ironclaw_reborn_traces::contribution so the facade can fetch server
traces without coupling to RuntimeHttpEgress.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(reborn): GET /api/webchat/v2/traces/account handler + contract test

* feat(reborn-ui): render submitted Trace Commons traces in settings

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* chore(traces): document flush-gate limitation, hermetic account-traces contract test, annotate sink scaffold

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* style(traces): cargo fmt across Trace Commons slice changes

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(traces): resolver-aware flush gate (instance-enrolled users can contribute)

The autonomous trace-flush gate read only the per-scope (personal-invite)
policy and aborted when it was disabled, so instance-enrolled users (whose
enrollment lives at scope None) could never contribute traces — and the
per-user scope_dir would also fail to load the instance device key.

Introduce a single EffectiveFlushTarget resolver (resolve_effective_flush_target,
mirroring resolve_trace_credentials but keyed on the already-composed scope
string) that returns the policy, device-key dir, and per-user subject in one
policy-read/path pass. The flush gate now proceeds for instance-only enrollment,
loads the device key from the instance (None) dir, and attributes uploads via
the per-user pseudonymous subject. The redundant subject_for_scope helper (which
re-read the same policies with silent .ok() error swallowing) is removed and its
logic folded into the new helper with proper error propagation.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(traces): include per-user subject in upload-claim cache key

Under instance enrollment every user shares the same instance device-key
dir (scope None), so the upload-claim cache key — which keyed on scope_dir
but not subject — collided across users. A bearer minted for one subject
could be served from cache to another, mis-attributing traces / leaking
across users. Add a hashed subject component to the DeviceKey cache key and
a regression test proving two subjects sharing a scope_dir get distinct keys
(and a no-subject personal-invite context stays distinct from both).

Found by Codex review of PR nearai#5280.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(traces): address CodeRabbit review on PR nearai#5280

- account_login_link manifest: declare ReadFilesystem effect (it reads
  local enrollment/policy/device-key state before egress), matching
  profile_token. (CR #2)
- account-traces fetch: always send a bounded, clamped limit
  ([1, 500], default 200) so None never triggers an unbounded server
  history fetch. (CR #3)
- direct fetch path: bound the response body with a hard byte ceiling
  (256 KiB) via a chunked bounded reader, instead of buffering unbounded. (CR #5)
- account-traces fetch (both sink + direct): stop swallowing every
  non-2xx as an empty list — 404 = legitimate empty (no account yet),
  all other non-2xx surface as Err so the WebUI renders a sanitized
  unavailable state. Add regression tests (500 -> err, 404 -> empty). (CR #6)
- trace-commons-tab.js: render missing final_credit as "—" not "0.00";
  surface useAccountTraces() query errors instead of collapsing them to
  "no traces". (CR #7, #8)
- handlers contract test: capture the forwarded caller in the
  trace_account_traces stub and assert the route threads the
  authenticated user id (test-through-the-caller). (CR #9)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(ci): cover trace_commons.account_login_link + backfill trace i18n keys

PR nearai#5280 added the builtin.trace_commons.account_login_link capability and a
submitted-traces UI section, but left three guardrail/parity tests un-updated,
turning CI red:

- ironclaw_host_runtime builtin_first_party_package_declares_expected_capabilities:
  register account_login_link in the expected id list and its Ask-permission arm.
- reborn_builtin_first_party_capability_e2e_coverage_is_complete: add genuine
  e2e coverage by exercising account_login_link in the existing trace_commons
  parity test (confirmed=true on a not-enrolled scope returns a deterministic
  NotEnrolled, no network), grant it in the harness allow-set, and add it to the
  model-visible surface test and the covered-capability list.
- ironclaw_webui_v2_static all_locales_share_the_en_key_set: backfill the six new
  traceCommons.* submitted-traces keys into all ten non-en locales.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(traces): address CodeRabbit review — withhold login URL, type errors, wire i18n

- Security (host_runtime): dispatch_account_login_link returned the one-time
  login `url` (a code-bearing account-access credential) on the model-visible
  surface, persisting it into the LLM transcript and any downstream logging.
  Follow the profile_token pattern: persist the URL to a 0600 private file
  (atomic temp+rename) and return an opaque `link_delivery` marker instead.
  E2e test now asserts the URL/code never appears in the result and is
  delivered out-of-band to the private file.

- Typed error (product_workflow): account_traces_for_user flattened backend
  errors into String before the WebUI boundary. Introduce AccountTracesError
  (thiserror) that names the failing operation and preserves the full cause
  chain ({:#}); the boundary keeps returning a sanitized, diagnosable 500. Also
  document that fetch_account_traces(None) is already server-bounded (default
  200, clamp 500, 256 KiB response cap) — the "unbounded fetch" concern was
  resolved by prior hardening.

- i18n (webui_v2_static): the traceStatus and traceReceivedAt keys backfilled
  for locale parity were unused by the consumer. Wire traceStatus as the status
  badge's accessible title/aria-label and render traceReceivedAt as the
  timestamp label, so all six submitted-traces keys are now consumed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(traces): address CodeRabbit re-review — async persist + typed identifiers

- Blocking I/O (host_runtime): persist_account_login_link does mkdir/write/
  fsync/rename with std::fs on the async dispatch path. Wrap the persist call in
  tokio::task::spawn_blocking so it never stalls a Tokio worker (coding guideline:
  all I/O async). Atomic temp+rename behavior is unchanged; a join failure maps to
  the same sanitized "could not write" result.
- Typed identifiers (product_workflow): account_traces_for_user took bare &str
  tenant/user; the caller already holds TenantId/UserId newtypes. Take
  &TenantId/&UserId and only cross to &str at the ironclaw_reborn_traces boundary.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(traces): instance-aware enrollment across trace_commons dispatch + UI

Instance-only-enrolled users (admin-provisioned instance policy, no personal
invite) were falsely rejected across the Trace Commons surface: the dispatch
gates and the profile mints read only the personal per-scope policy, and the
submitted-traces UI was gated behind the personal-credits branch. Addresses
CodeRabbit re-review (#3, #4, #5) on PR nearai#5280.

reborn_traces:
- Add instance-aware entry points mint_profile_attribution_token_for_user_via_sink
  and set_community_profile_for_user_via_sink that resolve enrollment via
  resolve_trace_credentials (personal OR instance) and build the claim context
  with the instance scope_dir + per-user pseudonymous subject, mirroring
  mint_account_login_link_inner. Refactor the token mint to share a
  context-based core. New tests assert the per-user subject reaches the issuer.

host_runtime (trace_commons dispatch):
- Route the enrollment gates in dispatch_status, dispatch_profile_token,
  dispatch_profile_set, and dispatch_account_login_link through
  resolve_trace_credentials so instance-only contributors pass. status now
  reports the resolved (instance or personal) policy. profile_token/profile_set
  call the new instance-aware mints.
- #4: preserve the stage_secret_material_once failure cause (log it) instead of
  discarding it with map_err(|_|); wire message stays sanitized.

webui_v2_static (#3):
- Lift the submitted-traces section out of the credits/empty-state branch so
  instance-enrolled users with no personal credits still see their traces and
  any tracesQuery errors.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(traces): isolated dispatch-layer e2e for instance-only enrollment

Extract the trace_commons dispatch e2e helpers into a shared
tests/support/trace_commons_dispatch.rs module (base-dir setup, mock issuer,
runtime/dispatch helpers, find_persisted_login_link, test_jwt_eddsa) so a second
test binary can reuse them.

Add trace_commons_instance_dispatch_e2e.rs — a SEPARATE binary (fresh process =
private IRONCLAW_BASE_DIR) that provisions the process-global instance policy
(scope None) without bleeding into the personal-invite suite. It pins the
CodeRabbit #5 fix at the layer it manifests: an instance-only-enrolled user
(no personal invite) passes dispatch_status and dispatch_account_login_link and
mints under the shared instance device key with a per-user pseudonymous subject
(asserted via the subject on the login-links POST).

No production changes; trace_commons_dispatch_e2e.rs behavior is unchanged
(5 tests still pass) — only its helpers moved to the shared module.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(traces): sanitize bearer-staging log + typed IDs on mint entry points

Addresses CodeRabbit overnight review on PR nearai#5280.

- Security (#1): the trace-bearer staging error was debug-logged via `?error`,
  which can leak secret-store/backend detail on the credential path. The
  host_runtime logging guideline forbids backend error detail here — log only
  the safe fact of failure; the wire message stays sanitized. (Supersedes the
  earlier "preserve cause" change specifically on this bearer-material path.)

- Typed identities (#2): the three agent-facing Trace Commons mint entry points
  (mint_account_login_link_via_sink, mint_profile_attribution_token_for_user_via_sink,
  set_community_profile_for_user_via_sink) now take &TenantId/&UserId instead of
  adjacent &str, so callers can't transpose tenant/user and misattribute a
  contributor. Identity stays typed to the public boundary and is stringified
  only when handing off to the dir-parameterised `_inner` cores / resolver
  (the storage edge). Adds ironclaw_host_api as a reborn_traces dependency
  (no cycle: host_api does not depend on reborn_traces). Dispatch callers pass
  the typed scope ids directly.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(traces): sanitize persist-path logs, preserve handle-validation cause

Two follow-up CodeRabbit findings on PR nearai#5280:

- Security (Major): dispatch_account_login_link's spawn_blocking persist arms
  logged %error / %join_error at debug. Filesystem errors (mkdir/write/fsync/
  rename) can carry raw host paths, which the host_runtime guideline forbids in
  logs. Drop the interpolation; log only the generic fact, keep the message
  sanitized — same treatment as the bearer-staging path.

- Maintainability (Minor): SecretHandle::new(TRACE_COMMONS_BEARER_HANDLE) used
  map_err(|_| ...), discarding the cause (non-exemptible per the guideline). The
  handle name is a compile-time constant, so its validation error carries no
  secret/path — bind and log it.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(traces): typed login-link errors, per-request bearer handle, doc accuracy

Addresses the third CodeRabbit review round on PR nearai#5280.

- Security (Major): the trace-bearer staging used a constant SecretHandle
  (TRACE_COMMONS_BEARER_HANDLE). The injection store is a HashMap keyed by
  (scope, capability, handle) with overwrite-on-insert, so two concurrent
  same-scope Trace Commons egresses could race and stage/consume the wrong
  bearer. Suffix the handle with a per-request uuid so every staged bearer key
  is distinct. Localized to the shared HostEgressContributionSink, so all
  trace_commons flows benefit.

- Correctness (Major): account_login_link_error_value classified failures by
  substring-matching upstream error wording, coupling the public error_code
  contract to phrasing. Introduce a typed AccountLoginLinkError (thiserror) in
  reborn_traces; mint_account_login_link_via_sink returns it, producing the
  specific variant at each failure site. The host maps variants -> error_code
  with no substring checks. NotEnrolled (the only tested code) is preserved;
  the two bearer-derived codes collapse into EnrollmentIncomplete (both meant
  "re-run onboarding"), and persist failures get a distinct LocalStateWrite.

- Docs (Minor): the persist_account_login_link comments promised 0600 across
  platforms though only Unix enforces it. Softened to "private local file
  (0600 on Unix)".

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(traces): type profile_token/profile_set error mappers (systemic)

Follow-up to the account_login_link typed-error change: convert the remaining
substring-based error mappers so all four trace_commons dispatch flows derive
the public error_code contract from typed variants instead of matching upstream
error wording. (onboard was already typed via OnboardError.)

- reborn_traces: add ProfileAttributionError (shared by the profile_token and
  profile_set token mints) and CommunityProfileError (profile_set wrapper adding
  InvalidProfile). mint_profile_attribution_token_for_user_via_sink and
  set_community_profile_for_user_via_sink now return these; each failure site
  produces the specific variant (NotEnrolled / PolicyRead / EnrollmentIncomplete
  / Backend / LocalStateWrite, plus InvalidProfile for profile_set).

- host_runtime: profile_token_error_value / profile_set_error_value now match on
  the typed variants — no error.contains(...) anywhere in the file. NotEnrolled
  and InvalidProfile (the tested codes) are preserved; the issuer/device/refused
  substrings collapse into EnrollmentIncomplete, consistent with the
  account_login_link mapping. Also sanitized the profile_token persist-failure
  log (host-path leak class), matching the login-link path.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(traces): split enrollment precondition from backend in token mints

CodeRabbit re-review: collapsing every mint_profile_attribution_token_with_context
(and bearer_token) failure into EnrollmentIncomplete mislabels transient
transport/status/serde failures as "re-run onboarding".

Split the local precondition (upload-claim issuer URL configured) from
post-resolution failures: a missing issuer URL maps to EnrollmentIncomplete via
an explicit upload_claim_issuer_missing() check (typed, no substring), while the
claim mint / bearer fetch / PUT failures now map to Backend. Applied
consistently across profile_token, profile_set, and account_login_link so the
error_code contract reflects the real failure class. URL-derivation
preconditions (ingest/login-links URL) stay EnrollmentIncomplete.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(traces): check login-link URL precondition before minting bearer

Fail-closed ordering: the local account_login_links_url derivation ran after
bearer_token, so a malformed/absent login-links URL would mint a device-key
bearer and hit the issuer before failing. Move that local precondition ahead of
all secret/egress work so incomplete enrollment fails closed with no side
effects. (profile_token/profile_set already order local preconditions first.)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(traces): make upload-claim cache key match issuer payload exactly

The subject cache-key component trimmed/collapsed context.subject, but the
DeviceKey issuer request sends it unchanged — so None, Some(""), and
whitespace variants could share a cache key while minting different payloads,
letting one user's claim be served from cache to another (cross-user trace
mis-attribution). This is nearai#5280's per-user-subject cache-key path.

Hash the exact optional bytes the request sends (DeviceKey → subject,
WorkloadTokenEnv → None) with a None/Some discriminator. Extend the cache-key
test with the Some("")-vs-None and whitespace-variant collision cases.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(traces): check response-size cap before growing the buffer

Both bounded response readers (upload-claim and account-traces) enforced the
hard byte ceiling only after extend_from_slice, so a single oversized chunk
could push the buffer past the advertised limit before the error returned.
Compute bytes.len() + chunk.len() (checked_add) and validate before appending.

Pre-existing pattern (from nearai#4559), fixed here per review.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(traces): fail loud when trace policy cannot be statted

read_trace_policy_for_scope_at used Path::exists(), which maps stat/permission
errors to false — silently treating an unreadable policy as missing and
default-disabled, flipping enrollment/flush behavior. Use try_exists() and
propagate the stat error with context; only a confirmed non-existent path
returns the not-enrolled default.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(traces): capture traces for instance-only enrolled users

Codex P1: capture_turn_trace gated on the per-user scope policy
(read_trace_policy_for_scope(Some(scope)) + policy.enabled), so an instance-only
enrolled user — whose per-user policy is absent/disabled — had every turn
dropped before an envelope was queued, leaving the instance-aware flush gate
nothing to submit. The headline instance-enrollment feature never captured for
exactly the users it targets.

Gate capture on the effective enrollment instead, mirroring the flush gate: add
resolve_effective_capture_policy (personal-invite policy if enabled, else the
admin-provisioned instance policy at scope None, else None) and prepare the
envelope under that governing policy. Add a resolver test covering the
personal / instance-only / neither cases.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Remove accidentally committed frontend node_modules, restore .gitignore

The merge commit f34dfa7 dropped crates/ironclaw_webui_v2_static/frontend/.gitignore
and swept 1065 node_modules files into the index. Untrack them and restore the ignore.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Address PR review feedback: egress hardening, effect declaration, instance status sync

- fetch_account_traces_direct now uses a pinned-DNS, private-IP-filtered
  HTTP client (shared pinned_trace_commons_http_client) instead of an
  unrestricted reqwest lookup, closing the DNS-rebinding window between
  claim validation and the bearer-authenticated account-traces GET.
- account_login_link capability manifest declares EffectKind::WriteFilesystem
  for the local delivery-file write.
- Queue-flush status sync (and the public sync entry point) now run off the
  resolved effective flush target (policy, device-key dir, per-user subject)
  instead of re-reading the per-scope policy, so instance-enrolled users get
  final credit status after submission; subject is threaded into the
  status-sync claim context.
- Each login-link mint persists to a unique account_login_link.<uuid>.url
  file so concurrent mints cannot clobber each other; stale link files are
  pruned best-effort after one hour.
- resolve_trace_credentials takes typed &TenantId/&UserId at the public
  boundary; call sites drop their .as_str() conversions.
- Login-link/account-traces requests honor the policy-configured issuer
  timeout; the sink-path traces fetch uses ACCOUNT_TRACES_MAX_RESPONSE_BYTES.
- Removed the AdminScope::enroll_instance_trace_commons wrapper from the v1
  monolith (crate-side entry point is onboard_instance_with_sink; noted in
  the slice1 plan).
- Tests: direct account-traces path covered for 500/404; new regression test
  pins instance-target status sync (subject + instance device-key dir).
- Plan docs: server login-link contract callout, no developer-local paths,
  resolver errors propagate, 404-only zero-state, scope_dir threading.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Pin DNS resolution on the background trace submit/status/revoke lane

The background lane (queue flush submission, status sync, revocation)
previously relied only on enrollment-time endpoint validation
(validate_trace_commons_ingest_url); the per-request client did a fresh
unrestricted DNS lookup. Replace trace_remote_http_client with
pinned_trace_remote_http_client: per-request host resolution through
resolve_trace_upload_claim_issuer_host (private/internal IPs rejected,
literal-loopback local-dev exception) pinned via resolve_to_addrs, so an
endpoint host that passed validation at enrollment cannot later rebind to
an internal address and receive bearer-authenticated requests. Timeout
behavior (env/test task-local override) is unchanged.

Regression test: pinned_trace_remote_client_rejects_private_endpoint_hosts.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Address CodeRabbit follow-up: sanitize status log, sync plan snippets

- trace commons status dispatcher no longer formats the resolver error into
  the log (it can embed the policy file's host path); logs the safe fact only,
  matching the sibling dispatchers.
- slice4 plan: AccountTraceItem snippet derives Deserialize (matches shipped
  code, which parses the response).
- slice3 plan: login-link parsing snippet fails loud on missing account_id/url
  instead of unwrap_or_default (matches shipped code).
- slice1 plan: the AdminScope wrapper task is marked SUPERSEDED up front so the
  plan no longer gives conflicting guidance.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Address round-2 review: opt-out precedence, salted subjects, UI branch tests

- Explicit per-user opt-out (scoped policy present with enabled=false, as
  written by 'traces opt-out') now blocks the instance-enrollment fallback in
  resolve_trace_credentials and resolve_effective_flush_target (and thus
  capture) — only a never-configured scope falls through to the instance
  policy. Regression test covers all three resolution surfaces.
- Instance-enrollment subjects are now salted: a per-instance random salt
  (persisted 0600 at the instance trace dir, create_new race-safe) feeds
  sha256(salt:scope), so the server or ledger holders cannot dictionary-match
  guessable tenant/user ids against an unsalted scope hash. Unsalted
  local_pseudonymous_contributor_id remains for local state keying/log refs.
- contribution.rs carries the architecture-rule file-size justification
  referencing decomposition tracking issue nearai#4088; state_scope field docs now
  say which state it does (and does not) locate.
- Submitted-traces UI: extracted the pure tracesSectionMode decision (error
  wins over list; list needs enrolled + non-empty) and covered it plus the
  row formatters in trace-commons-tab.test.mjs.
- Docs: slice4 plan points at crates/ironclaw_webui_v2 (static crate was
  folded in), slice3 signature snippet matches the typed contract, and the
  webui_v2 CLAUDE.md route table gains the three trace routes.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Update crossbeam-epoch 0.9.18 -> 0.9.20 for RUSTSEC-2026-0204

Lockfile-only patch bump of a transitive dep (via termimad/crossbeam) to
clear the new advisory failing cargo-deny; verified locally with
cargo deny check advisories.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* Route login-link/account-traces claim mint through the caller's sink

The sink-based entry points (mint_account_login_link_via_sink,
fetch_account_traces_via_sink) used the sink for the final POST/GET but
minted the upload-claim bearer via DefaultTraceUploadCredentialProvider,
whose issuer request takes the direct reqwest path — so an agent-invoked
account_login_link performed a network call outside RuntimeHttpEgress.
New trace_upload_bearer_token_via threads Option<sink> into the claim
mint (cache behavior unchanged; the default provider passes None), and
both sink paths pass Some(sink), matching the profile-token/profile-set
flows. Tests now use a RecordingSink to pin the invariant that both the
claim mint and the follow-up request route through the sink.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: regular Contributor history classification risk: medium Risk classification scope: ci size: XS Changed-line size classification

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant