[codex] restrict Docker publishing to upstream - #8
Merged
Merged
Conversation
theredspoon
marked this pull request as ready for review
June 17, 2026 23:18
theredspoon
pushed a commit
that referenced
this pull request
Jun 18, 2026
Consolidated fixes for serrrfirat's 10 unresolved review threads plus zmanian's CHANGES_REQUESTED review (3 blockers + 7 mediums + 3 follow-up items + title typo). ## Blockers 1. CI now runs the package tests (zmanian blocker 2 / serrrfirat #7). The matrix `cargo test ${{ matrix.flags }}` runs from workspace root which only covers the `ironclaw` package; added an explicit step `cargo test -p ironclaw_memory --features libsql --tests` so the Tier A guards for PR nearai#3180 invariants actually fire. 2. `#[ignore]` markers converted to `#[cfg_attr(not(feature = "pr3180-ready"), ignore = ...)]` (zmanian blocker 1 / serrrfirat #1). Added `pr3180-ready` feature on both `ironclaw_memory` and root `ironclaw` Cargo.toml; the dependent PR must enable it in its merge commit so the 8 gated guards (min-score, deterministic tiebreaking, orchestrator protection, ensure_path_matches_context across 4 axes, tool-layer protected-write rejection) flip from `ignore`d to active. 3. Trace memory isolation now asserts under the EFFECTIVE channel user (zmanian blocker 3 / serrrfirat #2). Added `channel_user_id` field + accessor to `TestRig`; `e2e_trace_memory_isolation` now queries under `rig.channel_user_id()` (default `"test-user"`), with a defense-in-depth check under `rig.owner_id()` for mis-routing regressions. ## Test-correctness mediums 4. Min-score test pins `with_query_embedding([1,0,0])` to favor hybrid.md (serrrfirat #3 / zmanian #4). Removed the permissive `* 0.99` fallback — FTS-only-below-hybrid is now a hard assertion. 5. Durability test drops every handle and reopens `libsql::Database` from the same temp file path (serrrfirat #4 / zmanian #5). Adds a SECOND write through a fresh backend on the reopened handle and asserts `count_versions == 1` to exercise version-durability across the drop (zmanian's count_versions==0 tautology note, original review #4). 6. Append versioning asserts exact row count `== 1`, not `!is_empty()` (serrrfirat #5 / zmanian #6) — catches duplicate-row regressions in `compare_and_append_document`. 7. Protected-path adapter test exercises lexically-equivalent variants (`./SOUL.md`, `//SOUL.md`) in addition to canonical (serrrfirat #6 / zmanian #7). VirtualPath rejects `..` so no `..` variant; in-loop `count_documents_total == 0` after EACH variant. 8. Hybrid search isolation now varies all four scope axes (serrrfirat #8 / zmanian #8): tenant, user, agent, project. 5 documents seeded; search from caller scope must return exactly one. 9. Tool round-trip asserts EXACT persisted content via direct DB read (serrrfirat #9 / zmanian #9). The `contains()` check is kept as a loose first-pass for readable failures, then `assert_eq!` on the exact byte string is the load-bearing assertion. 10. Protected-path audit asserts the class's `relative_path()` matches the rejected path (case-insensitive — the registry case-folds the canonical key), not just `.is_some()` (serrrfirat #10 / zmanian #10). A regression that emits the wrong path class now fails. ## zmanian follow-ups Z1. Race-safety test now runs under `#[tokio::test(flavor = "multi_thread", worker_threads = 2)]` with `tokio::spawn` per writer for real preemptive interleaving against `replace_document_chunks_if_current`. Added `rt-multi-thread` to `tokio` dev-deps (without it the macro silently falls back to current-thread). Z2. `write_to_protected_path_rejected.json` trace fixture sets `all_tools_succeeded: false` explicitly. Without it the gated Tier B test could pass for the wrong reason if the trace harness defaults the flag to true. Z3. Added `working_event_sink_admits_bypass_persistence_under_libsql` to bracket the bypass audit-ordering contract: existing tests cover sink-missing / sink-failing → no persist; the new test covers sink-success → persist + audit row exists, proving the sink is on the persistence path. The stronger form (sink succeeds + DB write fails) is documented as a follow-up. ## Cleanup - Removed `_link_in_memory_repo_for_unused_imports` shim and the `InMemoryMemoryDocumentRepository` import that only existed to feed it (zmanian original-review #3). - Fixed PR title typo `momery` → `memory` via gh. Helper-consolidation into `tests/common/libsql_helpers.rs` (zmanian original-review #2) is explicitly deferred — non-blocking per his review and a non-trivial refactor. ## Verified - `cargo fmt --all -- --check` clean - `cargo clippy -p ironclaw_memory --features libsql --all-targets -- -D warnings` zero warnings - `cargo test -p ironclaw_memory --features libsql` all suites green (gated tests stay `ignored` without `--features pr3180-ready`)
theredspoon
pushed a commit
that referenced
this pull request
Jun 18, 2026
…view Four high-level architectural decisions from human review (serrrfirat, zmanian) on PR nearai#3544 land as spec amendments before WS-0 seals the trait shapes: 1. LoopFamily as first-class abstraction (zmanian #2.2). New WS-3.5 brief: LoopFamilyId + ComponentIdentity + LoopFamily + LoopFamilyRegistry (Guice-style singleton built once at startup). 2. Strategies sealed pub(crate) (zmanian #1.2, #2.5, #2.6; serrrfirat #1). AgentLoopPlanner is pub but uses sealed-trait pattern; strategy access lives on pub(crate) AgentLoopPlannerInternal extension trait. 3. Hooks as middleware (zmanian #1.1). Adopt the four-scenario design in PR nearai#3523-comment-4435808547. Two concrete follow-ups: composition seam architecture test in WS-9's brief; LoopFailureKind::PolicyDenied variant in WS-0. 4. Broad scope retained with stress-test discipline (zmanian #2.1, #2.7; serrrfirat #8). New §12.5 enumerates anticipated families (default + hypothetical routine/mission/coding/planning) and the strategies each would swap; trait shapes accommodate every row. Side effects: - Split ControlStrategyState → StopStrategyState + GateStrategyState in WS-0 so stop and gate strategies grow independently. - Drop generics on PlannedDriver — non-generic { family, executor }. WS-7 + downstream briefs (WS-9, WS-10, WS-11, WS-13, WS-14) updated. - Subsume PlannerId into LoopFamilyId + ComponentIdentity (one versioning primitive per zmanian #1.4, partly addressed). - Executor entry point becomes execute_family(&LoopFamily, host, state). - Document HostXxxContextSource as the pattern for family-specific durable context (mirrors WS-15's HostIdentityContextSource). 15 files changed (1 new brief: loop-family-registry.md). Spec-only; no code changes. Three remaining comment clusters (checkpoint durability + LoopExit witness; versioning + JSON canonicalization; serrrfirat's remaining 7 correctness seams) are noted as follow-up work. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
theredspoon
pushed a commit
that referenced
this pull request
Jun 18, 2026
…3920) * Implement installed WASM hook runtime Adds crates/ironclaw_hooks/docs/threat-model-wasm.md and follows the reviewed design ack: 1) module bytes are resolved, digest-cached, and compiled in the tool-WASM style while reusing its resource limiter; 2) each invocation gets a fresh wasmtime Store; 3) the ABI is a wasmtime::Linker surface, not wit-bindgen; 4) host-import sink shims enforce call, patch-byte, observer-fact, and decision budgets. * Harden WASM hook string and metadata budgets * fix(hooks): validate WASM hook ABI at install time (serrrfirat #3 on PR nearai#3634) Address serrrfirat MEDIUM finding #3: `WasmHookRuntime::prepare()` compiled and cached module bytes but did not validate imports or the requested export. ABI mismatches (unsupported import, missing export, wrong export signature) were deferred to first live dispatch — and the prior `wasm_unsupported_host_import_fails_closed` test codified that a bad-import module would install successfully and only fail closed at invocation. Malformed untrusted modules should never reach live traffic. Changes: - `prepare()` derives the target hook point from `request.kind`, then runs `validate_module_abi()`: scratch-instantiate the module against the point-specific linker (catches unsupported / wrong-type imports) and resolve the typed export `() -> ()` (catches missing export and wrong signature). Failures surface as new `WasmHookRuntimeError::InvalidImports` or existing `WasmHookRuntimeError::InvalidExport`, both of which bubble up as `HookError::RegistryConstruction` from the registrar. - `wasm_point_for_kind(HookManifestKind)` helper centralizes the kind → wasm-point mapping; the previous `execute_*` paths can share it in a follow-up but kept inline for now to minimize churn. Tests: - `wasm_unsupported_host_import_is_rejected_at_install_time`: replaces the prior test that codified late-failure behavior; asserts the registrar returns `RegistryConstruction` citing the bad import. - `wasm_missing_export_is_rejected_at_install_time`: new module that compiles but lacks the manifest-declared export; same install-time rejection. * fix(hooks): address henrypark133 must-fix #1, #2, #3 on PR nearai#3634 Three items from the 5-15 review: **#1 (must-fix) Extract ironclaw_wasm_limiter micro-crate** Replace `#[path = "../../../ironclaw_wasm/src/limiter.rs"]` cross-crate file import with a proper Cargo edge. The 111-line `WasmResourceLimiter` moves into a new `crates/ironclaw_wasm_limiter` micro-crate that both `ironclaw_wasm` and `ironclaw_hooks` depend on. The architecture rule forbidding `ironclaw_hooks -> ironclaw_wasm` is preserved (the new crate sits below both consumers and pulls in only `wasmtime` + `tracing`); `cargo check`, `cargo doc`, and architecture-linting tests now see the edge, and the file can't be moved out from under one of the consumers silently. Mechanical changes: - new `crates/ironclaw_wasm_limiter/` (Cargo.toml + src/lib.rs with the type exposed as `pub` instead of `pub(crate)`) - workspace `members` entry added - `crates/ironclaw_wasm/src/limiter.rs` deleted - `crates/ironclaw_wasm/src/lib.rs`: `mod limiter` removed - `crates/ironclaw_wasm/src/store.rs`: import switched to `ironclaw_wasm_limiter::WasmResourceLimiter` - `crates/ironclaw_wasm/Cargo.toml`: dep added - `crates/ironclaw_hooks/Cargo.toml`: dep added - `crates/ironclaw_hooks/src/wasm/runtime.rs`: `#[path = ...]` block removed; import switched to the crate **#2 + #3 (must-fix) Dead WASM arms in dispatch** `run_before_capability_hook`, `run_before_prompt_hook`, and `run_observer_hook` each had an early-return guard that dispatched WASM hooks with `catch_unwind` + timeout, then ALSO had a matching WASM arm in the inner `match` that ran without those protections. The prompt-path arm additionally swallowed `WasmHookFailure` via `|_| ()`, making the must-fix #2 problem worse on that path specifically. If a future refactor removed any of the early-return guards, those inner arms would silently take over and drop panic isolation, deadline enforcement, AND (for prompts) the failure category. Replaced each inner arm with `unreachable!()` carrying a comment that explains why the arm exists and references the early-return guard above it. A future refactor that removes the guard will now trip the `unreachable!` at first call instead of silently degrading. All 154 hooks lib + 29 reborn integration tests still pass. * fix(hooks): plumb context to WASM hooks + runtime hardening Critical #1 on PR nearai#3634: WASM hooks previously received no context. The `execute_*` entry points dropped the `&BeforeCapabilityHookContext` / `&BeforePromptHookContext` / `&ObserverHookContext` value and invoked the guest export with `()`, so a WASM gate could never decide based on the capability name, tenant, provider, or other dispatch-time facts. Add an `ic:hooks/context@1` host-import module exposing two read-only calls — `ctx_size() -> i32` and `ctx_read(ptr, len) -> i32` — backed by a JSON-serialized blob the dispatcher writes per-invocation into the fresh store. Modules that don't import these continue to link; modules that do import them get a stable, non-empty payload to read. An integration test (`wasm_before_capability_hook_reads_context_blob`) asserts the contract end-to-end: a guest that fails to read a non-empty blob traps before its `deny` call. Also rolls up the other reviewer-flagged WASM runtime issues, all of which touch `wasm/runtime.rs`: HIGH #2: epoch-tick background thread now holds a shutdown `AtomicBool` and joins on `Drop`. Previously it looped forever and leaked an Engine clone on every runtime drop. MED #4: compiled-module cache is now an `lru::LruCache` bounded by `MODULE_CACHE_CAPACITY = 128`. Replaces the unbounded `HashMap`. MED #7: `prepare()` no longer compiles under the cache lock. Fast path reads from LRU under a brief lock; slow path compiles outside the lock and re-checks on insert to avoid the TOCTOU window where two concurrent installs of the same module both compile. Bug #9: post-call `deadline_exceeded()` re-check on the Ok branch is gone. wasmtime epoch-interrupt is the authoritative wall-clock signal; an Ok return is no longer reclassified as a timeout because the wall ticked over during host-side return. Bug #10: `add_milestone_metadata` returns a distinct "metadata value exceeds the u32 byte-length ceiling" error when the guest-supplied `value.len()` overflows u32, instead of misreporting it as "exceeded total prompt-patch byte budget". Existing integration tests for WASM hooks are also re-wired through `HookRegistrar::with_verified_grants` so the grants-store gate added in the foundation-01 merge stops failing the pre-existing fixtures. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(hooks): run WASM hooks on the blocking pool HIGH #3 on PR nearai#3634: `tokio::time::timeout` does NOT cancel synchronous wasmtime execution. The previous code awaited a `catch_unwind(async { h.evaluate(ctx) })` future whose body completed in one poll, so the timeout could only fire *around* the WASM call rather than against it; a hook that wedged inside wasmtime simply pinned the calling tokio task. Route gate, prompt, and observer WASM dispatch paths through `tokio::task::spawn_blocking` via a shared `run_wasm_blocking` helper. The outer `tokio::time::timeout` now governs the JoinHandle, so a stuck blocking task stops blocking the dispatcher's caller; the wasmtime epoch interrupt configured in the runtime (10 ms tick) is the authoritative in-WASM wall-clock cancel signal. JoinError (panic in the blocking task) maps to `FailureCategory::Panic`, matching the pre-existing semantics for synchronous panics caught via `catch_unwind`. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * perf(hooks): O(1) hook-id lookup via side index Finding #8 on PR nearai#3634: `set_priority`, `poison`, `is_poisoned`, and `contains_hook` all did full-registry scans over every binding at every point. Each is called per-dispatch (poison-checks on the snapshot loop in particular), so the cost is `O(registered_hooks)` per `(installed_hook, registered_hook)` pair. Maintain a denormalized `HashMap<HookId, (HookPointSpec, usize)>` side index in lock-step with `by_point` so every per-hook-id operation becomes a single hash lookup + a direct vec indexed access. The duplicate-id rejection in `insert` now reads from the side index too, turning what used to be a flat-map scan into a `HashMap::contains_key`. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(hooks): wall-clock timeout, observer memory, limiter rollback, registrar happy path Round out the test set for the WASM hook execution path: #11 / #12: gate + observer wall-clock timeout. The pre-fix dispatcher ran wasmtime synchronously on the executor, so the outer `tokio::time::timeout` `Err(_elapsed)` arm was effectively unreachable. Now that WASM execution runs on the blocking pool, the timeout actually fires; the new tests give the wasm budget headroom (1B fuel, 5s wall) and the dispatcher a 20 ms timeout, then assert the failure classification (FailClosed for gate, FailIsolated for observer). #13: observer memory exhaustion. Mirrors `wasm_memory_exhaustion_fails_closed_for_gate` against the observer dispatch path so the FailIsolated branch of the failure matrix has explicit memory coverage, not just fuel/wall. #15: `WasmResourceLimiter::memory_grow_failed` rollback. Stages an approved grow, simulates the OS-level grow failing, and asserts a subsequent grow of the full ceiling succeeds — the inflated `memory_used` from the failed attempt must be released. #16: registrar WASM happy path. Companion to the existing `install_wasm_body_requires_runtime` negative case: a valid module installs, the binding is visible via the public registry accessor, and is not pre-poisoned. #14 (`add_milestone_metadata` happy path) is intentionally omitted — the BeforePrompt dispatch path is currently unreachable due to a pre-existing manifest-vs-registry scope conflict (`OwnCapabilities` is the only valid `BeforePrompt` scope per manifest validation, but the registry rejects `OwnCapabilities` at `BeforePrompt` because the point has no provider context). That contradiction sits outside this PR's scope; flagging for a follow-up. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(hooks): typed WASM version material, reconcile design doc LOW #20 on PR nearai#3634: extract the `{extension_version}+wasm:{module_digest_hex}` concatenation into a `WasmVersionMaterial` newtype with a single `Display` impl. The identity material no longer floats free as a stringly-typed argument inside the registrar. Reconcile `docs/successors/02-wasm-runtime.md` with the implementation: - Spell out that wall-clock cancellation depends on the `tokio::time::timeout(tokio::task::spawn_blocking(...))` pair, and explain why a bare timeout over a synchronous wasmtime call cannot actually cancel. - Define `FailIsolated` and `FailClosed` as `FailureDisposition` values, distinct from the older `HookFailureMode::{FailOpen, FailClosed}` policy switch that applies to predicates. - Clarify the generic `evaluate` export contract — name is whatever the manifest declares, signature is `(): ()`, context arrives through the new `ic:hooks/context@1` host imports — and note the intentional divergence from `WitToolRuntime`'s hardcoded interface. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(hooks): drop .expect() in WASM module cache capacity Pre-commit no-panics CI flagged the .expect() on the LruCache capacity. Move the validity check to a const match, so the NonZeroUsize is fixed at compile time and the no-panics regex is satisfied. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(hooks): use HookLocalId::new after newtype privatization The newtype-privatization landed in reborn-integration after the hooks-fu-wasm-runtime branch's WASM scaffolding tests were written; update the affected test/registrar sites to use HookLocalId::new instead of the now-private tuple constructor. * style: cargo fmt after newtype-privatization fixups * test(hooks): ignore 3 BeforePrompt WASM tests with manifest/registry conflict These tests were failing on the original branch tip too (verified against origin/hooks-fu-wasm-runtime @ 571efdf). The Installed-tier BeforePrompt WASM install path has no valid scope today: - OwnCapabilities is rejected by the registry C3 check (finding #2 on PR nearai#3573) since BeforePrompt has no per-capability invocation context. - SameTenant is rejected by manifest validation ("cannot combine scope = same_tenant with kind = before_prompt"). The budget-overflow paths these tests exercise are point-agnostic; the follow-up is to either rewrite the helper to install through BeforeCapability or add a Global manifest scope. Tracked as a deferred item on the new PR. --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
theredspoon
pushed a commit
that referenced
this pull request
Jun 18, 2026
…earai#4559) * docs: trace commons agent onboarding design spec Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: address spec review findings (trust anchoring, key staging, consumption atomicity, replay validation) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: spec review round 2 nits (server-anchored tenant wording, pending-key cleanup) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: implementation plan for trace commons agent onboarding Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: address plan review findings (scope threading refactor, dispatch model, dev-deps, LazyLock hazard) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: plan review round 2 fixes (literal dep versions, context constructor threading depth) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: incorporate server-agent coordination feedback (optional community/profile/leaderboard URLs) From TraceCommons/trace-commons#136-#141 comments: onboard response gains optional browser-surface navigation hints, sanitized client-side (HTTPS or dropped), never part of issuer trust anchoring. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(traces): onboarding wire types matching trace-commons-server contract Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(traces): invite URL parsing with origin trust anchoring Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(traces): device keypair lifecycle with pending staging and self-signed workload JWTs Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(traces): auth_mode and device_key_id policy fields with legacy-compatible defaults Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(traces): onboard() orchestration with trust anchoring and retry-safe key staging Wire invite parsing, device key staging, onboard POST, issuer origin trust anchoring, ingest_url HTTPS enforcement, keypair promotion, and policy write into onboard_at_dir(). Refactors invite.rs to extract pub(crate) is_https_or_loopback, origin_of, and host_only helpers shared with mod.rs (one source of truth for origin/bracket handling). Adds axum mock-issuer tests covering the happy path, mismatch rejection, terminal vs transient error key retention, insecure ingest URL, loopback ingest allowance, community URL sanitisation, and retry key reuse. Partial-failure lockout fix (spec §2.2): promote() no longer deletes the pending file. The flow now writes the tenant key file, then the policy, and only discards the pending file after BOTH durably succeed. If the policy write fails the pending key survives, so a retry reloads the same key (server idempotency returns the original registration) and harmlessly overwrites the tenant file — no permanent lockout from a consumed invite with a regenerated keypair. Regression test simulates a policy-write failure (policy.json pre-created as a non-empty dir so the atomic rename fails), asserts Err(Persist) with the pending key intact, then asserts a retry succeeds reusing the same device_key_id. Response validation (defense-in-depth): reject schema_version != the v1 response constant as MalformedResponse, and cross-check the response device_key_id against the locally derived id (we never trust the response value for policy; a disagreement is now treated as a tamper signal and rejected). Both covered by tests. The onboard response body is read with the 64 KB cap enforced per-chunk during streaming (mirroring read_bounded_trace_upload_claim_response) rather than buffering the whole body first, so a hostile server cannot force a large allocation. Also fixes a pre-existing test-isolation defect surfaced by the added load: the remote-request timeout test configured a 50ms timeout via the process-global IRONCLAW_TRACE_REMOTE_REQUEST_TIMEOUT_MS env var. set_var is process-global, so under parallel execution the 50ms value leaked into other tests' trace HTTP clients, producing spurious `operation timed out` failures against fast local mocks. Replace the env mutation with a task-scoped TEST_REMOTE_REQUEST_TIMEOUT_OVERRIDE task-local (visible only within the awaiting test's own task tree, zero production change; documents the spawn caveat), and decouple the timing assertion from a tight wall-clock race so it no longer flakes when reqwest's timer is delayed under an oversubscribed runtime. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(traces): device-key self-signed workload JWT branch in upload-claim refresh Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(engine): trace_commons onboard and status first-party tools with agent guidance Add two model-visible first-party capabilities to the Reborn engine: - builtin.trace_commons.onboard: drives operator-invite enrollment flow with explicit per-conversation consent gate (confirmed=true required before any network call); maps OnboardOutcome/OnboardError to clean agent-readable JSON - builtin.trace_commons.status: read-only enrollment state inspector Wires ironclaw_reborn_traces into ironclaw_host_runtime, creates schema files (schemas/builtin/trace-commons-{onboard,status}.{input,output}.v1.json) and prompt doc files (prompts/builtin/trace-commons-{onboard,status}.md) at the manifest-derived paths. Includes 11 unit tests covering input parsing, consent refusal, success/error value formatting, and status formatting. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: add Task 11 — credits visibility (console display + agent-queryable balance) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(engine): e2e trace commons onboarding through capability dispatch Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(traces): document agent onboarding flow in trace-commons internal doc Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: correct Task 11 console scope (credit endpoint already exists; frontend = coordinate with designer) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(traces): trace_commons.credits agent-queryable balance tool Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(gateway): minimal Trace Commons credits card in settings Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(traces): store upload-claim endpoint in policy; preserve primary onboard error; block metadata/link-local/multicast issuers Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(traces): route agent onboarding HTTP through host network-egress policy (nearai#4560) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * build: update Cargo.lock for trace-commons onboarding dev-deps Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(traces): drop orphaned schema/prompt files (main resolves builtin schemas inline; prompt_doc_ref dropped) Post-merge cleanup: main's first_party_tools now resolves builtin input schemas via the inline schemas.rs match (trace_commons arms added during the merge) and sets prompt_doc_ref: None for all builtins, so the physical trace-commons-*.json schema files and trace-commons-*.md prompt docs are no longer referenced. The onboard consent contract remains in the capability description and is enforced in dispatch_onboard. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Fix Trace Commons invite hash contract * fix(traces): grant trace_commons capabilities in local-dev policy The three builtin.trace_commons.* capabilities were declared in the first-party package but had no [[grants]] entries in local_dev_capability_policy.toml, so local-dev runs (repl/serve) filtered them out of the model-visible tool surface entirely. The provider-level authority_effects ceiling had external_write, but the per-capability grants were never added. onboard gets the local_dev_wildcard egress profile (invite origins are operator-chosen; private/metadata IP ranges stay blocked by the shared enforcer). status/credits are read-only. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(traces): add Reborn e2e coverage for trace_commons first-party tools Closes the coverage gate failure: builtin.trace_commons.{onboard,status, credits} were declared in the first-party package but missing from REBORN_FIRST_PARTY_E2E_COVERED_CAPABILITIES, failing reborn_builtin_first_party_capability_e2e_coverage_is_complete on both the Reborn root tests and all-features CI jobs. Adds a trace_commons host-runtime harness (network policy populated so the onboard Network-effect obligation passes) and a parity test driving all three capabilities through the scripted model loop: onboard with confirmed=false exercises the deterministic consent gate with no network, status and credits return the unenrolled/zero-credit defaults. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(traces): community profile second opt-in (token mint + profile set) After device-key enrollment, public leaderboard attribution is a second, separate opt-in: IronClaw mints a short-lived profile token from the claim issuer with consent_scopes=[public_attribution] and empty allowed_uses (such a claim cannot submit traces), then either prints it for the web profile page or performs the profile update itself. The browser cannot sign device-key requests, so the token must be minted by IronClaw — previously this step was impossible and agent guidance invented flows. - ConsentScope::PublicAttribution mirrors the server protocol enum; default_allowed_uses_for_scope returns empty for it. - mint_profile_attribution_token_for_scope / set_community_profile_for_scope / withdraw_community_profile_for_scope reuse the hardened issuer HTTP path (allowlist validation, pinned DNS, no redirects, bounded reads, token never in errors). PUT/DELETE /v1/community/profile per the server contract; handle (3-32 ASCII alnum/-/_) and bio (<=280 bytes) validated client-side. - CLI: ironclaw-reborn traces profile token|set|withdraw. - Onboard tool next_steps now describes the profile second opt-in so agent guidance stops inventing browser login flows. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(traces): autonomous turn-end trace capture in the Reborn runtime The Reborn binary could onboard, report status/credits, and manage profiles, but never captured or submitted traces — the autonomous pipeline existed only in the v1 agent loop. This wires it into the Reborn runtime composition: - TraceCaptureTurnEventSink subscribes best-effort to the turn lifecycle bus (the existing turn_event_sink injection seam). On Completed/Failed events with an explicit owner it spawns a detached task that reads the owner's standing policy (one file read for non-enrolled users), loads the recent thread history (last 24 messages, 5 turns — v1 parity), adapts user/assistant text rows into the neutral ConversationMessage shape, redacts + scores locally, and queues + immediately flushes eligible envelopes. All failures are debug!-logged and never touch the turn lifecycle path. - A periodic flush worker (300s, 25/scope — v1 parity) retries queued envelopes for the runtime owner plus every scope observed since boot, with CancellationToken shutdown alongside the other workers. - TraceClientAutonomousCaptureRequest gains outcome_override so the lifecycle event's terminal status (authoritative in Reborn, where transcripts carry no structured outcome payload) marks failed turns as TaskSuccess::Failure; v1 passes None (no behavior change). - Tool-result rows and credit-notice delivery are documented follow-ups (refs-only records; no composition-level outbound channel surface). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(traces): end-to-end auto-capture through send_user_message Proves the full Reborn auto-submission chain with a real runtime: a completed turn for an enrolled owner scope lands a redacted envelope in that scope's submission queue with no manual trace command — turn completion -> lifecycle bus -> capture sink -> thread-history read -> redact/score -> eligibility -> queue (+ local-failing immediate flush leaves the entry for the retry worker). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Expose Trace Commons profile token tool * Expose Trace Commons profile set tool * Allow Trace Commons profile setup from agent * feat(webui-v2): Trace Commons credits card in WebChat v2 settings Adds GET /api/webchat/v2/traces/credit and a read-only Trace Commons settings tab to the v2 SPA, giving webui-v2-beta parity with the v1 console's credits card. - Route follows the descriptor system end to end: bearer-auth required, NoBody, 120/60 per-caller read rate limit; descriptor-driven body/rate-limit enforcement applies automatically. - RebornServicesApi::trace_credits derives the trace scope exclusively from the authenticated caller's user id (never from query/body) and reads contributor-local state via ironclaw_reborn_traces (policy + trace_credit_report), soft-falling back to an unenrolled zero-state on missing/unreadable local state, mirroring builtin.trace_commons.credits. - SPA: Trace Commons subtab (enrollment, pending/final credit, delayed ledger delta, submission counts, last submission/sync, recent credit explanations) with the server-authoritative framing and a not-enrolled empty state pointing at agent onboarding. - Tests: descriptor contract row, handler oneshot, and three composed- router serve tests (200 zero-state, 401 without bearer, enrolled policy reporting with per-test scope isolation). - Drive-by: cfg-gate openai_user_id in webui_serve.rs to clear a pre-existing unused-variable warning under default features. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Exempt Trace Commons profile setup from local-dev gate * Route Trace Commons profile writes to ingest * review(4559): address serrrfirat feedback - Drop stray working-note markdown files from the repo root (they rode in via an early origin/main merge and are not this PR's documentation). - trace_commons_dispatch_e2e: setup_base_dir is now a OnceLock that every test calls first — the previous 'single-threaded during init' claim was wrong under tokio's multi-threaded test runtime, and two of three tests skipped the setup entirely. - settings.js: extract shared appendDisplayGroup + declarative row defs; loadTraceCommonsCredits drops from ~120 lines of manual DOM to a rows array; also removes a double-escape (textContent + escapeHtml) on explanation lines. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(webui-v2): add traceCommons i18n keys to all locales The credits card added the traceCommons.* key set to en.js only; the i18n consistency test (all_locales_share_the_en_key_set) requires every locale to carry the same key set. Adds translated entries to ar, de, es, fr, hi, ja, ko, pt-BR, uk, and zh-CN. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(reborn): fail loud with source context on malformed local-dev master key The local-dev secret store resolver read the cached key file (and the SECRETS_MASTER_KEY env fallback) and passed the material straight into SecretsCrypto::new several layers deep. A corrupt or low-entropy key (e.g. a 64-char all-zeros value, which passes the length floor but has one distinct byte) surfaced only as the opaque "Invalid master key", with no pointer to the file the operator must fix. - Add ironclaw_secrets::validate_master_key_material as the single source of truth for master-key rules; SecretsCrypto::new delegates to it. - resolve_local_dev_secret_master_key now validates at the source (cached file vs SECRETS_MASTER_KEY env) and returns a RebornBuildError::InvalidConfig naming the offending path/env var and the actual constraint, before any crypto is constructed. - A malformed env value is now rejected before being persisted to the cached key file (no more poisoned-cache state). Tests: malformed-file path-context rejection, malformed-env source-context rejection, valid cached file accepted. Refs nearai#4741 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(webui-v2): add Trace Commons credits card to chat sidebar Surface trace contribution credits at a glance in the chat sidebar, above the conversation list. Previously credits were only visible under Settings -> Trace Commons. - New SidebarTraceCredits component reuses the existing useTraceCredits hook (/api/webchat/v2/traces/credit) — no new endpoint. Renders only when enrolled; loading/error/not-enrolled render nothing to keep the sidebar clean. Shows final credit and accepted/submitted counts and clicks through to Settings -> Trace Commons for the full ledger. - useTraceCredits now refetches (60s interval + on window focus) so the card and the Settings tab reflect newly-accepted submissions live. - Add one compact i18n key (traceCommons.cardAccepted) across all 11 locales; reuse existing keys for the rest. - Source-shape regression test in assets.rs. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(reborn-traces): reconstruct tool calls in turn-end trace capture The Reborn capture adapter dropped every tool-result row, so captured trace envelopes were text-only. That left the two highest-value scoring levers — replayability (0.20) and tool coverage (0.15) — permanently at zero, so even agentic tool-using turns scored as plain chat and stayed below the 0.35 submission gate. Nothing ever submitted. conversation_messages_from_records now reconstructs a `tool_calls` message from each run of ToolResultReference rows that carry `tool_result_provider_call` replay metadata, collapsing consecutive rows into one message positioned between the user message and the assistant response (the shape capture_turns_from_conversation_messages' per-turn lookahead consumes). Tool names always flow through so the value scorecard sees required_tools/replayable; raw tool payloads stay consent-gated downstream by include_tool_payloads. Rows without provider metadata remain dropped. TDD: - adapter unit tests: single tool call -> tool_calls message; consecutive calls collapse into one; ref without provider metadata still dropped. - integration guard: a captured tool-using turn's queued envelope carries replay.required_tools + replayable=true. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reborn-traces): read capture history from context window, not display projection Tool-call reconstruction (previous commit) had no data to work with: the capture history source read SessionThreadService::list_thread_history, whose product-display projection (history_message) hard-nulls tool_result_provider_call. So even though tool calls persist with full provider metadata, the adapter received None on every tool row, dropped them, and produced a text-only envelope that scored below the 0.35 submission gate. Nothing ever submitted. SessionThreadHistorySource now reads load_context_window (the model-context/replay view, which preserves tool_result_provider_call) and maps ContextMessage -> ThreadMessageRecord via context_window_to_records. This is the semantically correct source for trace capture anyway: the replay transcript, not the display transcript. TDD: a caller-level test (per .claude/rules/testing.md "test through the caller") drives SessionThreadHistorySource against a real InMemorySessionThreadService with an appended tool result, asserting the returned tool row keeps provider_call. Failed on list_thread_history (None), passes on load_context_window. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(reborn-traces): auto-submit traces with PII risk below High Previously any non-Low residual PII risk was blocked from auto-submission two ways: the manual-approval eligibility gate held everything != Low, and the value scorecard halved the score (privacy_gate Medium 0.5) and subtracted a 0.60-weighted penalty. A minimal tool trace scores ~0.36 at Low (barely over the 0.35 gate), so any Medium penalty collapsed it to 0 — nothing below High could ever submit. Treat below-High residual risk as clean for auto-submission (the deterministic redactor has already scrubbed detected PII): - trace_autonomous_eligibility manual-approval gate now holds only High (== High, was != Low). - privacy_gate: Low|Medium => 1.0 (was Medium 0.5); High => 0.0. - privacy_risk_score: Low|Medium => 0.0 (was Medium 0.5); High => 1.0. High remains fully blocked: privacy_gate zeros its score and the gate holds it for manual review. The 0.35 submission gate leaves no headroom for a partial Medium discount on a minimal trace, so below-High is clean rather than partially penalized. TDD: medium_pii_tool_trace_auto_submits_while_high_is_held asserts a Medium-risk tool trace clears 0.35 and auto-submits while High is held. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(reborn-traces): design for Trace Commons held-trace review Held traces are currently dropped on the autonomous capture path with no visibility or authorize path. This plan reuses the existing hold-sidecar machinery (TraceQueueHold / .held.json / read_trace_queue_holds_for_scope / ManualReview) and adds: retain held traces, surface a held count+list on the /traces/credit response, a card/tab UI, and a promote-as-is authorize endpoint. Four independently-shippable TDD slices. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(reborn-traces): retain manual-review held traces instead of dropping (slice 1) Autonomous turn-end capture dropped every held trace (logged at debug, envelope discarded), so PII-gated traces were unrecoverable and invisible. Slice 1 of the held-review feature retains manual-review holds: - TraceQueueEligibility::Hold now carries a typed TraceQueueHoldKind (ManualReview for the High residual-PII gate; PolicyGate for score / tool-allowlist / submission-class gates), replacing reason-string classification at the flush call site. - TraceClientAutonomousCaptureOutcome::Held carries the built envelope and its kind so callers can persist it. - New queue_trace_envelope_as_held_for_scope: queues the envelope plus a ManualReview .held.json sidecar under one scope lock; the flush worker already skips held sidecars, so it is retained but not submitted. - capture_turn_trace retains ManualReview holds and still drops PolicyGate holds (low-value traces never pollute the review surface). TDD: held-retain function (RED on missing sidecar -> GREEN), eligibility kind classification, and caller-level capture tests (an AWS-key message forces High PII -> retained ManualReview hold; a sub-threshold trace is dropped, not retained). Refs docs/plans/2026-06-10-trace-commons-held-review.md Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(webui-v2): surface manual-review held count + list on /traces/credit (slice 2) Held traces retained by slice 1 were invisible to the UI. Slice 2 surfaces them on the existing trace-credits response so one fetch powers the whole card/tab. - ironclaw_reborn_traces: manual_review_holds_for_scope() returns only ManualReview holds (excludes PolicyGate value-gates and transient RetryableSubmissionFailure retry holds), via an extracted retain_manual_review_holds filter. - RebornTraceCreditsResponse gains manual_review_hold_count + holds[] ({ submission_id, reason }). Sanitized: submission id and the already privacy-safe hold reason only, never raw trace content. TDD: retain_manual_review_holds filter unit test (excludes policy/retry), disk-level manual_review_holds_for_scope test, and the facade zero-state test asserts the new fields default empty. webui_v2 handler contract tests (42) still pass with the propagated fields. Refs docs/plans/2026-06-10-trace-commons-held-review.md Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(webui-v2): show held-for-review traces on card + Settings tab (slice 3) Surface the manual-review held count/list from slice 2 in the UI. Both render only when there are holds, so the common (nothing-held) state is unchanged. - Sidebar card: "{count} held for review" line when manual_review_hold_count > 0. - Settings -> Trace Commons tab: a "Held for review" section listing each held trace's sanitized reason + submission id from holds[]. - No hook/api change: fetchTraceCredits already returns the raw response, so credits.holds / credits.manual_review_hold_count are available. - Three i18n keys (cardHeld, heldTitle, heldDescription) across all 11 locales. The per-trace Authorize action ships with its endpoint in slice 4 (so the UI never offers a button that 404s). Source-shape assertions extended. Refs docs/plans/2026-06-10-trace-commons-held-review.md Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(webui-v2): authorize held traces for submission (slice 4) Complete the held-review feature with a promote-as-is authorize action across the stack. ironclaw_reborn_traces: - TraceContributionEnvelope gains `manual_review_authorized`; an authorized envelope submits past every gate in trace_autonomous_eligibility (the flush re-evaluates eligibility each pass, so removing the hold sidecar alone is not enough to promote). - authorize_manual_review_hold_for_scope: stamps the envelope (durable consent record) BEFORE removing the .held.json sidecar, so a crash between the two leaves the trace held (fail closed). Only ManualReview holds are authorizable; unknown submissions return Ok(false), not an error. ironclaw_product_workflow: - RebornServicesApi::authorize_trace_hold derives scope from the authenticated caller (the path submission id is never cross-scope authority), validates the id, and returns RebornTraceHoldAuthorizeResponse. ironclaw_webui_v2: - POST /api/webchat/v2/traces/holds/{submission_id}/authorize — NoBody, mutation rate limit, bearer auth. Descriptor + handler + router + contract table (now 46 routes). Frontend: - authorizeTraceHold api, an authorize mutation in useTraceCredits that invalidates the credits query on success, and a per-hold Authorize button on the Settings tab. `authorize`/`authorizing` i18n in all 11 locales. TDD: authorize promotes a High-PII held envelope past all gates; facade zero-state; webui_v2 descriptor/handler contracts; composition serve (47); source-shape assertions. clippy/fmt clean across crates. Refs docs/plans/2026-06-10-trace-commons-held-review.md Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(traces): loopback dev claim exception + profile_set consent gate Address the two codex P2 findings from review: - Preserve loopback claim uploads after onboarding: the loopback-HTTP dev invite form stores a loopback claim/ingest endpoint in the policy, but the claim/ingest validators required https and rejected loopback hosts, so a successful loopback onboarding could never mint a claim or submit credits. The validators and the pinned DNS resolution now honor the same literal-loopback exception as invite parsing (shared is_loopback_host predicate); for loopback hosts the pinned resolution additionally requires all resolved addresses to be loopback. Non-loopback http, internal hostnames, and private ranges stay rejected, and the issuer allowlist still applies. - Require explicit confirmation before community profile updates: trace_commons.profile_set now has the same hard confirmed=true input gate as onboarding — it short-circuits with consent_required before the enrollment check and any network write, since the capability is approval-gate-exempt in local-dev policy. Schema and manifest document the field. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(merge): thread attachments field through trace-capture record construction main added ThreadMessageRecord.attachments (Vec<AttachmentRef>); the trace-capture reconstruction path and its two test helpers construct records and must set it. The capture path reconstructs records from a context window for redaction/scoring and carries no attachment refs of its own, so Vec::new() is correct. * fix(traces): adapt v1 autonomous capture to new Held variant shape The merge brought in slice 1 of the held-trace-review feature, which changed TraceClientAutonomousCaptureOutcome::Held from { submission_id, reason } to { kind, reason, envelope } so manual-review holds can be retained instead of dropped. The v1 autonomous-capture path in thread_ops.rs still matched the old shape, breaking the `--no-default-features --features libsql` build (and default build). Adapt the v1 path to the new shape and give it the same retain-or-drop parity as the Reborn capture path (ironclaw_reborn_composition::trace_capture): ManualReview holds are retained via queue_held_envelope_for_scope (the on-disk held queue is shared, so a v1-captured hold surfaces in the v2 review UI); policy/value gates are dropped as before, just logged. Behavior mirrors the tested Reborn path (send_user_message_auto_queues_trace_for_enrolled_scope); the v1 autonomous-capture path is a detached tokio::spawn with no unit-testable seam, so no focused regression test is added. [skip-regression-check] * fix(traces): set manual_review_authorized in reborn-cli test envelope fixture The merge brought in the held-trace-review manual_review_authorized field on TraceContributionEnvelope. The reborn-cli trace_queue test fixture constructs the envelope directly and missed the field, breaking `cargo clippy --all-features --tests` and `Tests (all-features)` (the fixture is test-only, so the libsql binary build did not surface it). Fresh queued envelopes are not yet authorized, so false is correct. [skip-regression-check] * test(traces): pass confirmed=true in profile_set parity step The trace_commons first-party-tools parity test invoked profile_set without confirmed=true and asserted the NotEnrolled enrollment-gate result. Commit 6bc776d added the public-attribution consent gate to dispatch_profile_set, which now short-circuits to consent_required before the enrollment check when confirmed is unset — so the test's NotEnrolled assertion failed (the gate output carries no error_code). Pass confirmed=true so the call clears the consent gate and reaches the enrollment check, deterministically returning NotEnrolled with no network (the scope never onboarded). Matches the unit-test pattern established for the other profile_set tests in the same change. [skip-regression-check] * fix(traces): onboarding-security + contribution correctness (coderabbit batch 1) Addresses 6 coderabbit findings in ironclaw_reborn_traces: - device_key.rs: re-assert 0o700 on pre-existing key dirs (not just on create), so broader perms on an existing device_keys/ or pending/ can't leave invite/tenant hashes enumerable. - device_key.rs: fail closed on load when on-disk public_key/device_key_id don't match the loaded private key (tampered/partial files no longer load an inconsistent identity that only fails later at remote auth). - invite.rs: scope the staged pending-key filename by invite ORIGIN, not just code, so two issuers reusing one invite code can't share a device key (invite_hash stays code-only as the server allowlist subject). - onboarding/mod.rs: reject ingest_url values with embedded userinfo before persisting, so a malicious onboarding response can't smuggle credentials into policy.json + outbound requests. - contribution.rs: preserve mount path prefixes when deriving the community-profile endpoint (mirrors trace_submission_status_endpoint); a prefixed deployment no longer 404s on profile PUT/DELETE. - contribution.rs: fail closed in trace_autonomous_eligibility on envelopes with no allowed-uses (public_attribution-only) instead of relying on the remote to bounce them. Updated two retry tests that encoded the cross-issuer key-sharing bug now fixed: they retried against a second mock on a different port; a new spawn_flaky_mock_issuer keeps the retry on the same origin so it exercises genuine same-issuer pending-key reuse. Added regression tests for each fix. * fix(trace-commons): address coderabbit review findings on nearai#4559 - index.html: add type="button" to the Trace Commons settings subtab to prevent accidental form submission. - settings.js + i18n/en.js: route the Trace Commons credits copy through I18n.t(...) and register the matching locale keys (matches the existing surface pattern; en-only like settings.traceCommons, fallback covers rest). - factory.rs: drive the malformed SECRETS_MASTER_KEY env case through the real caller resolve_local_dev_secret_master_key (via an env-parameterized inner) and assert the rejected key is never persisted to the cached file. - trace_commons_dispatch_e2e.rs: give each test a distinct user/extension scope so onboarding state can no longer bleed across tests. - local_dev_capability_policy.toml: exempt builtin.trace_commons.onboard from the REPL approval gate (it has its own confirmed=true consent gate, mirroring profile_set). - docs: fix the onboard prompt-file reference, match the held-trace JSON shape to RebornTraceHold (submission_id + reason only), and resolve the wire-protocol ownership split (types live locally in onboarding/protocol.rs, no shared trace-commons-protocol crate). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(traces): tenant-scoping + token leak + read-failure + unbounded scopes (coderabbit batch 2) Addresses the coupled backend findings: - Tenant-scope Trace Commons local state across the Reborn paths: new trace_scope_key(tenant, user) helper keys policy / device-key / credit / profile / capture state by tenant+user, so the same user id in two tenants no longer shares state. Applied in host_runtime trace_commons dispatchers, product_workflow credits/hold, and composition trace-capture (v1 stays user-only — legacy single-tenant). Updated the affected runtime/sink tests and added a non-owner attribution assertion. - Do not return the raw profile token from the model-visible profile_token capability: persist it to a 0600 <scope>/profile_token.jwt and return the file path + instructions instead, keeping the bearer credential off the LLM transcript. - Stop masking genuine local-state read failures as zero/not-enrolled: the status capability and the WebUI credits path now propagate a read/parse failure (NotFound is already softened inside read_*_for_scope) so an enrolled user with a corrupt policy file is not told they have nothing. - Bound ObservedTraceScopes: the periodic flush worker now prunes drained scopes (new trace_scope_has_pending_queue) after each tick, so the set is bounded by actual pending backlog instead of growing one entry per caller ever seen. Note: a v1 caller-level test for the ManualReview hold-retention path is not included — v1 ingress blocks secrets outright and the outbound leak detector redacts them, so the High-residual-PII condition that produces a ManualReview hold cannot be reproduced through process_user_input. The retention logic is identical to and covered by the Reborn-side capture_retains_manual_review_hold_for_high_pii_trace. * test(traces): enroll under tenant-scoped key in webui_v2_serve credits test trace_credits_reports_enrolled_for_caller_with_enabled_policy wrote the policy under the bare user id, but the credits route now keys local state by trace_scope_key(tenant, user). Enroll (and clean up) under the composite TENANT/user scope so the route sees the enrollment. * fix(factory): fail closed on explicit-but-unusable SECRETS_MASTER_KEY An explicitly-set-but-unusable local-dev master key silently fell through to generating + persisting a fresh key, leaving local-dev secrets encrypted under an unintended master key the operator never chose: - resolve_local_dev_secret_master_key used std::env::var(...).ok(), which drops VarError::NotUnicode -> treated as absent. Now only NotPresent is absent; a non-Unicode value returns InvalidConfig. - resolve_local_dev_secret_master_key_with_env collapsed a set-but-empty (or whitespace-only) value to None via .filter(). Now a set-but-empty value returns InvalidConfig instead of generating a key. Added resolve_local_dev_secret_master_key_rejects_set_but_empty_env_without_persisting asserting empty/whitespace env values fail closed and persist nothing. (coderabbit follow-up on nearai#3794) * fix(factory): reject empty SECRETS_MASTER_KEY before the cached-file read Follow-up to the prior fix: the empty-env rejection lived in the env branch, which only runs when no cached key file exists. On a rebuild where .reborn-local-dev-secrets-master-key already exists, the cached key was returned first, so an explicitly-set-but-empty SECRETS_MASTER_KEY was still silently ignored. Hoist the empty/whitespace rejection (and env normalization) above the cached-file read so it fails closed regardless of cached state. Added resolve_local_dev_secret_master_key_rejects_empty_env_even_with_cached_file asserting the empty env is rejected and the cached key is left unchanged. * fix(traces): address 14:54 coderabbit re-review (tenant-seed, IO errors, effects, test) Four outside-diff findings from the re-review: - runtime.rs: seed ObservedTraceScopes with the runtime owner's trace_scope_key(tenant, owner) composite, not the bare owner id, so startup pending-queue discovery matches how capture keys state; the enrolled-scope test cleanup now removes the composite scope dir too. - runtime.rs: the trace-queue polling test helper no longer swallows read_dir errors via unwrap_or_default() — only NotFound is the expected pre-capture fallback; any other IO error panics instead of masking as 'no queued traces'. - trace_commons.rs manifests + local_dev grants: onboard (device-key material) and profile_token (0600 token file) now declare Read/WriteFilesystem effects, and the local-dev grants allow them, so the effect model accurately models the local secret-material writes. - local_dev_authorization test: added local_dev_trace_commons_onboard_skips_approval_gate (the onboard exemption was the actual fix; the profile_set-only test would pass even if the onboard TOML exemption were dropped). * fix(factory): validate non-empty SECRETS_MASTER_KEY before the cached-file read Follow-up: the prior fix rejected an *empty* env value before the cached read but still validated a non-empty *malformed* value only after it. So a valid cache + SECRETS_MASTER_KEY=0000... silently ignored the explicit bad secret config on rebuilds. Move validate_resolved_master_key into the up-front env normalization so any explicit-but-unusable env key (empty OR malformed) fails closed regardless of cached state. Added resolve_local_dev_secret_master_key_rejects_malformed_env_even_with_cached_file. * fix(traces): address 15:41 coderabbit re-review (credits read-failure + 2 test guards) - trace_commons.rs dispatch_credits: stop masking genuine records read/parse failures as 'no records' (NotFound is already softened inside read_local_trace_records_for_scope); report RecordsReadFailed, mirroring dispatch_status. - runtime.rs trace-queue polling helper: fail loud on per-ENTRY read_dir IO errors too (map + unwrap_or_else panic) instead of filter_map(e.ok()), so a broken entry can't be silently dropped while claiming the queue holds one. - local_dev_authorization approval-gate test: assert the effects DO require approval without the exemption (local_dev_effects_require_approval), so the test can't pass via a non-gating default policy if the TOML exemption were dropped. * fix(traces): address Henri review — backend findings (atomic token, error mapping, validation, egress test) - persist_profile_token now writes atomically (unique 0600 temp + fsync + rename) so a reader never observes a half-written or overwritten bearer credential under overlapping mints (Henri perf/security Medium). - dispatch_onboard error mapping: OnboardError::DeviceKey is reported as a distinct DeviceKeyError (re-run onboarding) instead of being collapsed into PersistError's check-disk-and-permissions guidance (Henri bugs Medium). - parse_profile_set_input enforces the manifest's declared schema at parse time: handle 3-32 ASCII letters/digits/-/_, bio <= 280 bytes (Henri conventions Medium). Added schema-limit test. - Added dispatch_onboard_confirmed_without_host_egress_is_network_denied covering the NetworkDenied host-egress-miswiring branch (Henri tests Medium). * fix(traces): address Henri review — frontend findings (enrolled empty-state + polling) - v1 credits: TraceCreditResponse now carries `enrolled` (read from the standing policy), and settings.js keys the opt-in empty state on `!data.enrolled` instead of `!submissions_total` — an enrolled user with zero submissions now sees their zero-credit view, not the not-enrolled prompt (Henri bugs Medium). - useTraceCredits: each fetch rebuilds the full server-side credit view, so the aggressive 60s poll made an open tab steady O(history) work. Relaxed to a 5-min interval + staleTime + no background polling, keeping a focus refetch for liveness; mutation invalidation still updates promptly. Added a TODO to incrementalize the server-side view (Henri perf Medium). * perf(traces): memoize server-side credit view by on-disk input signature Bounds the trace-credits polling cost to O(new submissions) instead of O(total history). New scoped_credit_view(scope) caches the computed credit report + manual-review holds keyed by a cheap change signature (submissions file mtime+len, plus a hash of the held-trace sidecars). On the steady-state polling case (unchanged history) a request is a couple of stat()s + a clone rather than reading/parsing the full submissions file and re-aggregating. On any change the signature differs and it recomputes once. Cache is bounded (4096 scopes, cleared on overflow). Wired through the polled WebUI path (local_trace_credits_for_user) and the model-visible credits capability (dispatch_credits). Added scoped_credit_view_reflects_record_changes_via_signature covering the cache-hit path and signature-based invalidation on record changes. Completes the TODO from the Henri perf-review follow-up (#5). * fix(traces): gate profile_set behind runtime approval (Henri #1 High) profile_set publishes a public community profile (an external write to a public surface). Its `confirmed=true` input is model-controlled, so a prompt-injected or confused model could supply it. Make the runtime approval gate the primary, user-controlled consent control: - Drop `builtin.trace_commons.profile_set` from the local-dev approval-gate exemption list (keep `onboard`, which runs its own in-turn confirmed=true consent before the network POST). - Set profile_set's manifest default_permission to Ask (was Allow). - Split the local-dev authorization test into `local_dev_trace_commons_profile_set_requires_approval_gate` (asserts Decision::RequireApproval) and `local_dev_trace_commons_onboard_skips_approval_gate` (asserts Decision::Allow), via a shared `trace_commons_authorize_decision` helper that first asserts the effects would gate without an exemption. Also fix a pre-existing trace_commons harness gap: onboard + profile_token gained a WriteFilesystem effect (device-key persistence) but the `trace_commons_tools` harness allow-set was never updated, so those capabilities were filtered out of the model-visible surface and the parity/visibility tests failed with driver_unavailable. Grant WriteFilesystem in the harness allow-set. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * refactor(traces): extract onboarding test harness to sibling file (Henri #8) The onboarding module's ~840-line `#[cfg(test)] mod tests` block (mock issuer harness, retry/idempotency coverage, URL-validation tests) made `onboarding/mod.rs` a 1319-line file dominated by test scaffolding. Move the module body into `onboarding/tests.rs` declared `#[cfg(test)] mod tests;`, leaving mod.rs focused on production logic (now 480 lines). No test behavior changes; `use super::*;` still resolves to the onboarding module. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(webui-v2): update embedded-asset assertion for incrementalized credits poll The Henri #5 polling fix changed useTraceCredits.js from refetchInterval 60_000 to 300_000 (plus refetchIntervalInBackground: false and staleTime: 60_000), but the embedded-asset test in assets.rs still asserted the old 60_000 value and failed in CI. Update the assertion to lock the new infrequent-poll + paused-while-hidden + focus-refetch shape. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(traces): address CodeRabbit review + stale capability-policy test CodeRabbit findings on the gating/refactor commits: - Major: format_profile_token returned the absolute host path of the token file (token_file) on the model-visible surface, which violates the "never expose absolute paths" guideline. Replace with an opaque token_delivery marker; the token is still persisted 0600 for out-of-band retrieval by a bearer-auth UI/CLI. Update the message + test accordingly. - Major (fail-loud): profile_token_error_value and profile_set_error_value collapsed "could not read policy" into NotEnrolled, sending enrolled users back through onboarding on unreadable/corrupt state. Split into a distinct PolicyReadFailed result in both formatters (matches dispatch_status). - Minor: stale comment claiming profile_set is approval-gate-exempt (it is now PermissionMode::Ask and NOT exempt) — corrected. - Minor: inaccurate harness comments (profile_token writes profile_token.jwt not device-key material; yolo auto-approves all Trace Commons Ask-gated tools, not just onboard) — corrected. Also fix bundled_local_dev_capability_policy_parses, which still asserted the pre-gating policy shape: profile_set as exempt (now onboard exempt / profile_set NOT exempt), onboard's grant missing the read/write filesystem effects, and profile_token/profile_set sharing one effect-set assertion even though profile_token now carries WriteFilesystem and profile_set does not. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * style(traces): collapse single-line use block after Path import removal rustfmt collapses `use std::{panic, path::PathBuf, sync::Arc}` to one line once Path was dropped; the prior commit skipped re-running fmt after that edit, reddening the Formatting CI check. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(traces): consent-gate profile_token + drop fixed-origin profile URL (CodeRabbit) Two Major CodeRabbit security findings on the profile tools: - profile_token minted and persisted a bearer credential with no in-turn consent gate. PermissionMode::Ask can be auto-approved under local-yolo, so a model call could mint a credential without explicit per-conversation consent. Add a hard confirmed=true gate (schema + parse + consent_required short-circuit) before minting, mirroring dispatch_onboard / dispatch_profile_set. - format_profile_token and profile_set_success_value hardcoded https://tracecommons.ai/profile. The token is scoped to the user's ENROLLED issuer (which may be self-hosted or loopback), so steering the user to paste a bearer profile-management token at a fixed origin could leak it to the wrong host. Drop the fixed profile_url; route through the enrolled profile flow / local UI/CLI out of band. Tests: new dispatch_profile_token_without_confirmed_returns_consent_required_no_mint; existing without-enrollment test now passes confirmed=true; profile_set success test asserts no fixed origin; parity step mints with confirmed=true. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(traces): route agent-invoked profile writes through host egress (CodeRabbit #3) profile_token (upload-claim mint) and profile_set (community-profile PUT/DELETE) previously made network writes via the ironclaw_reborn_traces crate-local reqwest client, bypassing the host RuntimeHttpEgress pipeline (private-IP filtering, redaction, byte accounting) that onboard already uses. Add a `ContributionHttpSink` port (mirroring `OnboardingHttpSink`): when a sink is injected, the mint POST and the profile PUT/DELETE run through host egress; when `None`, the existing hardened crate-local client is used unchanged. host_runtime supplies `HostEgressContributionSink` (wraps RuntimeHttpEgress, sanitizes errors via stable_runtime_reason, never leaks URL/token), and dispatch_profile_token / dispatch_profile_set fail closed with NetworkDenied if egress is absent (after the enrollment pre-check, so a not-enrolled user still gets NotEnrolled guidance). The background trace-upload / status-sync worker and the CLI keep the crate-local client (pass `None`): that lane is a durable, model-input-free internal task that sends only already-redacted envelopes to the operator-enrolled endpoint and does its own SSRF/private-IP validation, so host egress adds complexity without security benefit. Justification recorded in a comment on `trace_remote_http_client`. New public surface: ContributionHttpSink/Request/Response/Error/Method, mint_profile_attribution_token_for_scope_via_sink, set_community_profile_for_scope_via_sink. Existing public fns keep their signatures (None path) so CLI/worker/tests are unchanged. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
theredspoon
added a commit
that referenced
this pull request
Jun 18, 2026
Only allow the Docker Image publishing job to run from nearai/ironclaw so the control branch copy cannot publish from theredspoon.
theredspoon
pushed a commit
that referenced
this pull request
Jun 21, 2026
* Add WebSocket gateway and control plane endpoint Adds bidirectional WebSocket transport to the web gateway alongside the existing SSE stream. Clients can send messages, approvals, and pings over a single persistent connection at /api/chat/ws. - Enable axum `ws` feature for built-in WebSocket support - Add WsClientMessage/WsServerMessage types with tagged JSON protocol - Add subscribe_raw() to SseManager for non-SSE consumers - Create ws.rs with connection handler (split sender/receiver tasks) - Add WsConnectionTracker for active connection counting - Add /api/gateway/status control plane endpoint (SSE + WS counts) - 35 new tests covering message types, broadcast, and handler logic https://claude.ai/code/session_01KEaLN6Xq2j5EeV3SGHQT6b * Add e2e WebSocket gateway integration tests - Add tokio-tungstenite dev-dependency for WebSocket client in tests - Update start_server to return actual bound SocketAddr (enables port 0) - Add 10 e2e tests covering full HTTP upgrade → WebSocket → message flow: ping/pong, message routing to agent, broadcast event delivery, connection tracking, invalid message handling, auth rejection, gateway status endpoint, and multi-event sequencing https://claude.ai/code/session_01KEaLN6Xq2j5EeV3SGHQT6b --------- Co-authored-by: Claude <noreply@anthropic.com>
theredspoon
pushed a commit
that referenced
this pull request
Jun 21, 2026
…tion (nearai#238) * feat: add extension registry with metadata catalog, CLI, and onboarding integration Adds a central registry that catalogs all 14 available extensions (10 tools, 4 channels) with their capabilities, auth requirements, and artifact references. The onboarding wizard now shows installable channels from the registry and offers tool installation as a new Step 7. - registry/ folder with per-extension JSON manifests and bundle definitions - src/registry/ module: manifest structs, catalog loader, installer - `ironclaw registry list|info|install|install-defaults` CLI commands - Setup wizard enhanced: channels from registry, new extensions step (8 steps) Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(setup): resolve workspace errors for tool crates and channels-only onboarding Tool crates in tools-src/ and channels-src/ failed `cargo metadata` during onboard install because Cargo resolved them as part of the root workspace. Add `[workspace]` table to each standalone crate and extend the root `workspace.exclude` list so they build independently. Channels-only mode (`onboard --channels-only`) failed with "Secrets not configured" and "No database connection" because it skipped database and security setup. Add `reconnect_existing_db()` to establish the DB connection and load saved settings before running channel configuration. Also improve the tunnel "already configured" display to show full provider details (domain, mode, command) instead of just the provider name. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(registry): address PR review feedback on installer and catalog - Use manifest.name (not crate_name) for installed filenames so discovery, auth, and CLI commands all agree on the stem (#1) - Add AlreadyInstalled error variant instead of misleading ExtensionNotFound (#2) - Add DownloadFailed error variant with URL context instead of stuffing URLs into PathBuf (#3) - Validate HTTP status with error_for_status() before reading response bytes in artifact downloads (#4) - Switch build_wasm_component to tokio::process::Command with status() so build output streams to the terminal (#6) - Find WASM artifact by crate_name specifically instead of picking the first .wasm file in the release directory (#7) - Add is_file() guard in catalog loader to skip directories (#8) - Detect ambiguous bare-name lookups when both tools/<name> and channels/<name> exist, with get_strict() returning an error (#9) - Fix wizard step_extensions to check tool.name for installed detection, consistent with the new naming (#11, #12) - Fix redundant closures and map_or clippy warnings in changed files Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(setup): restore DB connection fields after settings reload reconnect_postgres() and reconnect_libsql() called Settings::from_db_map() which overwrote database_url / libsql_path / libsql_url set from env vars. Also use get_strict() in cmd_info to surface ambiguous bare-name errors. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * style: fix clippy collapsible_if and print_literal warnings Collapse nested if-let chains and inline string literals in format macros to satisfy CI clippy lint checks (deny warnings). Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(registry): prefer artifacts for install-defaults and improve dir lookup - InstallDefaults now defaults to downloading pre-built artifacts (matching `registry install` behavior), with --build flag for source builds. - find_registry_dir() walks up 3 ancestor levels from the exe and adds a CARGO_MANIFEST_DIR fallback, matching load_registry_catalog() logic. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
theredspoon
pushed a commit
that referenced
this pull request
Jun 21, 2026
* refactor: extract shared assertion helpers to support/assertions.rs Move 5 assertion helpers from e2e_spot_checks.rs to a shared module. Add assert_all_tools_succeeded and assert_tool_succeeded for eliminating false positives in E2E tests. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: add tool output capture via tool_results() accessor Extract (name, preview) from ToolResult status events in TestChannel and TestRig, enabling content assertions on tool outputs. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: correct tool parameters in 3 broken trace fixtures - tool_time.json: add missing "operation": "now" for time tool - robust_correct_tool.json: same fix - memory_full_cycle.json: change "path" to "target" for memory_write Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: add tool success and output assertions to eliminate false positives Every E2E test that exercises tools now calls assert_all_tools_succeeded. Added tool output content assertions where tool results are predictable (time year, read_file content, memory_read content). Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: capture per-tool timing from ToolStarted/ToolCompleted events Record Instant on ToolStarted and compute elapsed duration on ToolCompleted, wiring real timing data into collect_metrics() instead of hardcoded zeros. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * refactor: add RAII CleanupGuard for temp file/dir cleanup in tests Replace manual cleanup_test_dir() calls and inline remove_file() with Drop-based CleanupGuard that ensures cleanup even if a test panics. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: add Drop impl and graceful shutdown for TestRig Wrap agent_handle in Option so Drop can abort leaked tasks. Signal the channel shutdown before aborting for future cooperative shutdown. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: replace agent startup sleep with oneshot ready signal Use a oneshot channel fired in Channel::start() instead of a fixed 100ms sleep, eliminating the race condition on slow systems. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: replace fragile string-matching iteration limit with count-based detection Use tool completion count vs max_tool_iterations instead of scanning status messages for "iteration"/"limit" substrings. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: use assert_all_tools_succeeded for memory_full_cycle test Remove incorrect comment about memory_tree failing with empty path (it actually succeeds). Omit empty path from fixture and use the standard assert_all_tools_succeeded instead of per-tool assertions. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * refactor: promote benchmark metrics types to library code Move TraceMetrics, ScenarioResult, RunResult, MetricDelta, and compare_runs() from tests/support/metrics.rs to src/benchmark/metrics.rs. Existing tests use re-export for backward compatibility. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: add Scenario and Criterion types for agent benchmarking Scenario defines a task with input, success criteria, and resource limits. Criterion is an enum of programmatic checks (tool_used, response_contains, etc.) evaluated without LLM judgment. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: add initial benchmark scenario suite (12 scenarios across 5 categories) Scenarios cover tool_selection, tool_chaining, error_recovery, efficiency, and memory_operations. All loaded from JSON with deserialization validation test. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: add benchmark runner with BenchChannel and InstrumentedLlm BenchChannel is a minimal Channel implementation for benchmarks. InstrumentedLlm wraps any LlmProvider to capture per-call metrics. Runner creates a fresh agent per scenario, evaluates success criteria, and produces RunResult with timing, token, and cost metrics. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: add baseline management, reports, and benchmark entry point - baseline.rs: load/save/promote benchmark results - report.rs: format comparison reports with regression detection - benchmark_runner.rs: integration test with real LLM (feature-gated) - Add benchmark feature flag to Cargo.toml Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * style: apply cargo fmt to benchmark module Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat(benchmark): add multi-turn scenario types with setup, judge, ResponseNotContains Add BenchScenario, Turn, TurnAssertions, JudgeConfig, ScenarioSetup, WorkspaceSetup, SeedDocument types for multi-turn benchmark scenarios. Add ResponseNotContains criterion variant. Add TurnAssertions::to_criteria() converter for backward compat with existing evaluation engine. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat(benchmark): add JSON scenario loader with recursive discovery and tag filter Add load_bench_scenarios() for the new BenchScenario format with recursive directory traversal and tag-based filtering. Create 4 initial trajectory scenarios across tool-selection, multi-turn, and efficiency categories. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat(benchmark): multi-turn runner with workspace seeding and per-turn metrics Add run_bench_scenario() that loops over BenchScenario turns, seeds workspace documents, collects per-turn metrics (tokens, tool calls, wall time), and evaluates per-turn assertions. Add TurnMetrics to metrics.rs and clear_for_next_turn() to BenchChannel. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat(benchmark): add LLM-as-judge scoring with prompt formatting and score parsing Create judge.rs with format_judge_prompt, parse_judge_score, and judge_turn. Wire into run_bench_scenario for turns with judge config -- scores below min_score fail the turn. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat(benchmark): add CLI subcommand (ironclaw benchmark) Add BenchmarkCommand with --tags, --scenario, --no-judge, --timeout, --update-baseline flags. Wire into Command enum and main.rs dispatch. Feature-gated behind benchmark flag. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat(benchmark): per-scenario JSON output with full trajectory Add save_scenario_results() that writes per-scenario JSON files alongside the run summary. Each scenario gets its own file with turn_metrics trajectory. Update CLI to use new output format. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat(benchmark): add ToolRegistry::retain_only and wire tool filtering in scenarios Add a retain_only() method to ToolRegistry that filters tools down to a given allowlist. Wire this into run_bench_scenario() so that when a scenario specifies a tools list in its setup, only those tools are available during the benchmark run. Includes two tests for the new method: one verifying filtering works and one verifying empty input is a no-op. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat(benchmark): wire identity overrides into workspace before agent start Add seed_identity() helper that writes identity files (IDENTITY.md, USER.md, etc.) into the workspace before the agent starts, so that workspace.system_prompt() picks them up. Wire it into run_bench_scenario() after workspace seeding. Include a test that verifies identity files are written and readable. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat(benchmark): add --parallel and --max-cost CLI flags Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(benchmark): use feature-conditional snapshot names for CLI help tests Prevents snapshot conflicts between default (no benchmark) and all-features (with benchmark) builds by using separate snapshot names per feature set. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat(benchmark): parallel execution with JoinSet and budget cap enforcement Replace sequential loop in run_all_bench() with parallel execution using JoinSet + semaphore when config.parallel > 1. Add budget cap enforcement that skips remaining scenarios when max_total_cost_usd is exceeded. Track skipped count in RunResult.skipped_scenarios and display it in format_report(). Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat(benchmark): add tool restriction and identity override test scenarios Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * chore: fix formatting for Phase 3 Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat(benchmark): add SkillRegistry::retain_only and wire skill filtering in scenarios Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat(benchmark): add --json flag for machine-readable output Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * ci: add GitHub Actions benchmark workflow (manual trigger) Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * refactor(benchmark): remove in-tree benchmark harness, keep retain_only utilities Move benchmark-specific code out of ironclaw in preparation for the nearai/benchmarks trajectory adapter. This removes: - src/benchmark/ (runner, scenarios, metrics, judge, report, etc.) - src/cli/benchmark.rs and the Benchmark CLI subcommand - benchmarks/ data directory (scenarios + trajectories) - .github/workflows/benchmark.yml - The "benchmark" Cargo feature flag What remains: - ToolRegistry::retain_only() and SkillRegistry::retain_only() - Test support types (TraceMetrics, InstrumentedLlm) inlined into tests/support/ instead of re-exporting from the deleted module Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * docs: add README for LLM trace fixture format Documents the trajectory JSON format, response types, request hints, directory structure, and how to write new traces. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat(test): unify trace format around turns, add multi-turn support Introduce TraceTurn type that groups user_input with LLM response steps, making traces self-contained conversation trajectories. Add run_trace() to TestRig for automatic multi-turn replay. Backward-compatible: flat "steps" JSON is deserialized as a single turn transparently. Includes all trace fixtures (spot, coverage, advanced), plan docs, and new e2e tests for steering, error recovery, long chains, memory, and prompt injection resilience. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(test): fix CI failures after merging main - Fix tool_json fixture: use "data" parameter (not "input") to match JsonTool schema - Fix status_events test: remove assertion for "time" tool that isn't in the fixture (only "echo" calls are used) - Allow dead_code in test support metrics/instrumented_llm modules (utilities for future benchmark tests) [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * Working on recording traces and testing them * feat(test): add declarative expects to trace fixtures, split infra tests Add TraceExpects struct with 9 optional assertion fields (response_contains, tools_used, all_tools_succeeded, etc.) that can be declared in fixture JSON instead of hand-written Rust. Add verify_expects() and run_recorded_trace() so recorded trace tests become one-liners. Split trace infra tests (deserialization, backward compat) into tests/trace_format.rs which doesn't require the libsql feature gate. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * refactor(test): add expects to all trace fixtures, simplify e2e tests Add declarative expects blocks to all 19 trace fixture JSONs across spot/, coverage/, advanced/, and root directories. Update all 8 e2e test files to use verify_trace_expects() / run_and_verify_trace(), replacing ~270 lines of hand-written assertions with fixture-driven verification. Tests that check things beyond expects (file content on disk, metrics, event ordering) keep those extra assertions alongside the declarative ones. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(test): adapt tests to AppBuilder refactor, fix formatting Update test files to work with refactored TestRigBuilder that uses AppBuilder::build_all() (removing with_tools/with_workspace methods). Update telegram_check fixture to use tool_list instead of echo. Fix cargo fmt issues in src/llm/mod.rs and src/llm/recording.rs. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * refactor(test): deduplicate support unit tests into single binary Support modules (assertions, cleanup, test_channel, test_rig, trace_llm) had #[cfg(test)] mod tests blocks that were compiled and run 12 times — once per e2e test binary that declares `mod support;`. Extracted all 29 support unit tests into a dedicated `tests/support_unit_tests.rs` so they run exactly once. [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * style: fix trailing newlines in support files Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * refactor(test): unify trace types and fix recorded multi-turn replay Import shared types (TraceStep, TraceResponse, TraceToolCall, RequestHint, ExpectedToolResult, MemorySnapshotEntry, HttpExchange*) from ironclaw::llm::recording instead of redefining them in trace_llm.rs. Fix the flat-steps deserializer to split at UserInput boundaries into multiple turns, instead of filtering them out and wrapping everything into a single turn. This enables recorded multi-turn traces to be replayed as proper multi-turn conversations via run_trace(). [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(test): fix CI failures - unused imports and missing struct fields - Add #[allow(unused_imports)] on pub use re-exports in trace_llm.rs (types are re-exported for downstream test files, not used locally) - Add `..` to ToolCompleted pattern in test_channel.rs to match new `error` and `parameters` fields Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(test): fix CI failures after merging main - Add missing `error` and `parameters` fields to ToolCompleted constructors in support_unit_tests.rs - Add `..` to ToolCompleted pattern match in support_unit_tests.rs - Add #[allow(dead_code)] to CleanupGuard, LlmTrace impl, and TraceLlm impl (only used behind #[cfg(feature = "libsql")]) Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * Adding coverage running script * fix(test): address review feedback on E2E test infrastructure - Increase wait_for_responses polling to exponential backoff (50ms-500ms) and raise default timeout from 15s to 30s to reduce CI flakiness (#1) - Strengthen prompt_injection_resilience test with positive safety layer assertion via has_safety_warnings(), enable injection_check (#2) - Add assert_tool_order() helper and tools_order field in TraceExpects for verifying tool execution ordering in multi-step traces (#3) - Document TraceLlm sequential-call assumption for concurrency (#6) - Clean up CleanupGuard with PathKind enum instead of shotgun remove_file + remove_dir_all on every path (#8) - Fix coverage.sh: default to --lib only, fix multi-filter syntax, add COV_ALL_TARGETS option - Add coverage/ to .gitignore - Remove planning docs from PR [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: address PR review - use HashSet in retain_only, improve skill test - Use HashSet for O(N+M) lookup in SkillRegistry::retain_only and ToolRegistry::retain_only instead of linear scan - Strengthen test_retain_only_empty_is_noop in SkillRegistry to pre-populate with a skill before asserting the no-op behavior [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(test): revert incorrect safety layer assertion in injection test The safety layer sanitizes tool output, not user input. The injection test sends a malicious user message with no tools called, so the safety layer never fires. Reverted to the original test which correctly validates the LLM refuses via trace expects. Also fixed case-sensitive request hint ("ignore" -> "Ignore") to suppress noisy warning. [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: clean stale profdata before coverage run Adds `cargo llvm-cov clean` before each run to prevent "mismatched data" warnings from stale instrumentation profiles. [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * style: fix formatting in retain_only test [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> Co-authored-by: Illia Polosukhin <ilblackdragon@gmail.com>
theredspoon
pushed a commit
that referenced
this pull request
Jun 21, 2026
) * feat: add inbound attachment support to WASM channel system Add attachment record to WIT interface and implement inbound media parsing across all four channel implementations (Telegram, Slack, WhatsApp, Discord). Attachments flow from WASM channels through EmittedMessage to IncomingMessage with validation (size limits, MIME allowlist, count caps) at the host boundary. - Add `attachment` record to `emitted-message` in wit/channel.wit - Add `IncomingAttachment` struct to channel.rs and re-export - Add host-side validation (20MB total, 10 max, MIME allowlist) - Telegram: parse photo, document, audio, video, voice, sticker - Slack: parse file attachments with url_private - WhatsApp: parse image, audio, video, document with captions - Discord: backward-compatible empty attachments - Update FEATURE_PARITY.md section 7 - Add fixture-based tests per channel and host integration tests [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: integrate outbound attachment support and reconcile WIT types (nearai#409) Reconcile PR nearai#409's outbound attachment work with our inbound attachment support into a unified design: WIT type split: - `inbound-attachment` in channel-host: metadata-only (id, mime_type, filename, size_bytes, source_url, storage_key, extracted_text) - `attachment` in channel: raw bytes (filename, mime_type, data) on agent-response for outbound sending Outbound features (from PR nearai#409): - `on-broadcast` WIT export for proactive messages without prior inbound - Telegram: multipart sendPhoto/sendDocument with auto photo→document fallback for files >10MB - wrapper.rs: `call_on_broadcast`, `read_attachments` from disk, attachment params threaded through `call_on_respond` - HTTP tool: `save_to` param for binary downloads to /tmp/ (50MB limit, path traversal protection, SSRF-safe redirect following) - Message tool: allow /tmp/ paths for attachments alongside base_dir - Credential env var fallback in inject_channel_credentials Channel updates: - All 4 channels implement on_broadcast (Telegram full, others stub) - Telegram: polling_enabled config, adjusted poll timeout - Inbound attachment types renamed to InboundAttachment in all channels Tests: 1965 passing (9 new), 0 clippy warnings [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: add audio transcription pipeline and extensible WIT attachment design Add host-side transcription middleware (OpenAI Whisper) that detects audio attachments with inline data on incoming messages and transcribes them automatically. Refactor WIT inbound-attachment to use extras-json and a store-attachment-data host function instead of typed fields, so future attachment properties (dimensions, codec, etc.) don't require WIT changes that invalidate all channel plugins. - Add src/transcription/ module: TranscriptionProvider trait, TranscriptionMiddleware, AudioFormat enum, OpenAI Whisper provider - Add src/config/transcription.rs: TRANSCRIPTION_ENABLED/MODEL/BASE_URL - Wire middleware into agent message loop via AgentDeps - WIT: replace data + duration-secs with extras-json + store-attachment-data - Host: parse extras-json for well-known keys, merge stored binary data - Telegram: download voice files via store-attachment-data, add duration to extras-json, add /file/bot to HTTP allowlist, voice-only placeholder - Add reqwest multipart feature for Whisper API uploads - 5 regression tests for transcription middleware Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: wire attachment processing into LLM pipeline with multimodal image support Attachments on incoming messages are now augmented into user text via XML tags before entering the turn system, and images with data are passed as multimodal content parts (base64 data URIs) to LLM providers. This enables audio transcripts, document text, and image content to reach the LLM without changes to ChatMessage serialization or provider interfaces. - Add src/agent/attachments.rs with augment_with_attachments() and 9 unit tests - Add ContentPart/ImageUrl types to llm::provider with OpenAI-compatible serde - Carry image_content_parts transiently on Turn (skipped in serialization) - Update nearai_chat and rig_adapter to serialize multimodal content - Add 3 e2e tests verifying attachments flow through the full agent loop Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: CI failures — formatting, version bumps, and Telegram voice test - Fix cargo fmt formatting in attachments.rs, nearai_chat.rs, rig_adapter.rs, e2e_attachments.rs - Bump channel registry versions 0.1.0 → 0.2.0 (discord, slack, telegram, whatsapp) to satisfy version-bump CI check - Fix Telegram test_extract_attachments_voice: add missing required `duration` field to voice fixture JSON Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: bump WIT channel version to 0.3.0, fix Telegram voice test, add pre-commit hook - Bump wit/channel.wit package version 0.2.0 → 0.3.0 (interface changed with store-attachment-data) - Update WIT_CHANNEL_VERSION constant and registry wit_version fields to match - Fix Telegram test_extract_attachments_voice: gate voice download behind #[cfg(target_arch = "wasm32")] so host functions aren't called in native tests, update assertions for generated filename and extras_json duration - Add @0.3.0 linker stubs in wit_compat.rs - Add .githooks/pre-commit hook that runs scripts/check-version-bumps.sh when WIT or extension sources are staged - Symlink commit-msg regression hook into .githooks/ [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * refactor: extract voice download from extract_attachments into handle_message Move download_voice_file + store_attachment_data calls out of extract_attachments into a separate download_and_store_voice function called from handle_message. This keeps extract_attachments as a pure data-mapping function with no host calls, making it fully testable in native unit tests without #[cfg(target_arch)] gates. [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: address PR review comments — security, correctness, and code quality Security fixes: - Add path validation to read_attachments (restrict to /tmp/) preventing arbitrary file reads from compromised tools - Escape XML special characters in attachment filenames, MIME types, and extracted text to prevent prompt injection via tag spoofing - Percent-encode file_id in Telegram getFile URL to prevent query injection - Clone SecretString directly instead of expose_secret().to_string() Correctness fixes: - Fix store_attachment_data overwrite accounting: subtract old entry size before adding new to prevent inflated totals and false rejections - Use max(reported, stored_size) for attachment size accounting to prevent WASM channels from under-reporting size_bytes to bypass limits - Add application/octet-stream to MIME allowlist (channels default unknown types to this) Code quality: - Extract send_response helper in Telegram, deduplicating on_respond and on_broadcast - Rename misleading Discord test to test_parse_slash_command_interaction - Fix .githooks/commit-msg to use relative symlink (portable across machines) [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: add tool_upgrade command + fix TOCTOU in save_to path validation Add `tool_upgrade` — a new extension management tool that automatically detects and reinstalls WASM extensions with outdated WIT versions. Preserves authentication secrets during upgrade. Supports upgrading a single extension by name or all installed WASM tools/channels at once. Fix TOCTOU in `validate_save_to_path`: validate the path *before* creating parent directories, so traversal paths like `/tmp/../../etc/` cannot cause filesystem mutations outside /tmp before being rejected. [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: unify WIT package version to 0.3.0 across tool.wit and all capabilities tool.wit and channel.wit share the `near:agent` package namespace, so they must declare the same version. Bumps tool.wit from 0.2.0 to 0.3.0 and updates all capabilities files and registry entries to match. Fixes `cargo component build` failure: "package identifier near:agent@0.2.0 does not match previous package name of near:agent@0.3.0" [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: move WIT file comments after package declaration WIT treats `//` comments before `package` as doc comments. When both tool.wit and channel.wit had header comments, the parser rejected them as "doc comments on multiple 'package' items". Move comments after the package declaration in both files. Also bumps tool registry versions to 0.2.0 to match the WIT 0.3.0 bump. [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: display extension versions in gateway Extensions tab Add version field to InstalledExtension and RegistryEntry types, pipe through the web API (ExtensionInfo, RegistryEntryInfo), and render as a badge in the gateway UI for both installed and available extensions. For installed WASM extensions, version is read from the capabilities file with a fallback to the registry entry when the local file has no version (old installations). Bump all extension Cargo.toml and registry JSON versions from 0.1.0 to 0.2.0 to keep them in sync. [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: add document text extraction middleware for PDF, Office, and text files Extract text from document attachments (PDF, DOCX, PPTX, XLSX, RTF, plain text, code files) so the LLM can reason about uploaded documents. Uses pdf-extract for PDFs, zip+XML parsing for Office XML formats, and UTF-8 decode for text files. Wired into the agent loop after transcription middleware. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: download document files in Telegram channel for text extraction The DocumentExtractionMiddleware needs file bytes in the attachment `data` field, but only voice files were being downloaded. Document attachments (PDFs, DOCX, etc.) had empty `data` and a source_url with a credential placeholder that only works inside the WASM host's http_request. Add `download_and_store_documents()` that downloads non-voice, non-image, non-audio attachments via the existing two-step getFile→download flow and stores bytes via `store_attachment_data` for host-side extraction. Also rename `download_voice_file` → `download_telegram_file` since it's generic for any file_id. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: allow Office MIME types and increase file download limit for Telegram Two issues preventing document extraction from Telegram: 1. PPTX/DOCX/XLSX MIME types (application/vnd.*) were dropped by the WASM host attachment allowlist — add application/vnd., application/msword, and application/rtf prefixes. 2. Telegram file downloads over 10 MB failed with "Response body too large" — set max_response_bytes to 20 MB in Telegram capabilities. [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: report document extraction errors back to user instead of silently skipping - Bump max_response_bytes to 50 MB for Telegram file downloads - When document extraction fails (too large, download error, parse error), set extracted_text to a user-friendly error message instead of leaving it None. This ensures the LLM tells the user what went wrong. - On Telegram download failure, set extracted_text with the error so the user sees feedback even when the file never reaches the extraction middleware. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * feat: store extracted document text in workspace memory for search/recall After document extraction succeeds, write the extracted text to workspace memory at `documents/{date}/{filename}`. This enables: - Full-text and semantic search over past uploaded documents - Cross-conversation recall ("what did that PDF say?") - Automatic chunking and embedding via the workspace pipeline Documents are stored with metadata header (uploader, channel, date, MIME type). Error messages (extraction failures) are not stored — only successful extractions. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: CI failures — formatting, unused assignment warning - Run cargo fmt on document_extraction and agent_loop modules - Suppress unused_assignments warning on trace_llm_ref (used only behind #[cfg(feature = "libsql")]) [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: address PR review comments — security, correctness, and code quality Security fixes: - Remove SSRF-prone download() from DocumentExtractionMiddleware (#13) - Sanitize filenames in workspace path to prevent directory traversal (#11) - Pre-check file size before reading in WASM wrapper to prevent OOM (#2) - Percent-encode file_id in Telegram source URLs (#7) Correctness fixes: - Clear image_content_parts on turn end to prevent memory leak (#1) - Find first *successful* transcription instead of first overall (#3) - Enforce data.len() size limit in document extraction (#10) - Use UTF-8 safe truncation with char_indices() (#12) Robustness & code quality: - Add 120s timeout to OpenAI Whisper HTTP client (#5) - Trim trailing slash from Whisper base_url (#6) - Allow ~/.ironclaw/ paths in WASM wrapper (#8) - Return error from on_broadcast in Slack/Discord/WhatsApp (#9) - Fix doc comment in HTTP tool (#4) Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: formatting — cargo fmt Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: address latest PR review — doc comments, error messages, version bumps - Fix DocumentExtractionMiddleware doc comment (no longer downloads from source_url) - Fix error message: "no inline data" instead of "no download URL" - Log error + fallback instead of silent unwrap_or_default on Whisper HTTP client - Bump all capabilities.json versions from 0.1.0 to 0.2.0 to match Cargo.toml Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: remove unsupported profile: minimal from CI workflows [skip-regression-check] dtolnay/rust-toolchain@stable does not accept the 'profile' input (it was a parameter for the deprecated actions-rs/toolchain action). Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: merge with latest main — resolve compilation errors and PR review nits - Add version: None to RegistryEntry/InstalledExtension test constructors - Fix MessageContent type mismatches in nearai_chat tests (String → MessageContent::Text) - Fix .contains() calls on MessageContent — use .as_text().unwrap() - Remove redundant trace_llm_ref = None assignment in test_rig - Check data size before clone in document extraction to avoid unnecessary allocation [skip-regression-check] Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
theredspoon
pushed a commit
that referenced
this pull request
Jun 21, 2026
* feat(agent): queue and merge messages during active turns
Replace the hard rejection ("Turn in progress") when messages arrive
during an active turn with a bounded queue (max 10) that auto-drains
after the turn completes.
Queued messages are merged with newlines into a single turn so the LLM
receives full context from rapid consecutive inputs instead of producing
fragmented responses from partial context.
Key changes:
- Thread.pending_messages (VecDeque) with queue_message/drain_pending_messages
- Drain loop in agent_loop.rs merges all queued messages per iteration
- interrupt() and /clear both clear the pending queue
- MAX_PENDING_MESSAGES constant with cap enforced inside queue_message()
- Drain loop continues on soft errors, stops on NeedApproval/Interrupted
- Drain loop logs respond() failures instead of silently swallowing them
Fixes nearai#259 — debounces rapid inbound messages during processing
Fixes nearai#826 — drain loop is bounded by MAX_PENDING_MESSAGES cap
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: address PR review — drain loop busy-loop guard and stale state re-check
- Add Ok(SubmissionResult::Ok) to drain loop break conditions to prevent
a tight busy-loop if process_user_input returns a queued-ack (e.g. from
a corrupted/hydrated session stuck in Processing state)
- Re-check thread.state under the mutable lock in the Processing arm to
guard against the turn completing between the snapshot read and the
queue operation
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: clear attachments on drain-loop queued message processing
Queued messages are text-only (queued as strings during Processing
state). The drain loop was reusing the original IncomingMessage
reference which carried the first message's attachments, causing
augment_with_attachments to incorrectly re-apply them to unrelated
queued text. Clone the message with cleared attachments for drain-loop
turns.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: address PR review round 2 — stale state fallthrough and thread-not-found guard
- Processing arm: when re-checked state is no longer Processing, fall
through to normal processing instead of dropping user input
- Processing arm: return error when thread not found instead of false
"queued" ack
- Document intermediate drain-loop responses as best-effort for one-shot
channels (HttpChannel)
- Add regression tests for both edge cases
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: address PR review feedback for message queue drain loop
[skip-regression-check] — test modifications present but hook has
SIGPIPE/pipefail false negative when awk exits early on match
- Replace wildcard match in drain loop with explicit `while let
Ok(Response)` guard — stops on Error variant too, preventing
confusing interleaved output after soft errors (review issue #1)
- Reject queueing messages with attachments during Processing state
instead of silently dropping them (review issue #2)
- Document response routing limitation: all drain-loop responses
route via original message identity (review issue #3)
- Document why SubmissionResult::Ok is correct for queued ack and
how it interacts with drain loop break condition (review issue #4)
- Rewrite two dead regression tests to assert actual behavior:
thread-gone returns error, state-changed does not queue (review #5)
- Document MAX_PENDING_MESSAGES=10 as acceptable for personal
assistant use case (review issue #6)
- Fix misleading one-shot channel comment — HttpChannel consumes
sender on first call, subsequent calls are dropped (review issue #8)
- Simplify drain loop intermediate response since while-let guard
guarantees Response variant
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: add missing extension_manager field in webhook EngineContext
The fire_webhook method's EngineContext initializer was missing the
extension_manager field added in staging, causing CI compilation failure.
[skip-regression-check]
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: gate TestRig::session_manager() behind libsql feature flag
The field is #[cfg(feature = "libsql")] so the accessor must match.
All callers are already inside #[cfg(feature = "libsql")] blocks.
[skip-regression-check]
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: re-queue drained messages on drain loop failure
If process_user_input fails after drain_pending_messages() removed
all queued content, that user input was permanently lost. Now the
merged content is re-queued at the front of pending_messages on any
non-Response result so it will be processed on the next successful
turn.
Adds Thread::requeue_drained() helper and unit test.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: remove unreachable!() from drain loop, add lock-drop comments
- Extract content binding in `while let` pattern instead of using a
separate match with unreachable!() — satisfies the no-panic-in-
production convention (zmanian review item #1)
- Add comment clarifying session lock is dropped at Processing arm
boundary before fall-through (zmanian review item #5)
- Document bounded cap overshoot on requeue_drained (review item #2)
[skip-regression-check]
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(security): validate queued messages and touch updated_at on queue ops
- Run safety validation, policy checks, and secret scanning on
messages before queueing during Processing state. Previously,
content with leaked secrets could be stored in pending_messages
and serialized without hitting the inbound scanner.
- Touch updated_at in queue_message(), drain_pending_messages(),
and requeue_drained() so thread timestamps reflect queue activity.
[skip-regression-check] — safety validation requires full Agent;
updated_at is a data-level fix on existing tested methods
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
theredspoon
pushed a commit
that referenced
this pull request
Jun 21, 2026
…i-tenant isolation (nearai#1626) * feat: complete multi-tenant isolation — per-user budgets, model selection, heartbeat cycling Finishes the remaining isolation work from phases 2–4 of #59: Phase 2 (DB scoping): Fix /status and /list commands to use _for_user DB variants instead of global queries that leaked cross-user job data. Phase 3 (Runtime isolation): Per-user workspace in routine engine's spawn_fire so lightweight routines run in the correct user context. Per-user daily cost tracking in CostGuard with configurable budget via MAX_COST_PER_USER_PER_DAY_CENTS. Multi-user heartbeat that cycles through all users with routines, auto-detected from GATEWAY_USER_TOKENS. Phase 4 (Provider/tools): Per-user model selection via preferred_model setting — looked up from SettingsStore on first iteration, threaded through ReasoningContext.model_override to CompletionRequest. Works with providers that support per-request model overrides (NearAI). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: use selected_model setting key to match /model command persistence The dispatcher was reading "preferred_model" but the /model command (merged from staging) persists to "selected_model". Since set_setting is already per-user scoped, using the same key makes /model work as the per-user model override in multi-tenant mode. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: heartbeat hygiene, /model multi-tenant guard, RigAdapter model override Three follow-up fixes for multi-tenant isolation: 1. Multi-user heartbeat now runs memory hygiene per user before each heartbeat check, matching single-user heartbeat behavior. 2. /model command in multi-tenant mode only persists to per-user settings (selected_model) without calling set_model() on the shared LlmProvider. The per-request model_override in the dispatcher reads from the same setting. Added multi_tenant flag to AgentConfig (auto-detected from GATEWAY_USER_TOKENS). 3. RigAdapter now supports per-request model overrides by injecting the model name into rig-core's additional_params. OpenAI/Anthropic/Ollama API servers use last-key-wins for duplicate JSON keys, so the override takes effect via serde's flatten serialization order. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address PR review — cost model attribution, heartbeat concurrency, pruning Fixes from review comments on nearai#1614: - Cost tracking now uses the override model name (not active_model_name) when a per-user model override is active, for accurate attribution. - Multi-user heartbeat runs per-user checks concurrently via JoinSet instead of sequentially, preventing one slow user from blocking others. - Per-user failure counts tracked independently; users exceeding max_failures are skipped (matching single-user semantics). - per_user_daily_cost HashMap pruned on day rollover to prevent unbounded growth in long-lived deployments. - Doc comment fixed: says "routines" not "active routines". Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: /status ownership, model persistence scoping, heartbeat robustness Addresses second round of PR review on nearai#1614: - /status <job_id> DB path now validates job.user_id == requesting user before returning data (was missing ownership check, security fix). - persist_selected_model takes user_id param instead of owner_id, and skips .env/TOML writes in multi-tenant mode (these are shared global files). handle_system_command now receives user_id from caller. - JoinSet collection handles Err(JoinError) explicitly instead of silently dropping panicked tasks. - Notification forwarder extracts owner_id from response metadata in multi-tenant mode for per-user routing instead of broadcasting to the agent owner. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: cost pricing, fire_manual workspace, heartbeat concurrency cap Round 3 review fixes: - Cost tracking passes None for cost_per_token when model override is active, letting CostGuard look up pricing by model name instead of using the default provider's rates (serrrfirat). - fire_manual() now uses per-user workspace, matching spawn_fire() pattern (serrrfirat). - Removed MULTI_TENANT env var — multi-tenant mode is auto-detected solely from GATEWAY_USER_TOKENS presence (serrrfirat + Copilot). - Multi-user heartbeat capped at 8 concurrent tasks to avoid flooding the LLM provider (serrrfirat + Copilot). - Fixed inject_model_override doc comment accuracy (Copilot). - Added comment explaining multi-tenant notification routing priority (Copilot). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat: user-scoped webhook endpoint for multi-tenant isolation Adds POST /api/webhooks/u/{user_id}/{path} — a user-scoped webhook endpoint that filters the routine lookup by user_id, preventing cross-user webhook triggering when paths collide. The existing /api/webhooks/{path} endpoint remains unchanged for backward compatibility in single-user deployments. Changes: - get_webhook_routine_by_path gains user_id: Option<&str> param - Both postgres and libsql implementations add AND user_id = ? filter when user_id is provided - New webhook_trigger_user_scoped_handler extracts (user_id, path) from URL and passes to shared fire_webhook_inner logic - Route registered on public router (webhooks are called by external services that can't send bearer tokens) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(db): add UserStore trait with users, api_tokens, invitations tables Foundation for DB-backed user management (nearai#1605): - UserRecord, ApiTokenRecord, InvitationRecord types in db/mod.rs - UserStore sub-trait (17 methods) added to Database supertrait - PostgreSQL migration V14__users.sql (users, api_tokens, invitations) - libSQL schema + incremental migration V14 - Full implementations for both PgBackend (via Store delegation) and LibSqlBackend (direct SQL in libsql/users.rs) - authenticate_token JOINs api_tokens+users with active/non-revoked checks; has_any_users for bootstrap detection Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(web): DB-backed auth, user/token/invitation API handlers Adds the web gateway layer for DB-backed user management (nearai#1605): Auth refactor: - CombinedAuthState wraps env-var tokens (MultiAuthState) + optional DbAuthenticator for DB-backed token lookup with LRU cache (60s TTL, 1024 max entries) - auth_middleware tries env-var tokens first, then DB fallback - From<MultiAuthState> impl for backward compatibility - main.rs wires with_db_auth when database is available API handlers (12 new endpoints): - /api/admin/users — CRUD: create, list, detail, update, suspend, activate - /api/tokens — create (returns plaintext once), list, revoke - /api/invitations — create, list, accept (creates user + first token) Token creation: 32 random bytes → hex plaintext, SHA-256 hash stored. Invitation accept: validates hash + pending + not expired, creates user record and first API token atomically. All test files updated for CombinedAuthState type change. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat: startup env-var user migration + UserStore integration tests Completes the DB-backed user management feature (nearai#1605): - Startup migration: when GATEWAY_USER_TOKENS is set and the users table is empty, inserts env-var users + hashed tokens into DB. Logs deprecation notice when DB already has users. - hash_token made pub for reuse in migration code. - 10 integration tests for UserStore (libsql file-backed): - has_any_users bootstrap detection - create/get/get_by_email/list/update user lifecycle - token create → authenticate → revoke → reject cycle - suspended user tokens rejected - wrong-user token revoke returns false - invitation create → accept → user created - record_login and record_token_usage timestamps - libSQL migration: removed FK constraints from V14 (incompatible with execute_batch inside transactions). Tables in both base SCHEMA and incremental migration for fresh and existing databases. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * refactor: remove GATEWAY_USER_TOKENS, fix review feedback GATEWAY_USER_TOKENS never went to production — replaced entirely by DB-backed user management via /api/admin/users and /api/tokens. Removed: - UserTokenConfig struct and GATEWAY_USER_TOKENS env var parsing - user_tokens field from GatewayConfig - GatewayChannel::new_multi_auth() constructor - Env-var user migration block in main.rs (~90 lines) - multi_tenant auto-detection from GATEWAY_USER_TOKENS (now runtime via db.has_any_users() in app.rs) Review fixes (zmanian): - User ID generation: UUID instead of display-name derivation (#1) - Invitation accept moved to public router (no auth needed) (#3) - libSQL get_invitation_by_hash aligned with postgres: filters status='pending' AND expires_at > now (#4) - UUID parse: returns DatabaseError::Serialization instead of unwrap_or_default (#7) - PostgreSQL SELECT * replaced with explicit column lists (#8) - Sort order aligned (both backends use DESC) (#6) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat: add role-based access control (admin/member) Adds a `role` field (admin|member) to user management: Schema: - `role TEXT NOT NULL DEFAULT 'member'` added to users table in both PostgreSQL V14 migration and libSQL schema/incremental migration - UserRecord gains `role: String` field - UserIdentity gains `role: String` field, populated from DB in DbAuthenticator and defaulting to "admin" for single-user mode Access control: - AdminUser extractor: returns 403 Forbidden if role != "admin" - /api/admin/users/* handlers: require AdminUser (create, list, detail, update, suspend, activate) - POST /api/invitations: requires AdminUser (only admins can invite) - User creation accepts optional "role" param (defaults to "member") - Invitation acceptance creates users with "member" role Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(web): add Users admin tab to web UI Adds a Users tab to the web gateway UI for managing users, tokens, and roles without needing direct API calls. Features: - User list table with ID, name, email, role, status, created date - Create user form with display name, email, role selector - Suspend/activate actions per user - Create API token for any user (shows plaintext once with copy button) - Role badges (admin highlighted, member muted) - Non-admin users see "Admin access required" message - Keyboard shortcut: Cmd/Ctrl+5 switches to Users tab CSS: - Reuses routines-table styles for the user list - Badge, token-display, btn-small, btn-danger, btn-primary components Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: move Users to Settings subtab, bootstrap admin user on first run - Moved Users from top-level tab to Settings sidebar subtab (under Skills, before Theme toggle) - On first startup with empty users table, automatically creates an admin user from GATEWAY_USER_ID config with a corresponding API token from GATEWAY_AUTH_TOKEN. This ensures the owner appears in the Users panel immediately. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: user creation shows token, + Token works, no password save popup Three UI/UX fixes: 1. Create user now generates an initial API token and shows it in a copy-able banner instead of triggering the browser's password save dialog. Uses autocomplete="off" and type="text" for email field. 2. "+ Token" button works: exposed createTokenForUser/suspendUser/ activateUser on window for inline onclick handlers in dynamically generated table rows. Token creation uses showTokenBanner helper. 3. Admin token creation: POST /api/tokens now accepts optional "user_id" field when the requesting user is admin, allowing token creation for other users from the Users panel. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: use event delegation for user action buttons (CSP compliance) Inline onclick handlers are blocked by the Content-Security-Policy (script-src 'self' without 'unsafe-inline'). Switched to data-action attributes with a delegated click listener on the users table. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: add i18n for Users subtab, show login link on user creation - Added 'settings.users' i18n key for English and Chinese - Token banner now shows a full login link (domain/?token=xxx) with a Copy Link button, plus the raw token below - Login link works automatically via existing ?token= auto-auth Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: token hash mismatch — hash hex string, not raw bytes Critical auth bug: token creation hashed the raw 32 bytes (hasher.update(token_bytes)) but authentication hashed the hex-encoded string (hash_token(candidate) where candidate is the hex string the user sends). This meant newly created tokens could never authenticate. Fixed all 4 token creation sites (users, tokens, invitations create, invitations accept) to use hash_token(&plaintext_token) which hashes the hex string consistently with the auth lookup path. Removed now-unused sha2::Digest imports from handlers. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * refactor: remove invitation system The invitation flow is redundant — admin create user already generates a token and shows a login link. Invitations add complexity without value until email integration exists. Removed: - InvitationRecord struct and 4 UserStore trait methods - invitations table from V14 migration (postgres + both libsql schemas) - PostgreSQL Store methods (create/get/accept/list invitations) - libSQL UserStore invitation methods + row_to_invitation helper - invitations.rs handler file (212 lines) - /api/invitations routes (create, list, accept) - test_invitation_lifecycle test Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat: user deletion, self-service profile, per-user job limits, usage API Four multi-tenancy improvements: 1. User deletion cascade (DELETE /api/admin/users/{id}): Deletes user and all data across 11 user-scoped tables (settings, secrets, routines, memory, jobs, conversations, etc.). Admin only. 2. Self-service profile (GET/PATCH /api/profile): Users can read and update their own display_name and metadata without admin privileges. 3. Per-user job concurrency (MAX_JOBS_PER_USER env var): Scheduler checks active_jobs_for(user_id) before dispatch. Prevents one user from exhausting all job slots. 4. Usage reporting (GET /api/admin/usage?user_id=X&period=day|week|month): Aggregates LLM costs from llm_calls via agent_jobs.user_id. Returns per-user, per-model breakdown of calls, tokens, and cost. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat: add TenantCtx for compile-time tenant isolation Implements zmanian's architectural proposal from nearai#1614 review: two-tier scoped database access (TenantScope/AdminScope) so handler code cannot accidentally bypass tenant scoping. TenantScope (default): wraps user_id + Arc<dyn Database>, auto-binds user_id on every operation. ID-based lookups return None for cross- tenant resources. No escape hatch — forgetting to scope is a compile error. AdminScope (explicit opt-in): cross-tenant access for system-level components (heartbeat, routine engine, self-repair, scheduler, worker). TenantCtx bundles TenantScope + workspace + cost guard + per-user rate limiting. Constructed once per request in handle_message, threaded through all command handlers and ChatDelegate. Key changes: - New src/tenant.rs (~920 lines): TenantScope, AdminScope, TenantCtx, TenantRateState, TenantRateRegistry - All command handlers: user_id: &str → ctx: &TenantCtx - ChatDelegate: cost check/record/settings via self.tenant - System components: store field changed to AdminScope - Config: TENANT_MAX_LLM_CONCURRENT, TENANT_MAX_JOBS_CONCURRENT env vars - Fixes bug: /status <job_id> cross-tenant leak (now auto-filtered) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address PR nearai#1626 review feedback — bounded LRU cache, admin auth, FK cleanup - Replace HashMap with lru::LruCache in DbAuthenticator so the token cache is hard-bounded at 1024 entries (evicts LRU, not just expired) - Gate admin user endpoints (list/detail/update/suspend/activate) with AdminUser extractor so members get 403 instead of full access - Add api_tokens to libSQL delete_user cleanup list to prevent orphaned tokens (libSQL has no FK cascade) - Add regression tests for all three fixes Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: update CA certificates in runtime Docker image Ensures the root certificate bundle is current so TLS handshakes to services like Supabase succeed on Railway. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: resolve CI failures — formatting, no-panics check - Run cargo fmt on test code - Replace .expect() with const NonZeroUsize in DbAuthenticator - Add // safety: comments for test-only code in multi_tenant.rs Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: switch PostgreSQL TLS from rustls to native-tls rustls with rustls-native-certs fails TLS handshake on Railway's slim container (empty or stale root cert store). native-tls delegates to OpenSSL on Linux which handles system certs more reliably. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * Adding user management api * feat: admin secrets provisioning API + API documentation - Add PUT/GET/DELETE /api/admin/users/{id}/secrets/{name} endpoints for application backends to provision per-user secrets (AES-256-GCM encrypted) - Add secrets_store field to GatewayState with builder wiring - Create docs/USER_MANAGEMENT_API.md with full API spec covering users, secrets, tokens, profile, and usage endpoints - Update web gateway CLAUDE.md route table Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: add CatchPanicLayer to capture handler panics Without this, panics in async handlers silently drop the connection and the edge proxy returns a generic 503. Now panics are caught, logged, and returned as 500 with the panic message. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address second-round review — transactional delete, overflow, error logging - C1: Wrap PostgreSQL delete_user() in a transaction so partial cleanup can't leave users in a half-deleted state - M2: Add job_events to delete cleanup (both backends) — FK to agent_jobs without CASCADE would cause FK violation - H1/M4: Cap expires_in_days to 36500 before i64 cast (tokens + secrets) - H2: Validate target user exists before creating admin token to prevent orphan tokens on libSQL - H3: Log DB errors in DbAuthenticator::authenticate() instead of silently swallowing them as 401 Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: revert to rustls with webpki-roots fallback for PostgreSQL TLS native-tls/OpenSSL caused silent crashes (segfaults in C code) during DB writes on Railway containers. Switch back to rustls but add webpki-roots as a fallback when system certs are missing, which was the original TLS handshake failure on slim container images. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * chore: update Cargo.lock for rustls + webpki-roots Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * debug: add /api/debug/db-write endpoint to diagnose user insert failure Temporary diagnostic endpoint that tests DB INSERT to users table with full error logging. No auth required. Will be removed after debugging. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * perf: use cargo-chef in Dockerfile for dependency caching Splits the build into planner/deps/builder stages. Dependencies are only recompiled when Cargo.toml or Cargo.lock change. Source-only changes skip straight to the final build stage. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * debug: add tracing to users_create_handler Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: guard created_by FK in user creation handler The auth identity user_id (from owner_id scope) may not match any user row in the DB, causing a FK violation on the created_by column. Check that the referenced user exists before setting created_by. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * refactor: collapse GATEWAY_USER_ID into IRONCLAW_OWNER_ID Remove the separate GATEWAY_USER_ID config. The gateway now uses IRONCLAW_OWNER_ID (config.owner_id) directly for auth identity, bootstrap user creation, and workspace scoping. Previously, with_owner_scope() rebinds the auth identity to owner_id while keeping default_sender_id as the gateway user_id. This caused a FK constraint violation when creating users because the auth identity ("default") didn't match any user in the DB ("nearai"). Changes: - Remove GATEWAY_USER_ID env var and gateway_user_id from settings - Remove user_id field from GatewayConfig - Add owner_id parameter to GatewayChannel::new() - Remove with_owner_scope() method - Remove default_sender_id from GatewayState - Remove sender override logic in chat/approval handlers - Remove debug endpoint and tracing from prior debugging - Update all tests and E2E fixtures Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: hide Users tab for non-admins, remove auth hint text - Fetch /api/profile after login and hide the Users settings tab when the user's role is not admin - Remove the "Enter the GATEWAY_AUTH_TOKEN" hint from the login page since tokens are now managed via the admin panel, not .env files Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address review feedback (auth 503, token expiry, CORS PATCH) - DB auth errors now return 503 instead of 401 so outages are distinguishable from invalid tokens (serrrfirat H3) - Cap expires_in_days to 36500 before i64 cast to prevent negative duration from u64 overflow (serrrfirat H1) - Add PATCH to CORS allowed methods for profile/user update endpoints (Copilot) - Stop leaking panic details in CatchPanicLayer response body Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: harden multi-tenant isolation — review fixes from nearai#1614 - Add conversation ownership checks in TenantScope: add_conversation_message, touch_conversation, list_conversation_messages (+ paginated), update_conversation_metadata_field, get_conversation_metadata now return NotFound for conversations not owned by the tenant (cross-tenant data leak) - Fix multi-user heartbeat: clear notify_user_id per runner so notifications persist to the correct user, not the shared config target - Move hygiene tasks into bounded JoinSet instead of unbounded tokio::spawn - Revert send_notification to private visibility (only used within module) - Use effective_model_name() for cost attribution in dispatcher so providers that ignore per-request model overrides report the actual model used - Fix inject_model_override doc comment; add 3 unit tests - Fix heartbeat doc comment ("routines" not "active routines") Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat: add Jobs, Cost, Last Active columns to admin Users table Add UserSummaryStats struct and user_summary_stats() batch query to the UserStore trait (both PostgreSQL and libSQL backends). The admin users list endpoint now fetches per-user aggregates (job count, total LLM spend, most recent activity) in a single query and includes them inline in the response. The frontend Users table displays three new columns. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address review comments and CI formatting failures CI fixes: - cargo fmt fixes in cli/mod.rs and db/tls.rs Security/correctness (from Copilot + serrrfirat + pranavraja99 reviews): - Token create: reject expires_in_days > 36500 with 400 instead of silent clamp - Token create: return 404 when admin targets non-existent user - User create: map duplicate email constraint violations to 409 Conflict - User create: remove unnecessary DB roundtrip for created_by (use AdminUser directly) - DB auth: log warn on DB lookup failures instead of silently swallowing errors - libSQL: add FK constraints on users.created_by and api_tokens.user_id Config fixes: - agent.multi_tenant: resolve from AGENT_MULTI_TENANT env var instead of hardcoding false - heartbeat.multi_tenant: fix doc comment to match actual env-var-based behavior UI fix: - showTokenBanner: pass correct title ("Token created!" vs "User created!") Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address remaining review comments (round 2) - Secrets handlers: normalize name to lowercase before store operations, validate target user_id exists (returns 404 if not found) - libSQL: propagate cost parsing errors instead of unwrap_or_default() in both user_usage_stats and user_summary_stats - users_list_handler: propagate user_summary_stats DB errors (was silently swallowed with unwrap_or_default) - loadUsers: distinguish 401/403 (admin required) from other errors - Docs: fix users.id type (TEXT not UUID), remove "invitation flow" from V14 migration comment Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat: i18n for Users tab, atomic user+token creation, transactional delete_user i18n: - Add 31 translation keys for all Users tab strings (en + zh-CN) - Wire data-i18n attributes on HTML elements (headings, buttons, inputs, table headers, empty state) - Replace all hard-coded strings in app.js with I18n.t() calls Atomic user+token creation: - Add create_user_with_token() to UserStore trait - PostgreSQL: wraps both INSERTs in conn.transaction() with auto-rollback - libSQL: wraps in explicit BEGIN/COMMIT with ROLLBACK on error - Handler uses single atomic call instead of two separate operations Transactional delete_user for libSQL: - Wrap multi-table DELETE cascade in BEGIN/COMMIT transaction - ROLLBACK on any error to prevent partial cleanup / inconsistent state - Matches the PostgreSQL implementation which already used transactions Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: revert V14 migration to match deployed checksum [skip-regression-check] Refinery checksums applied migrations — editing V14__users.sql after it was already applied causes deployment failures. Revert the cosmetic comment changes (added in df40b22) to restore the original checksum. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: bootstrap onboarding flow for multi-tenant users The bootstrap greeting and workspace seeding only ran for the owner workspace at startup, so new users created via the admin API never received the welcome message or identity files (BOOTSTRAP.md, SOUL.md, AGENTS.md, USER.md, etc.). Three fixes: - tenant_ctx(): seed per-user workspace on first creation via seed_if_empty(), which writes identity files and sets bootstrap_pending when the workspace is truly fresh - handle_message(): check take_bootstrap_pending() on the tenant workspace (not the owner workspace) and persist the greeting to the user's own assistant conversation + broadcast via SSE - WorkspacePool: seed new per-user workspaces in the web gateway so memory tools also see identity files immediately The existing single-user bootstrap in Agent::run() is preserved for non-multi-tenant deployments. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address remaining PR review comments (round 3) - Docs: fix metadata description from "merge patch" to "full replacement" - Secrets: reject expires_in_days > 36500 with 400 (was silently clamped) - libSQL: CAST(SUM(cost) AS TEXT) in user_usage_stats and user_summary_stats to prevent SQLite numeric coercion from crashing get_text() — this was the root cause of the Copilot "SUM returns numeric type" comments - Add 3 regression tests: user_summary_stats (empty + with data) and user_usage_stats (multi-model aggregation) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat: add role change support for users (admin/member toggle) - Add update_user_role() to UserStore trait + both backends (PostgreSQL and libSQL) - Extend PATCH /api/admin/users/{id} to accept optional "role" field with validation (must be "admin" or "member") - Add "Make Admin" / "Make Member" toggle button in Users table actions - Add i18n keys for role change (en + zh-CN) - Update API docs to document the role field on PATCH - Fix test helpers to use fmt_ts() for timestamps (was using SQLite datetime('now') which produces incompatible format for string comparison) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: show live LLM spend in Users table instead of only DB-recorded costs [skip-regression-check] Chat turns record LLM cost in CostGuard (in-memory) but don't create agent_jobs/llm_calls DB rows — those are only written for background jobs. The Users table was querying only from DB, so it showed $0.00 for users who only chatted. Now supplements DB stats with CostGuard.daily_spend_for_user() — the same source displayed in the status bar token counter. Shows whichever is larger (DB historical total vs live daily spend). Also falls back to last_login_at for "Last Active" when no DB job activity exists. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: persist chat LLM calls to DB and fix usage stats query Two root causes for zero usage stats: 1. ChatDelegate only recorded LLM costs to CostGuard (in-memory) — never to the llm_calls DB table. Added DB persistence via TenantScope.record_llm_call() after each chat LLM call, with job_id=NULL and conversation_id=thread_id. 2. user_summary_stats query only joined agent_jobs→llm_calls, missing chat calls (which have job_id=NULL). Redesigned query to start from llm_calls and resolve user_id via COALESCE(agent_jobs.user_id, conversations.user_id) — covers both job and chat LLM calls. Both PostgreSQL and libSQL queries updated. TenantScope gets record_llm_call() method. Tests updated for new query semantics. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address review comments — input validation, cost semantics, panic safety [skip-regression-check] - Validate display_name: trim whitespace, reject empty strings (create + update) - Validate metadata: must be a JSON object, return 400 if not (admin + profile) - secrets_list_handler: verify target user_id exists before listing - Cost display: use DB total directly (chat calls now persist to DB), remove confusing max(db,live) CostGuard fallback - CatchPanicLayer: truncate panic payload to 200 chars in log to limit potential sensitive data exposure Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address Copilot round 5 — docs, secrets consistency, token name, provider field [skip-regression-check] - Docs: users.id note updated to "typically UUID v4 strings (bootstrap admin may use a custom ID)" - secrets_list_handler: return 503 when DB store is None (was falling through to list secrets without user validation) - tokens_create: trim + reject empty token name (matching display_name pattern) - LlmCallRecord.provider: use llm_backend ("nearai","openai") instead of model_name() which returns the model identifier - user_summary_stats zero-LLM users: acceptable — handler already falls back to 0 cost and last_login_at for missing entries Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: DB auth returns 503 on outage, scheduler counts only blocking jobs From serrrfirat review: - DB auth: return Err(()) on database errors so middleware returns 503 instead of silently returning Ok(None) → 401 (auth miss) - Scheduler: add parallel_blocking_count_for() that uses is_parallel_blocking() (Pending/InProgress/Stuck) instead of is_active() for per-user concurrency — Completed/Submitted jobs no longer count against MAX_JOBS_PER_USER From Copilot: - CLAUDE.md: fix secrets route paths from {id} to {user_id} - token_hash: use .as_slice() instead of .to_vec() to avoid heap allocation on every token auth/creation call Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: immediate auth cache invalidation on security-critical actions (zmanian review #6) Add DbAuthenticator::invalidate_user() that evicts all cached entries for a user. Called after: - Suspend user (immediate lockout, was 60s delay) - Activate user (immediate access restoration) - Role change (admin↔member takes effect immediately) - Token revocation (revoked token can't be reused from cache) The DbAuthenticator is shared (via Clone, which Arc-clones the cache) between the auth middleware and GatewayState, so handlers can evict entries from the same cache the middleware reads. Also from zmanian's review: - Items 1-5, 7-11 were already resolved in prior commits - Item 12 (String→enum for status/role) is deferred as a broader refactor Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: last-admin protection, usage stats for chat calls, UTF-8 safe panic truncation Last-admin protection: - Suspend, delete, and role-demotion of the last active admin now return 409 Conflict instead of succeeding and locking out the admin API - Helper is_last_admin() checks active admin count before destructive ops Usage stats: - user_usage_stats() now includes chat LLM calls (job_id=NULL) by joining via conversations.user_id, matching user_summary_stats() - Both PostgreSQL and libSQL queries updated Panic handler: - Use floor_char_boundary(200) instead of byte-index [..200] to prevent panic on multi-byte UTF-8 characters in panic messages Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: workspace seed race, bootstrap atomicity, email trim, secrets upsert response [skip-regression-check] - WorkspacePool: await seed_if_empty() synchronously after inserting into cache (drop lock first to avoid blocking), so callers see identity files immediately instead of racing a background task - Bootstrap admin: use create_user_with_token() for atomic user+token creation, matching the admin create endpoint - Email: trim whitespace, treat empty as None to prevent " " being stored and breaking uniqueness - Secrets PUT: report "updated" vs "created" based on prior existence - Last token_hash.to_vec() → .as_slice() in authenticate_token Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: disable unscoped webhook endpoint in multi-tenant mode [skip-regression-check] The original /api/webhooks/{path} endpoint looks up routines across all users. In multi-tenant mode, anyone who knows the webhook path + secret could trigger another user's routine. Now returns 410 Gone with a message pointing to the scoped endpoint /api/webhooks/u/{user_id}/{path}. Detection uses state.db_auth.is_some() — present only when DB-backed auth is enabled (multi-tenant). Single-user deployments are unaffected. From: standardtoaster review comment Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: webhook multi-tenant check, secrets error propagation, stale doc comment [skip-regression-check] - Webhook: use workspace_pool.is_some() instead of db_auth.is_some() for multi-tenant detection — db_auth is set for any DB deployment, workspace_pool is only set when has_any_users() was true at startup - Secrets: propagate exists() errors instead of unwrap_or(false) so backend outages surface as 500 rather than incorrect "created" status - Config: fix stale workspace_read_scopes comment referencing user_id Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
theredspoon
pushed a commit
that referenced
this pull request
Jun 21, 2026
…earai#2617) * feat(common): add CredentialName and ExtensionName newtypes Introduce typed identifiers for the backend-secret vs user-facing extension identity split that the Extension/Auth Invariants section of CLAUDE.md describes. Four recent PRs (nearai#2561, nearai#2473, nearai#2512, nearai#2574) have been identity- confusion bugs with the same shape: a stringly-typed value passed through multiple layers with each layer meaning a different thing. Newtypes make each of those a compile error. This is PR 1 of 2. PR 1 lands the newtypes and migrates the core auth seam (ResumeKind::Authentication, MissingCredential, ToolReadiness::NeedsAuth, LatentActionExecution::NeedsAuth, extensions/naming.rs). PR 2 will migrate AppEvent.extension_name, OAuth/pending-flow stores, TUI events, and the remaining extension_name: String fields. Wire format is unchanged — both newtypes use #[serde(transparent)] so on- wire and on-disk representations stay plain strings and legacy persisted rows keep deserializing. Validation runs at explicit construction (::new / ::try_from / ::from_str), not at deserialize time. Also adds .claude/rules/types.md codifying the "no stringly-typed internals" rule. Regression coverage: 17 new unit tests in identity.rs; existing auth_manager, router, and gate tests (130+ cases) all pass unchanged. * fix(common): address PR nearai#2611 review feedback Four fixes from Copilot, Gemini, and Claude reviews: - **identity.rs docs**: drop reference to a non-existent `validate()` re-validation API. Document that instances represent "passed validation at some point in history" rather than "guaranteed valid right now" — by design. - **effect_adapter.rs**: the `awaiting_authorization` / `awaiting_token` gate path was using `CredentialName::from_trusted` to wrap a value read straight out of a tool's JSON output. Tool output is external/untrusted; use `CredentialName::new` (validating) with a cascade: external → tool name → `from_trusted(tool_name)` as final fallback. Closes a credential-name shape-injection vector. - **canonicalize()**: reorder checks cheapest-first against the trimmed slice so invalid inputs reject without allocating a canonicalized `String`. `replace('-', "_")` is deferred until after the structural checks pass; since `-`/`_` are both one byte, the earlier length check stays valid. - **Remove `Deref<Target = str>`** from identity newtypes, keep `AsRef<str>`. Auto-deref let `&cred_name` silently coerce to `&str`, which is exactly the implicit-conversion pattern these newtypes exist to prevent. Callers that had a `&CredentialName` where `&str` was expected now write `.as_str()` explicitly. Added a regression test for the accessor contract and updated the rule template in `.claude/rules/types.md` to document the decision. Declined one review item (Claude): the remaining `to_string()` calls inside `IdentityError` variants are on the exception path; the common invalid-input case no longer allocates twice after the canonicalize reorder, and errors must carry owned strings so they can escape the function. Regression coverage: 5035 lib tests + 18 identity tests (one new — `explicit_accessors_work`) pass. Zero clippy warnings. * feat(common): apply ExtensionName newtype to fan-out sites (PR 2/2) Follow-up to nearai#2611. Migrates the remaining stringly-typed extension_name and credential_name fields to use the ExtensionName and CredentialName newtypes introduced in ironclaw_common::identity. Fields now typed: - AppEvent::{OnboardingState, GateRequired, ExtensionStatus}.extension_name (serde transparent — wire format unchanged) - StatusUpdate::{AuthRequired, AuthCompleted}.extension_name - TuiEvent::{AuthRequired, AuthCompleted}.extension_name (adds ironclaw_common dep to ironclaw_tui) - PendingOAuthLaunchParams.extension_name - PendingOAuthFlow.extension_name - PendingAuth.extension_name, PendingAuthPrompt.extension_name - ParsedAuthData.extension_name, selected_auth_prompt tuple - emit_auth_required_status() and Session::enter_auth_mode() parameters - event_from_configure_result() parameter - resolve_extension_for_action() and resolve_auth_gate_display_name() return types - normalize_extension_name() return type PendingAuthPrompt::new is now infallible (accepts ExtensionName directly) since the identity validator carries the non-empty invariant the constructor used to re-check. The "blank extension name" rejection test moved out — that logic lives in ironclaw_common::identity tests. Test updates use `ExtensionName::new("...").unwrap()` at construction sites and `from_trusted(...)` where a trusted upstream string is being adapted. Every site is a compile-time audit of where the type was crossing a boundary untyped. Regression coverage: existing 5034 lib tests + 26 engine_v2_gate integration tests + 40 ironclaw_common tests all pass. Zero clippy warnings across all features. * fix(web): return ExtensionName from pending_gate_extension_name Addresses Claude's review comment on nearai#2611: the function was doing `Some(credential_name.as_str().to_string())` in the fallback branch, defeating the newtype's purpose by re-stringifying the identity. Return `Option<ExtensionName>` instead. Plumbs through `PendingGateInfo. extension_name` (wire format unchanged — `#[serde(transparent)]`). The fallback path's cross-identity conversion (credential name → extension name) is now an explicit `ExtensionName::from_trusted` call, making the boundary crossing visible at the call site. Also fixes the `Deref<Target = str>` removal fallout that followed the rebase onto the updated PR 1: call sites that relied on auto-deref (`ext.contains(...)`, `auth_manager.submit_auth_token(&cred_name, ...)`) now explicitly call `.as_str()`. * fix(router,web): address PR nearai#2617 review feedback Four Gemini review comments, all on the boundary between credential/ extension identifiers and user input. 1. [HIGH, security] extensions_setup_submit_handler was wrapping the URL path segment in ExtensionName::from_trusted, which skips the newtype's path-traversal / invalid-character validation. That path is user-controlled (`/api/extensions/{name}/setup`). Validate with ExtensionName::new at the handler entry and return 400 on failure; downstream uses switch to .as_str() or .clone() of the validated value, and the three in-handler from_trusted sites disappear. 2. Rename resolve_auth_gate_display_name -> resolve_auth_gate_extension_name. The function returns an identifier/slug, not a human-readable display name — the old name was a leftover from when the value was a String. 3. Return Option<ExtensionName> from the renamed function. Previously the non-Authentication gate branch fabricated an ExtensionName::from_trusted(pending.action_name), which was semantically wrong (an action name is not an extension identifier) and silently defeated the type's invariants. Now it returns None for Approval/External gates, and callers thread an Option through. send_pending_gate_status accepts Option<&ExtensionName> and only uses it on the Authentication arm, with a warn! log if upstream plumbing ever reaches the arm with None. The GateRequired SSE event's extension_name is now a clean .clone() of the Option. 4. Rename auth_display_name -> extension_name on send_pending_gate_status so the parameter name matches both its type and the StatusUpdate::AuthRequired.extension_name field it feeds. Regression: new test_extensions_setup_submit_rejects_path_traversal_name at the handler tier (per .claude/rules/testing.md "Test Through the Caller, Not Just the Helper") drives the handler with malformed path segments and asserts 400 before the value reaches extension lookup or any from_trusted wrap. 5035 lib tests pass, zero clippy warnings. * docs(identity): codify web-boundary rules + add static check Three rule additions + one enforcement hook covering the identity boundary that PR nearai#2617 review uncovered: - src/channels/web/CLAUDE.md — extend "Unified Extension Onboarding" with explicit rules: * Setup/configure/activate routes MUST validate `{name}` via `ExtensionName::new` at handler entry (return 400 on failure). * Web DTOs and handlers MUST NOT reference `CredentialName` — credential identity is backend-only; the dispatcher/auth_manager resolves it from the ExtensionName server-side. * Auth-flow extension resolution happens in *one* place (`AuthManager::resolve_extension_name_for_auth_flow`). Wrappers are thin and delegate; they must not duplicate the precedence logic or re-derive from credential prefixes. The four recent identity bugs (nearai#2561, nearai#2473, nearai#2512, nearai#2574) were duplicate- resolution drift. - src/bridge/CLAUDE.md — new module spec documenting auth_manager.rs as the single authority for auth-flow extension resolution, with the resolver's four-step precedence order and the approved wrapper call sites. - scripts/pre-commit-safety.sh — new check #8 (CREDNAME): flags `CredentialName` references in newly-added production lines under `src/channels/web/**`. Test-mod code is excluded via the existing `strip_test_mod_lines` filter. Suppression via `// web-identity-exempt: <reason>` for the rare legitimate case of reading an already-typed value off a backend struct. Smoke-tested: * baseline (current branch) — no warnings * injected violation — fires with CREDNAME warning * injected violation + `// web-identity-exempt:` — suppressed The rules and the check live at the same level — humans read the rule, CI enforces it. * fix(auth): validate user-influenced names at the resolver boundary Addresses four Copilot review comments on PR nearai#2617 that all pointed at the same seam: the canonical `AuthManager::resolve_extension_name_for_auth_flow` returned a raw `String` whose first branch (the LLM-supplied `name` parameter on `tool_install` / `tool_activate` / `tool_auth` actions) passed through without `ExtensionName` validation. Both call sites then wrapped the result in `ExtensionName::from_trusted`, promoting an unvalidated user-influenced value to a typed identity. - **Resolver now returns `ExtensionName`.** Branch 1 validates the user-controlled name via `ExtensionName::new` and falls through on failure; branches 2–4 use `from_trusted` because their sources (tool registry hint, canonicalizer, typed credential fallback) are already trusted upstream. This consolidates validation in the single "resolve once" site documented in `src/bridge/CLAUDE.md`. - **router.rs and server.rs drop their wraps.** `resolve_extension_for_action` (router) and `pending_gate_extension_name` (server) return the resolver's typed output directly. The tool-registry fallback in router.rs (no-auth-manager path) keeps its `from_trusted` wrap since it operates on the same trusted sources as branch 2. - **`restore_selected_auth_prompt` re-validates rehydrated prompts.** `PendingAuthPrompt` is `#[serde(transparent)]`, so deserialize does not re-check the inner `ExtensionName` string. A legacy-persisted invalid name would previously have been dropped by the old `PendingAuthPrompt::new(String, ...)` empty-string rejection; now `restore_selected_auth_prompt` re-runs `ExtensionName::new` and drops + warns on failure, upgrading the old non-empty-only check to the full identity invariant. New test `test_restore_selected_auth_prompt_rejects_invalid_legacy_row` forges three invalid rows (empty / uppercase / path-traversal) straight through serde and asserts each is dropped. - **Docstring on `PendingAuthPrompt` refreshed.** The old comment claimed `::new` "trims and validates extension_name is non-empty", which is no longer true — `::new` is infallible and the invariant lives in `ExtensionName` itself. The new comment documents the split: validation runs at `ExtensionName::new` construction and at restore-from-persistence, not inside `PendingAuthPrompt`. Regression: 5063 lib tests pass (+1 new). Clippy zero warnings. * fix(ci): adapt post-merge-from-staging sites to ExtensionName Staging shipped nearai#2640 (repl unlock) and gateway refactor commits after my last merge. The CI build picked them up via auto-merge and hit three type mismatches my branch hadn't seen: - src/channels/repl.rs:908 — new test constructs `StatusUpdate::AuthRequired { extension_name: "google_oauth_token" .to_string(), ... }`. Typed field; now `ExtensionName::new(...).unwrap()`. - src/channels/web/server.rs:1405-1424 — staging added a no-auth-manager fallback chain to `pending_gate_extension_name` that returned raw `Some(String)` on three branches. Aligned with `AuthManager::resolve_extension_name_for_auth_flow`: branch 1 (user-influenced `tool_install`/`tool_activate`/`tool_auth` `name` param) validates via `ExtensionName::new` and falls through on failure; branches 2-3 (provider-extension hint, credential-name fallback) use `from_trusted` because they're sourced from typed upstream state. Mirrors the fix applied to the canonical resolver in c813caa. - src/channels/web/server.rs:3831 — test used `.as_deref()` on the function's Option<ExtensionName> return; switched to `.as_ref().map(|n| n.as_str())` matching the pattern from the adjacent test. No new logic — just adapting two staging landings to the typed surface PR nearai#2617 introduces. The validation behaviour for the fallback path is already locked in by the identity-layer tests in `ironclaw_common::identity` (rejects_path_traversal, rejects_uppercase, etc.) and by the regression test added in c813caa (test_restore_selected_auth_prompt_rejects_invalid_legacy_row). [skip-regression-check] — type adaptation to unblock CI, no behaviour change needing its own regression test. Clippy with `-D warnings` clean, 5074 lib tests pass. * fix(auth): extract shared resolver; wrapper delegates instead of duplicating Addresses two Copilot comments on PR nearai#2617 that surfaced the same architectural issue: the no-auth-manager fallback in `pending_gate_extension_name` had grown a three-branch copy of the resolver's precedence that quietly skipped branch 3 (canonicalize action_name + check `ExtensionManager::extension_info`). Exactly the duplicate-resolution drift the "one resolver" rule in `src/bridge/CLAUDE.md` warns against — four prior identity bugs (nearai#2561, nearai#2473, nearai#2512, nearai#2574) were the same pattern. - Extracted `pub(crate) async fn resolve_auth_flow_extension_name` to `src/bridge/auth_manager.rs` as the single site of the four-branch precedence. Takes `Option<&ToolRegistry>` + `Option<&ExtensionManager>` so both the `AuthManager` method (which passes its own fields) and the web wrapper (which passes `state.tool_registry` / `state.extension_manager`) share identical logic. - `AuthManager::resolve_extension_name_for_auth_flow` is now a 1-block delegator. - `pending_gate_extension_name` in `web/server.rs` drops its inline fallback entirely and calls the shared free function from both branches. The bare-test-harness path now runs branch 3 (canonicalize + installed-extension check) that it previously missed. - Updated `src/bridge/CLAUDE.md` to document the free function as the single authority, the three approved wrappers as thin delegators, and the return type as `ExtensionName` (was stale `String` from the pre-c813caa9 era). Regression coverage: the existing `resolve_extension_name_for_auth_flow_prefers_installed_channel_name` test passes unchanged — it exercises branch 3 through the method, which now reaches it via the extracted free function. * Merge remote-tracking branch 'origin/staging' into feat/identity-newtypes-pr2 Picks up nearai#2644 (platform/ extraction) and nearai#2645 (features/oauth/ move). Manual resolutions: - src/channels/web/server.rs: staging removed 720 lines of OAuth callback code (moved to features/oauth/mod.rs in nearai#2645). My PR 2 ExtensionName changes to two of those functions (oauth_callback_handler, slack_relay_oauth_callback_handler) ported to the new location. - src/bridge/auth_manager.rs: extended the shared resolver's branch-1 action pattern to include 'tool-activate' and 'tool-auth' variants, matching staging's new pending_gate_extension_name_uses_install_parameters_for_hyphenated_activate_tool test expectation. Underscore + hyphen variants for all three actions. No new PR 2 logic — just aligning the type surface with two staging refactors. 5074 lib tests pass (+1 vs previous — the new staging hyphenated-tool test). Clippy -D warnings clean. * fix(web): address PR nearai#2617 round-3 review feedback Two Copilot findings from the 2026-04-18 review: 1. `/api/extensions/{name}/{activate,remove,setup}` handlers accepted `Path<String>` and forwarded it to the extension manager without validating path-traversal, invalid characters, or case — only `extensions_setup_submit_handler` had the `ExtensionName::new` guard. Applied the same boundary validation to all three siblings. 2. `restore_pending_auth_mode` took `extension_name: &str` and re-wrapped it with `ExtensionName::from_trusted`, re-introducing an unvalidated string boundary even though every caller already held an `ExtensionName` (`pending_auth.extension_name`). Changed the helper to accept `&ExtensionName` so the identity stays typed end-to-end; `from_trusted` is no longer needed here. Regression: added `test_extensions_sibling_handlers_reject_path_traversal_name` covering activate / remove / setup-GET with the same malformed slugs the setup-submit test already locks in (path traversal, slash in segment, uppercase, space, trailing underscore). Drives the handlers through axum routing so the boundary is exercised end-to-end. * fix(ci): adapt replay_outcome to ExtensionName after staging merge Staging nearai#2621 added `tests/support/replay_outcome.rs`, which destructures `StatusUpdate::{AuthRequired,AuthCompleted}.extension_name` into a `String` field of `EventSummary`. This PR made those `StatusUpdate` fields `ExtensionName`, so the post-merge build breaks in the replay snapshot gate and all-features clippy jobs. Convert to `String` at the destructure via `ExtensionName::into()` so the `EventSummary` shape (and the persisted `.snap` files) stay unchanged. The test-support / snapshot wire format is a legitimate String boundary per `.claude/rules/types.md`.
theredspoon
pushed a commit
that referenced
this pull request
Jun 21, 2026
* feat(engine-v2): mount-backend abstraction for per-project sandbox (Phase 1) Adds the engine-side `MountBackend` trait + minimal `WorkspaceMounts` registry and a host-side bridge interceptor that routes sandbox-eligible tool calls (`file_read`, `file_write`, `list_dir`, `apply_patch`, `shell`) through a backend when their path argument starts with `/project/`. Default behavior is unchanged: until `EffectBridgeAdapter::set_workspace_mounts(Some(...))` is called (Phase 6), the interception path is dormant. This is the first phase of the per-project sandbox plan (`docs/plans/2026-04-10-engine-v2-sandbox.md`) and a deliberately small subset of the unified Workspace VFS proposed in nearai#1894 — just enough abstraction so the sandbox can be a `MountBackend` rather than a special case in the bridge. When nearai#1894's full mount table lands, the sandbox backend slots in unchanged. Engine crate (`crates/ironclaw_engine/src/workspace/`): - `mount.rs` — `MountBackend` trait, `MountError` (NotFound / InvalidPath / PermissionDenied / Io / Tool / Backend / Unsupported), `DirEntry`, `EntryKind`, `ShellOutput` - `filesystem.rs` — `FilesystemBackend`: passthrough host-fs implementation with two-layer path validation (lexical reject of absolute / `..`, then symlink-escape canonicalization). `read`/`write`/`list` fully implemented; `patch`/`shell` return `Unsupported` so the bridge falls through to the host tool until Phase 5 - `registry.rs` — `WorkspaceMounts` per-project registry with lazy `ProjectMountFactory`, longest-prefix-match resolution, cached and invalidatable Bridge (`src/bridge/sandbox/`): - `intercept.rs` — `maybe_intercept` and `SANDBOX_TOOL_NAMES`. Returns `Handled(json)` on a successful backend dispatch, `FellThrough` for non-sandbox tools, host paths, missing path params, or `Unsupported` backend ops - `effect_adapter.rs` — `workspace_mounts` field + `set_workspace_mounts` setter; interception block in `execute_action_internal` right before `execute_tool_with_safety`, gated on the optional mount table Tests (31 new): - 17 engine workspace unit tests covering trait error mapping, path safety (lexical + symlink), longest-prefix routing, and lazy factory caching - 9 bridge sandbox unit tests including `intercept_actually_dispatches_into_backend` (counting backend) which proves the interceptor reaches the backend - 5 integration tests in `tests/engine_v2_sandbox_integration.rs` driving `EffectBridgeAdapter::execute_action()` end-to-end per the "Test Through the Caller" rule (`.claude/rules/testing.md`), including a host-path-falls-through test that asserts the sandbox tempdir was not touched, and a `..`-escape test that verifies no `/etc/passwd` content leaks even after safety-layer redaction Drive-by: feature-gate two pre-existing dead-code helpers in `crates/ironclaw_skills/src/parser.rs` on `#[cfg(feature = "registry")]` to match their only call site, fixing a pre-existing clippy warning that blocked the workspace's `-D warnings` policy when `ironclaw_skills` is built with `default-features = false` (as the engine crate does). Verification: - `cargo fmt --check` clean - `cargo clippy --all --benches --tests --examples --all-features` zero warnings - 31 / 31 new tests passing; no existing tests broken Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine-v2): per-project sandbox — Phases 2–7 + live Docker e2e test Completes the per-project sandbox plan (docs/plans/2026-04-10-engine-v2-sandbox.md Phases 2–7), building on Phase 1's mount-backend abstraction (nearai#2211). Phase 2 — Project workspace folder: - `Project.workspace_path: Option<PathBuf>` field + `with_workspace_path()` - Host-side `project_workspace_path()`, `ensure_project_workspace_dir()` (creates `~/.ironclaw/projects/<id>/` mode 0700, idempotent) - `FilesystemMountFactory` taking a `ProjectPathResolver` closure (decoupled from `Store`); wired into `EffectBridgeAdapter` via `set_workspace_mounts()` Phase 3 — Standalone daemon binary: - `src/bin/sandbox_daemon.rs` — NDJSON over stdin/stdout, health/shutdown/execute_tool - Constructs ReadFileTool/WriteFileTool/ListDirTool/ApplyPatchTool/ShellTool with `base_dir=/project` (override via `IRONCLAW_SANDBOX_BASE_DIR`) Phase 4 — Dockerfile.sandbox: - Multi-stage build: rust-slim builder (+ python3 for pyo3) compiles sandbox_daemon; debian-slim runtime with tini PID 1, common build tools, `/project` mount target Phase 5 — ProjectSandboxManager + ContainerizedFilesystemBackend: - protocol.rs: Request/Response/RpcError matching daemon wire format - transport.rs: `SandboxTransport` trait (seam for testing without Docker) - containerized_backend.rs: `ContainerizedFilesystemBackend` impls `MountBackend`, translates relative→`/project/<rel>`, maps tool-error→MountError - docker_transport.rs: real bollard exec session, serialized Mutex, lazy reconnect - lifecycle.rs: deterministic `ironclaw-sandbox-<pid>` naming, ensure_running/stop/remove - manager.rs: `ProjectSandboxManager` per-project transport cache Phase 6 — Router gating on ENGINE_V2_SANDBOX: - `engine_v2_sandbox_enabled()` helper (truthy: 1/true/yes/on) - Router selects `ContainerizedMountFactory` when enabled + Docker reachable; falls back to `FilesystemMountFactory` with warning otherwise Live e2e bugs caught and fixed: - Shell without explicit `workdir` defaulted to host (not sandbox); fixed by defaulting to `/project/` in `extract_path_param` - `ContainerizedFilesystemBackend::shell` parsed `stdout`/`stderr` but host ShellTool returns merged `output` field; fixed with fallback key lookup - SANDBOX_TOOL_NAMES only had v2 names (`file_read`/`file_write`) but host registry uses v1 names (`read_file`/`write_file`); added both aliases Tests (62 sandbox-related, all green): - 27 bridge sandbox unit tests (intercept, workspace_path, factory, protocol, lifecycle, containerized_backend with ScriptedTransport mock) - 7 containerized-backend tests (including 2 regression tests for the shell bugs) - 5 engine v2 sandbox integration tests (EffectBridgeAdapter end-to-end) - 5 daemon binary smoke tests (real subprocess + NDJSON I/O) - 17 engine workspace unit tests - 1 live Docker e2e test: agent clones nearai/ironclaw into sandbox, renames to megaclaw via sed, verifies with grep — 70s, $0.09, recorded trace committed Verification: - `cargo fmt --check` clean - `cargo clippy --all --benches --tests --examples --all-features` zero warnings - All 62 sandbox tests passing; no existing tests broken Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: replace .expect() with Result in DockerTransport::ensure_session CI's no-panics checker flagged the .expect("just inserted") in production code. Replace with .ok_or_else() returning MountError::Backend. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: multi-tenant project paths + unify sandbox env var with v1 Two issues addressed: 1. Project workspace paths now namespace by user_id: `~/.ironclaw/projects/<user_id>/<project_id>/` instead of `~/.ironclaw/projects/<project_id>/`. Prevents filesystem collisions in multi-tenant deployments where two users could theoretically have the same project UUID. 2. Sandbox enablement now reads `SANDBOX_ENABLED` (same env var as v1 sandbox) in addition to `ENGINE_V2_SANDBOX`. Either being truthy enables the per-project sandbox. This means a single flag governs sandbox behavior regardless of engine version, while the v2-specific override remains available for transitional setups. Tests: 30 bridge sandbox unit tests passing (added multi-tenant path tests + env var combination tests). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address PR review — TOCTOU race, shell env passthrough, canonicalize guard Three issues flagged by the code review bot on nearai#2211: 1. TOCTOU race in WorkspaceMounts::resolve (HIGH): Added double-checked locking — re-check the cache after acquiring the write lock so two threads racing on the same project's first access don't both call factory.build(). The second thread finds the insert from the first. 2. Shell intercept ignores env parameter (MEDIUM): The shell arm in maybe_intercept was passing HashMap::new() instead of forwarding the tool call's env map. Fixed to parse parameters["env"] and pass it through to backend.shell(). 3. Canonicalization fails when root doesn't exist (MEDIUM): When self.root hasn't been created yet (first write to a new project), canonicalize_under_root would walk up to a real ancestor and the starts_with check against the non-existent root would always fail. Now skips canonicalization entirely when root doesn't exist — lexical safety is already guaranteed by safe_join. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address PR review round 2 — apply_patch schema, content validation, dir perms, docs - Fix apply_patch schema mismatch: MountBackend::patch now takes (old_string, new_string, replace_all) matching ApplyPatchTool's actual contract. Previously sent {patch: diff} which would fail with invalid_params in the containerized daemon. - Validate file_write content param: return error instead of silently writing empty string when content is missing. - Log stderr frames from sandbox daemon at debug! instead of silently discarding them in docker_transport StreamReader. - Tighten permissions on intermediate directories created by ensure_project_workspace_dir (projects/, <user_id>/) to 0o700, not just the leaf. - Fix stale module doc in sandbox/mod.rs (referenced "Phase 5 will add" but all phases shipped). - Fix doc path mismatch: workspace path is <user_id>/<project_id>/, not <project_id>/ (workspace_path.rs, CLAUDE.md, design plan). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address PR review round 3 — symlink safety, visibility, debug logging - Close TOCTOU window in canonicalize_under_root: re-canonicalize and verify containment when the reassembled path exists on disk - Fix list_dir_recursive: use symlink_metadata (lstat) so symlinks are detected instead of followed; validate directories against root before recursive traversal - Tighten is_mountable_path to /project/, /memory/, /home/ prefixes instead of any absolute path (defense-in-depth) - Narrow sandbox module visibility to pub(crate) and remove unused pub use re-exports - Remove concrete types (FilesystemBackend, DirEntry, EntryKind, ShellOutput) from engine crate top-level re-exports; access via ironclaw_engine::workspace:: module path - Add debug! tracing to sandbox intercept routing decisions - Add read_file/write_file v1 aliases to daemon SUPPORTED_TOOLS health response - Remove developer-local path from sandbox mod.rs doc comment - Merge staging to fix CI (user_timezone field on ThreadExecutionContext) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address PR review round 4 — safety validation, network isolation, binary writes - Add pre-intercept safety param validation so sandbox-dispatched calls go through the same checks as host-dispatched calls (#1) - Set network_mode: "none" on sandbox containers to prevent outbound network access (#3) - Reject binary content in containerized write instead of silently corrupting via from_utf8_lossy (#5) - Cap list_dir depth to 10 to prevent unbounded traversal (#8) - Change container creation log from info! to debug! to avoid breaking REPL/TUI output (#10) - Make is_truthy case-insensitive so SANDBOX_ENABLED=True works (#11) - Return error instead of unwrap_or_default for missing container ID (#12) - Propagate set_permissions errors instead of silently ignoring (#13) - Return error for missing daemon output key instead of defaulting to empty object (#14) - Add env mutex guard in sandbox_live_e2e test (#15) - Fix rustfmt formatting for let-chain in canonicalize_under_root Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address review round 5 — path traversal, error types, tests Security fixes: - Sanitize user_id in workspace path to prevent directory traversal via malicious user IDs containing `..` or `/` - Add Component::ParentDir check in ContainerizedFilesystemBackend::container_path matching the defense-in-depth approach of FilesystemBackend::safe_join Correctness: - Use MountError::Tool instead of MountError::InvalidPath for missing tool parameters (content, old_string, new_string) — fixes confusing LLM-visible error messages - Fix clippy sort_by_key suggestion in registry.rs Cleanup: - Remove spurious Notify import and dead _notify_link function New tests: - ContainerizedFilesystemBackend path traversal rejection (read + write) - container_path unit tests for safe and unsafe paths - Adversarial user_id test in workspace_path - Daemon-side path traversal test in sandbox_daemon_smoke Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address review round 6 — param normalization, error types, edge cases - Normalize sandbox params via prepare_tool_params() before validation, matching the host execution path (fixes inconsistent validation) - Return ToolError::InvalidParameters instead of EngineError::Effect for sandbox param validation failures (consistent error surface) - ensure_dir checks path.is_dir() not path.exists() (rejects files) - Empty user_id returns "_anonymous" sentinel instead of empty hex string that would drop the tenant namespace via PathBuf::join("") - Restore ENGINE_V2_SANDBOX env var after sandbox live E2E test - Tighten is_mountable_path to /project/ only (no mounts for /memory/ or /home/ yet) - Add v1 tool name aliases (read_file, write_file) to SUPPORTED_TOOLS Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * refactor: unify sandbox env var — remove ENGINE_V2_SANDBOX, use SANDBOX_ENABLED only Single env var controls sandboxing for both engine versions. The transitional ENGINE_V2_SANDBOX override is removed from code, tests, docs, and Dockerfile. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: double-checked locking in transport_for, explicit stdin close in smoke test - ProjectSandboxManager::transport_for no longer holds the mutex across the Docker ensure_running await. Uses double-checked locking so concurrent projects initialize in parallel. - sandbox_daemon_smoke: explicitly take() stdin before wait_with_output so EOF is sent even without a shutdown request. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address review — network mode, error types, race, protocol dedup - Change sandbox container network_mode from "none" to default bridge so git clone / cargo build / pip install work inside the container - Fix binary content rejection to use MountError::Tool instead of MountError::InvalidPath (semantic mismatch) - Fix list depth: use actual depth value instead of depth.max(1) - Fix orphan container race in transport_for by holding lock across container creation instead of double-checked locking - Deduplicate protocol types: daemon now imports from shared bridge::sandbox::protocol instead of defining its own copies - Make bridge::sandbox pub (narrow exposure: only protocol and workspace_path sub-modules are pub) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * docs: update plan doc — sandbox uses bridge networking, not network_mode=none Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
theredspoon
pushed a commit
that referenced
this pull request
Jun 21, 2026
…earai#2680) Biggest single-slice migration in the epic: ten chat handlers + every chat-private helper + all chat helper-tests leave `server.rs` for a new `src/channels/web/features/chat/` module. Routes carried over end-to-end (no behavior change): - POST /api/chat/send, /api/chat/approval, /api/chat/gate/resolve - POST /api/chat/auth-token, /api/chat/auth-cancel (legacy v1 shims) - GET /api/chat/ws, /api/chat/events - GET /api/chat/history, /api/chat/threads - POST /api/chat/thread/new Chat-private helpers that moved along with the handlers: - `is_local_origin` (CSRF-gate for the WS upgrade) - `pending_gate_extension_name` → routes through the canonical `AuthManager::resolve_auth_flow_extension_name` (the identity invariant called out in `src/channels/web/CLAUDE.md` + check #8 in `scripts/pre-commit-safety.sh`); the wrapper is preserved byte-identical so the "one resolver" rule holds after the move. - In-progress reconciliation chain: `reconcile_in_progress_with_turns`, `in_progress_matches_turn`, `in_progress_from_metadata`, `is_stale_in_progress`, `completed_turn_is_newer_than_in_progress`, `in_progress_from_thread`, `summary_live_state`. - `turn_info_from_in_memory_turn`, `thread_state_label`, `turn_state_label`, `IN_PROGRESS_STALE_AFTER_MINUTES`. - `HistoryQuery`, `ChatEventsQuery` request DTOs and `extract_last_event_id` helper. - `engine_pending_gate_info` / `history_pending_gate_info` gate-info hydrators. Tests: 15 helper-level tests (5 reconcile, 2 in-memory-turn-info, 2 summary-live-state, 1 thread-state-label, 5 is-local-origin) move with the helpers into `features/chat/mod.rs::tests`. 8 caller-level tests (chat_history × 3, chat_approval, chat_auth_token × 2, chat_auth_cancel, chat_gate_resolve) stay in `server.rs::tests` for now because they rely on shared `GatewayState` builders (`test_gateway_state`, `test_gateway_state_with_store_and_session_manager`, `test_gateway_state_with_dependencies`) that construct state for multiple slices — promoting those builders to `src/channels/web/test_helpers.rs` is a follow-up. Also dropped 3 `test_build_turns_from_db_messages_*` tests in `server.rs` that were redundant with the 14 already in `util.rs::tests`. Cleanup of dead code: `src/channels/web/handlers/chat.rs` deleted entirely. The file held live `chat_events_handler` + `extract_last_event_id` + `ChatEventsQuery` (absorbed into `features/chat/`), plus three zombie duplicate handler definitions (`chat_ws_handler`, `chat_threads_handler`, `chat_new_thread_handler`) that predated the `server.rs` canonicals but were never deleted — one of them (`chat_ws_handler`) used a weaker `is_local_origin` heuristic that skipped IPv6-literal parsing, so accidentally wiring through it would have been a silent security degradation. The router.rs imports and `handlers/mod.rs` declaration are updated accordingly. Router updates: nine `server::` imports swapped for `features::chat::`, plus the `handlers::chat::chat_events_handler` import removed (now `features::chat::chat_events_handler`). The migration docstring above the feature-handler imports lists chat as extracted alongside logs / oauth / pairing / status. Quality gate: fmt clean, clippy clean on `--all --tests --examples --all-features`, 424 `channels::web` tests pass, boundary checker reports no back-edges with the existing empty allowlist. Net shape: `server.rs` shrinks from ~5,770 → ~3,620 lines (-2,150 lines). `handlers/chat.rs` goes from 432 → 0. Explicit non-scope (noted in the migration plan): - Four handlers (`chat_send`, `chat_ws`, `chat_threads`, `chat_new_thread`) gate side effects and currently have no caller-level test. Adding them is genuine new coverage, not regression preservation — a follow-up PR. The helper-level tests that DO cover things (`is_local_origin`, `reconcile_*`, `turn_info_from_in_memory_turn`, `summary_live_state`, `thread_state_label`) move with their helpers. - The pre-existing `engine_v2` / `engine_v2_enabled` duplicate in `GatewayStatusResponse` (flagged on PRs nearai#2665 and earlier) still needs a coordinated frontend fix and isn't touched here. Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
theredspoon
pushed a commit
that referenced
this pull request
Jun 21, 2026
…easoning-augmented recall (nearai#2336) * feat(memory): configurable insights interval, session summary hook, reasoning-augmented recall Three memory enrichment features: 1. Configurable conversation insights interval via MISSION_INSIGHTS_INTERVAL env var (default: 5, min: 1) with MissionsConfig + MissionSettings wiring 2. SessionSummaryHook that writes LLM-generated conversation summaries to workspace daily logs on session end (fail-open, 30s timeout) 3. Optional reasoning parameter on memory_search that synthesizes raw chunks via cheap LLM before returning, controlled by SEARCH_REASONING_ENABLED Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(memory): address PR nearai#2336 review feedback and CI failures Critical fixes: - Use DB-first config system for MissionsConfig instead of raw std::env::var in router.rs (issue #1) - SessionSummaryHook now uses thread_ids from HookEvent::SessionEnd to summarize the correct conversation instead of guessing via recency; falls back to most-recent for backward compatibility (#2) - Add per-user rate limiter (10/min, 60/hr) and 15s timeout on reasoning LLM calls in MemorySearchTool to prevent unbounded usage (#3) Test coverage: - Caller-level tests for reasoning-augmented recall (LLM wiring, disabled config, and failure fallback paths) (#4) - SessionSummaryHook LLM failure path test confirming fail-open behavior (#5) - reasoning_enabled config field tests (default, env, DB override) (#6) - MissionSettings and SearchSettings round-trip assertions in comprehensive_db_map_round_trip (#11) Convention fixes: - Remove double env-var parsing in MissionsConfig::resolve (#7) - Use ChatMessage::system()/user() constructors in SessionSummaryHook (#8) - Add TODO comments for inline prompt strings (#9) - Add timeout on reasoning LLM call (#10) CI fixes: - Remove 4 stale wasmtime advisory entries from deny.toml - Add RUSTSEC-2026-0097 (rand 0.8.5) to advisory ignore list Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(memory): address henrypark133 + ilblackdragon review — safety, concurrency, prompts (nearai#2336) - Move inline prompt templates to prompts/*.md per project convention (session_summary.md, memory_reasoning_synthesis.md) — resolves TODOs - Add Arc<Semaphore> to SessionSummaryHook to cap concurrent LLM calls on mass session expiry (follows OutboundWebhookHook pattern) - Sanitize LLM-generated summaries via ironclaw_safety::Sanitizer before writing to workspace (mitigates stored prompt injection vector) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(memory): CI compile fix + reasoning sanitizer parity + harden test - Add live_state / live_state_started_at fields to ConversationSummary literals in three session-summary test sites; staging added these fields after the branch was created and clippy/test builds were failing on missing-field errors. - Replace silent unwrap_or_default on MissionsConfig::resolve in bridge::router::init_engine with an explicit warn-and-default match, so a misconfigured MISSION_INSIGHTS_INTERVAL surfaces in logs instead of being absorbed into the default. - Run the reasoning-synthesis output through ironclaw_safety::Sanitizer before persisting it to the tool result, matching the parity already applied in SessionSummaryHook. Memory chunks fed into synthesis can carry attacker-controlled text and the synthesis flows back into future LLM contexts via memory_search results. - Strengthen reasoning_enabled_fires_llm_and_returns_synthesis: add a preflight assertion that FTS returns the seeded doc, then unconditionally assert the LLM was called once and that synthesis matches the mocked response. Removes the prior `if llm.calls() > 0` guard that made the synthesis assertions vacuous when search returned empty. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> Co-authored-by: Illia Polosukhin <ilblackdragon@gmail.com>
theredspoon
pushed a commit
that referenced
this pull request
Jun 21, 2026
Consolidated fixes for serrrfirat's 10 unresolved review threads plus zmanian's CHANGES_REQUESTED review (3 blockers + 7 mediums + 3 follow-up items + title typo). ## Blockers 1. CI now runs the package tests (zmanian blocker 2 / serrrfirat #7). The matrix `cargo test ${{ matrix.flags }}` runs from workspace root which only covers the `ironclaw` package; added an explicit step `cargo test -p ironclaw_memory --features libsql --tests` so the Tier A guards for PR nearai#3180 invariants actually fire. 2. `#[ignore]` markers converted to `#[cfg_attr(not(feature = "pr3180-ready"), ignore = ...)]` (zmanian blocker 1 / serrrfirat #1). Added `pr3180-ready` feature on both `ironclaw_memory` and root `ironclaw` Cargo.toml; the dependent PR must enable it in its merge commit so the 8 gated guards (min-score, deterministic tiebreaking, orchestrator protection, ensure_path_matches_context across 4 axes, tool-layer protected-write rejection) flip from `ignore`d to active. 3. Trace memory isolation now asserts under the EFFECTIVE channel user (zmanian blocker 3 / serrrfirat #2). Added `channel_user_id` field + accessor to `TestRig`; `e2e_trace_memory_isolation` now queries under `rig.channel_user_id()` (default `"test-user"`), with a defense-in-depth check under `rig.owner_id()` for mis-routing regressions. ## Test-correctness mediums 4. Min-score test pins `with_query_embedding([1,0,0])` to favor hybrid.md (serrrfirat #3 / zmanian #4). Removed the permissive `* 0.99` fallback — FTS-only-below-hybrid is now a hard assertion. 5. Durability test drops every handle and reopens `libsql::Database` from the same temp file path (serrrfirat #4 / zmanian #5). Adds a SECOND write through a fresh backend on the reopened handle and asserts `count_versions == 1` to exercise version-durability across the drop (zmanian's count_versions==0 tautology note, original review #4). 6. Append versioning asserts exact row count `== 1`, not `!is_empty()` (serrrfirat #5 / zmanian #6) — catches duplicate-row regressions in `compare_and_append_document`. 7. Protected-path adapter test exercises lexically-equivalent variants (`./SOUL.md`, `//SOUL.md`) in addition to canonical (serrrfirat #6 / zmanian #7). VirtualPath rejects `..` so no `..` variant; in-loop `count_documents_total == 0` after EACH variant. 8. Hybrid search isolation now varies all four scope axes (serrrfirat #8 / zmanian #8): tenant, user, agent, project. 5 documents seeded; search from caller scope must return exactly one. 9. Tool round-trip asserts EXACT persisted content via direct DB read (serrrfirat #9 / zmanian #9). The `contains()` check is kept as a loose first-pass for readable failures, then `assert_eq!` on the exact byte string is the load-bearing assertion. 10. Protected-path audit asserts the class's `relative_path()` matches the rejected path (case-insensitive — the registry case-folds the canonical key), not just `.is_some()` (serrrfirat #10 / zmanian #10). A regression that emits the wrong path class now fails. ## zmanian follow-ups Z1. Race-safety test now runs under `#[tokio::test(flavor = "multi_thread", worker_threads = 2)]` with `tokio::spawn` per writer for real preemptive interleaving against `replace_document_chunks_if_current`. Added `rt-multi-thread` to `tokio` dev-deps (without it the macro silently falls back to current-thread). Z2. `write_to_protected_path_rejected.json` trace fixture sets `all_tools_succeeded: false` explicitly. Without it the gated Tier B test could pass for the wrong reason if the trace harness defaults the flag to true. Z3. Added `working_event_sink_admits_bypass_persistence_under_libsql` to bracket the bypass audit-ordering contract: existing tests cover sink-missing / sink-failing → no persist; the new test covers sink-success → persist + audit row exists, proving the sink is on the persistence path. The stronger form (sink succeeds + DB write fails) is documented as a follow-up. ## Cleanup - Removed `_link_in_memory_repo_for_unused_imports` shim and the `InMemoryMemoryDocumentRepository` import that only existed to feed it (zmanian original-review #3). - Fixed PR title typo `momery` → `memory` via gh. Helper-consolidation into `tests/common/libsql_helpers.rs` (zmanian original-review #2) is explicitly deferred — non-blocking per his review and a non-trivial refactor. ## Verified - `cargo fmt --all -- --check` clean - `cargo clippy -p ironclaw_memory --features libsql --all-targets -- -D warnings` zero warnings - `cargo test -p ironclaw_memory --features libsql` all suites green (gated tests stay `ignored` without `--features pr3180-ready`)
theredspoon
pushed a commit
that referenced
this pull request
Jun 21, 2026
…view Four high-level architectural decisions from human review (serrrfirat, zmanian) on PR nearai#3544 land as spec amendments before WS-0 seals the trait shapes: 1. LoopFamily as first-class abstraction (zmanian #2.2). New WS-3.5 brief: LoopFamilyId + ComponentIdentity + LoopFamily + LoopFamilyRegistry (Guice-style singleton built once at startup). 2. Strategies sealed pub(crate) (zmanian #1.2, #2.5, #2.6; serrrfirat #1). AgentLoopPlanner is pub but uses sealed-trait pattern; strategy access lives on pub(crate) AgentLoopPlannerInternal extension trait. 3. Hooks as middleware (zmanian #1.1). Adopt the four-scenario design in PR nearai#3523-comment-4435808547. Two concrete follow-ups: composition seam architecture test in WS-9's brief; LoopFailureKind::PolicyDenied variant in WS-0. 4. Broad scope retained with stress-test discipline (zmanian #2.1, #2.7; serrrfirat #8). New §12.5 enumerates anticipated families (default + hypothetical routine/mission/coding/planning) and the strategies each would swap; trait shapes accommodate every row. Side effects: - Split ControlStrategyState → StopStrategyState + GateStrategyState in WS-0 so stop and gate strategies grow independently. - Drop generics on PlannedDriver — non-generic { family, executor }. WS-7 + downstream briefs (WS-9, WS-10, WS-11, WS-13, WS-14) updated. - Subsume PlannerId into LoopFamilyId + ComponentIdentity (one versioning primitive per zmanian #1.4, partly addressed). - Executor entry point becomes execute_family(&LoopFamily, host, state). - Document HostXxxContextSource as the pattern for family-specific durable context (mirrors WS-15's HostIdentityContextSource). 15 files changed (1 new brief: loop-family-registry.md). Spec-only; no code changes. Three remaining comment clusters (checkpoint durability + LoopExit witness; versioning + JSON canonicalization; serrrfirat's remaining 7 correctness seams) are noted as follow-up work. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
theredspoon
pushed a commit
that referenced
this pull request
Jun 21, 2026
…3920) * Implement installed WASM hook runtime Adds crates/ironclaw_hooks/docs/threat-model-wasm.md and follows the reviewed design ack: 1) module bytes are resolved, digest-cached, and compiled in the tool-WASM style while reusing its resource limiter; 2) each invocation gets a fresh wasmtime Store; 3) the ABI is a wasmtime::Linker surface, not wit-bindgen; 4) host-import sink shims enforce call, patch-byte, observer-fact, and decision budgets. * Harden WASM hook string and metadata budgets * fix(hooks): validate WASM hook ABI at install time (serrrfirat #3 on PR nearai#3634) Address serrrfirat MEDIUM finding #3: `WasmHookRuntime::prepare()` compiled and cached module bytes but did not validate imports or the requested export. ABI mismatches (unsupported import, missing export, wrong export signature) were deferred to first live dispatch — and the prior `wasm_unsupported_host_import_fails_closed` test codified that a bad-import module would install successfully and only fail closed at invocation. Malformed untrusted modules should never reach live traffic. Changes: - `prepare()` derives the target hook point from `request.kind`, then runs `validate_module_abi()`: scratch-instantiate the module against the point-specific linker (catches unsupported / wrong-type imports) and resolve the typed export `() -> ()` (catches missing export and wrong signature). Failures surface as new `WasmHookRuntimeError::InvalidImports` or existing `WasmHookRuntimeError::InvalidExport`, both of which bubble up as `HookError::RegistryConstruction` from the registrar. - `wasm_point_for_kind(HookManifestKind)` helper centralizes the kind → wasm-point mapping; the previous `execute_*` paths can share it in a follow-up but kept inline for now to minimize churn. Tests: - `wasm_unsupported_host_import_is_rejected_at_install_time`: replaces the prior test that codified late-failure behavior; asserts the registrar returns `RegistryConstruction` citing the bad import. - `wasm_missing_export_is_rejected_at_install_time`: new module that compiles but lacks the manifest-declared export; same install-time rejection. * fix(hooks): address henrypark133 must-fix #1, #2, #3 on PR nearai#3634 Three items from the 5-15 review: **#1 (must-fix) Extract ironclaw_wasm_limiter micro-crate** Replace `#[path = "../../../ironclaw_wasm/src/limiter.rs"]` cross-crate file import with a proper Cargo edge. The 111-line `WasmResourceLimiter` moves into a new `crates/ironclaw_wasm_limiter` micro-crate that both `ironclaw_wasm` and `ironclaw_hooks` depend on. The architecture rule forbidding `ironclaw_hooks -> ironclaw_wasm` is preserved (the new crate sits below both consumers and pulls in only `wasmtime` + `tracing`); `cargo check`, `cargo doc`, and architecture-linting tests now see the edge, and the file can't be moved out from under one of the consumers silently. Mechanical changes: - new `crates/ironclaw_wasm_limiter/` (Cargo.toml + src/lib.rs with the type exposed as `pub` instead of `pub(crate)`) - workspace `members` entry added - `crates/ironclaw_wasm/src/limiter.rs` deleted - `crates/ironclaw_wasm/src/lib.rs`: `mod limiter` removed - `crates/ironclaw_wasm/src/store.rs`: import switched to `ironclaw_wasm_limiter::WasmResourceLimiter` - `crates/ironclaw_wasm/Cargo.toml`: dep added - `crates/ironclaw_hooks/Cargo.toml`: dep added - `crates/ironclaw_hooks/src/wasm/runtime.rs`: `#[path = ...]` block removed; import switched to the crate **#2 + #3 (must-fix) Dead WASM arms in dispatch** `run_before_capability_hook`, `run_before_prompt_hook`, and `run_observer_hook` each had an early-return guard that dispatched WASM hooks with `catch_unwind` + timeout, then ALSO had a matching WASM arm in the inner `match` that ran without those protections. The prompt-path arm additionally swallowed `WasmHookFailure` via `|_| ()`, making the must-fix #2 problem worse on that path specifically. If a future refactor removed any of the early-return guards, those inner arms would silently take over and drop panic isolation, deadline enforcement, AND (for prompts) the failure category. Replaced each inner arm with `unreachable!()` carrying a comment that explains why the arm exists and references the early-return guard above it. A future refactor that removes the guard will now trip the `unreachable!` at first call instead of silently degrading. All 154 hooks lib + 29 reborn integration tests still pass. * fix(hooks): plumb context to WASM hooks + runtime hardening Critical #1 on PR nearai#3634: WASM hooks previously received no context. The `execute_*` entry points dropped the `&BeforeCapabilityHookContext` / `&BeforePromptHookContext` / `&ObserverHookContext` value and invoked the guest export with `()`, so a WASM gate could never decide based on the capability name, tenant, provider, or other dispatch-time facts. Add an `ic:hooks/context@1` host-import module exposing two read-only calls — `ctx_size() -> i32` and `ctx_read(ptr, len) -> i32` — backed by a JSON-serialized blob the dispatcher writes per-invocation into the fresh store. Modules that don't import these continue to link; modules that do import them get a stable, non-empty payload to read. An integration test (`wasm_before_capability_hook_reads_context_blob`) asserts the contract end-to-end: a guest that fails to read a non-empty blob traps before its `deny` call. Also rolls up the other reviewer-flagged WASM runtime issues, all of which touch `wasm/runtime.rs`: HIGH #2: epoch-tick background thread now holds a shutdown `AtomicBool` and joins on `Drop`. Previously it looped forever and leaked an Engine clone on every runtime drop. MED #4: compiled-module cache is now an `lru::LruCache` bounded by `MODULE_CACHE_CAPACITY = 128`. Replaces the unbounded `HashMap`. MED #7: `prepare()` no longer compiles under the cache lock. Fast path reads from LRU under a brief lock; slow path compiles outside the lock and re-checks on insert to avoid the TOCTOU window where two concurrent installs of the same module both compile. Bug #9: post-call `deadline_exceeded()` re-check on the Ok branch is gone. wasmtime epoch-interrupt is the authoritative wall-clock signal; an Ok return is no longer reclassified as a timeout because the wall ticked over during host-side return. Bug #10: `add_milestone_metadata` returns a distinct "metadata value exceeds the u32 byte-length ceiling" error when the guest-supplied `value.len()` overflows u32, instead of misreporting it as "exceeded total prompt-patch byte budget". Existing integration tests for WASM hooks are also re-wired through `HookRegistrar::with_verified_grants` so the grants-store gate added in the foundation-01 merge stops failing the pre-existing fixtures. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(hooks): run WASM hooks on the blocking pool HIGH #3 on PR nearai#3634: `tokio::time::timeout` does NOT cancel synchronous wasmtime execution. The previous code awaited a `catch_unwind(async { h.evaluate(ctx) })` future whose body completed in one poll, so the timeout could only fire *around* the WASM call rather than against it; a hook that wedged inside wasmtime simply pinned the calling tokio task. Route gate, prompt, and observer WASM dispatch paths through `tokio::task::spawn_blocking` via a shared `run_wasm_blocking` helper. The outer `tokio::time::timeout` now governs the JoinHandle, so a stuck blocking task stops blocking the dispatcher's caller; the wasmtime epoch interrupt configured in the runtime (10 ms tick) is the authoritative in-WASM wall-clock cancel signal. JoinError (panic in the blocking task) maps to `FailureCategory::Panic`, matching the pre-existing semantics for synchronous panics caught via `catch_unwind`. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * perf(hooks): O(1) hook-id lookup via side index Finding #8 on PR nearai#3634: `set_priority`, `poison`, `is_poisoned`, and `contains_hook` all did full-registry scans over every binding at every point. Each is called per-dispatch (poison-checks on the snapshot loop in particular), so the cost is `O(registered_hooks)` per `(installed_hook, registered_hook)` pair. Maintain a denormalized `HashMap<HookId, (HookPointSpec, usize)>` side index in lock-step with `by_point` so every per-hook-id operation becomes a single hash lookup + a direct vec indexed access. The duplicate-id rejection in `insert` now reads from the side index too, turning what used to be a flat-map scan into a `HashMap::contains_key`. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(hooks): wall-clock timeout, observer memory, limiter rollback, registrar happy path Round out the test set for the WASM hook execution path: #11 / #12: gate + observer wall-clock timeout. The pre-fix dispatcher ran wasmtime synchronously on the executor, so the outer `tokio::time::timeout` `Err(_elapsed)` arm was effectively unreachable. Now that WASM execution runs on the blocking pool, the timeout actually fires; the new tests give the wasm budget headroom (1B fuel, 5s wall) and the dispatcher a 20 ms timeout, then assert the failure classification (FailClosed for gate, FailIsolated for observer). #13: observer memory exhaustion. Mirrors `wasm_memory_exhaustion_fails_closed_for_gate` against the observer dispatch path so the FailIsolated branch of the failure matrix has explicit memory coverage, not just fuel/wall. #15: `WasmResourceLimiter::memory_grow_failed` rollback. Stages an approved grow, simulates the OS-level grow failing, and asserts a subsequent grow of the full ceiling succeeds — the inflated `memory_used` from the failed attempt must be released. #16: registrar WASM happy path. Companion to the existing `install_wasm_body_requires_runtime` negative case: a valid module installs, the binding is visible via the public registry accessor, and is not pre-poisoned. #14 (`add_milestone_metadata` happy path) is intentionally omitted — the BeforePrompt dispatch path is currently unreachable due to a pre-existing manifest-vs-registry scope conflict (`OwnCapabilities` is the only valid `BeforePrompt` scope per manifest validation, but the registry rejects `OwnCapabilities` at `BeforePrompt` because the point has no provider context). That contradiction sits outside this PR's scope; flagging for a follow-up. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * refactor(hooks): typed WASM version material, reconcile design doc LOW #20 on PR nearai#3634: extract the `{extension_version}+wasm:{module_digest_hex}` concatenation into a `WasmVersionMaterial` newtype with a single `Display` impl. The identity material no longer floats free as a stringly-typed argument inside the registrar. Reconcile `docs/successors/02-wasm-runtime.md` with the implementation: - Spell out that wall-clock cancellation depends on the `tokio::time::timeout(tokio::task::spawn_blocking(...))` pair, and explain why a bare timeout over a synchronous wasmtime call cannot actually cancel. - Define `FailIsolated` and `FailClosed` as `FailureDisposition` values, distinct from the older `HookFailureMode::{FailOpen, FailClosed}` policy switch that applies to predicates. - Clarify the generic `evaluate` export contract — name is whatever the manifest declares, signature is `(): ()`, context arrives through the new `ic:hooks/context@1` host imports — and note the intentional divergence from `WitToolRuntime`'s hardcoded interface. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(hooks): drop .expect() in WASM module cache capacity Pre-commit no-panics CI flagged the .expect() on the LruCache capacity. Move the validity check to a const match, so the NonZeroUsize is fixed at compile time and the no-panics regex is satisfied. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(hooks): use HookLocalId::new after newtype privatization The newtype-privatization landed in reborn-integration after the hooks-fu-wasm-runtime branch's WASM scaffolding tests were written; update the affected test/registrar sites to use HookLocalId::new instead of the now-private tuple constructor. * style: cargo fmt after newtype-privatization fixups * test(hooks): ignore 3 BeforePrompt WASM tests with manifest/registry conflict These tests were failing on the original branch tip too (verified against origin/hooks-fu-wasm-runtime @ 571efdf). The Installed-tier BeforePrompt WASM install path has no valid scope today: - OwnCapabilities is rejected by the registry C3 check (finding #2 on PR nearai#3573) since BeforePrompt has no per-capability invocation context. - SameTenant is rejected by manifest validation ("cannot combine scope = same_tenant with kind = before_prompt"). The budget-overflow paths these tests exercise are point-agnostic; the follow-up is to either rewrite the helper to install through BeforeCapability or add a Global manifest scope. Tracked as a deferred item on the new PR. --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
theredspoon
pushed a commit
that referenced
this pull request
Jun 21, 2026
…earai#4559) * docs: trace commons agent onboarding design spec Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: address spec review findings (trust anchoring, key staging, consumption atomicity, replay validation) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: spec review round 2 nits (server-anchored tenant wording, pending-key cleanup) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: implementation plan for trace commons agent onboarding Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: address plan review findings (scope threading refactor, dispatch model, dev-deps, LazyLock hazard) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: plan review round 2 fixes (literal dep versions, context constructor threading depth) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: incorporate server-agent coordination feedback (optional community/profile/leaderboard URLs) From TraceCommons/trace-commons#136-#141 comments: onboard response gains optional browser-surface navigation hints, sanitized client-side (HTTPS or dropped), never part of issuer trust anchoring. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(traces): onboarding wire types matching trace-commons-server contract Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(traces): invite URL parsing with origin trust anchoring Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(traces): device keypair lifecycle with pending staging and self-signed workload JWTs Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(traces): auth_mode and device_key_id policy fields with legacy-compatible defaults Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(traces): onboard() orchestration with trust anchoring and retry-safe key staging Wire invite parsing, device key staging, onboard POST, issuer origin trust anchoring, ingest_url HTTPS enforcement, keypair promotion, and policy write into onboard_at_dir(). Refactors invite.rs to extract pub(crate) is_https_or_loopback, origin_of, and host_only helpers shared with mod.rs (one source of truth for origin/bracket handling). Adds axum mock-issuer tests covering the happy path, mismatch rejection, terminal vs transient error key retention, insecure ingest URL, loopback ingest allowance, community URL sanitisation, and retry key reuse. Partial-failure lockout fix (spec §2.2): promote() no longer deletes the pending file. The flow now writes the tenant key file, then the policy, and only discards the pending file after BOTH durably succeed. If the policy write fails the pending key survives, so a retry reloads the same key (server idempotency returns the original registration) and harmlessly overwrites the tenant file — no permanent lockout from a consumed invite with a regenerated keypair. Regression test simulates a policy-write failure (policy.json pre-created as a non-empty dir so the atomic rename fails), asserts Err(Persist) with the pending key intact, then asserts a retry succeeds reusing the same device_key_id. Response validation (defense-in-depth): reject schema_version != the v1 response constant as MalformedResponse, and cross-check the response device_key_id against the locally derived id (we never trust the response value for policy; a disagreement is now treated as a tamper signal and rejected). Both covered by tests. The onboard response body is read with the 64 KB cap enforced per-chunk during streaming (mirroring read_bounded_trace_upload_claim_response) rather than buffering the whole body first, so a hostile server cannot force a large allocation. Also fixes a pre-existing test-isolation defect surfaced by the added load: the remote-request timeout test configured a 50ms timeout via the process-global IRONCLAW_TRACE_REMOTE_REQUEST_TIMEOUT_MS env var. set_var is process-global, so under parallel execution the 50ms value leaked into other tests' trace HTTP clients, producing spurious `operation timed out` failures against fast local mocks. Replace the env mutation with a task-scoped TEST_REMOTE_REQUEST_TIMEOUT_OVERRIDE task-local (visible only within the awaiting test's own task tree, zero production change; documents the spawn caveat), and decouple the timing assertion from a tight wall-clock race so it no longer flakes when reqwest's timer is delayed under an oversubscribed runtime. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(traces): device-key self-signed workload JWT branch in upload-claim refresh Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(engine): trace_commons onboard and status first-party tools with agent guidance Add two model-visible first-party capabilities to the Reborn engine: - builtin.trace_commons.onboard: drives operator-invite enrollment flow with explicit per-conversation consent gate (confirmed=true required before any network call); maps OnboardOutcome/OnboardError to clean agent-readable JSON - builtin.trace_commons.status: read-only enrollment state inspector Wires ironclaw_reborn_traces into ironclaw_host_runtime, creates schema files (schemas/builtin/trace-commons-{onboard,status}.{input,output}.v1.json) and prompt doc files (prompts/builtin/trace-commons-{onboard,status}.md) at the manifest-derived paths. Includes 11 unit tests covering input parsing, consent refusal, success/error value formatting, and status formatting. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: add Task 11 — credits visibility (console display + agent-queryable balance) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(engine): e2e trace commons onboarding through capability dispatch Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(traces): document agent onboarding flow in trace-commons internal doc Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: correct Task 11 console scope (credit endpoint already exists; frontend = coordinate with designer) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(traces): trace_commons.credits agent-queryable balance tool Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(gateway): minimal Trace Commons credits card in settings Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(traces): store upload-claim endpoint in policy; preserve primary onboard error; block metadata/link-local/multicast issuers Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(traces): route agent onboarding HTTP through host network-egress policy (nearai#4560) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * build: update Cargo.lock for trace-commons onboarding dev-deps Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * chore(traces): drop orphaned schema/prompt files (main resolves builtin schemas inline; prompt_doc_ref dropped) Post-merge cleanup: main's first_party_tools now resolves builtin input schemas via the inline schemas.rs match (trace_commons arms added during the merge) and sets prompt_doc_ref: None for all builtins, so the physical trace-commons-*.json schema files and trace-commons-*.md prompt docs are no longer referenced. The onboard consent contract remains in the capability description and is enforced in dispatch_onboard. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Fix Trace Commons invite hash contract * fix(traces): grant trace_commons capabilities in local-dev policy The three builtin.trace_commons.* capabilities were declared in the first-party package but had no [[grants]] entries in local_dev_capability_policy.toml, so local-dev runs (repl/serve) filtered them out of the model-visible tool surface entirely. The provider-level authority_effects ceiling had external_write, but the per-capability grants were never added. onboard gets the local_dev_wildcard egress profile (invite origins are operator-chosen; private/metadata IP ranges stay blocked by the shared enforcer). status/credits are read-only. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(traces): add Reborn e2e coverage for trace_commons first-party tools Closes the coverage gate failure: builtin.trace_commons.{onboard,status, credits} were declared in the first-party package but missing from REBORN_FIRST_PARTY_E2E_COVERED_CAPABILITIES, failing reborn_builtin_first_party_capability_e2e_coverage_is_complete on both the Reborn root tests and all-features CI jobs. Adds a trace_commons host-runtime harness (network policy populated so the onboard Network-effect obligation passes) and a parity test driving all three capabilities through the scripted model loop: onboard with confirmed=false exercises the deterministic consent gate with no network, status and credits return the unenrolled/zero-credit defaults. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(traces): community profile second opt-in (token mint + profile set) After device-key enrollment, public leaderboard attribution is a second, separate opt-in: IronClaw mints a short-lived profile token from the claim issuer with consent_scopes=[public_attribution] and empty allowed_uses (such a claim cannot submit traces), then either prints it for the web profile page or performs the profile update itself. The browser cannot sign device-key requests, so the token must be minted by IronClaw — previously this step was impossible and agent guidance invented flows. - ConsentScope::PublicAttribution mirrors the server protocol enum; default_allowed_uses_for_scope returns empty for it. - mint_profile_attribution_token_for_scope / set_community_profile_for_scope / withdraw_community_profile_for_scope reuse the hardened issuer HTTP path (allowlist validation, pinned DNS, no redirects, bounded reads, token never in errors). PUT/DELETE /v1/community/profile per the server contract; handle (3-32 ASCII alnum/-/_) and bio (<=280 bytes) validated client-side. - CLI: ironclaw-reborn traces profile token|set|withdraw. - Onboard tool next_steps now describes the profile second opt-in so agent guidance stops inventing browser login flows. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * feat(traces): autonomous turn-end trace capture in the Reborn runtime The Reborn binary could onboard, report status/credits, and manage profiles, but never captured or submitted traces — the autonomous pipeline existed only in the v1 agent loop. This wires it into the Reborn runtime composition: - TraceCaptureTurnEventSink subscribes best-effort to the turn lifecycle bus (the existing turn_event_sink injection seam). On Completed/Failed events with an explicit owner it spawns a detached task that reads the owner's standing policy (one file read for non-enrolled users), loads the recent thread history (last 24 messages, 5 turns — v1 parity), adapts user/assistant text rows into the neutral ConversationMessage shape, redacts + scores locally, and queues + immediately flushes eligible envelopes. All failures are debug!-logged and never touch the turn lifecycle path. - A periodic flush worker (300s, 25/scope — v1 parity) retries queued envelopes for the runtime owner plus every scope observed since boot, with CancellationToken shutdown alongside the other workers. - TraceClientAutonomousCaptureRequest gains outcome_override so the lifecycle event's terminal status (authoritative in Reborn, where transcripts carry no structured outcome payload) marks failed turns as TaskSuccess::Failure; v1 passes None (no behavior change). - Tool-result rows and credit-notice delivery are documented follow-ups (refs-only records; no composition-level outbound channel surface). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * test(traces): end-to-end auto-capture through send_user_message Proves the full Reborn auto-submission chain with a real runtime: a completed turn for an enrolled owner scope lands a redacted envelope in that scope's submission queue with no manual trace command — turn completion -> lifecycle bus -> capture sink -> thread-history read -> redact/score -> eligibility -> queue (+ local-failing immediate flush leaves the entry for the retry worker). Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Expose Trace Commons profile token tool * Expose Trace Commons profile set tool * Allow Trace Commons profile setup from agent * feat(webui-v2): Trace Commons credits card in WebChat v2 settings Adds GET /api/webchat/v2/traces/credit and a read-only Trace Commons settings tab to the v2 SPA, giving webui-v2-beta parity with the v1 console's credits card. - Route follows the descriptor system end to end: bearer-auth required, NoBody, 120/60 per-caller read rate limit; descriptor-driven body/rate-limit enforcement applies automatically. - RebornServicesApi::trace_credits derives the trace scope exclusively from the authenticated caller's user id (never from query/body) and reads contributor-local state via ironclaw_reborn_traces (policy + trace_credit_report), soft-falling back to an unenrolled zero-state on missing/unreadable local state, mirroring builtin.trace_commons.credits. - SPA: Trace Commons subtab (enrollment, pending/final credit, delayed ledger delta, submission counts, last submission/sync, recent credit explanations) with the server-authoritative framing and a not-enrolled empty state pointing at agent onboarding. - Tests: descriptor contract row, handler oneshot, and three composed- router serve tests (200 zero-state, 401 without bearer, enrolled policy reporting with per-test scope isolation). - Drive-by: cfg-gate openai_user_id in webui_serve.rs to clear a pre-existing unused-variable warning under default features. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * Exempt Trace Commons profile setup from local-dev gate * Route Trace Commons profile writes to ingest * review(4559): address serrrfirat feedback - Drop stray working-note markdown files from the repo root (they rode in via an early origin/main merge and are not this PR's documentation). - trace_commons_dispatch_e2e: setup_base_dir is now a OnceLock that every test calls first — the previous 'single-threaded during init' claim was wrong under tokio's multi-threaded test runtime, and two of three tests skipped the setup entirely. - settings.js: extract shared appendDisplayGroup + declarative row defs; loadTraceCommonsCredits drops from ~120 lines of manual DOM to a rows array; also removes a double-escape (textContent + escapeHtml) on explanation lines. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(webui-v2): add traceCommons i18n keys to all locales The credits card added the traceCommons.* key set to en.js only; the i18n consistency test (all_locales_share_the_en_key_set) requires every locale to carry the same key set. Adds translated entries to ar, de, es, fr, hi, ja, ko, pt-BR, uk, and zh-CN. Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com> * fix(reborn): fail loud with source context on malformed local-dev master key The local-dev secret store resolver read the cached key file (and the SECRETS_MASTER_KEY env fallback) and passed the material straight into SecretsCrypto::new several layers deep. A corrupt or low-entropy key (e.g. a 64-char all-zeros value, which passes the length floor but has one distinct byte) surfaced only as the opaque "Invalid master key", with no pointer to the file the operator must fix. - Add ironclaw_secrets::validate_master_key_material as the single source of truth for master-key rules; SecretsCrypto::new delegates to it. - resolve_local_dev_secret_master_key now validates at the source (cached file vs SECRETS_MASTER_KEY env) and returns a RebornBuildError::InvalidConfig naming the offending path/env var and the actual constraint, before any crypto is constructed. - A malformed env value is now rejected before being persisted to the cached key file (no more poisoned-cache state). Tests: malformed-file path-context rejection, malformed-env source-context rejection, valid cached file accepted. Refs nearai#4741 Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(webui-v2): add Trace Commons credits card to chat sidebar Surface trace contribution credits at a glance in the chat sidebar, above the conversation list. Previously credits were only visible under Settings -> Trace Commons. - New SidebarTraceCredits component reuses the existing useTraceCredits hook (/api/webchat/v2/traces/credit) — no new endpoint. Renders only when enrolled; loading/error/not-enrolled render nothing to keep the sidebar clean. Shows final credit and accepted/submitted counts and clicks through to Settings -> Trace Commons for the full ledger. - useTraceCredits now refetches (60s interval + on window focus) so the card and the Settings tab reflect newly-accepted submissions live. - Add one compact i18n key (traceCommons.cardAccepted) across all 11 locales; reuse existing keys for the rest. - Source-shape regression test in assets.rs. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(reborn-traces): reconstruct tool calls in turn-end trace capture The Reborn capture adapter dropped every tool-result row, so captured trace envelopes were text-only. That left the two highest-value scoring levers — replayability (0.20) and tool coverage (0.15) — permanently at zero, so even agentic tool-using turns scored as plain chat and stayed below the 0.35 submission gate. Nothing ever submitted. conversation_messages_from_records now reconstructs a `tool_calls` message from each run of ToolResultReference rows that carry `tool_result_provider_call` replay metadata, collapsing consecutive rows into one message positioned between the user message and the assistant response (the shape capture_turns_from_conversation_messages' per-turn lookahead consumes). Tool names always flow through so the value scorecard sees required_tools/replayable; raw tool payloads stay consent-gated downstream by include_tool_payloads. Rows without provider metadata remain dropped. TDD: - adapter unit tests: single tool call -> tool_calls message; consecutive calls collapse into one; ref without provider metadata still dropped. - integration guard: a captured tool-using turn's queued envelope carries replay.required_tools + replayable=true. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(reborn-traces): read capture history from context window, not display projection Tool-call reconstruction (previous commit) had no data to work with: the capture history source read SessionThreadService::list_thread_history, whose product-display projection (history_message) hard-nulls tool_result_provider_call. So even though tool calls persist with full provider metadata, the adapter received None on every tool row, dropped them, and produced a text-only envelope that scored below the 0.35 submission gate. Nothing ever submitted. SessionThreadHistorySource now reads load_context_window (the model-context/replay view, which preserves tool_result_provider_call) and maps ContextMessage -> ThreadMessageRecord via context_window_to_records. This is the semantically correct source for trace capture anyway: the replay transcript, not the display transcript. TDD: a caller-level test (per .claude/rules/testing.md "test through the caller") drives SessionThreadHistorySource against a real InMemorySessionThreadService with an appended tool result, asserting the returned tool row keeps provider_call. Failed on list_thread_history (None), passes on load_context_window. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(reborn-traces): auto-submit traces with PII risk below High Previously any non-Low residual PII risk was blocked from auto-submission two ways: the manual-approval eligibility gate held everything != Low, and the value scorecard halved the score (privacy_gate Medium 0.5) and subtracted a 0.60-weighted penalty. A minimal tool trace scores ~0.36 at Low (barely over the 0.35 gate), so any Medium penalty collapsed it to 0 — nothing below High could ever submit. Treat below-High residual risk as clean for auto-submission (the deterministic redactor has already scrubbed detected PII): - trace_autonomous_eligibility manual-approval gate now holds only High (== High, was != Low). - privacy_gate: Low|Medium => 1.0 (was Medium 0.5); High => 0.0. - privacy_risk_score: Low|Medium => 0.0 (was Medium 0.5); High => 1.0. High remains fully blocked: privacy_gate zeros its score and the gate holds it for manual review. The 0.35 submission gate leaves no headroom for a partial Medium discount on a minimal trace, so below-High is clean rather than partially penalized. TDD: medium_pii_tool_trace_auto_submits_while_high_is_held asserts a Medium-risk tool trace clears 0.35 and auto-submits while High is held. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs(reborn-traces): design for Trace Commons held-trace review Held traces are currently dropped on the autonomous capture path with no visibility or authorize path. This plan reuses the existing hold-sidecar machinery (TraceQueueHold / .held.json / read_trace_queue_holds_for_scope / ManualReview) and adds: retain held traces, surface a held count+list on the /traces/credit response, a card/tab UI, and a promote-as-is authorize endpoint. Four independently-shippable TDD slices. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(reborn-traces): retain manual-review held traces instead of dropping (slice 1) Autonomous turn-end capture dropped every held trace (logged at debug, envelope discarded), so PII-gated traces were unrecoverable and invisible. Slice 1 of the held-review feature retains manual-review holds: - TraceQueueEligibility::Hold now carries a typed TraceQueueHoldKind (ManualReview for the High residual-PII gate; PolicyGate for score / tool-allowlist / submission-class gates), replacing reason-string classification at the flush call site. - TraceClientAutonomousCaptureOutcome::Held carries the built envelope and its kind so callers can persist it. - New queue_trace_envelope_as_held_for_scope: queues the envelope plus a ManualReview .held.json sidecar under one scope lock; the flush worker already skips held sidecars, so it is retained but not submitted. - capture_turn_trace retains ManualReview holds and still drops PolicyGate holds (low-value traces never pollute the review surface). TDD: held-retain function (RED on missing sidecar -> GREEN), eligibility kind classification, and caller-level capture tests (an AWS-key message forces High PII -> retained ManualReview hold; a sub-threshold trace is dropped, not retained). Refs docs/plans/2026-06-10-trace-commons-held-review.md Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(webui-v2): surface manual-review held count + list on /traces/credit (slice 2) Held traces retained by slice 1 were invisible to the UI. Slice 2 surfaces them on the existing trace-credits response so one fetch powers the whole card/tab. - ironclaw_reborn_traces: manual_review_holds_for_scope() returns only ManualReview holds (excludes PolicyGate value-gates and transient RetryableSubmissionFailure retry holds), via an extracted retain_manual_review_holds filter. - RebornTraceCreditsResponse gains manual_review_hold_count + holds[] ({ submission_id, reason }). Sanitized: submission id and the already privacy-safe hold reason only, never raw trace content. TDD: retain_manual_review_holds filter unit test (excludes policy/retry), disk-level manual_review_holds_for_scope test, and the facade zero-state test asserts the new fields default empty. webui_v2 handler contract tests (42) still pass with the propagated fields. Refs docs/plans/2026-06-10-trace-commons-held-review.md Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(webui-v2): show held-for-review traces on card + Settings tab (slice 3) Surface the manual-review held count/list from slice 2 in the UI. Both render only when there are holds, so the common (nothing-held) state is unchanged. - Sidebar card: "{count} held for review" line when manual_review_hold_count > 0. - Settings -> Trace Commons tab: a "Held for review" section listing each held trace's sanitized reason + submission id from holds[]. - No hook/api change: fetchTraceCredits already returns the raw response, so credits.holds / credits.manual_review_hold_count are available. - Three i18n keys (cardHeld, heldTitle, heldDescription) across all 11 locales. The per-trace Authorize action ships with its endpoint in slice 4 (so the UI never offers a button that 404s). Source-shape assertions extended. Refs docs/plans/2026-06-10-trace-commons-held-review.md Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(webui-v2): authorize held traces for submission (slice 4) Complete the held-review feature with a promote-as-is authorize action across the stack. ironclaw_reborn_traces: - TraceContributionEnvelope gains `manual_review_authorized`; an authorized envelope submits past every gate in trace_autonomous_eligibility (the flush re-evaluates eligibility each pass, so removing the hold sidecar alone is not enough to promote). - authorize_manual_review_hold_for_scope: stamps the envelope (durable consent record) BEFORE removing the .held.json sidecar, so a crash between the two leaves the trace held (fail closed). Only ManualReview holds are authorizable; unknown submissions return Ok(false), not an error. ironclaw_product_workflow: - RebornServicesApi::authorize_trace_hold derives scope from the authenticated caller (the path submission id is never cross-scope authority), validates the id, and returns RebornTraceHoldAuthorizeResponse. ironclaw_webui_v2: - POST /api/webchat/v2/traces/holds/{submission_id}/authorize — NoBody, mutation rate limit, bearer auth. Descriptor + handler + router + contract table (now 46 routes). Frontend: - authorizeTraceHold api, an authorize mutation in useTraceCredits that invalidates the credits query on success, and a per-hold Authorize button on the Settings tab. `authorize`/`authorizing` i18n in all 11 locales. TDD: authorize promotes a High-PII held envelope past all gates; facade zero-state; webui_v2 descriptor/handler contracts; composition serve (47); source-shape assertions. clippy/fmt clean across crates. Refs docs/plans/2026-06-10-trace-commons-held-review.md Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(traces): loopback dev claim exception + profile_set consent gate Address the two codex P2 findings from review: - Preserve loopback claim uploads after onboarding: the loopback-HTTP dev invite form stores a loopback claim/ingest endpoint in the policy, but the claim/ingest validators required https and rejected loopback hosts, so a successful loopback onboarding could never mint a claim or submit credits. The validators and the pinned DNS resolution now honor the same literal-loopback exception as invite parsing (shared is_loopback_host predicate); for loopback hosts the pinned resolution additionally requires all resolved addresses to be loopback. Non-loopback http, internal hostnames, and private ranges stay rejected, and the issuer allowlist still applies. - Require explicit confirmation before community profile updates: trace_commons.profile_set now has the same hard confirmed=true input gate as onboarding — it short-circuits with consent_required before the enrollment check and any network write, since the capability is approval-gate-exempt in local-dev policy. Schema and manifest document the field. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(merge): thread attachments field through trace-capture record construction main added ThreadMessageRecord.attachments (Vec<AttachmentRef>); the trace-capture reconstruction path and its two test helpers construct records and must set it. The capture path reconstructs records from a context window for redaction/scoring and carries no attachment refs of its own, so Vec::new() is correct. * fix(traces): adapt v1 autonomous capture to new Held variant shape The merge brought in slice 1 of the held-trace-review feature, which changed TraceClientAutonomousCaptureOutcome::Held from { submission_id, reason } to { kind, reason, envelope } so manual-review holds can be retained instead of dropped. The v1 autonomous-capture path in thread_ops.rs still matched the old shape, breaking the `--no-default-features --features libsql` build (and default build). Adapt the v1 path to the new shape and give it the same retain-or-drop parity as the Reborn capture path (ironclaw_reborn_composition::trace_capture): ManualReview holds are retained via queue_held_envelope_for_scope (the on-disk held queue is shared, so a v1-captured hold surfaces in the v2 review UI); policy/value gates are dropped as before, just logged. Behavior mirrors the tested Reborn path (send_user_message_auto_queues_trace_for_enrolled_scope); the v1 autonomous-capture path is a detached tokio::spawn with no unit-testable seam, so no focused regression test is added. [skip-regression-check] * fix(traces): set manual_review_authorized in reborn-cli test envelope fixture The merge brought in the held-trace-review manual_review_authorized field on TraceContributionEnvelope. The reborn-cli trace_queue test fixture constructs the envelope directly and missed the field, breaking `cargo clippy --all-features --tests` and `Tests (all-features)` (the fixture is test-only, so the libsql binary build did not surface it). Fresh queued envelopes are not yet authorized, so false is correct. [skip-regression-check] * test(traces): pass confirmed=true in profile_set parity step The trace_commons first-party-tools parity test invoked profile_set without confirmed=true and asserted the NotEnrolled enrollment-gate result. Commit 6bc776d added the public-attribution consent gate to dispatch_profile_set, which now short-circuits to consent_required before the enrollment check when confirmed is unset — so the test's NotEnrolled assertion failed (the gate output carries no error_code). Pass confirmed=true so the call clears the consent gate and reaches the enrollment check, deterministically returning NotEnrolled with no network (the scope never onboarded). Matches the unit-test pattern established for the other profile_set tests in the same change. [skip-regression-check] * fix(traces): onboarding-security + contribution correctness (coderabbit batch 1) Addresses 6 coderabbit findings in ironclaw_reborn_traces: - device_key.rs: re-assert 0o700 on pre-existing key dirs (not just on create), so broader perms on an existing device_keys/ or pending/ can't leave invite/tenant hashes enumerable. - device_key.rs: fail closed on load when on-disk public_key/device_key_id don't match the loaded private key (tampered/partial files no longer load an inconsistent identity that only fails later at remote auth). - invite.rs: scope the staged pending-key filename by invite ORIGIN, not just code, so two issuers reusing one invite code can't share a device key (invite_hash stays code-only as the server allowlist subject). - onboarding/mod.rs: reject ingest_url values with embedded userinfo before persisting, so a malicious onboarding response can't smuggle credentials into policy.json + outbound requests. - contribution.rs: preserve mount path prefixes when deriving the community-profile endpoint (mirrors trace_submission_status_endpoint); a prefixed deployment no longer 404s on profile PUT/DELETE. - contribution.rs: fail closed in trace_autonomous_eligibility on envelopes with no allowed-uses (public_attribution-only) instead of relying on the remote to bounce them. Updated two retry tests that encoded the cross-issuer key-sharing bug now fixed: they retried against a second mock on a different port; a new spawn_flaky_mock_issuer keeps the retry on the same origin so it exercises genuine same-issuer pending-key reuse. Added regression tests for each fix. * fix(trace-commons): address coderabbit review findings on nearai#4559 - index.html: add type="button" to the Trace Commons settings subtab to prevent accidental form submission. - settings.js + i18n/en.js: route the Trace Commons credits copy through I18n.t(...) and register the matching locale keys (matches the existing surface pattern; en-only like settings.traceCommons, fallback covers rest). - factory.rs: drive the malformed SECRETS_MASTER_KEY env case through the real caller resolve_local_dev_secret_master_key (via an env-parameterized inner) and assert the rejected key is never persisted to the cached file. - trace_commons_dispatch_e2e.rs: give each test a distinct user/extension scope so onboarding state can no longer bleed across tests. - local_dev_capability_policy.toml: exempt builtin.trace_commons.onboard from the REPL approval gate (it has its own confirmed=true consent gate, mirroring profile_set). - docs: fix the onboard prompt-file reference, match the held-trace JSON shape to RebornTraceHold (submission_id + reason only), and resolve the wire-protocol ownership split (types live locally in onboarding/protocol.rs, no shared trace-commons-protocol crate). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * fix(traces): tenant-scoping + token leak + read-failure + unbounded scopes (coderabbit batch 2) Addresses the coupled backend findings: - Tenant-scope Trace Commons local state across the Reborn paths: new trace_scope_key(tenant, user) helper keys policy / device-key / credit / profile / capture state by tenant+user, so the same user id in two tenants no longer shares state. Applied in host_runtime trace_commons dispatchers, product_workflow credits/hold, and composition trace-capture (v1 stays user-only — legacy single-tenant). Updated the affected runtime/sink tests and added a non-owner attribution assertion. - Do not return the raw profile token from the model-visible profile_token capability: persist it to a 0600 <scope>/profile_token.jwt and return the file path + instructions instead, keeping the bearer credential off the LLM transcript. - Stop masking genuine local-state read failures as zero/not-enrolled: the status capability and the WebUI credits path now propagate a read/parse failure (NotFound is already softened inside read_*_for_scope) so an enrolled user with a corrupt policy file is not told they have nothing. - Bound ObservedTraceScopes: the periodic flush worker now prunes drained scopes (new trace_scope_has_pending_queue) after each tick, so the set is bounded by actual pending backlog instead of growing one entry per caller ever seen. Note: a v1 caller-level test for the ManualReview hold-retention path is not included — v1 ingress blocks secrets outright and the outbound leak detector redacts them, so the High-residual-PII condition that produces a ManualReview hold cannot be reproduced through process_user_input. The retention logic is identical to and covered by the Reborn-side capture_retains_manual_review_hold_for_high_pii_trace. * test(traces): enroll under tenant-scoped key in webui_v2_serve credits test trace_credits_reports_enrolled_for_caller_with_enabled_policy wrote the policy under the bare user id, but the credits route now keys local state by trace_scope_key(tenant, user). Enroll (and clean up) under the composite TENANT/user scope so the route sees the enrollment. * fix(factory): fail closed on explicit-but-unusable SECRETS_MASTER_KEY An explicitly-set-but-unusable local-dev master key silently fell through to generating + persisting a fresh key, leaving local-dev secrets encrypted under an unintended master key the operator never chose: - resolve_local_dev_secret_master_key used std::env::var(...).ok(), which drops VarError::NotUnicode -> treated as absent. Now only NotPresent is absent; a non-Unicode value returns InvalidConfig. - resolve_local_dev_secret_master_key_with_env collapsed a set-but-empty (or whitespace-only) value to None via .filter(). Now a set-but-empty value returns InvalidConfig instead of generating a key. Added resolve_local_dev_secret_master_key_rejects_set_but_empty_env_without_persisting asserting empty/whitespace env values fail closed and persist nothing. (coderabbit follow-up on nearai#3794) * fix(factory): reject empty SECRETS_MASTER_KEY before the cached-file read Follow-up to the prior fix: the empty-env rejection lived in the env branch, which only runs when no cached key file exists. On a rebuild where .reborn-local-dev-secrets-master-key already exists, the cached key was returned first, so an explicitly-set-but-empty SECRETS_MASTER_KEY was still silently ignored. Hoist the empty/whitespace rejection (and env normalization) above the cached-file read so it fails closed regardless of cached state. Added resolve_local_dev_secret_master_key_rejects_empty_env_even_with_cached_file asserting the empty env is rejected and the cached key is left unchanged. * fix(traces): address 14:54 coderabbit re-review (tenant-seed, IO errors, effects, test) Four outside-diff findings from the re-review: - runtime.rs: seed ObservedTraceScopes with the runtime owner's trace_scope_key(tenant, owner) composite, not the bare owner id, so startup pending-queue discovery matches how capture keys state; the enrolled-scope test cleanup now removes the composite scope dir too. - runtime.rs: the trace-queue polling test helper no longer swallows read_dir errors via unwrap_or_default() — only NotFound is the expected pre-capture fallback; any other IO error panics instead of masking as 'no queued traces'. - trace_commons.rs manifests + local_dev grants: onboard (device-key material) and profile_token (0600 token file) now declare Read/WriteFilesystem effects, and the local-dev grants allow them, so the effect model accurately models the local secret-material writes. - local_dev_authorization test: added local_dev_trace_commons_onboard_skips_approval_gate (the onboard exemption was the actual fix; the profile_set-only test would pass even if the onboard TOML exemption were dropped). * fix(factory): validate non-empty SECRETS_MASTER_KEY before the cached-file read Follow-up: the prior fix rejected an *empty* env value before the cached read but still validated a non-empty *malformed* value only after it. So a valid cache + SECRETS_MASTER_KEY=0000... silently ignored the explicit bad secret config on rebuilds. Move validate_resolved_master_key into the up-front env normalization so any explicit-but-unusable env key (empty OR malformed) fails closed regardless of cached state. Added resolve_local_dev_secret_master_key_rejects_malformed_env_even_with_cached_file. * fix(traces): address 15:41 coderabbit re-review (credits read-failure + 2 test guards) - trace_commons.rs dispatch_credits: stop masking genuine records read/parse failures as 'no records' (NotFound is already softened inside read_local_trace_records_for_scope); report RecordsReadFailed, mirroring dispatch_status. - runtime.rs trace-queue polling helper: fail loud on per-ENTRY read_dir IO errors too (map + unwrap_or_else panic) instead of filter_map(e.ok()), so a broken entry can't be silently dropped while claiming the queue holds one. - local_dev_authorization approval-gate test: assert the effects DO require approval without the exemption (local_dev_effects_require_approval), so the test can't pass via a non-gating default policy if the TOML exemption were dropped. * fix(traces): address Henri review — backend findings (atomic token, error mapping, validation, egress test) - persist_profile_token now writes atomically (unique 0600 temp + fsync + rename) so a reader never observes a half-written or overwritten bearer credential under overlapping mints (Henri perf/security Medium). - dispatch_onboard error mapping: OnboardError::DeviceKey is reported as a distinct DeviceKeyError (re-run onboarding) instead of being collapsed into PersistError's check-disk-and-permissions guidance (Henri bugs Medium). - parse_profile_set_input enforces the manifest's declared schema at parse time: handle 3-32 ASCII letters/digits/-/_, bio <= 280 bytes (Henri conventions Medium). Added schema-limit test. - Added dispatch_onboard_confirmed_without_host_egress_is_network_denied covering the NetworkDenied host-egress-miswiring branch (Henri tests Medium). * fix(traces): address Henri review — frontend findings (enrolled empty-state + polling) - v1 credits: TraceCreditResponse now carries `enrolled` (read from the standing policy), and settings.js keys the opt-in empty state on `!data.enrolled` instead of `!submissions_total` — an enrolled user with zero submissions now sees their zero-credit view, not the not-enrolled prompt (Henri bugs Medium). - useTraceCredits: each fetch rebuilds the full server-side credit view, so the aggressive 60s poll made an open tab steady O(history) work. Relaxed to a 5-min interval + staleTime + no background polling, keeping a focus refetch for liveness; mutation invalidation still updates promptly. Added a TODO to incrementalize the server-side view (Henri perf Medium). * perf(traces): memoize server-side credit view by on-disk input signature Bounds the trace-credits polling cost to O(new submissions) instead of O(total history). New scoped_credit_view(scope) caches the computed credit report + manual-review holds keyed by a cheap change signature (submissions file mtime+len, plus a hash of the held-trace sidecars). On the steady-state polling case (unchanged history) a request is a couple of stat()s + a clone rather than reading/parsing the full submissions file and re-aggregating. On any change the signature differs and it recomputes once. Cache is bounded (4096 scopes, cleared on overflow). Wired through the polled WebUI path (local_trace_credits_for_user) and the model-visible credits capability (dispatch_credits). Added scoped_credit_view_reflects_record_changes_via_signature covering the cache-hit path and signature-based invalidation on record changes. Completes the TODO from the Henri perf-review follow-up (#5). * fix(traces): gate profile_set behind runtime approval (Henri #1 High) profile_set publishes a public community profile (an external write to a public surface). Its `confirmed=true` input is model-controlled, so a prompt-injected or confused model could supply it. Make the runtime approval gate the primary, user-controlled consent control: - Drop `builtin.trace_commons.profile_set` from the local-dev approval-gate exemption list (keep `onboard`, which runs its own in-turn confirmed=true consent before the network POST). - Set profile_set's manifest default_permission to Ask (was Allow). - Split the local-dev authorization test into `local_dev_trace_commons_profile_set_requires_approval_gate` (asserts Decision::RequireApproval) and `local_dev_trace_commons_onboard_skips_approval_gate` (asserts Decision::Allow), via a shared `trace_commons_authorize_decision` helper that first asserts the effects would gate without an exemption. Also fix a pre-existing trace_commons harness gap: onboard + profile_token gained a WriteFilesystem effect (device-key persistence) but the `trace_commons_tools` harness allow-set was never updated, so those capabilities were filtered out of the model-visible surface and the parity/visibility tests failed with driver_unavailable. Grant WriteFilesystem in the harness allow-set. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * refactor(traces): extract onboarding test harness to sibling file (Henri #8) The onboarding module's ~840-line `#[cfg(test)] mod tests` block (mock issuer harness, retry/idempotency coverage, URL-validation tests) made `onboarding/mod.rs` a 1319-line file dominated by test scaffolding. Move the module body into `onboarding/tests.rs` declared `#[cfg(test)] mod tests;`, leaving mod.rs focused on production logic (now 480 lines). No test behavior changes; `use super::*;` still resolves to the onboarding module. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(webui-v2): update embedded-asset assertion for incrementalized credits poll The Henri #5 polling fix changed useTraceCredits.js from refetchInterval 60_000 to 300_000 (plus refetchIntervalInBackground: false and staleTime: 60_000), but the embedded-asset test in assets.rs still asserted the old 60_000 value and failed in CI. Update the assertion to lock the new infrequent-poll + paused-while-hidden + focus-refetch shape. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(traces): address CodeRabbit review + stale capability-policy test CodeRabbit findings on the gating/refactor commits: - Major: format_profile_token returned the absolute host path of the token file (token_file) on the model-visible surface, which violates the "never expose absolute paths" guideline. Replace with an opaque token_delivery marker; the token is still persisted 0600 for out-of-band retrieval by a bearer-auth UI/CLI. Update the message + test accordingly. - Major (fail-loud): profile_token_error_value and profile_set_error_value collapsed "could not read policy" into NotEnrolled, sending enrolled users back through onboarding on unreadable/corrupt state. Split into a distinct PolicyReadFailed result in both formatters (matches dispatch_status). - Minor: stale comment claiming profile_set is approval-gate-exempt (it is now PermissionMode::Ask and NOT exempt) — corrected. - Minor: inaccurate harness comments (profile_token writes profile_token.jwt not device-key material; yolo auto-approves all Trace Commons Ask-gated tools, not just onboard) — corrected. Also fix bundled_local_dev_capability_policy_parses, which still asserted the pre-gating policy shape: profile_set as exempt (now onboard exempt / profile_set NOT exempt), onboard's grant missing the read/write filesystem effects, and profile_token/profile_set sharing one effect-set assertion even though profile_token now carries WriteFilesystem and profile_set does not. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * style(traces): collapse single-line use block after Path import removal rustfmt collapses `use std::{panic, path::PathBuf, sync::Arc}` to one line once Path was dropped; the prior commit skipped re-running fmt after that edit, reddening the Formatting CI check. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(traces): consent-gate profile_token + drop fixed-origin profile URL (CodeRabbit) Two Major CodeRabbit security findings on the profile tools: - profile_token minted and persisted a bearer credential with no in-turn consent gate. PermissionMode::Ask can be auto-approved under local-yolo, so a model call could mint a credential without explicit per-conversation consent. Add a hard confirmed=true gate (schema + parse + consent_required short-circuit) before minting, mirroring dispatch_onboard / dispatch_profile_set. - format_profile_token and profile_set_success_value hardcoded https://tracecommons.ai/profile. The token is scoped to the user's ENROLLED issuer (which may be self-hosted or loopback), so steering the user to paste a bearer profile-management token at a fixed origin could leak it to the wrong host. Drop the fixed profile_url; route through the enrolled profile flow / local UI/CLI out of band. Tests: new dispatch_profile_token_without_confirmed_returns_consent_required_no_mint; existing without-enrollment test now passes confirmed=true; profile_set success test asserts no fixed origin; parity step mints with confirmed=true. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(traces): route agent-invoked profile writes through host egress (CodeRabbit #3) profile_token (upload-claim mint) and profile_set (community-profile PUT/DELETE) previously made network writes via the ironclaw_reborn_traces crate-local reqwest client, bypassing the host RuntimeHttpEgress pipeline (private-IP filtering, redaction, byte accounting) that onboard already uses. Add a `ContributionHttpSink` port (mirroring `OnboardingHttpSink`): when a sink is injected, the mint POST and the profile PUT/DELETE run through host egress; when `None`, the existing hardened crate-local client is used unchanged. host_runtime supplies `HostEgressContributionSink` (wraps RuntimeHttpEgress, sanitizes errors via stable_runtime_reason, never leaks URL/token), and dispatch_profile_token / dispatch_profile_set fail closed with NetworkDenied if egress is absent (after the enrollment pre-check, so a not-enrolled user still gets NotEnrolled guidance). The background trace-upload / status-sync worker and the CLI keep the crate-local client (pass `None`): that lane is a durable, model-input-free internal task that sends only already-redacted envelopes to the operator-enrolled endpoint and does its own SSRF/private-IP validation, so host egress adds complexity without security benefit. Justification recorded in a comment on `trace_remote_http_client`. New public surface: ContributionHttpSink/Request/Response/Error/Method, mint_profile_attribution_token_for_scope_via_sink, set_community_profile_for_scope_via_sink. Existing public fns keep their signatures (None path) so CLI/worker/tests are unchanged. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
theredspoon
added a commit
that referenced
this pull request
Jun 21, 2026
Only allow the Docker Image publishing job to run from nearai/ironclaw so the control branch copy cannot publish from theredspoon.
personal-upstream-sync Bot
pushed a commit
that referenced
this pull request
Jul 8, 2026
… inspection (nearai#5280) * docs: spec for Trace Commons instance enrollment, profiles, and trace inspection Cross-repo design (ironclaw + trace-commons-server) for three coexisting capabilities: instance-wide enrollment, per-user contributor accounts via login-links, and submitted-trace inspection. Introduces a trace-credential resolver so the existing user-invite model and the new instance-wide model both function on one instance, with personal-invite enrollment taking precedence. Server change is additive (optional per-user subject through claim issuance + login-link + account resolution). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: Slice 0 plan — trace-commons-server per-user subject TDD plan for the one server change the whole effort depends on: accept an optional opaque subject in the upload-claim request and derive a per-user, tenant-namespaced principal at device-key issuance. Submission attribution, login-link account resolution, and trace readback all become per-user automatically from the shared bearer principal; absent subject reproduces today's behavior. Targets trace-commons-server (contributor-account-slice1). Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * docs: IronClaw plans for Trace Commons slices 1-4 Slice 1: trace-credential resolver (personal-invite wins, instance fallback with per-user subject) + admin-gated instance enrollment. Slice 2: per-user subject plumbing through upload-claim request + submission. Slice 3: trace_commons.account_login_link first-party capability (profiles). Slice 4: per-user submitted-trace inspection across reborn_traces → product_workflow facade → webui_v2 handler → frontend. Each plan is bite-sized TDD against verbatim-extracted current code. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(traces): trace-credential resolver (personal invite wins, instance fallback w/ subject) * refactor(traces): single dir-parameterized policy-path site (remove resolver duplication) Extract `trace_contribution_dir_for_scope_at`, `trace_policy_path_at`, `read_trace_policy_for_scope_at`, and `write_trace_policy_for_scope_at` as the canonical base-dir-parameterized path helpers. All public functions (`trace_contribution_dir_for_scope`, `read_trace_policy_for_scope`, `write_trace_policy_for_scope`) now delegate to the `_at` variants with `ironclaw_base_dir()` — signatures unchanged. The inline `read_policy` closure in `resolve_trace_credentials_at` that re-implemented path layout is deleted; it now calls `read_trace_policy_for_scope_at` directly. The test `write_policy_at` helper's bespoke path construction is replaced with a call to `write_trace_policy_for_scope_at`. The now-dead `trace_policy_path` function is removed. Path layout is encoded in exactly one place. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(traces): instance-level enrollment write path (scope None) * test(traces): make instance-enrollment test hermetic (tempdir, no global base) Rework `instance_onboard_writes_instance_level_policy` to operate entirely under a `tempfile::tempdir()`: - Compute instance_dir as base.path().join("trace_contributions") (scope=None layout, no users/<hash> segment) rather than calling the global LazyLock. - Call `onboard_at_dir_with_sink` directly against the tempdir so the test never touches the real ~/.ironclaw tree. - Assert policy.json by reading and deserializing it from the tempdir. - Remove all manual std::fs::remove_* cleanup lines; tempdir drops automatically. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(admin): AdminScope::enroll_instance_trace_commons (admin-gated instance enrollment) Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(traces): carry optional per-user subject in upload-claim request Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(traces): thread resolver subject into submission claim context * test(traces): claim request carries per-user subject end-to-end * feat(traces): mint_account_login_link_via_sink (POST /v1/account/login-links) Add `mint_account_login_link_via_sink` to ironclaw_reborn_traces: - `TraceUploadClaimContext::for_account(subject)` constructor for account-management call contexts (no trace/submission ids, no consent scopes). - `AccountLoginLink { account_id, url }` return type. - `account_login_links_url(policy)` helper that derives the login-links URL from the upload-claim issuer URL (strip /v1/trace-upload-claim, append /v1/account/login-links). - `mint_account_login_link_inner(base_dir, ...)` private dir-parameterised core: resolves credentials, selects correct scope_dir for DeviceKey auth (instance enrollment → instance scope dir; personal → user scope dir), mints bearer, POSTs subject, parses response. - `mint_account_login_link_via_sink(tenant_id, user_id, sink)` public entry point wrapping the inner function with the real base dir. Tests (hermetic, tempdir-isolated): - `mint_account_login_link_posts_subject_and_returns_url`: verifies the posted subject equals `local_pseudonymous_contributor_id(trace_scope_key(...))` for instance-enrolled users via an axum mock serving both the upload-claim issuer and the login-links endpoint. - `mint_account_login_link_errors_when_not_enrolled`: verifies error path. - `ReqwestContributionSink` test helper added to the test module. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(traces): error instead of silent misroute in account_login_links_url Replace the unwrap_or_else fallback (which silently used the full issuer URL as a base when the /v1/trace-upload-claim suffix was absent) with an explicit anyhow error. Add two unit tests: one asserting an Err on a wrong-suffix URL, one asserting the correct .../v1/account/login-links URL on a valid issuer. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(host_runtime): add consent-gated trace_commons.account_login_link capability Mints a Trace Commons browser login URL via host network egress, mirroring dispatch_profile_token. Includes consent gate, enrollment pre-check, HostEgressContributionSink routing, and two e2e tests. Also fixes a sanitizer bug: validate_runtime_request was rejecting authorization headers on all requests, including RuntimeKind::FirstParty. FirstParty requests are host-internal and trusted to carry bearer tokens; the sensitive-header and manual-credentials guards now only apply to untrusted plugin runtimes (WASM/MCP/Script). Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * fix(host_runtime): route trace bearer via credential injection; restore FirstParty sensitive-header guard Commit 9e25d99 blanket-exempted all RuntimeKind::FirstParty requests from the egress sensitive-header and manual-credentials guards so the host-minted Trace Commons bearer could pass. builtin.http is also FirstParty but forwards model-supplied headers, so this let the model smuggle Authorization/Cookie/ x-api-key headers (or user:pass@ URLs) to allowlisted hosts. Revert the sanitize.rs exemption (guards now apply to ALL runtimes again) and deliver the trace bearer through the staged credential-injection path instead: the HostEgressContributionSink stages the minted token one-shot via RuntimeSecretMaterialStager and declares a StagedObligation Authorization-header injection, mirroring the SlackProtocolHttpEgress pattern. The stager is now exposed to first-party handlers via InvocationServices. Covers the profile_token, profile_set/community-profile, and account_login_link bearer paths. Regression tests: FirstParty + raw authorization header -> denied; FirstParty + user:pass@ URL -> denied. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(traces): fetch_account_traces_via_sink (GET /v1/account/traces, per-user) - Add ContributionHttpMethod::Get variant; update all exhaustive match sites in ironclaw_reborn_traces and HostEgressContributionSink in ironclaw_host_runtime. - Extract account_api_base_url() shared helper; account_login_links_url and new account_traces_url both delegate to it (DRY). - Add AccountTraceItem (Debug, Clone, Serialize, Deserialize; serde defaults). - Add fetch_account_traces_via_sink / fetch_account_traces_inner mirroring mint_account_login_link pattern: unenrolled -> Ok(vec![]), non-2xx -> Ok(vec![]), transport error -> Err. - Tests: hermetic axum mock (GET /v1/account/traces), unenrolled empty-list, URL shape with/without limit. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(reborn): trace_account_traces facade method + wire types Adds RebornAccountTrace / RebornAccountTracesResponse wire types and a trace_account_traces default method on RebornServicesApi, mirroring the trace_credits egress pattern (crate-local hardened reqwest, no host-egress sink). Also adds fetch_account_traces (direct path) to ironclaw_reborn_traces::contribution so the facade can fetch server traces without coupling to RuntimeHttpEgress. Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * feat(reborn): GET /api/webchat/v2/traces/account handler + contract test * feat(reborn-ui): render submitted Trace Commons traces in settings Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * chore(traces): document flush-gate limitation, hermetic account-traces contract test, annotate sink scaffold Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com> * style(traces): cargo fmt across Trace Commons slice changes Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * feat(traces): resolver-aware flush gate (instance-enrolled users can contribute) The autonomous trace-flush gate read only the per-scope (personal-invite) policy and aborted when it was disabled, so instance-enrolled users (whose enrollment lives at scope None) could never contribute traces — and the per-user scope_dir would also fail to load the instance device key. Introduce a single EffectiveFlushTarget resolver (resolve_effective_flush_target, mirroring resolve_trace_credentials but keyed on the already-composed scope string) that returns the policy, device-key dir, and per-user subject in one policy-read/path pass. The flush gate now proceeds for instance-only enrollment, loads the device key from the instance (None) dir, and attributes uploads via the per-user pseudonymous subject. The redundant subject_for_scope helper (which re-read the same policies with silent .ok() error swallowing) is removed and its logic folded into the new helper with proper error propagation. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(traces): include per-user subject in upload-claim cache key Under instance enrollment every user shares the same instance device-key dir (scope None), so the upload-claim cache key — which keyed on scope_dir but not subject — collided across users. A bearer minted for one subject could be served from cache to another, mis-attributing traces / leaking across users. Add a hashed subject component to the DeviceKey cache key and a regression test proving two subjects sharing a scope_dir get distinct keys (and a no-subject personal-invite context stays distinct from both). Found by Codex review of PR nearai#5280. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(traces): address CodeRabbit review on PR nearai#5280 - account_login_link manifest: declare ReadFilesystem effect (it reads local enrollment/policy/device-key state before egress), matching profile_token. (CR #2) - account-traces fetch: always send a bounded, clamped limit ([1, 500], default 200) so None never triggers an unbounded server history fetch. (CR #3) - direct fetch path: bound the response body with a hard byte ceiling (256 KiB) via a chunked bounded reader, instead of buffering unbounded. (CR #5) - account-traces fetch (both sink + direct): stop swallowing every non-2xx as an empty list — 404 = legitimate empty (no account yet), all other non-2xx surface as Err so the WebUI renders a sanitized unavailable state. Add regression tests (500 -> err, 404 -> empty). (CR #6) - trace-commons-tab.js: render missing final_credit as "—" not "0.00"; surface useAccountTraces() query errors instead of collapsing them to "no traces". (CR #7, #8) - handlers contract test: capture the forwarded caller in the trace_account_traces stub and assert the route threads the authenticated user id (test-through-the-caller). (CR #9) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(ci): cover trace_commons.account_login_link + backfill trace i18n keys PR nearai#5280 added the builtin.trace_commons.account_login_link capability and a submitted-traces UI section, but left three guardrail/parity tests un-updated, turning CI red: - ironclaw_host_runtime builtin_first_party_package_declares_expected_capabilities: register account_login_link in the expected id list and its Ask-permission arm. - reborn_builtin_first_party_capability_e2e_coverage_is_complete: add genuine e2e coverage by exercising account_login_link in the existing trace_commons parity test (confirmed=true on a not-enrolled scope returns a deterministic NotEnrolled, no network), grant it in the harness allow-set, and add it to the model-visible surface test and the covered-capability list. - ironclaw_webui_v2_static all_locales_share_the_en_key_set: backfill the six new traceCommons.* submitted-traces keys into all ten non-en locales. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(traces): address CodeRabbit review — withhold login URL, type errors, wire i18n - Security (host_runtime): dispatch_account_login_link returned the one-time login `url` (a code-bearing account-access credential) on the model-visible surface, persisting it into the LLM transcript and any downstream logging. Follow the profile_token pattern: persist the URL to a 0600 private file (atomic temp+rename) and return an opaque `link_delivery` marker instead. E2e test now asserts the URL/code never appears in the result and is delivered out-of-band to the private file. - Typed error (product_workflow): account_traces_for_user flattened backend errors into String before the WebUI boundary. Introduce AccountTracesError (thiserror) that names the failing operation and preserves the full cause chain ({:#}); the boundary keeps returning a sanitized, diagnosable 500. Also document that fetch_account_traces(None) is already server-bounded (default 200, clamp 500, 256 KiB response cap) — the "unbounded fetch" concern was resolved by prior hardening. - i18n (webui_v2_static): the traceStatus and traceReceivedAt keys backfilled for locale parity were unused by the consumer. Wire traceStatus as the status badge's accessible title/aria-label and render traceReceivedAt as the timestamp label, so all six submitted-traces keys are now consumed. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(traces): address CodeRabbit re-review — async persist + typed identifiers - Blocking I/O (host_runtime): persist_account_login_link does mkdir/write/ fsync/rename with std::fs on the async dispatch path. Wrap the persist call in tokio::task::spawn_blocking so it never stalls a Tokio worker (coding guideline: all I/O async). Atomic temp+rename behavior is unchanged; a join failure maps to the same sanitized "could not write" result. - Typed identifiers (product_workflow): account_traces_for_user took bare &str tenant/user; the caller already holds TenantId/UserId newtypes. Take &TenantId/&UserId and only cross to &str at the ironclaw_reborn_traces boundary. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(traces): instance-aware enrollment across trace_commons dispatch + UI Instance-only-enrolled users (admin-provisioned instance policy, no personal invite) were falsely rejected across the Trace Commons surface: the dispatch gates and the profile mints read only the personal per-scope policy, and the submitted-traces UI was gated behind the personal-credits branch. Addresses CodeRabbit re-review (#3, #4, #5) on PR nearai#5280. reborn_traces: - Add instance-aware entry points mint_profile_attribution_token_for_user_via_sink and set_community_profile_for_user_via_sink that resolve enrollment via resolve_trace_credentials (personal OR instance) and build the claim context with the instance scope_dir + per-user pseudonymous subject, mirroring mint_account_login_link_inner. Refactor the token mint to share a context-based core. New tests assert the per-user subject reaches the issuer. host_runtime (trace_commons dispatch): - Route the enrollment gates in dispatch_status, dispatch_profile_token, dispatch_profile_set, and dispatch_account_login_link through resolve_trace_credentials so instance-only contributors pass. status now reports the resolved (instance or personal) policy. profile_token/profile_set call the new instance-aware mints. - #4: preserve the stage_secret_material_once failure cause (log it) instead of discarding it with map_err(|_|); wire message stays sanitized. webui_v2_static (#3): - Lift the submitted-traces section out of the credits/empty-state branch so instance-enrolled users with no personal credits still see their traces and any tracesQuery errors. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * test(traces): isolated dispatch-layer e2e for instance-only enrollment Extract the trace_commons dispatch e2e helpers into a shared tests/support/trace_commons_dispatch.rs module (base-dir setup, mock issuer, runtime/dispatch helpers, find_persisted_login_link, test_jwt_eddsa) so a second test binary can reuse them. Add trace_commons_instance_dispatch_e2e.rs — a SEPARATE binary (fresh process = private IRONCLAW_BASE_DIR) that provisions the process-global instance policy (scope None) without bleeding into the personal-invite suite. It pins the CodeRabbit #5 fix at the layer it manifests: an instance-only-enrolled user (no personal invite) passes dispatch_status and dispatch_account_login_link and mints under the shared instance device key with a per-user pseudonymous subject (asserted via the subject on the login-links POST). No production changes; trace_commons_dispatch_e2e.rs behavior is unchanged (5 tests still pass) — only its helpers moved to the shared module. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(traces): sanitize bearer-staging log + typed IDs on mint entry points Addresses CodeRabbit overnight review on PR nearai#5280. - Security (#1): the trace-bearer staging error was debug-logged via `?error`, which can leak secret-store/backend detail on the credential path. The host_runtime logging guideline forbids backend error detail here — log only the safe fact of failure; the wire message stays sanitized. (Supersedes the earlier "preserve cause" change specifically on this bearer-material path.) - Typed identities (#2): the three agent-facing Trace Commons mint entry points (mint_account_login_link_via_sink, mint_profile_attribution_token_for_user_via_sink, set_community_profile_for_user_via_sink) now take &TenantId/&UserId instead of adjacent &str, so callers can't transpose tenant/user and misattribute a contributor. Identity stays typed to the public boundary and is stringified only when handing off to the dir-parameterised `_inner` cores / resolver (the storage edge). Adds ironclaw_host_api as a reborn_traces dependency (no cycle: host_api does not depend on reborn_traces). Dispatch callers pass the typed scope ids directly. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(traces): sanitize persist-path logs, preserve handle-validation cause Two follow-up CodeRabbit findings on PR nearai#5280: - Security (Major): dispatch_account_login_link's spawn_blocking persist arms logged %error / %join_error at debug. Filesystem errors (mkdir/write/fsync/ rename) can carry raw host paths, which the host_runtime guideline forbids in logs. Drop the interpolation; log only the generic fact, keep the message sanitized — same treatment as the bearer-staging path. - Maintainability (Minor): SecretHandle::new(TRACE_COMMONS_BEARER_HANDLE) used map_err(|_| ...), discarding the cause (non-exemptible per the guideline). The handle name is a compile-time constant, so its validation error carries no secret/path — bind and log it. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(traces): typed login-link errors, per-request bearer handle, doc accuracy Addresses the third CodeRabbit review round on PR nearai#5280. - Security (Major): the trace-bearer staging used a constant SecretHandle (TRACE_COMMONS_BEARER_HANDLE). The injection store is a HashMap keyed by (scope, capability, handle) with overwrite-on-insert, so two concurrent same-scope Trace Commons egresses could race and stage/consume the wrong bearer. Suffix the handle with a per-request uuid so every staged bearer key is distinct. Localized to the shared HostEgressContributionSink, so all trace_commons flows benefit. - Correctness (Major): account_login_link_error_value classified failures by substring-matching upstream error wording, coupling the public error_code contract to phrasing. Introduce a typed AccountLoginLinkError (thiserror) in reborn_traces; mint_account_login_link_via_sink returns it, producing the specific variant at each failure site. The host maps variants -> error_code with no substring checks. NotEnrolled (the only tested code) is preserved; the two bearer-derived codes collapse into EnrollmentIncomplete (both meant "re-run onboarding"), and persist failures get a distinct LocalStateWrite. - Docs (Minor): the persist_account_login_link comments promised 0600 across platforms though only Unix enforces it. Softened to "private local file (0600 on Unix)". Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * refactor(traces): type profile_token/profile_set error mappers (systemic) Follow-up to the account_login_link typed-error change: convert the remaining substring-based error mappers so all four trace_commons dispatch flows derive the public error_code contract from typed variants instead of matching upstream error wording. (onboard was already typed via OnboardError.) - reborn_traces: add ProfileAttributionError (shared by the profile_token and profile_set token mints) and CommunityProfileError (profile_set wrapper adding InvalidProfile). mint_profile_attribution_token_for_user_via_sink and set_community_profile_for_user_via_sink now return these; each failure site produces the specific variant (NotEnrolled / PolicyRead / EnrollmentIncomplete / Backend / LocalStateWrite, plus InvalidProfile for profile_set). - host_runtime: profile_token_error_value / profile_set_error_value now match on the typed variants — no error.contains(...) anywhere in the file. NotEnrolled and InvalidProfile (the tested codes) are preserved; the issuer/device/refused substrings collapse into EnrollmentIncomplete, consistent with the account_login_link mapping. Also sanitized the profile_token persist-failure log (host-path leak class), matching the login-link path. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(traces): split enrollment precondition from backend in token mints CodeRabbit re-review: collapsing every mint_profile_attribution_token_with_context (and bearer_token) failure into EnrollmentIncomplete mislabels transient transport/status/serde failures as "re-run onboarding". Split the local precondition (upload-claim issuer URL configured) from post-resolution failures: a missing issuer URL maps to EnrollmentIncomplete via an explicit upload_claim_issuer_missing() check (typed, no substring), while the claim mint / bearer fetch / PUT failures now map to Backend. Applied consistently across profile_token, profile_set, and account_login_link so the error_code contract reflects the real failure class. URL-derivation preconditions (ingest/login-links URL) stay EnrollmentIncomplete. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(traces): check login-link URL precondition before minting bearer Fail-closed ordering: the local account_login_links_url derivation ran after bearer_token, so a malformed/absent login-links URL would mint a device-key bearer and hit the issuer before failing. Move that local precondition ahead of all secret/egress work so incomplete enrollment fails closed with no side effects. (profile_token/profile_set already order local preconditions first.) Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(traces): make upload-claim cache key match issuer payload exactly The subject cache-key component trimmed/collapsed context.subject, but the DeviceKey issuer request sends it unchanged — so None, Some(""), and whitespace variants could share a cache key while minting different payloads, letting one user's claim be served from cache to another (cross-user trace mis-attribution). This is nearai#5280's per-user-subject cache-key path. Hash the exact optional bytes the request sends (DeviceKey → subject, WorkloadTokenEnv → None) with a None/Some discriminator. Extend the cache-key test with the Some("")-vs-None and whitespace-variant collision cases. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(traces): check response-size cap before growing the buffer Both bounded response readers (upload-claim and account-traces) enforced the hard byte ceiling only after extend_from_slice, so a single oversized chunk could push the buffer past the advertised limit before the error returned. Compute bytes.len() + chunk.len() (checked_add) and validate before appending. Pre-existing pattern (from nearai#4559), fixed here per review. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(traces): fail loud when trace policy cannot be statted read_trace_policy_for_scope_at used Path::exists(), which maps stat/permission errors to false — silently treating an unreadable policy as missing and default-disabled, flipping enrollment/flush behavior. Use try_exists() and propagate the stat error with context; only a confirmed non-existent path returns the not-enrolled default. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * fix(traces): capture traces for instance-only enrolled users Codex P1: capture_turn_trace gated on the per-user scope policy (read_trace_policy_for_scope(Some(scope)) + policy.enabled), so an instance-only enrolled user — whose per-user policy is absent/disabled — had every turn dropped before an envelope was queued, leaving the instance-aware flush gate nothing to submit. The headline instance-enrollment feature never captured for exactly the users it targets. Gate capture on the effective enrollment instead, mirroring the flush gate: add resolve_effective_capture_policy (personal-invite policy if enabled, else the admin-provisioned instance policy at scope None, else None) and prepare the envelope under that governing policy. Add a resolver test covering the personal / instance-only / neither cases. Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com> * Remove accidentally committed frontend node_modules, restore .gitignore The merge commit f34dfa7 dropped crates/ironclaw_webui_v2_static/frontend/.gitignore and swept 1065 node_modules files into the index. Untrack them and restore the ignore. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Address PR review feedback: egress hardening, effect declaration, instance status sync - fetch_account_traces_direct now uses a pinned-DNS, private-IP-filtered HTTP client (shared pinned_trace_commons_http_client) instead of an unrestricted reqwest lookup, closing the DNS-rebinding window between claim validation and the bearer-authenticated account-traces GET. - account_login_link capability manifest declares EffectKind::WriteFilesystem for the local delivery-file write. - Queue-flush status sync (and the public sync entry point) now run off the resolved effective flush target (policy, device-key dir, per-user subject) instead of re-reading the per-scope policy, so instance-enrolled users get final credit status after submission; subject is threaded into the status-sync claim context. - Each login-link mint persists to a unique account_login_link.<uuid>.url file so concurrent mints cannot clobber each other; stale link files are pruned best-effort after one hour. - resolve_trace_credentials takes typed &TenantId/&UserId at the public boundary; call sites drop their .as_str() conversions. - Login-link/account-traces requests honor the policy-configured issuer timeout; the sink-path traces fetch uses ACCOUNT_TRACES_MAX_RESPONSE_BYTES. - Removed the AdminScope::enroll_instance_trace_commons wrapper from the v1 monolith (crate-side entry point is onboard_instance_with_sink; noted in the slice1 plan). - Tests: direct account-traces path covered for 500/404; new regression test pins instance-target status sync (subject + instance device-key dir). - Plan docs: server login-link contract callout, no developer-local paths, resolver errors propagate, 404-only zero-state, scope_dir threading. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Pin DNS resolution on the background trace submit/status/revoke lane The background lane (queue flush submission, status sync, revocation) previously relied only on enrollment-time endpoint validation (validate_trace_commons_ingest_url); the per-request client did a fresh unrestricted DNS lookup. Replace trace_remote_http_client with pinned_trace_remote_http_client: per-request host resolution through resolve_trace_upload_claim_issuer_host (private/internal IPs rejected, literal-loopback local-dev exception) pinned via resolve_to_addrs, so an endpoint host that passed validation at enrollment cannot later rebind to an internal address and receive bearer-authenticated requests. Timeout behavior (env/test task-local override) is unchanged. Regression test: pinned_trace_remote_client_rejects_private_endpoint_hosts. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Address CodeRabbit follow-up: sanitize status log, sync plan snippets - trace commons status dispatcher no longer formats the resolver error into the log (it can embed the policy file's host path); logs the safe fact only, matching the sibling dispatchers. - slice4 plan: AccountTraceItem snippet derives Deserialize (matches shipped code, which parses the response). - slice3 plan: login-link parsing snippet fails loud on missing account_id/url instead of unwrap_or_default (matches shipped code). - slice1 plan: the AdminScope wrapper task is marked SUPERSEDED up front so the plan no longer gives conflicting guidance. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Address round-2 review: opt-out precedence, salted subjects, UI branch tests - Explicit per-user opt-out (scoped policy present with enabled=false, as written by 'traces opt-out') now blocks the instance-enrollment fallback in resolve_trace_credentials and resolve_effective_flush_target (and thus capture) — only a never-configured scope falls through to the instance policy. Regression test covers all three resolution surfaces. - Instance-enrollment subjects are now salted: a per-instance random salt (persisted 0600 at the instance trace dir, create_new race-safe) feeds sha256(salt:scope), so the server or ledger holders cannot dictionary-match guessable tenant/user ids against an unsalted scope hash. Unsalted local_pseudonymous_contributor_id remains for local state keying/log refs. - contribution.rs carries the architecture-rule file-size justification referencing decomposition tracking issue nearai#4088; state_scope field docs now say which state it does (and does not) locate. - Submitted-traces UI: extracted the pure tracesSectionMode decision (error wins over list; list needs enrolled + non-empty) and covered it plus the row formatters in trace-commons-tab.test.mjs. - Docs: slice4 plan points at crates/ironclaw_webui_v2 (static crate was folded in), slice3 signature snippet matches the typed contract, and the webui_v2 CLAUDE.md route table gains the three trace routes. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Update crossbeam-epoch 0.9.18 -> 0.9.20 for RUSTSEC-2026-0204 Lockfile-only patch bump of a transitive dep (via termimad/crossbeam) to clear the new advisory failing cargo-deny; verified locally with cargo deny check advisories. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> * Route login-link/account-traces claim mint through the caller's sink The sink-based entry points (mint_account_login_link_via_sink, fetch_account_traces_via_sink) used the sink for the final POST/GET but minted the upload-claim bearer via DefaultTraceUploadCredentialProvider, whose issuer request takes the direct reqwest path — so an agent-invoked account_login_link performed a network call outside RuntimeHttpEgress. New trace_upload_bearer_token_via threads Option<sink> into the claim mint (cache behavior unchanged; the default provider passes None), and both sink paths pass Some(sink), matching the profile-token/profile-set flows. Tests now use a RecordingSink to pin the invariant that both the claim mint and the follow-up request route through the sink. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
personal-upstream-sync Bot
pushed a commit
that referenced
this pull request
Jul 23, 2026
* Unify runtime store graph selection
* Shrink runtime graph mode branching
* Unify scoped filesystem graph resources
* Route identity and secrets through runtime graph
* Use composite filesystem for production graph
* Collapse production runtime graph wrapper
* Store one runtime graph in services
* Rename runtime auxiliary surfaces
* Consolidate Reborn runtime composition
* Unify Reborn runtime storage assembly
* Unify reborn runtime assembly
* refactor(reborn-composition): fix branch build + advance composition refactor (F2–F7, config/pool/DEL-7 phases)
Repairs the broken branch build and advances the composition cleanup. All
production build paths are green workspace-wide; the test suite is
intentionally RED pending an end-functionality rewrite (see below) — this is a
draft/WIP push.
Build fix + earlier cleanups:
- Finish the build_reborn_services→build_runtime / RebornServices→RebornRuntime
consumer migration that HEAD left half-done (composition lib, product_workflow
+ architecture tests, root integration harness).
- F2: rename RebornRuntimeSubstrate→RebornRuntimeStores; RebornRuntime carries
production-only fields (test-only handles removed).
- F3: dedup resource-governor guards (filesystem_resource_governor) + the
libsql/postgres store-assembly tail (finish_production_backend).
- F4: shared HostRuntimeServices setter chain → with_shared_host_runtime_wiring!.
- F5: single-source the /system subroots (SYSTEM_SUBROOTS).
- F7: make runtime_store_parts infallible.
Deployment-config / bindings refactor:
- Phase A: move declarative shape (required_runtime_backends + require_* flags)
into DeploymentConfig.
- Phase B: relocate DB pool/db acquisition from RebornBuildInput construction
into build_runtime build-time via declarative connection config
(PostgresPoolSource{Config|Prebuilt}, LibsqlConnectionConfig + prebuilt test
hatch); 4-way pool sharing + rg-singleton gates preserved; no new deps.
- Phase C (1–3): manifest-declared network egress — new per-capability
max_egress_bytes; gmail/calendar/web-access manifests declare their exact
host allowlists + caps (egress reproduced byte-for-byte); composition's
gsuite/web-access network-policy special-cases removed.
Deferred (WIP): the test suite is red by design — internal-peeking *_for_test
accessors + the new max_egress_bytes field need the end-functionality test
rewrite; DEL-7 dep-removal (drop ironclaw_first_party_extensions from
composition prod deps) is blocked on that test net + injecting the gsuite
account-visibility policy.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* style+test: cargo fmt the composition refactor + fix max_egress_bytes test literals + regen pub-use snapshot
- Run `cargo fmt` over the prior commit's composition changes (they were
committed without formatting — 30 fmt diffs across factory.rs/runtime.rs/
test_support.rs/etc.). `cargo fmt --check` now clean.
- Add `max_egress_bytes: None` to 14 `#[cfg(test)]` CapabilityDescriptor/
CapabilityDeclV2 literals across authorization/approvals/runner/host_runtime/
runtime_policy/loop_host/composition that the new manifest-egress field left
uncompilable; those crates' test targets compile again.
- Regenerate docs/plans/composition-pubuse.snapshot to the current lib.rs
surface (drops stale RebornServices/build_reborn_services, adds build_runtime);
composition_public_pub_use_surface_matches_snapshot passes.
Still WIP: composition `--features test-support` + the integration suite remain
red pending the end-functionality test rewrite (orphaned *_for_test accessors);
and reborn_generic_code_names_no_concrete_extension stays red until DEL-7
Changes 4-7 remove composition's concrete-extension naming (pre-existing —
already failed at the branch's prior HEAD).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(reborn-composition): rewrite budget/auto-approve peek tests to end-functionality; un-red composition test-support
First bounded unit of the internal-peek test rewrite. Removes orphaned
#[cfg(test/test-support)] accessors on RebornRuntime that read the deleted
test-only fields, and re-expresses their tests through real seams.
- Budget: delete budget_resource_governor/budget_event_sink/budget_gate_store/
budget_gate_scope_for_conversation/apply_resolved_budget_gate. Rewrite
budget_e2e + budget_approval_e2e to assert enforcement via production seams
(with_budget_defaults setup, send_user_message_until_gate -> BlockedOnGate,
reconcile receipt for spend, broadcast_budget_event_sink()/observer for
events) — e.g. a pause-threshold run now asserts a real blocked-gate outcome
AND zero model calls (short-circuit), not internal store contents.
- local_dev_approval_test_parts / attachment reader: re-sourced over the
runtime's own composite root / read_write view (byte-identical durable rows).
- local_dev_triggered_run_delivery_for_test: deleted (no callers).
- local_dev_storage_root_for_test: deleted; outbound_store_durability caller
derives the root via std::fs::canonicalize (behaviorally identical).
- runtime/tests/core.rs attachment test: read port built over the read-write
workspace view (reader never writes).
- harness_mcp.rs: add max_egress_bytes: None to its two capability literals.
cargo build -p ironclaw_reborn_composition --features test-support and
cargo build --workspace are clean; budget tests pass (12/12).
Dropped-coverage (verified no production seam, reported for review): budget-gate
RESOLUTION (approve-with-raised-limit/cancel/expire) has no production caller
outside the deleted test hooks; per-agent-account caps and reading a seeded
cap's exact USD value had no config/event seam. Core enforcement + pause + event
coverage retained.
Still WIP: the integration suite (channel/pairing substrate-accessor peeks in
extension_delivery.rs etc.) is the next bounded rewrite unit; DEL-7 Changes 4-7
remain.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(integration): defer 2 channel/pairing suites; integration suite now compiles
The channel delivery/pairing integration tests drive their scenarios through
RebornRuntime test-support accessors (channel_config_facade,
start_channel_host_assembly_for_test, pairing_*, delivery_coordinator,
outbound_delivery_stores_for_test, register_static_channel_egress_credentials_for_test)
that were removed in the RebornRuntime shape cleanup. Per the chosen path, defer
these two reversibly (tracked, task #8) rather than restore the gated accessors now:
- tests/integration/extension_delivery.rs: `#![cfg(any())]` (whole bin inactive).
- group_extensions slack lifecycle scenario: `mod` + `run(&g)` commented.
With these deferred, `cargo test -p ironclaw_reborn_integration_tests --no-run
--all-features` compiles cleanly — the integration net is restored (needed to
verify DEL-7 Changes 4-7's catalog injection). Re-enable by restoring the gated
channel/pairing accessors or building a channel-delivery harness seam.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(reborn-composition): finish DEL-7 — drop ironclaw_first_party_extensions from composition prod deps
Composition no longer links a concrete extension crate in production;
ironclaw_first_party_extensions is now a dev-dependency only. The binary
(ironclaw) injects gsuite + web-access via the same seam slack/telegram use.
Behavior preserved exactly across all three security axes.
- Account visibility: RuntimeCredentialAccountVisibilityPolicy made pub;
injected on the build input; production uses the injected policy else a
promoted fail-closed Default (generic is_authorized_for_requester — strictly
more restrictive). GsuiteRuntimeCredentialAccountVisibilityPolicy moved to the
CLI, which injects it (production behavior identical).
- Trust effects: builtin_first_party_trust_policy() no longer called at
input-construction; production_first_party_trust_policy(bundles) composes the
policy at build time from the injected first-party bundles' trust_effects
(identical per-extension grants). No-arg fn kept #[cfg(test)].
- Catalog: composition-neutral FirstPartyPackageBundle (extension_host/first_party.rs)
injected on the build input; catalog/trust/search-alias sites take &[bundle];
is_gsuite_extension_id dropped (folded into per-bundle metadata). CLI injects
the real inventory; a non-empty assert guards against vanishing extensions.
- Handlers: composition-owned FirstPartyHandlerRegistrar seam; factory registers
via injected registrars. Deleted composition web_access.rs + extension_host/gsuite.rs;
the gsuite/web-access handlers moved to CLI src/first_party/. The 2 test files
that imported the old bundling (composition tests/gsuite.rs, integration
harness_web_access.rs) rewritten onto local wrappers.
- Egress: unchanged (manifest-declared from the prior commit); the test-support
network-policy helper now delegates to the production builder.
- Cargo.toml: first_party → [dev-dependencies] (ports crate untouched).
reborn_dependency_boundaries.rs: first_party added to the ironclaw allow-list.
reborn_extension_specificity.rs + pub-use snapshot updated; both pass.
Verified: composition (default + test-support), ironclaw (CLI) build clean;
integration suite + composition tests compile; cargo fmt clean; ironclaw_architecture
passes except the pre-existing reborn_localdev_typename ratchet (LocalDevApprovalHarness,
unrelated, fails identically on HEAD).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(reborn-composition): rename RebornBuildInput -> RebornHostBindings (Phase A step 1)
Mechanical rename of the composition build-input carrier to RebornHostBindings
ahead of splitting its DATA fields into DeploymentConfig. No behavior change;
updates the pub-use snapshot and doc references.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(reborn-composition): move declarative DATA into DeploymentConfig (Phase A step 2)
The owner_id, local_runtime_identity, resolved runtime_policy,
turn_state_store_limits, oauth provider/DCR configs, nearai bootstrap config,
account-setup descriptors, and first-party bundle inventory now live on
DeploymentConfig as declarative data; RebornHostBindings keeps only the
irreducible code (trait objects, factories, registrars, pre-opened handles,
test hooks). Builder setters/readers delegate to self.deployment, so call
sites are unchanged. DeploymentConfig drops PartialEq/Eq (it now carries
non-Eq secret/config data); the one equality test compares observable axes.
FirstPartyPackageBundle and its asset/onboarding/oauth sub-types gain
Debug/Clone so the config can derive them.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(reborn-composition): surface DeploymentConfig as first-class config on RebornRuntimeInput (Phase A step 3)
Adds RebornRuntimeInput::config() (the authoritative declarative deployment
config, read separately from the code-carrying bindings) and with_config() /
RebornHostBindings::with_deployment_config() to install an accurately-resolved
config after construction. The build entry stays build_reborn_runtime(input),
so all existing callers are unchanged; the config is sourced from the bindings
the caller supplied, giving the runtime layer a bindings-independent view of
'what deployment is this' now that all declarative DATA lives on
DeploymentConfig.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(reborn-composition): resolve runtime policy in local-dev construction
RebornHostBindings::local_dev / local_dev_with_profile routed through
local_dev_from_deployment, which built an input with an UNRESOLVED runtime
policy (runtime_policy left None) — so every build_runtime_substrate on that
bare fixture failed 'MissingRuntimePolicy'. Fold the policy_request resolution
(previously only done by the local_runtime_build_input* bridge) into
local_dev_from_deployment so the local-dev fixture is buildable without the
caller separately calling .with_runtime_policy(); this removes the reason a
bare, unresolved-policy local-dev constructor existed. Un-reds 35 composition
lib tests (70 -> 35 failing) with no regressions.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* refactor(reborn-composition): remove RebornHostBindings::local_dev[_with_profile]
A local-dev deployment is data (DeploymentConfig::local_dev) plus generic
bindings, not a bindings-typed constructor. Replace the two associated
constructors with test-support-gated free helpers local_dev_build_input /
local_dev_build_input_with_profile (behaviour-identical, routed through
local_dev_from_deployment) and migrate all ~200 call sites. Reroute the one
production caller (hosted_single_tenant_volume_build_input) to
local_dev_from_deployment directly. No behavior change; composition lib tests
unchanged at 1356 passed / 35 failed (the pre-existing task-#8 red set).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(reborn-composition): inject first-party surface into local-dev test constructors
DEL-7 removed composition's internal first-party bundling; the binary now
injects catalog bundles + capability-handler registrars on the build input.
Composition's own unit tests lost that surface and failed to install/activate
first-party extensions ("available extension was not found").
Mirror the binary's assembly from the dev-dependency inventory in a
\#[cfg(test)] test-support module and inject it by default from the local-dev
test constructors (local_dev_build_input[_with_profile]) — restoring the
pre-refactor state where composition unit tests always saw first-party
extensions. Fixes 24 of the 35 deferred failures with no regressions.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(reborn-composition): restore 3 local-dev behaviors dropped by unify-assembly (975bcd2)
The unify-runtime-store-graph refactor (commit 975bcd2) collapsed the
separate local-dev builder into the single build_production_shaped ->
build_backend_production path and silently dropped three local-dev behaviors
the deleted branch had been providing (the 'removing a redundant layer
un-masks behavior' hazard):
1. Product-auth flow_record_source: the durable local-dev product-auth service
was no longer wired as the AuthFlowRecordSource, so flow_record_source() /
as_auth_challenge_provider() returned None and the WebUI auth interaction
surface composed as Unavailable instead of routing gates into the durable
service. Re-wired via a new flow_record_source field on
ProductAuthServicesCompositionInput, set only for the builder's own durable
service (None arm) — the caller-supplied-bundle path stays unwired so its
surface remains explicitly unavailable per the crate contract.
2. Test host-HTTP-egress override: RebornHostBindings::network_http_egress_for_test
was set but never read, so every local-dev test reached the real network
(hosted-MCP discovery, deployment-channel DM provisioning). Re-introduced the
TestNetworkHttpEgress adapter and thread the override through
RebornProductionBuildContext into the host-egress wiring (cfg(test|test-support)).
3. Local-dev host-runtime selection: the choice keyed only on a wired LocalHost
process port, so a local-dev deployment using a non-LocalHost process backend
(e.g. an injected TenantSandbox port) wrongly routed through
host_runtime_for_production, whose validate_production_wiring rejects the
LocalSingleUser deployment mode. Also select the non-validating local-testing
host runtime whenever the resolved policy is LocalSingleUser (exactly the
shape production validation rejects; production never resolves to it).
Regression coverage: the pre-existing composition tests that exercised these
paths (factory manual-token-rebuild + DCR challenge-provider, notion hosted-MCP
activate+publish, runtime channel-identity DM provisioning, webui auth-gate
routing, local-dev approved-shell tenant-sandbox) go green again unchanged.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* test(reborn-composition): update runtime-unification expectations (unify-runtime-store-graph)
Runtime-store unification made the runtime store graph unconditional for every
build (extension_lifecycle_surface_context non-optional, local_runtime_for_test
-> Some(self), local_runtime -> Some(&services)). Three assertions pinned the
now-obsolete split-runtime premise; update them to the new-but-correct behavior
(the reject-branches they exercised are now unreachable):
- production_libsql_turn_state_uses_configured_runtime_identity: local_runtime
is_none() -> is_some() (its real subject, turn_state keyed by configured
identity, is unchanged).
- rejects_trajectory_observer_for_production -> wires_trajectory_observer_through_unified_runtime:
the observer is now wired through the shared capability path (no more silent
empty trajectory), so a production runtime accepts and starts with one.
- rejects_enabled_hooks_without_local_runtime -> wires_enabled_hooks_through_unified_runtime:
the hook framework is wired for production builds, so enabling hooks builds and
validates readiness instead of failing MalformedConfig.
Also hosted_single_tenant_rejects_local_dev_storage_input: the dedicated
storage-shape guard string was removed in 975bcd2 (substrate/storage are now
deployment-derived, making the mismatch unreachable in production). The test
now asserts the surviving fail-closed behavior — the artificially mismatched
pairing fails closed on the runtime-policy guard rather than composing hosted
over local storage.
PRODUCTION-BEHAVIOR NOTE for review: the observer/hooks tests now lock that a
production (HostedMultiTenant) runtime ACCEPTS a trajectory observer and enabled
hooks. This behavior change was made by the branch's unification (local_runtime
always Some), not by this test edit; flagging it explicitly.
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* style(reborn-composition): cargo fmt the task-8 edits
rustfmt normalization of the first-party test-support module, the local-dev
test constructors, and the with_bundled_first_party_for_test builder added in
this task's earlier commits (import ordering + line wrapping only; no semantic
change).
Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* Strengthen runtime observer and hook coverage
* Fix code style CI gates
* Trigger PR CI sync
* Refresh PR checks
* Cover runtime observers and CI wiring
* Wire bundled extensions into lifecycle CLI
* Refresh PR sync
* Fix runtime composition CI ratchets
* Address extension runtime review coverage
* Fix Reborn adapters CI gates
* Fix Reborn runtime wiring CI coverage
* Wire bundled extensions in product API integration
---------
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
Restrict the Docker Image workflow's publishing job to
nearai/ironclawonly.This keeps the control branch's copy of the workflow consistent with the deployment branch path and prevents Docker publishing from
theredspoon/ironclaw.Validation
.github/workflows/docker.yml.