Skip to content

Update - #6

Merged
elliotBraem merged 58 commits into
mainfrom
v/barcelona
Jun 17, 2026
Merged

elliotBraem merged 58 commits into
mainfrom
v/barcelona

Conversation

@elliotBraem

Copy link
Copy Markdown

Summary

Change Type

  • Bug fix
  • New feature
  • Refactor
  • Documentation
  • CI/Infrastructure
  • Security
  • Dependencies

Linked Issue

Validation

  • cargo fmt --all -- --check
  • cargo clippy --all --benches --tests --examples --all-features -- -D warnings
  • cargo build
  • Relevant tests pass:
  • cargo test --features integration if database-backed or integration behavior changed
  • Manual testing:
  • If a coding agent was used and supports it, review-pr or pr-shepherd --fix was run before requesting review

Security Impact

Reborn Trust-Boundary Checklist

  • Public policy/evidence/trust-bearing types: who can construct them?
  • Untrusted content enters prompts only through an envelope/escaping primitive.
  • Hashes declare purpose; trust/binding/authenticity uses SHA-256/BLAKE3 or separate authenticity check.
  • New/changed status, exit, policy, runtime, or error variants: downstream match sites audited. Command/output:
  • Security/durability serde(default) fields fail closed or have migration tests.
  • Queues/maps/buffers/counters have bounds and overflow-safe arithmetic.
  • Driver/operator-visible errors have stable class semantics (Transient, Permanent, Misconfigured, PolicyDenied or equivalent).
  • Sandbox/native/host names accurately describe trust boundary.

Database Impact

Blast Radius

Rollback Plan

Review Follow-Through


Review track:

henrypark133 and others added 30 commits June 14, 2026 22:03
…Error gaps (nearai#4895)

Addresses post-merge review findings on nearai#4836.

- Bug (Medium): the communication-context provider keyed outbound
  delivery preferences by the *actor* instead of the run *owner*. Product
  inbound and trusted-trigger runs can carry an explicit thread owner
  (subject/creator) distinct from the actor; the stored preference belongs
  to the owner. Resolve the caller's user_id via
  `scope.explicit_owner_user_id()` with actor fallback — matching
  `TurnScope::to_resource_scope` — so shared/channel inbound and trigger
  runs render the owner's delivery target, not the actor's. Adds two
  regression tests asserting the lookup is keyed by owner vs actor through
  a caller-capturing facade.

- Tests (Medium): cover `CommunicationContextFetch::resolve`'s JoinError
  branches that were previously unexercised — actorless failure degrades
  to `None`, actor-present failure degrades to `Some(Unknown)`.

- Docs (Low): the product-context-factory plan still specified the
  rejected 256-byte `RunOriginAdapter` bound; update to the as-built
  512-byte cap (mirroring `AdapterKind`) so follow-up work does not
  reintroduce the narrowing.

Not addressed (verified out of scope): the `model_safe_label` "injection"
finding is a false positive — label sources are admin/system-set and the
sanitizer strips structural characters (existing hostile-input tests pass);
the shared-route surface assertion finding is already covered by
`shared_user_message_records_channel_surface_type`; the string-based
trigger-helper finding is owned by issue nearai#4851's plan; the
RunOriginAdapter byte-mirror is already mitigated by a named const + docs
+ tests.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…earai#4840)

* fix(host_runtime): surface AuthRequired before approval gate on missing credentials (Fix B)

Extracts capability_credential_requirements() as the single source of truth
for credential requirements derived from the capability manifest descriptor.
Both the new credential pre-flight check and the existing dispatch-time
obligation check call this function — no second computation added.

Adds credential_preflight_check() on DefaultHostRuntime and calls it in
invoke_capability() and spawn_capability() BEFORE apply_persistent_approval_policy(),
so AuthRequired is returned without persisting an approval request when a
required credential is absent.

Wires the optional secret store from HostRuntimeServices.build_host_runtime()
into DefaultHostRuntime via with_credential_preflight_store(). The pre-flight
is skipped in minimal/test graphs that don't supply a secret store; the
dispatch-time obligation check remains the enforcement backstop regardless.

Tests (host_runtime_services_contract):
- invoke_capability_missing_credential_returns_auth_before_approval
- invoke_capability_present_credential_proceeds_to_approval
- invoke_capability_no_credential_requirement_proceeds_normally
- credential_requirements_preflight_and_dispatch_agree_on_same_handles

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(host_runtime): correct capability_credential_requirements docstring and strengthen test

The docstring claimed "both the pre-flight check and the dispatch-time
obligation check call this function" — that was false. The obligation
handler in BuiltinObligationHandler derives required handles by iterating
descriptor.runtime_credentials directly; it does not call this function.
Correct the docstring to accurately describe the agreement at source-data
level (both iterate the same required-true entries) and explain why gate-ID
divergence between pre-flight and backstop is moot in practice (pre-flight
fires first when a secret store is wired).

Rename `credential_requirements_preflight_and_dispatch_agree_on_same_handles`
to `credential_requirements_extraction_matches_descriptor_required_credentials`
with a scope note clarifying what the test actually verifies (canonical fn
vs descriptor, not vs obligation handler), add an assertion that
credential_requirements is empty for secret_handle source type, and reference
the caller-level test that covers the gate-ordering guarantee.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(host-runtime): skip credential pre-flight on store errors per review

On a transient SecretStore Err, credential_preflight_check previously
returned AuthRequired (treating backend failure as credential absent),
burning a user auth interaction. Now it returns None on Err so the
dispatch-time obligation check remains the sole enforcement backstop.

Also applied:
- FIX 2: add trust-class-agnostic comment at both pre-flight call sites
- FIX 3: rename field secret_store → credential_preflight_store to match
  the with_credential_preflight_store builder
- FIX 4: take registry.snapshot() once in invoke_capability and
  spawn_capability and pass it into credential_preflight_check, removing
  the redundant internal snapshot

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test(host-runtime): add credential pre-flight edge-case tests per review

- FIX 5: change section separator to triple-dash style matching
  memory_prompt_context.rs
- FIX 6: four new tests:
  a. spawn_capability_missing_credential_returns_auth_before_approval —
     spawn path mirrors invoke_capability pre-flight behavior
  b. invoke_capability_no_credential_requirement_with_wired_store_proceeds_normally —
     wired store + zero required credentials hits is_empty() branch, not
     no-store early exit
  c. invoke_capability_secret_store_error_skips_preflight — erroring
     store stub confirms pre-flight skips on Err (FIX 1) and flow
     reaches the approval gate
  d. credential_requirements_extraction_returns_empty_for_all_optional_credentials —
     descriptor with only required=false credentials yields empty vecs

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(host-runtime): address nearai#4840 review — product-auth preflight skip, scope validation, test wiring

- capability_credential_requirements no longer treats ProductAuthAccount injection
  slots as presence-checkable secrets, fixing a false-positive AuthRequired preflight
  for capabilities with already-connected product-auth accounts.
- Validate context/resource-scope consistency before the credential preflight queries
  the secret store, closing a forged-scope presence-probe window (invoke + spawn).
- Wire the preflight store in contract fixtures and seed credentials on the request's
  own ResourceScope; add product-auth and scope-validation regression tests.
- silent-ok annotation on the store-error skip; docs name both invoke and spawn;
  moved preflight tests to host_runtime_credential_preflight_contract.rs.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(host-runtime): nearai#4840 review round 2 — test fidelity + tighten public surface

- store-error regression now drives the manifest-backed credential backstop via
  ApprovalThenGrantAuthorizer (was masked by a test-authorizer-injected obligation).
- add spawn_capability present-credential happy-path test (proceeds to ApprovalRequired).
- make the credential-preflight-store setter test-only/crate-private; tests wire it
  through HostRuntimeServices::build_host_runtime.
- stop publishing capability_credential_requirements from the crate root; keep it
  crate-private with equivalent coverage.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(host-runtime): exercise the real dispatch-time obligation backstop on store error (nearai#4840)

The store-error regression now grants the required secret so dispatch authorization
passes and the resumed call reaches BuiltinObligationHandler::preflight_secret_injection,
where the AlwaysErrorSecretStore metadata() probe errors and the handler fails closed
(secret_obligation_failed). Previously the grant omitted the secret, so the block came
from grant-matching authorization — not the dispatch-time credential backstop the PR
contract relies on (obligations.rs preflight_secret_injection fails closed on store error).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(host-runtime): single owner for the secret-presence rule (nearai#4840)

Extract obligations::secret_present as the one definition of 'is this required
secret present in scope', shared by the credential pre-flight (ordering) and the
dispatch-time obligation backstop (enforcement). Removes the duplicated metadata()
presence rule across production.rs and obligations.rs so the two paths cannot drift.
Each caller still owns its store-error policy (pre-flight fails open and skips; the
backstop fails closed). The happy-path double-read remains (documented) — addresses
the split-ownership half of the review; collapsing the read is a separate follow-up.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(host-runtime): make store-error backstop test airtight + cover required product-auth (nearai#4840)

Addresses PR nearai#4840 review round 3:
- The store-error regression now uses a metadata()-call-counting error store and
  asserts the resume drove at least one further probe after a counter reset — proving
  the resume reaches BuiltinObligationHandler::preflight_secret_injection (authorization
  passed) and fails closed there, not at a premature authorization denial. Both paths map
  to RuntimeFailureKind::Authorization, so the probe count is the distinguishing signal.
  Removes the now-unused AlwaysErrorSecretStore and the stale test doc.
- Corrects the comment wording: the secret-injection obligation comes from grant
  evaluation against the manifest, not from ApprovalThenGrantAuthorizer.
- Adds credential_requirements_extraction_excludes_required_product_auth_account unit
  test: a required product_auth_account credential is excluded from required_secrets
  (no false-positive pre-flight AuthRequired) but still surfaces in credential_requirements.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(host-runtime): preserve secret-store failure cause when failing closed (nearai#4840)

The dispatch-time obligation backstop dropped SecretStoreError with map_err(|_|),
collapsing outages into an opaque secret-obligation failure with no server-side
trail. Bind the error and log it at debug! (SecretStoreError Display carries no raw
secret material) before returning the sanitized secret_obligation_failed(); the
caller still receives the opaque error. Mirror the same cause-logging in the
fail-open pre-flight skip arm.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
… dying as scope_mismatch (nearai#4899)

A 2nd+ Slack approval (or auth) resume of a capability could terminally
fail with scope_mismatch ("capability input ref is not scoped to this
loop run"). On a resume dispatch returning a transient Backend error,
handle_capability_error cleared the pending resume slot, then the
RecoveryOutcome::Retry path re-dispatched via
capability_invocation_from_candidate(call, None) — dropping the resume
context. The non-resume path resolved the original-run input_ref against
the resuming run, failed ensure_ref_scoped_to_run, and killed the run as
HostUnavailable.

Intercept Retry for approval-resume AND auth-resume origin failures:
surface the real backend error to the model as a tool result and
continue the loop (so the user can re-approve / re-auth) instead of
re-dispatching. Kills the scope_mismatch (S1) and avoids
double-executing a side effect whose lease is one-shot after a Backend
error (S2).

Adds regression tests for both resume origins.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(webui): keep code block overflow local

* fix(webui): keep code block typography readable after merge
…i#4910)

* fix auth resume input replay

* fix auth resume approval token carryover
* fix(reborn): normalize bare workspace tool paths

* Preserve scoped path URL validation for workspace aliases

* Normalize empty workspace alias segments

---------

Co-authored-by: Robert Yan <mstr.raphael@gmail.com>
* feat(reborn): declare Slack channel as extension manifest

* fix(reborn): address extension review feedback

* fix(reborn): address Slack product-adapter extension review feedback

- Reserve "slack" as host-bundled extension id (prevent filesystem shadowing)
- Add builtin_first_party_trust_policy regression test for Slack admin entry
- Activate manifest-backed channel packages from WebUI (suppress only wasm_channel)
- Preserve legacy Slack connect controls for pre-install deployments
- Project only ProductSurfaceKind::ExternalChannel to channel kind
- Parse each manifest once via ExtensionManifestRecord
- Rename product_adapter.host_beta section to stable product_adapter.inbound
- Add caller-level tests: list_extension_registry, extension_info, ChannelsTab render

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(runtime-context): enable connected-channel classification via surface_kinds

nearai#4778 lands the ProductAdapter surface projection, so the lifecycle summary
now carries surface_kinds. Flip CHANNEL_CLASSIFICATION_AVAILABLE to true and
make extension_is_channel_surface a real predicate (ExternalChannel), so
connected channel names render in the model runtime context instead of unknown.

Convert the two stubbed tests to positive cases: empty list -> Known([]),
mixed list -> only active channel-surface extensions reported.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* style(runtime-context): cargo fmt + drop unreachable classification branch

Remove the dead 'if !CHANNEL_CLASSIFICATION_AVAILABLE' arm inside the
Some(Ok(response)) match: lifecycle_fut only issues the ExtensionList call
when classification is enabled, so a present response always means it is on.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(webui-v2): repair Slack extension asset + locale checks after channel refactor

- assets.rs smoke test: assert showLegacySlackConnectActions (the refactor's
  built-in Slack status path) instead of the removed slackBuiltinStatus helper
- add extensions.kind.channel to all 10 non-en locales (en gained the key with
  the new channel surface kind; locale-parity test requires all locales match)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(webui-v2): cover ExtensionCard channel overflow + localize registry heading

Review findings (PR nearai#4778 review 4493764666):
- F2: add extension-card.test.mjs proving kind=channel/wasm_channel surface
  Setup (setup_required/failed) and Reconfigure (active/ready) overflow actions
  on the real component; channels-tab.test stubbed ExtensionCard so this was
  uncovered.
- F4: render the 'Available channels' registry heading via t(channels.availableChannels)
  instead of a hardcoded literal; add the key to all 11 locales. Update the
  channels-tab test to locate the registry section by the RegistryCard component
  (heading is now an interpolated value, not a template literal); drop the now
  unused renderedValueAfter helper.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(runtime-context): drop permanent classification flag, document surface_kinds cache

Review findings (PR nearai#4778 review 4493764666):
- F5: remove CHANNEL_CLASSIFICATION_AVAILABLE (permanently true after the stub
  flip) and run the lifecycle ExtensionList fetch unconditionally when a
  lifecycle facade is wired.
- F3: document AvailableExtensionPackage.surface_kinds as an intentional
  single-parse cache (re-deriving in summary() would re-run the manifest
  projection, undoing the parse-once optimization).

F1 (per-turn ExtensionList cost) accepted as-is: the fetch is spawned off the
critical path under a 500 ms budget with abort-on-drop; caching deferred to a
follow-up.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(webui-v2): suppress Activate for channel kinds during pairing

Review finding (PR nearai#4778 review 4494623523, finding 1): primaryExtensionAction
only suppressed the primary Activate button for legacy wasm_channel, so a
manifest-backed kind=channel Slack card fell through to 'activate' in
pairing_required/pairing states where the dedicated pairing section already
owns the flow. Suppress the primary action for channel-surface kinds in those
states (via isChannelExtensionKind); installed channels still return 'activate'.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(webui-v2): add channels.slack key + regression tests for surface-kind projection

Review findings (PR nearai#4778 review 06:28):
- Add missing channels.slack i18n key to all 11 locales (legacy Slack row
  rendered the raw key because the i18n helper returns the key on miss; the
  || "Slack" fallback never fired).
- Add filesystem-path test: a /system manifest with
  product_adapter.inbound.surface_kind = external_channel projects to
  ExternalChannel surface (previously only the bundled catalog path was covered).
- Add extension_kind regression test: non-channel summaries keep their runtime
  wire kind (wasm_tool, mcp_server) while channel surfaces map to "channel".

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reborn): gate Slack catalog entry behind slack-v2-host-beta; test wasm_channel registry

Review findings (PR nearai#4778 review 07:01):
- Gate the Slack first-party catalog entry, its only-Slack symbols (slack_package,
  slack_assets, SLACK_MANIFEST, slack_manifest_digest), the factory trust-policy
  Slack AdminEntry, and Slack-asserting tests behind slack-v2-host-beta. Without
  the feature the Slack route/runtime/WebUI mounts don't exist, so the catalog
  must not advertise an unrunnable Slack extension. Clean clippy + tests in both
  feature-on and feature-off configs.
- Add useExtensions hook test proving an uninstalled kind=wasm_channel registry
  entry lands in channelRegistry, not toolRegistry (isChannelExtensionKind covers
  both channel and wasm_channel).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(reborn): consolidate Slack trust policy tests (nearai#4778)

---------

Co-authored-by: Henry Park <henrypark133@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…endings) (nearai#4837)

* feat(agent-loop): gated final-answer nudge to avoid empty/canned turn endings

When the reborn loop would otherwise end a turn with no real assistant
answer — an empty/trailed-off reply, the model-call budget exhausted, or
NoProgressDetected (which emits a canned "I stopped repeating the same
step" reply) — issue ONE extra tool-free model call asking the model to
synthesize a closing answer from the work it already did. This is the
reborn equivalent of the legacy loop's on_tool_intent_nudge /
force-text-recovery.

Gated by the (previously unimplemented) SteeringPolicy
`allow_driver_specific_nudges` flag, which defaults to false — so
production behavior is unchanged. Capped at one nudge per run
(`LoopExecutionState.final_answer_nudges_used`) so it can't issue
unbounded extra model calls. Wired into all three exit modes
(assistant_reply empty/trailed, budget IterationLimit, NoProgressDetected);
falls back to existing behavior when disabled, capped, or the model still
declines to answer.

Mechanism note: the tool-free call uses an EMPTY capability_view
(visible_capability_ids: []), not surface_version=None — the reborn model
gateway attaches tools from the capability port regardless of
surface_version, so only an empty view yields a true text-only request.

Evidence (PinchBench, Qwen3.5-122B, claude-haiku judge, controlled A/B on
the same merged tree): nudge OFF 0.682 vs nudge ON 0.768 (+0.086), closing
reborn to v2 parity (0.793). It specifically rescues tasks that do the
work but trip the no-progress detector (stock, events, spreadsheet,
polymarket), turning the canned give-up into a real answer.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* review: route nudge through admission + accounting, drop dead branch, add tests

Addresses review feedback on the final-answer nudge:

- Remove the empty/trailed-reply nudge branch in AssistantReplyStage. As the
  reviewer noted, DefaultReplyAdmissionStrategy rejects empty/artifact replies
  before they reach AssistantReplyStage, so that branch was dead for the default
  family. Empty replies are rejected → the loop continues → NoProgressDetected,
  which is where the nudge still fires. assistant_reply.rs is back to main.
- The nudged reply now goes through the SAME admission policy as a normal reply
  (ctx.planner.reply_admission().admit_reply) instead of a bare !is_empty()
  check — so blank text and provider-transcript artifacts can't be finalized.
- Preserve canonical assistant-reply token accounting: record the nudge turn's
  output tokens (provider usage, else the same estimate AssistantReplyStage
  uses) into recent_output_token_counts so the diminishing-returns window isn't
  fed stale data.
- Add caller-level tests with the gate enabled at the boundaries the nudge
  affects: no-progress (synthesizes via one tool-free model call), budget
  iteration-limit (completes instead of failing closed), gate-disabled (no model
  call, canned fallback), and the one-shot cap (no second call).

All 302 ironclaw_agent_loop lib tests pass.

Note on the remaining structural point (model call still issued from the exit
boundary via the host primitive rather than ModelStage): see PR discussion —
the exit-boundary stages don't own Prompt/Model, so a full "typed stage outcome"
move means restructuring the terminal-exit flow to re-enter the loop for one
tool-free turn. Happy to do that if preferred over this factoring.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* review: address CodeRabbit findings on final-answer nudge

- Move FINAL_ANSWER_NUDGE prompt out of Rust source into
  prompts/final_answer_nudge.md, loaded via include_str! (repo prompt-template
  invariant).
- Fix stale tool-free comment: clarify that the empty capability view on the
  model request (not the surface_version/capability_view None assignments) is
  what suppresses provider tools.
- Make MockHost::with_driver_nudges_enabled flip the steering flag in-place so
  it composes with other context-level builders regardless of order; drop the
  now-unused test_run_context_with_driver_nudges fixture.
- Add legacy-checkpoint decode regression test asserting a payload missing
  final_answer_nudges_used decodes to 0 (#[serde(default)] contract).
- Apply rustfmt to the new stage tests.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

---------

Co-authored-by: Pranav Raja <pranav.raja@near.ai>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: serrrfirat <f@nuff.tech>
* feat(reborn): declare Slack channel as extension manifest

* fix(reborn): address extension review feedback

* fix(reborn): address Slack product-adapter extension review feedback

- Reserve "slack" as host-bundled extension id (prevent filesystem shadowing)
- Add builtin_first_party_trust_policy regression test for Slack admin entry
- Activate manifest-backed channel packages from WebUI (suppress only wasm_channel)
- Preserve legacy Slack connect controls for pre-install deployments
- Project only ProductSurfaceKind::ExternalChannel to channel kind
- Parse each manifest once via ExtensionManifestRecord
- Rename product_adapter.host_beta section to stable product_adapter.inbound
- Add caller-level tests: list_extension_registry, extension_info, ChannelsTab render

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(runtime-context): enable connected-channel classification via surface_kinds

nearai#4778 lands the ProductAdapter surface projection, so the lifecycle summary
now carries surface_kinds. Flip CHANNEL_CLASSIFICATION_AVAILABLE to true and
make extension_is_channel_surface a real predicate (ExternalChannel), so
connected channel names render in the model runtime context instead of unknown.

Convert the two stubbed tests to positive cases: empty list -> Known([]),
mixed list -> only active channel-surface extensions reported.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* style(runtime-context): cargo fmt + drop unreachable classification branch

Remove the dead 'if !CHANNEL_CLASSIFICATION_AVAILABLE' arm inside the
Some(Ok(response)) match: lifecycle_fut only issues the ExtensionList call
when classification is enabled, so a present response always means it is on.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(webui-v2): repair Slack extension asset + locale checks after channel refactor

- assets.rs smoke test: assert showLegacySlackConnectActions (the refactor's
  built-in Slack status path) instead of the removed slackBuiltinStatus helper
- add extensions.kind.channel to all 10 non-en locales (en gained the key with
  the new channel surface kind; locale-parity test requires all locales match)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(webui-v2): cover ExtensionCard channel overflow + localize registry heading

Review findings (PR nearai#4778 review 4493764666):
- F2: add extension-card.test.mjs proving kind=channel/wasm_channel surface
  Setup (setup_required/failed) and Reconfigure (active/ready) overflow actions
  on the real component; channels-tab.test stubbed ExtensionCard so this was
  uncovered.
- F4: render the 'Available channels' registry heading via t(channels.availableChannels)
  instead of a hardcoded literal; add the key to all 11 locales. Update the
  channels-tab test to locate the registry section by the RegistryCard component
  (heading is now an interpolated value, not a template literal); drop the now
  unused renderedValueAfter helper.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(runtime-context): drop permanent classification flag, document surface_kinds cache

Review findings (PR nearai#4778 review 4493764666):
- F5: remove CHANNEL_CLASSIFICATION_AVAILABLE (permanently true after the stub
  flip) and run the lifecycle ExtensionList fetch unconditionally when a
  lifecycle facade is wired.
- F3: document AvailableExtensionPackage.surface_kinds as an intentional
  single-parse cache (re-deriving in summary() would re-run the manifest
  projection, undoing the parse-once optimization).

F1 (per-turn ExtensionList cost) accepted as-is: the fetch is spawned off the
critical path under a 500 ms budget with abort-on-drop; caching deferred to a
follow-up.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(webui-v2): suppress Activate for channel kinds during pairing

Review finding (PR nearai#4778 review 4494623523, finding 1): primaryExtensionAction
only suppressed the primary Activate button for legacy wasm_channel, so a
manifest-backed kind=channel Slack card fell through to 'activate' in
pairing_required/pairing states where the dedicated pairing section already
owns the flow. Suppress the primary action for channel-surface kinds in those
states (via isChannelExtensionKind); installed channels still return 'activate'.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(webui-v2): add channels.slack key + regression tests for surface-kind projection

Review findings (PR nearai#4778 review 06:28):
- Add missing channels.slack i18n key to all 11 locales (legacy Slack row
  rendered the raw key because the i18n helper returns the key on miss; the
  || "Slack" fallback never fired).
- Add filesystem-path test: a /system manifest with
  product_adapter.inbound.surface_kind = external_channel projects to
  ExternalChannel surface (previously only the bundled catalog path was covered).
- Add extension_kind regression test: non-channel summaries keep their runtime
  wire kind (wasm_tool, mcp_server) while channel surfaces map to "channel".

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reborn): gate Slack catalog entry behind slack-v2-host-beta; test wasm_channel registry

Review findings (PR nearai#4778 review 07:01):
- Gate the Slack first-party catalog entry, its only-Slack symbols (slack_package,
  slack_assets, SLACK_MANIFEST, slack_manifest_digest), the factory trust-policy
  Slack AdminEntry, and Slack-asserting tests behind slack-v2-host-beta. Without
  the feature the Slack route/runtime/WebUI mounts don't exist, so the catalog
  must not advertise an unrunnable Slack extension. Clean clippy + tests in both
  feature-on and feature-off configs.
- Add useExtensions hook test proving an uninstalled kind=wasm_channel registry
  entry lands in channelRegistry, not toolRegistry (isChannelExtensionKind covers
  both channel and wasm_channel).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(reborn): consolidate Slack trust policy tests (nearai#4778)

* feat(reborn): expose outbound delivery targets to model

---------

Co-authored-by: Henry Park <henrypark133@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(reborn): unify extension registry flow

* fix(reborn): address extension registry review comments

* test(reborn): cover merged extension registry flow

* fix(reborn): clean google oauth configure copy

* fix(reborn): align notion oauth setup copy

* fix(reborn): clean token setup copy
* fix(reborn): prefer github extension for repository data

* refactor(reborn): centralize github http routing hint
* fix(reborn): filter pull requests from github issues

* fix(reborn): preserve github issue pagination

* fix(reborn): use issue-only github search pagination

* fix(reborn): update github wasm trace fixture
…css in git

- bos-dev.sh now generates BETTER_AUTH_SECRET via openssl when .env is created or has an empty value
- Track ui/src/styles.css in git (add ! exception in .gitignore)
- Prevents auth plugin crash and UI build failure on fresh clones
* fix(automations-ui): readable summary cards and NEXT RUN value

Reflow the summary strip to at most three cards per row so the detail
text no longer wraps one word per line, and let StatCard accept a
valueClassName override so the NEXT RUN date renders at a smaller size
instead of truncating to "Jun…". Default StatCard sizing is unchanged.

* fix(automations-ui): surface delivery save errors and gate Slack hint

The delivery-defaults panel swallowed save/clear failures and showed no
feedback; it now renders an inline error from the mutation and flashes the
"Saved" confirmation on Clear as well as Save. The "reply approve <code> in
Slack" footnote is hidden unless an external Slack-style target exists.

* fix(automations-ui): label sub-hourly cron schedules

Minute- and hour-level cadences such as "* * * * *", "*/15 * * * *", and
"0 * * * *" rendered as "Custom schedule" because they have no single clock
time. They now read as "Every minute", "Every 15 minutes", and "Hourly at
:00".

* fix(automations-ui): space the run-row action button icons

The "Open run" and "Logs" buttons in the recent-runs list rendered the
icon flush against the label because the non-primary Button variants don't
add a gap between children. Add the same icon margin the rest of the app
uses for icon+label buttons.

* test(automations): lock the panel UI fixes into the served bundle

Add static-asset assertions driving the composed router so each Automations
panel UX fix — sub-hourly cron labels, summary card reflow + smaller NEXT RUN
value, run-row icon spacing, and delivery save-error/Slack-hint gating — is
guarded against a regression that drops it from the shipped SPA source.

* fix(i18n): add automations.delivery.saveFailed to every locale pack

The new key was added only to en.js, which breaks the i18n consistency test
that requires all locale packs to share the English key set. Add it to the
ten other packs (English placeholder, matching the existing untranslated
automations strings there).
* feat(product): explicit gate-open feedback for busy threads, no parking

Design decision (supersedes the closed defer-and-drain PR nearai#4812): a message
arriving while another run holds the thread is recorded with the honest
terminal status RejectedBusy and the user gets an explicit notice — gate-aware
("an approval gate is open on this thread — resolve it before continuing,
then resend") when the blocking run is BlockedApproval/BlockedAuth, generic
otherwise. No background resubmission: the user is the retry actor.

- threads: MessageStatus::RejectedBusy + mark_message_rejected_busy (both
  backends); DeferredBusy kept as a legacy deserialization label, no longer
  written; RejectedBusy -> Submitted allowed so resends work
- product_workflow: ThreadBusy branches mark RejectedBusy; response variant
  renamed RejectedBusy with status-derived notice field
- webui: rejected_busy ack renders the notice as a system message in chat;
  wire-shape test asserts tag + notice
- slack: copy already honest from nearai#4811; stale nearai#4812 revisit comment removed
- e2e: runtime-level test proves ThreadBusy -> RejectedBusy with notice,
  NO resubmission on the blocking run's terminal event, and a fresh submit
  succeeds afterward

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(review): make RejectedBusy terminal across replay, ack, compaction, UI

Reviewers found RejectedBusy still inheriting auto-resubmit/defer semantics,
contradicting the no-parking contract. Fixes:

- inbound_turn: from_replay_parts returns a terminal AlreadyRejected handoff
  for RejectedBusy (re-rejects, never resubmits); to_ack now emits a settled
  ProductInboundAck::RejectedBusy so transport retries get Duplicate instead
  of resubmitting. Legacy DeferredBusy rows keep the resubmit path.
- reborn_services: replayed RejectedBusy returns RebornSubmitTurnResponse::
  RejectedBusy again (idempotent re-rejection) instead of building a fresh
  submission; status-to-notice mapping locked by BlockedApproval/BlockedAuth/
  generic tests.
- compaction: RejectedBusy (and frozen legacy DeferredBusy) map to
  SkipEphemeral(StableNonModelVisible), not DeferUntilStable, so busy threads
  can still compact instead of deferring forever.
- webui useChat: always clear processing on rejected_busy, mark the optimistic
  message failed, collision-free system-message id; added hook tests.
- tests: runtime no-resubmission assertion anchored on message identity;
  mark_message_rejected_busy negative coverage; webui handler test reuses
  StubServices via a queued response.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* refactor(threads): retire mark_message_deferred_busy writer

DeferredBusy is now a read-only legacy label — production writes RejectedBusy.
Remove the live writer from the SessionThreadService trait, both backends, the
Arc forwarder, and all test fakes. Legacy DeferredBusy read/replay coverage is
preserved via a doc-hidden inject_legacy_deferred_busy_for_test back-door on the
in-memory backend (never called from production). The DeferredBusy enum variant
and all read/replay/compaction handling are unchanged.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(review): RejectedBusy follow-ups — Slack hint, honest replay, fail-loud, gating

- slack_delivery: SlackFinalReplyDeliveryObserver now recognizes
  ProductInboundAck::RejectedBusy { active_run_id: Some(_) } and posts the
  gate-aware busy hint (was DeferredBusy-only, so Slack rejections settled
  silently); None active_run_id posts nothing. Tests added.
- reborn_services: RejectedBusy response run metadata (active_run_id, status,
  event_cursor) is now Option — fresh ThreadBusy returns Some(real values),
  idempotent replay returns None instead of fabricating a fake Running run at
  cursor 0. Wire-shape + replay tests updated.
- inbound_turn: RejectedBusy replay fails loud on a malformed stored
  turn_run_id instead of silently dropping it.
- compaction: RejectedBusy + frozen legacy DeferredBusy map to
  SkipEphemeral(StableNonModelVisible), not DeferUntilStable, so busy threads
  compact instead of deferring forever. Integration tests added.
- threads: inject_legacy_deferred_busy_for_test gated behind a test-support
  cargo feature (absent from production builds); filesystem contract coverage
  for mark_message_rejected_busy (happy + invalid transitions).
- webui useChat: always clear processing on rejected_busy, mark the optimistic
  message failed, collision-free system-message id; hook tests added.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* test/docs(review): RejectedBusy coverage, durable timeline rejection, contract docs

- inbound_turn: regression test for fail-loud malformed RejectedBusy turn_run_id
- product_workflow_contract: RejectedBusy(None) settles + transport retry = Duplicate
- webui wire test: assert status + event_cursor present on fresh RejectedBusy path
- thread contracts (both backends): RejectedBusy -> Submitted resend transition
- webui useChat: persisted rejected_busy/deferred_busy rows now render failed on
  history reload with durable resend copy (was a normal-looking sent message);
  history-messages tests added
- slack_delivery: shared busy-hint path renamed deferred_busy_* -> busy_hint_*
  (fns, call sites, logs, docs); DeferredBusy kept only in the legacy arm
- runtime.rs: inline arch justification above the large RejectedBusy e2e test
- docs/reborn/contracts: product-adapters.md + conversation-binding.md document
  RejectedBusy as a durable terminal outcome; DeferredBusy marked legacy

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(review): openai-compat build break, Slack duplicate hint, conversations doc

- openai-compat-beta: ProductInboundAck::RejectedBusy added to the four
  non-exhaustive match sites (ack_helpers, chat_workflow, responses_workflow x2),
  mapped to the same retryable 429 as DeferredBusy — fixes E0004 that broke any
  build enabling openai-compat-beta. RejectedBusy->429 tests added.
- slack_delivery: busy-hint run-id extraction now unwraps
  Duplicate { prior } recursively, so a transport retry arriving as
  Duplicate { prior: RejectedBusy { Some(run) } } still posts the busy-thread
  hint when the first was lost; the per-(conversation, run_id) throttle
  suppresses genuine repeat posts. Tests added.
- ironclaw_conversations/CLAUDE.md: the idempotency guardrail now distinguishes
  transient submit failures (retry, rotate key) from thread-busy admission
  (terminal RejectedBusy, no retry-until-submitted, user resends).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(review): RejectedBusy terminal at storage, compaction safety, ledger no-fake-run

- threads: RejectedBusy removed from ensure_user_accepted in both backends —
  a stored RejectedBusy row can no longer transition to Submitted; resend is a
  fresh message. Prior-pass RejectedBusy->Submitted tests inverted to assert the
  transition is now terminal (InvalidMessageTransition). DeferredBusy admission
  kept (legacy replay still resubmits).
- compaction: only RejectedBusy (terminal) is SkipEphemeral; DeferredBusy moved
  back to DeferUntilStable since legacy rows can still reach Submitted — prevents
  a summary silently omitting a message that later becomes model-visible.
- workflow ledger: RejectedBusy { active_run_id: None } maps to
  ActionDispatchKind::NoOp instead of minting a fresh TurnRunId; still settles
  durably, no fabricated run id.
- inbound_turn test: the misnamed legacy-DeferredBusy test now actually injects a
  legacy DeferredBusy row and asserts resubmission, distinct from the RejectedBusy
  re-rejection test.
- openai-compat: added cancel-path RejectedBusy->429 handler test.
- slack_delivery: comments corrected to describe the recursive Duplicate{prior}
  extraction (no longer claim Duplicate{DeferredBusy} returns None).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(review): non-retryable 429 for terminal RejectedBusy + busy rename + coverage

Address open PR review findings on the busy-thread rejection work:

- openai-compat: split the busy ack arm so terminal RejectedBusy maps to
  429 retryable=false (client must issue a new request), while legacy
  DeferredBusy keeps retryable=true. Covers chat create, responses create,
  and responses cancel paths; regression unit tests on the retryable flag.
- slack_delivery: rename SLACK_DEFERRED_BUSY_* constants to SLACK_BUSY_*
  (path now serves RejectedBusy + legacy DeferredBusy); refresh the stale
  "silently dropped (pending gate)" comment to cover generic RejectedBusy.
- compaction_task: correct the StableNonModelVisible doc comment — only
  terminal RejectedBusy is skipped; legacy DeferredBusy is DeferUntilStable
  (it can still transition to Submitted).
- tests: add filesystem legacy DeferredBusy on-disk round-trip coverage and
  a reborn_services RejectedBusy mark-failure reconcile-via-replay test.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(test): update root busy test to terminal RejectedBusy contract

The root integration test product_workflow_retries_after_filesystem_deferred_busy_release
still asserted the legacy DeferredBusy auto-resubmit contract (busy ->
DeferredBusy -> retry resubmits, submission_count 1->2). The live product
workflow now emits terminal RejectedBusy for busy user messages; the PR
updated crate-level tests but missed this root-level one, failing the
"Reborn root tests" CI job.

Rewrite + rename to product_workflow_rejects_busy_and_does_not_resubmit_on_filesystem_replay,
mirroring crates/ironclaw_product_workflow inbound_turn_contract's
rejected_busy_replay_is_re_rejected_not_resubmitted: first ack is
RejectedBusy (submission_count == 1); a same-event replay settles via the
idempotency ledger and returns Duplicate { prior: RejectedBusy } with no
resubmission (submission_count stays 1).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(compaction): terminal RejectedBusy cut point no longer blocks compaction

Review finding (High): the compaction terminal cut-point validation rejected
every SkipEphemeral disposition with InvalidCutPoint, but RejectedBusy now
classifies as SkipEphemeral(StableNonModelVisible). So a compaction range whose
drop_through_seq landed on a RejectedBusy message hard-failed — contradicting
this PR's goal that terminal RejectedBusy must never block compaction.

Allow a stable-non-model-visible terminal (RejectedBusy) as a legal cut point:
it is excluded from the compacted output like the in-range SkipEphemeral case
and compaction proceeds. Non-User Include and RejectInvalid still error.
Regression test: compaction_port_accepts_terminal_cut_point_that_is_rejected_busy.

Also add legacy_deferred_busy_mark_failure_reconciles_via_replay covering the
reconcile branch's legacy DeferredBusy replay path (RejectedBusy was already
covered).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(compaction): narrow terminal cut-point accept to StableNonModelVisible

Review follow-up: the terminal cut-point arm accepted SkipEphemeral(_) with a
wildcard, which would also admit CapabilityDisplayPreview as a valid terminal.
Only StableNonModelVisible (terminal RejectedBusy) should qualify. Match the
explicit variant so other ephemeral skip reasons fall through to
InvalidCutPoint and fail loud.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(review): reconcile only RejectedBusy as terminal + coverage/naming follow-ups

Address the latest review round (post origin/main merge):

- reborn_services: the mark-failure reconcile path treated legacy DeferredBusy
  as a terminal already-settled state. DeferredBusy is non-terminal (a later
  replay treats it like Accepted and can resubmit), so claiming terminal over
  it violated the no-resubmit guarantee. Drop DeferredBusy from the reconcile
  predicate — only RejectedBusy is terminal; a DeferredBusy row now surfaces
  the mark failure (503 retryable) instead of a false-terminal RejectedBusy.
  Flipped the legacy test to assert the surfaced error.
- compaction: add regression test that a terminal CapabilityDisplayPreview cut
  point returns InvalidCutPoint (only StableNonModelVisible is a legal terminal).
- fakes: FakeProductAdapter now records RejectedBusy in accepted_envelopes
  (durable like Accepted/DeferredBusy) so fake-based tests don't undercount.
- slack_delivery: arch-exempt comment updated from deferred-busy to busy-thread
  / RejectedBusy terminology.
- reborn_services_contract: retext a busy-submit test that asserts RejectedBusy
  but still labeled the path "deferred".

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(test): correct scripted-helper comments — DeferredBusy is non-terminal

Follow-up to the reconcile fix: the DeferredBusyMarkFails scripted helper and
its replay branch still documented the old behavior (DeferredBusy "settles"
reconciliation). reconcile_terminal_duplicate now accepts only RejectedBusy as
terminal, so a DeferredBusy replay surfaces the mark error (Unavailable/503)
instead of a false-terminal RejectedBusy. Comment-only; logic unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(test): align sibling probe-count comments with 3-call reconcile flow

Follow-up: the replay_call_count field doc and rejected_busy_mark_fails() doc
still described the old 2-call probe sequence. Both scripted mark-fail helpers
return None on the first two idempotency probes and Some(..) on the third
(reconcile) probe (count <= 2 guard). Comment-only; logic unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(compaction): reject empty cut-point range instead of summarizing nothing

Review finding (Medium): making a terminal RejectedBusy a legal cut point opened
an edge — a range whose only message is that rejection (or any all-skip-ephemeral
span) produced an empty validated_messages, then still ran inference on an empty
prompt and persisted a meaningless summary artifact.

Guard after the deferred-reason check: if validated_messages is empty, return
InvalidCutPoint before build_input — nothing model-visible to summarize. The
deferred-reason early-return stays first so legitimate deferrals are unaffected.
Regression test: a range whose only message is a terminal RejectedBusy returns
InvalidCutPoint and never calls inference.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test/docs(review): RejectedBusy command ack, ack_helpers, 429 spec, banner

Address review test/doc gaps on the busy-rejection work:

- product_command_workflow_contract: cover the command-dispatch path where
  command_service returns RejectedBusy -> UnsupportedActionKind -> terminal
  Rejected ack (previously only user-message RejectedBusy was tested).
- ack_helpers: unit-test that internal_refs_from_ack rejects RejectedBusy with
  the internal error (no internal refs bound for a terminal busy ack).
- docs/reborn/contracts/openai-compatible-api.md: document the busy 429 split —
  terminal RejectedBusy is non-retryable (client must issue a new request),
  legacy DeferredBusy stays retryable.
- reborn_services_contract: add the missing opening separator on the Legacy
  DeferredBusy test section banner.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(compaction): trim summary span to last visible message at a busy cut point

Review finding (Medium): when the compaction cut point is a terminal RejectedBusy,
the summary span ended at drop_through_seq, covering that non-visible message. The
thread backends' context builder skips any ReplaceRangeWhenSelected summary whose
span covers a non-model-context-visible message (summary_covers_hidden_content), so
the summary was persisted but never applied — a dead artifact.

Trim end_sequence to the last model-visible (Include'd) message's sequence so the
span excludes trailing non-visible terminals; the summary then applies. Folds the
empty-range guard into the same `validated_messages.last()` match (None => empty
range => InvalidCutPoint) — no production unwrap/expect. Regression test asserts a
[visible@1, RejectedBusy@2] range compacted through seq 2 yields a summary spanning
end_sequence=1, not 2.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test: pin command RejectedBusy error to ProductAdapterError::Internal

Review nit: the command-RejectedBusy test used a bare expect_err (any error).
Pin the concrete public variant: ProductAdapterError::Internal — which is what
ProductWorkflowError::UnsupportedActionKind maps to at the adapter boundary. The
kind string ("unsupported action kind: ...") is wrapped in RedactedString and
not exposed via Display, so Internal is the tightest assertable pin from the
public return type.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(threads): summary may span permanently-terminal non-visible messages

Review finding (Medium, egGm): summary_covers_hidden_content blocked a
ReplaceRangeWhenSelected summary whose span covered ANY non-model-context-visible
message. Compaction legitimately spans non-visible rows it skipped from the
summary content (e.g. an interior terminal RejectedBusy, or a capability preview),
so those summaries were silently dropped — the compaction-layer trailing trim
couldn't fix an interior hole.

Block the summary only when the span covers a non-visible message that can still
RESURFACE as model-visible (Draft / Interrupted / Superseded / DeferredBusy).
Permanently-terminal non-visible rows (RejectedBusy, CapabilityDisplayPreview
kind) never resurface, so spanning them is safe — the summary content already
excludes them and they are never shown in context. Identical change in both
in_memory and filesystem backends via a shared can_resurface_as_model_visible
helper; Redacted/Deleted keep blocking. Also corrects the pre-existing
capability-preview span behavior (two tests updated). Regression tests on both
backends: interior RejectedBusy summary is applied; interior Draft is not.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
* fix(reborn): allow read-only github capabilities

* test(reborn): future-proof github approval boundary

* fix(reborn): keep GitHub code search gated

---------

Co-authored-by: Robert Yan <mstr.raphael@gmail.com>
* fix(automations-ui): readable summary cards and NEXT RUN value

Reflow the summary strip to at most three cards per row so the detail
text no longer wraps one word per line, and let StatCard accept a
valueClassName override so the NEXT RUN date renders at a smaller size
instead of truncating to "Jun…". Default StatCard sizing is unchanged.

* fix(automations-ui): surface delivery save errors and gate Slack hint

The delivery-defaults panel swallowed save/clear failures and showed no
feedback; it now renders an inline error from the mutation and flashes the
"Saved" confirmation on Clear as well as Save. The "reply approve <code> in
Slack" footnote is hidden unless an external Slack-style target exists.

* fix(automations-ui): label sub-hourly cron schedules

Minute- and hour-level cadences such as "* * * * *", "*/15 * * * *", and
"0 * * * *" rendered as "Custom schedule" because they have no single clock
time. They now read as "Every minute", "Every 15 minutes", and "Hourly at
:00".

* fix(automations-ui): space the run-row action button icons

The "Open run" and "Logs" buttons in the recent-runs list rendered the
icon flush against the label because the non-primary Button variants don't
add a gap between children. Add the same icon margin the rest of the app
uses for icon+label buttons.

* fix(automations-ui): consistent summary counts and next-run

The Running/Failures summary cards counted individual runs while the
matching filter tabs counted automations, so the numbers disagreed; both
now count automations. The soonest "Next run" no longer includes paused
triggers, which keep a stored slot they will never actually fire.

* test(automations): lock the panel UI fixes into the served bundle

Add static-asset assertions driving the composed router so each Automations
panel UX fix — sub-hourly cron labels, summary card reflow + smaller NEXT RUN
value, run-row icon spacing, and delivery save-error/Slack-hint gating — is
guarded against a regression that drops it from the shipped SPA source.

* feat(automations): surface scheduler-off state and run it by default on serve

Scheduled automations never fired because the trigger poller is disabled by
default and nothing told the user. The list response now carries
scheduler_enabled (sourced from runtime readiness) and the panel shows a
"scheduling is turned off" notice when it is false. The local `ironclaw-reborn
serve` surface enables the poller by default; config and env still override it.

* fix(i18n): add automations.delivery.saveFailed to every locale pack

The new key was added only to en.js, which breaks the i18n consistency test
that requires all locale packs to share the English key set. Add it to the
ten other packs (English placeholder, matching the existing untranslated
automations strings there).

* fix(automations-ui): clear stale Saved flash before a new delivery write

The save-error alert is gated on !showSaved, so a "Saved" flash still
showing from a prior success would hide the error of a new failing
save/clear. Reset the flash (and its timer) at the start of every attempt.

* fix(automations): harden next-run filter and lock scheduler_enabled on the wire

Use loose `!= null` in the soonest-next-run filter so a missing
next_run_timestamp can't slip through, and assert scheduler_enabled in the
list-automations handler contract test so a serialization drift of the new
field is caught at the wire, not just in the facade.

* fix(i18n): add automations.schedulerOff keys to every locale pack

The scheduler-off notice keys were added only to en.js, which breaks the
i18n consistency test requiring all locale packs to share the English key
set. Add both keys to the ten other packs (English placeholder).

* i18n(automations): translate schedulerOff strings in all locale packs

The scheduler-off notice was English in every non-English pack, giving
Arabic/German/Spanish/French/Hindi/Japanese/Korean/pt-BR/Ukrainian/zh-CN
users a mixed-language UI. Provide real translations.

* i18n(automations): translate delivery.saveFailed in all locale packs

The save-failed delivery error was English in every non-English pack. Provide
real translations so users don't see mixed-language UI when a save fails.

* Localize automation summary counts

* Remove duplicate automation summary locale keys

---------

Co-authored-by: Robert Yan <mstr.raphael@gmail.com>
…d from persistent approval scope (nearai#4825) (nearai#4835)

* make 'always allow' approvals persist (tested on google suite)'

* test(approvals): lock criterion-5 backward-compat for project-scoped policies (nearai#4825)

* reduce slop

* fix(approvals): address approval scope review

* fix(approvals): preserve legacy approval lookup

* fix(approvals): find legacy grants across threads

* fix(approvals): drop legacy approval scope compatibility

* fix(approvals): simplify threadless policy lookup

---------

Co-authored-by: Emil Bogomolov <emil.bogomolov@near.ai>
Co-authored-by: Henry Park <henrypark133@gmail.com>
* Fix Railway WebUI OAuth callback origin

* fix(reborn-cli): address OAuth base URL review

* fix(reborn-cli): fail closed on hosted oauth base url
…rai#4827)

* fix(host-runtime): accept empty body/body_base64 in builtin.http

The HTTP tool's `body()` validator rejected any request that carried
*both* a `body` and a `body_base64` field, even when both were empty
strings. The JSON schema lists both fields, so models routinely emit
`{"body": "", "body_base64": "", ...}` as defaults for a bodyless GET.
That tripped the mutual-exclusion check and failed the call with
`InputEncode` before it ever dispatched — on every attempt — so the
agent could never make a successful request and would retry until its
loop gave up.

Treat a null or empty-string field as absent: only a non-empty value
counts as "set", so the mutual-exclusion check fires only when the
caller genuinely supplies two competing bodies. Behavior for a real
`body`, a real `body_base64`, or a genuine both-non-empty conflict is
unchanged.

Adds unit tests for body() covering empty-both, absent, string body,
base64 body, both-non-empty (rejected), and JSON-object body.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* style: rustfmt the body() unit tests

---------

Co-authored-by: Pranav Raja <pranav.raja@near.ai>
Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
…r injection (nearai#4588)

* feat(reborn): expose a trajectory observer hook on RebornRuntimeInput

The reborn runtime is sealed: build_reborn_runtime returns only the final
AssistantReply, and per-step capability (tool) calls + results live in internal
stores. Downstream consumers (benchmark harnesses, UI/debuggers) can't observe
the agent's trajectory.

Add `RebornTrajectoryObserver` (pub trait: on_capability_input(call_id, name,
args) / on_capability_result(call_id, output)) and
`RebornRuntimeInput::with_trajectory_observer`. The local-dev capability IO
(`LocalDevCapabilityIo`) forwards each tool call's name+args (at input staging)
and result (at result write) to the observer when present — reusing the same
data it already records for display previews. No-op when unset; best-effort.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* debug: trace observer hook firing (temporary)

* feat(reborn): trajectory observer — capability_id on result, reliable spine

Provider tool calls are staged by a lower decorator that bypasses the
LocalDevCapabilityIo input path, so on_capability_input does not fire for
them. on_capability_result fires for every completed capability — make it
carry the capability_id so consumers can reconstruct the trajectory (name +
output) from results alone. Input args capture is a follow-up.

* feat(reborn): capture capability input args at the host port chokepoint

Provider tool calls are staged by ProviderToolCallInputResolver, which keeps
args in a private map and bypasses the capability-IO input hook — so inputs
never reached the trajectory observer (only results did). Move the observer
trait down to ironclaw_loop_support (CapabilityTrajectoryObserver, re-exported
from composition as RebornTrajectoryObserver) and hook it in
HostRuntimeLoopCapabilityPort::invoke_capability right after the input
resolves — the one place the model's resolved arguments are visible. Threaded
through HostRuntimeLoopCapabilityPortFactory + the local-dev factory. Result
hook unchanged. Now name + args + output are all captured.

* feat(reborn): host LLM-provider injection seam

ResolvedRebornLlm::with_provider — drive the runtime with a caller-supplied
LlmProvider (e.g. an instrumented wrapper that counts tokens/cost and captures
reasoning) instead of always building one from config; build_llm_gateway honors
the override. The only viable observability path for reborn, whose model calls
run in spawned worker tasks a per-task tracing subscriber can't reach.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(reborn): cover trajectory observer + LLM provider override seams

Addresses Firat's two blocking review findings on nearai#4588 (both
missing-integration-test, per AGENTS.md "test through the caller"):

1. Trajectory observer callbacks — drive the real call sites with a
   recording CapabilityTrajectoryObserver:
   - host port: invoke_capability via HostRuntimeLoopCapabilityPortFactory
     ::with_trajectory_observer asserts on_capability_input fires with the
     resolved capability id + tool-call arguments.
   - local-dev IO: register_provider_tool_call_input + write_capability_result
     assert on_capability_input and on_capability_result fire and correlate by
     input ref.

2. LLM provider override — build_llm_gateway_drives_provider_override_not_config
   injects a counting mock via ResolvedRebornLlm::with_provider, points config
   at a dead endpoint, and asserts the gateway returns the mock's sentinel
   (proving the override is driven, not a config-built chain).

Also fixes pre-existing breakage this surfaced: 5 LocalDevLoopCapabilityPort
Factory test initializers (shell_tests.rs + tests.rs) were missing the
trajectory_observer field added by this PR, so the composition crate's tests
did not compile under --features root-llm-provider.

loop_support: 301 passed; composition (root-llm-provider): 520 passed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reborn): make trajectory observer input semantics consistent

Addresses Copilot's follow-up findings on the observer seam:

- Drop the `on_capability_input` callback from `LocalDevCapabilityIo::
  register_provider_tool_call_input`. It forwarded the raw provider tool
  name (`builtin_echo`) as the capability id — conflicting with the
  observer contract (resolved dotted `builtin.echo`) and the authoritative
  port-level hook — and `ProviderToolCallInputResolver` doesn't delegate
  here for provider tool calls, so it never fired in practice anyway.
  `HostRuntimeLoopCapabilityPort::invoke_capability` remains the single
  source of `on_capability_input` (resolved id); `LocalDevCapabilityIo`
  remains the source of `on_capability_result`.

- Clarify the trait doc: `arguments` is the raw model-emitted tool-call
  input resolved from the input ref (the callback fires before schema
  normalization), which is what the trajectory should record.

- Refocus the local-dev test on `on_capability_result` forwarding +
  correlation, and assert input staging does NOT emit `on_capability_input`
  from local-dev IO. Port-level input semantics stay covered by the
  capability_port.rs test.

loop_support: 301 passed; composition (root-llm-provider): 520 passed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* wire trajectory_observer through RefreshingLocalDevCapabilityPortConfig

Completes the main-merge conflict resolution: local_dev.rs passes
trajectory_observer into the refreshing-port config, so the config struct +
port struct must carry it and build_inner must apply it via
.with_trajectory_observer(). (Missed staging this file in the merge commit.)

* test(reborn): lock down the observability seams against regression

nearai#4588 exposes two seams a downstream harness relies on. Add tests so a
future refactor can't silently break either:

- capability_io_forwards_result_to_trajectory_observer: drives
  write_capability_result and asserts on_capability_result fires with the
  correct (call_id, capability_id, output) — the result half of the
  trajectory observer (tool-call outputs).
- build_llm_gateway_drives_provider_override_not_config: asserts the gateway
  drives a provider injected via ResolvedRebornLlm::with_provider (config
  points at a dead endpoint), proving the provider-injection seam works —
  this is how the bench captures reasoning / tokens / cost / system-prompt /
  tool-definitions. (Restores the test dropped during the main merge.)

The input half (on_capability_input) is already covered by
invoke_capability_forwards_resolved_input_to_trajectory_observer in
ironclaw_loop_support. All three pass.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(reborn): drop the false-confidence result-hook test

capability_io_forwards_result_to_trajectory_observer called
write_capability_result directly, so it stayed green even though the
result hook is unreachable end-to-end while capability dispatch fails
(the LocalDevYolo InputEncode regression) — i.e. it did not fail when
the feature it claimed to cover was actually broken. Remove it rather
than ship false confidence.

The result hook lives in LocalDevCapabilityIo and is only reached by a
real local-dev runtime turn, so an honest guard must drive the full
runtime and is red until the dispatch regression is fixed; that guard
belongs as an end-to-end test (PR, once green) or a bench pre-flight,
not a direct-call unit test.

Kept: invoke_capability_forwards_resolved_input_to_trajectory_observer
(input hook, real port code path) and
build_llm_gateway_drives_provider_override_not_config (provider seam) —
both genuinely fail if their seam regresses.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(reborn): address review on the trajectory-observer + provider seams

Resolves Henry + Firat review comments on nearai#4588:

- Composition-owned RebornTrajectoryObserver trait + adapter to the
  loop-support CapabilityTrajectoryObserver, instead of re-exporting the
  substrate trait directly (CLAUDE.md: facade-shaped handles only). Loop-support
  contract changes no longer break the public Reborn API. (Henry#8)

- Safe-preview by default: with_trajectory_observer now forwards bounded
  (truncated strings / capped arrays) payloads so a logs/UI/telemetry sink stays
  within the model-visible display boundary; a trusted in-process consumer that
  needs verbatim tool I/O opts in via the new with_raw_trajectory_observer.
  (Henry#5)

- catch_unwind around both observer call sites (input hook in capability_port,
  result hook in LocalDevCapabilityIo) so a panicking observer can't unwind the
  capability hot path; trait doc now states the never-block / panic-caught
  contract. (Henry#1/#6)

- e2e test local_dev_runtime_forwards_tool_call_trajectory_to_raw_observer:
  drives a real build_reborn_runtime turn dispatching builtin.echo and asserts
  BOTH input and result callbacks fire on the genuine dispatch path — honest
  coverage that replaces the dropped direct-call result-hook test, and proves
  the observer threads through build_reborn_runtime. (Firat#1, Henry#3/nearai#7)

- Strengthened provider-injection docs: the config-vs-override invariant and why
  the feature-gated seam takes the LlmProvider substrate trait. (Henry#4/nearai#9/nearai#11)

- Fixed the LocalDevCapabilityIo observer field comment to describe its actual
  result-only responsibility. (Henry#10)

Provider-override coverage (Firat#2/Henry#2) already landed in
build_llm_gateway_drives_provider_override_not_config.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(reborn): second-round review fixes on the trajectory/provider seams

Addresses Henry's review of the first round (nearai#4588):

- safe_preview_value now bounds objects (entry cap), recursion depth, and total
  nodes — not just strings/arrays — so a wide or deeply nested capability result
  can't force unbounded traversal/allocation on the hot path. (3405419089)

- Narrowed the loop-support CapabilityTrajectoryObserver to input-only:
  HostRuntimeLoopCapabilityPort never staged results through the port (results
  go via LoopCapabilityResultWriter), so advertising on_capability_result there
  was a contract a direct user could never see fire. Result observation stays on
  the composition path (LocalDevCapabilityIo). (3405419104)

- Synthetic capabilities (e.g. builtin.skill_activate) bypass the inner port's
  input hook, so the synthetic wrapper now emits on_capability_input itself after
  resolving input — otherwise consumers saw an unpaired result with no args.
  (3405419110)

- Provider injection no longer accepts a wholesale Arc<dyn LlmProvider> through
  the facade: with_provider is replaced by with_provider_factory, a decorator
  Fn(Arc<dyn LlmProvider>) -> Arc<dyn LlmProvider>. The composition always builds
  the provider from config (config stays the single construction source —
  collapses the old config-vs-override invariant too) and hands it to the factory
  to wrap. (3405419100, 3405419146)

- New caller-level test local_dev_runtime_safe_preview_observer_receives_bounded_payload:
  installs the default with_trajectory_observer, drives a real turn with a large
  echo payload, asserts the observer receives a truncated preview. (3405419095)

- Dropped the stale nearai#4588/main-rebase comment for a durable invariant. (3405419113)

cargo test (loop_support + reborn_composition, single-threaded) green; clippy
clean on touched files.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* style(reborn): rustfmt the trajectory/provider review changes

Formatting-only: import grouping + mod ordering in the two lib.rs re-export
blocks, and wrapping in runtime.rs / local_dev.rs / trajectory_observer.rs.
Fixes the Formatting + Code Style CI checks. No behaviour change.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* style(reborn): drop std-Mutex guard before await in observer e2e tests

clippy::await_holding_lock (-D warnings): the two trajectory-observer e2e
tests held the observer's std::sync::Mutex guard across runtime.shutdown().await.
Shut down before inspecting the recorded callbacks (the data is already captured
during the turn) so no guard is held across an await. Fixes Clippy (all-features).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* WIP(bench): http empty-body + multi-tool-call port reuse + final-answer nudge

Local checkpoint so the bench builds against a stable tree (uncommitted
edits were being reverted mid-session). Bundles: http body() empty-field
fix, RefreshingLocalDevCapabilityPort register reuse, the gated
final-answer nudge + interactive_profile gate flip, and the
trajectory-observer safe_preview borrow fix.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* style(reborn): wrap an over-long line for rustfmt 1.9.0

CI installs the latest stable rustfmt (1.9.0 / Rust 1.96), which wraps a
long eprintln! that older rustfmt left inline. Fixes the Formatting check.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* style(reborn): collapse nested if for clippy 1.96 collapsible_if

clippy 1.96 (CI's stable) flags the nested if-let in the final-answer-nudge
site as collapsible; fold it into a let-chain. No behaviour change. Fixes
Clippy (all-features).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* WIP(bench): multi-tool-call port reuse (matches main nearai#4790)

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* WIP(bench): nudge isolation - disable gate to measure marginal contribution

* Revert stray bench WIP accidentally committed onto this branch

Removes the http/nudge/multi-tool-call/diagnostic WIP commits
(c4bbb5f, 2c670b4, 6da818a) that were committed onto the
reborn-trajectory-observer branch by mistake during benchmarking and
swept to origin by a main-merge push. Restores the affected files to
origin/main (multi-tool-call is already fixed there by nearai#4790; the http
fix lives in PR nearai#4827). Observer-owned changes in state.rs,
refreshing_capability_port.rs, and local_dev.rs are preserved minus the
stray WIP additions. No history rewrite / force-push.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(reborn): preserve provider factory across reload + reject observer off local-dev

Addresses Firat's review of the trajectory/provider seams (nearai#4588):

- Provider factory now survives a live config reload. build_llm_gateway applied
  the factory to the bare config provider *before* wrapping it in the
  SwappableLlmProvider, so the first WebUI/settings reload (which swaps the
  swappable's inner) silently dropped the instrumentation wrapper. Invert the
  layering: build the config provider, put it behind the swappable + reload
  handle, then apply the factory *over the swappable* for the gateway-facing
  provider. Reloads swap the inner; the wrapper stays in the call path.
  Regression test provider_factory_survives_live_reload reloads and proves the
  wrapper still observes subsequent model calls.

- Reject a trajectory observer on profiles without a local runtime. The observer
  is wired only through the local-dev capability path; Production silently
  dropped it, so a caller got an empty trajectory with no error. Fail fast with
  InvalidArgument and document the seam as local-dev/bench-only. Test
  build_reborn_runtime_rejects_trajectory_observer_for_production.

cargo fmt + clippy (all-features, -D warnings) clean under rustfmt 1.9/clippy 1.96.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(reborn): note trajectory observer is local-dev/bench-only

Document the local-dev-only constraint + fail-fast behavior on the public
with_trajectory_observer / with_raw_trajectory_observer setters (Firat review).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Pranav Raja <pranav.raja@near.ai>
…reborn) (nearai#4936)

The dispatcher only parsed <suite> + --model, so every /benchmark comment ran
the legacy ironclaw runtime. nearai/benchmarks' reusable workflow already
accepts a 'framework' input (default ironclaw), so wire the comment through:

- parse an optional '--framework <ironclaw|ironclaw-reborn>' flag (either order
  with --model), validate it against an allowlist, and reject unrecognised
  --flags so typos fail loudly instead of silently defaulting;
- thread it through the parse job's outputs and forward it to
  bench-pr-reusable.yml (empty = unset → reusable workflow's ironclaw default);
- show the framework in the 'started' comment.

Usage: /benchmark <suite> [--model <id>] [--framework ironclaw-reborn]

Note: reborn runs require the built ironclaw to carry the reborn observability
seams (nearai#4588); point the PR/ironclaw-rev at a commit that has them.

Co-authored-by: Claude (rebase helper) <noreply@anthropic.com>
…earai#4559)

* docs: trace commons agent onboarding design spec

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs: address spec review findings (trust anchoring, key staging, consumption atomicity, replay validation)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs: spec review round 2 nits (server-anchored tenant wording, pending-key cleanup)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs: implementation plan for trace commons agent onboarding

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs: address plan review findings (scope threading refactor, dispatch model, dev-deps, LazyLock hazard)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs: plan review round 2 fixes (literal dep versions, context constructor threading depth)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs: incorporate server-agent coordination feedback (optional community/profile/leaderboard URLs)

From TraceCommons/trace-commons#136-nearai#141 comments: onboard response
gains optional browser-surface navigation hints, sanitized client-side
(HTTPS or dropped), never part of issuer trust anchoring.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(traces): onboarding wire types matching trace-commons-server contract

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(traces): invite URL parsing with origin trust anchoring

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(traces): device keypair lifecycle with pending staging and self-signed workload JWTs

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(traces): auth_mode and device_key_id policy fields with legacy-compatible defaults

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(traces): onboard() orchestration with trust anchoring and retry-safe key staging

Wire invite parsing, device key staging, onboard POST, issuer origin trust anchoring,
ingest_url HTTPS enforcement, keypair promotion, and policy write into onboard_at_dir().
Refactors invite.rs to extract pub(crate) is_https_or_loopback, origin_of, and host_only
helpers shared with mod.rs (one source of truth for origin/bracket handling). Adds axum
mock-issuer tests covering the happy path, mismatch rejection, terminal vs transient error
key retention, insecure ingest URL, loopback ingest allowance, community URL sanitisation,
and retry key reuse.

Partial-failure lockout fix (spec §2.2): promote() no longer deletes the pending file.
The flow now writes the tenant key file, then the policy, and only discards the pending
file after BOTH durably succeed. If the policy write fails the pending key survives, so a
retry reloads the same key (server idempotency returns the original registration) and
harmlessly overwrites the tenant file — no permanent lockout from a consumed invite with a
regenerated keypair. Regression test simulates a policy-write failure (policy.json
pre-created as a non-empty dir so the atomic rename fails), asserts Err(Persist) with the
pending key intact, then asserts a retry succeeds reusing the same device_key_id.

Response validation (defense-in-depth): reject schema_version != the v1 response constant
as MalformedResponse, and cross-check the response device_key_id against the locally derived
id (we never trust the response value for policy; a disagreement is now treated as a tamper
signal and rejected). Both covered by tests.

The onboard response body is read with the 64 KB cap enforced per-chunk during streaming
(mirroring read_bounded_trace_upload_claim_response) rather than buffering the whole body
first, so a hostile server cannot force a large allocation.

Also fixes a pre-existing test-isolation defect surfaced by the added load: the
remote-request timeout test configured a 50ms timeout via the process-global
IRONCLAW_TRACE_REMOTE_REQUEST_TIMEOUT_MS env var. set_var is process-global, so under
parallel execution the 50ms value leaked into other tests' trace HTTP clients, producing
spurious `operation timed out` failures against fast local mocks. Replace the env mutation
with a task-scoped TEST_REMOTE_REQUEST_TIMEOUT_OVERRIDE task-local (visible only within the
awaiting test's own task tree, zero production change; documents the spawn caveat), and
decouple the timing assertion from a tight wall-clock race so it no longer flakes when
reqwest's timer is delayed under an oversubscribed runtime.

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(traces): device-key self-signed workload JWT branch in upload-claim refresh

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(engine): trace_commons onboard and status first-party tools with agent guidance

Add two model-visible first-party capabilities to the Reborn engine:
- builtin.trace_commons.onboard: drives operator-invite enrollment flow with
  explicit per-conversation consent gate (confirmed=true required before any
  network call); maps OnboardOutcome/OnboardError to clean agent-readable JSON
- builtin.trace_commons.status: read-only enrollment state inspector

Wires ironclaw_reborn_traces into ironclaw_host_runtime, creates schema files
(schemas/builtin/trace-commons-{onboard,status}.{input,output}.v1.json) and
prompt doc files (prompts/builtin/trace-commons-{onboard,status}.md) at the
manifest-derived paths. Includes 11 unit tests covering input parsing, consent
refusal, success/error value formatting, and status formatting.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs: add Task 11 — credits visibility (console display + agent-queryable balance)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(engine): e2e trace commons onboarding through capability dispatch

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(traces): document agent onboarding flow in trace-commons internal doc

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs: correct Task 11 console scope (credit endpoint already exists; frontend = coordinate with designer)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(traces): trace_commons.credits agent-queryable balance tool

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* feat(gateway): minimal Trace Commons credits card in settings

Co-Authored-By: Claude Sonnet 4.6 <noreply@anthropic.com>

* fix(traces): store upload-claim endpoint in policy; preserve primary onboard error; block metadata/link-local/multicast issuers

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(traces): route agent onboarding HTTP through host network-egress policy (nearai#4560)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* build: update Cargo.lock for trace-commons onboarding dev-deps

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(traces): drop orphaned schema/prompt files (main resolves builtin schemas inline; prompt_doc_ref dropped)

Post-merge cleanup: main's first_party_tools now resolves builtin input schemas
via the inline schemas.rs match (trace_commons arms added during the merge) and
sets prompt_doc_ref: None for all builtins, so the physical trace-commons-*.json
schema files and trace-commons-*.md prompt docs are no longer referenced. The
onboard consent contract remains in the capability description and is enforced in
dispatch_onboard.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* Fix Trace Commons invite hash contract

* fix(traces): grant trace_commons capabilities in local-dev policy

The three builtin.trace_commons.* capabilities were declared in the
first-party package but had no [[grants]] entries in
local_dev_capability_policy.toml, so local-dev runs (repl/serve)
filtered them out of the model-visible tool surface entirely. The
provider-level authority_effects ceiling had external_write, but the
per-capability grants were never added.

onboard gets the local_dev_wildcard egress profile (invite origins are
operator-chosen; private/metadata IP ranges stay blocked by the shared
enforcer). status/credits are read-only.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(traces): add Reborn e2e coverage for trace_commons first-party tools

Closes the coverage gate failure: builtin.trace_commons.{onboard,status,
credits} were declared in the first-party package but missing from
REBORN_FIRST_PARTY_E2E_COVERED_CAPABILITIES, failing
reborn_builtin_first_party_capability_e2e_coverage_is_complete on both
the Reborn root tests and all-features CI jobs.

Adds a trace_commons host-runtime harness (network policy populated so
the onboard Network-effect obligation passes) and a parity test driving
all three capabilities through the scripted model loop: onboard with
confirmed=false exercises the deterministic consent gate with no
network, status and credits return the unenrolled/zero-credit defaults.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(traces): community profile second opt-in (token mint + profile set)

After device-key enrollment, public leaderboard attribution is a second,
separate opt-in: IronClaw mints a short-lived profile token from the
claim issuer with consent_scopes=[public_attribution] and empty
allowed_uses (such a claim cannot submit traces), then either prints it
for the web profile page or performs the profile update itself. The
browser cannot sign device-key requests, so the token must be minted by
IronClaw — previously this step was impossible and agent guidance
invented flows.

- ConsentScope::PublicAttribution mirrors the server protocol enum;
  default_allowed_uses_for_scope returns empty for it.
- mint_profile_attribution_token_for_scope / set_community_profile_for_scope /
  withdraw_community_profile_for_scope reuse the hardened issuer HTTP
  path (allowlist validation, pinned DNS, no redirects, bounded reads,
  token never in errors). PUT/DELETE /v1/community/profile per the
  server contract; handle (3-32 ASCII alnum/-/_) and bio (<=280 bytes)
  validated client-side.
- CLI: ironclaw-reborn traces profile token|set|withdraw.
- Onboard tool next_steps now describes the profile second opt-in so
  agent guidance stops inventing browser login flows.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* feat(traces): autonomous turn-end trace capture in the Reborn runtime

The Reborn binary could onboard, report status/credits, and manage
profiles, but never captured or submitted traces — the autonomous
pipeline existed only in the v1 agent loop. This wires it into the
Reborn runtime composition:

- TraceCaptureTurnEventSink subscribes best-effort to the turn
  lifecycle bus (the existing turn_event_sink injection seam). On
  Completed/Failed events with an explicit owner it spawns a detached
  task that reads the owner's standing policy (one file read for
  non-enrolled users), loads the recent thread history (last 24
  messages, 5 turns — v1 parity), adapts user/assistant text rows into
  the neutral ConversationMessage shape, redacts + scores locally, and
  queues + immediately flushes eligible envelopes. All failures are
  debug!-logged and never touch the turn lifecycle path.
- A periodic flush worker (300s, 25/scope — v1 parity) retries queued
  envelopes for the runtime owner plus every scope observed since
  boot, with CancellationToken shutdown alongside the other workers.
- TraceClientAutonomousCaptureRequest gains outcome_override so the
  lifecycle event's terminal status (authoritative in Reborn, where
  transcripts carry no structured outcome payload) marks failed turns
  as TaskSuccess::Failure; v1 passes None (no behavior change).
- Tool-result rows and credit-notice delivery are documented follow-ups
  (refs-only records; no composition-level outbound channel surface).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* test(traces): end-to-end auto-capture through send_user_message

Proves the full Reborn auto-submission chain with a real runtime: a
completed turn for an enrolled owner scope lands a redacted envelope in
that scope's submission queue with no manual trace command — turn
completion -> lifecycle bus -> capture sink -> thread-history read ->
redact/score -> eligibility -> queue (+ local-failing immediate flush
leaves the entry for the retry worker).

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* Expose Trace Commons profile token tool

* Expose Trace Commons profile set tool

* Allow Trace Commons profile setup from agent

* feat(webui-v2): Trace Commons credits card in WebChat v2 settings

Adds GET /api/webchat/v2/traces/credit and a read-only Trace Commons
settings tab to the v2 SPA, giving webui-v2-beta parity with the v1
console's credits card.

- Route follows the descriptor system end to end: bearer-auth required,
  NoBody, 120/60 per-caller read rate limit; descriptor-driven
  body/rate-limit enforcement applies automatically.
- RebornServicesApi::trace_credits derives the trace scope exclusively
  from the authenticated caller's user id (never from query/body) and
  reads contributor-local state via ironclaw_reborn_traces
  (policy + trace_credit_report), soft-falling back to an unenrolled
  zero-state on missing/unreadable local state, mirroring
  builtin.trace_commons.credits.
- SPA: Trace Commons subtab (enrollment, pending/final credit, delayed
  ledger delta, submission counts, last submission/sync, recent credit
  explanations) with the server-authoritative framing and a
  not-enrolled empty state pointing at agent onboarding.
- Tests: descriptor contract row, handler oneshot, and three composed-
  router serve tests (200 zero-state, 401 without bearer, enrolled
  policy reporting with per-test scope isolation).
- Drive-by: cfg-gate openai_user_id in webui_serve.rs to clear a
  pre-existing unused-variable warning under default features.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* Exempt Trace Commons profile setup from local-dev gate

* Route Trace Commons profile writes to ingest

* review(4559): address serrrfirat feedback

- Drop stray working-note markdown files from the repo root (they rode
  in via an early origin/main merge and are not this PR's documentation).
- trace_commons_dispatch_e2e: setup_base_dir is now a OnceLock that every
  test calls first — the previous 'single-threaded during init' claim was
  wrong under tokio's multi-threaded test runtime, and two of three tests
  skipped the setup entirely.
- settings.js: extract shared appendDisplayGroup + declarative row defs;
  loadTraceCommonsCredits drops from ~120 lines of manual DOM to a rows
  array; also removes a double-escape (textContent + escapeHtml) on
  explanation lines.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(webui-v2): add traceCommons i18n keys to all locales

The credits card added the traceCommons.* key set to en.js only; the
i18n consistency test (all_locales_share_the_en_key_set) requires
every locale to carry the same key set. Adds translated entries to
ar, de, es, fr, hi, ja, ko, pt-BR, uk, and zh-CN.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix(reborn): fail loud with source context on malformed local-dev master key

The local-dev secret store resolver read the cached key file (and the
SECRETS_MASTER_KEY env fallback) and passed the material straight into
SecretsCrypto::new several layers deep. A corrupt or low-entropy key
(e.g. a 64-char all-zeros value, which passes the length floor but has
one distinct byte) surfaced only as the opaque "Invalid master key",
with no pointer to the file the operator must fix.

- Add ironclaw_secrets::validate_master_key_material as the single
  source of truth for master-key rules; SecretsCrypto::new delegates
  to it.
- resolve_local_dev_secret_master_key now validates at the source
  (cached file vs SECRETS_MASTER_KEY env) and returns a
  RebornBuildError::InvalidConfig naming the offending path/env var and
  the actual constraint, before any crypto is constructed.
- A malformed env value is now rejected before being persisted to the
  cached key file (no more poisoned-cache state).

Tests: malformed-file path-context rejection, malformed-env
source-context rejection, valid cached file accepted.

Refs nearai#4741

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(webui-v2): add Trace Commons credits card to chat sidebar

Surface trace contribution credits at a glance in the chat sidebar,
above the conversation list. Previously credits were only visible under
Settings -> Trace Commons.

- New SidebarTraceCredits component reuses the existing useTraceCredits
  hook (/api/webchat/v2/traces/credit) — no new endpoint. Renders only
  when enrolled; loading/error/not-enrolled render nothing to keep the
  sidebar clean. Shows final credit and accepted/submitted counts and
  clicks through to Settings -> Trace Commons for the full ledger.
- useTraceCredits now refetches (60s interval + on window focus) so the
  card and the Settings tab reflect newly-accepted submissions live.
- Add one compact i18n key (traceCommons.cardAccepted) across all 11
  locales; reuse existing keys for the rest.
- Source-shape regression test in assets.rs.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(reborn-traces): reconstruct tool calls in turn-end trace capture

The Reborn capture adapter dropped every tool-result row, so captured
trace envelopes were text-only. That left the two highest-value scoring
levers — replayability (0.20) and tool coverage (0.15) — permanently at
zero, so even agentic tool-using turns scored as plain chat and stayed
below the 0.35 submission gate. Nothing ever submitted.

conversation_messages_from_records now reconstructs a `tool_calls`
message from each run of ToolResultReference rows that carry
`tool_result_provider_call` replay metadata, collapsing consecutive
rows into one message positioned between the user message and the
assistant response (the shape capture_turns_from_conversation_messages'
per-turn lookahead consumes). Tool names always flow through so the
value scorecard sees required_tools/replayable; raw tool payloads stay
consent-gated downstream by include_tool_payloads. Rows without provider
metadata remain dropped.

TDD:
- adapter unit tests: single tool call -> tool_calls message;
  consecutive calls collapse into one; ref without provider metadata
  still dropped.
- integration guard: a captured tool-using turn's queued envelope
  carries replay.required_tools + replayable=true.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reborn-traces): read capture history from context window, not display projection

Tool-call reconstruction (previous commit) had no data to work with: the
capture history source read SessionThreadService::list_thread_history,
whose product-display projection (history_message) hard-nulls
tool_result_provider_call. So even though tool calls persist with full
provider metadata, the adapter received None on every tool row, dropped
them, and produced a text-only envelope that scored below the 0.35
submission gate. Nothing ever submitted.

SessionThreadHistorySource now reads load_context_window (the
model-context/replay view, which preserves tool_result_provider_call)
and maps ContextMessage -> ThreadMessageRecord via context_window_to_records.
This is the semantically correct source for trace capture anyway: the
replay transcript, not the display transcript.

TDD: a caller-level test (per .claude/rules/testing.md "test through the
caller") drives SessionThreadHistorySource against a real
InMemorySessionThreadService with an appended tool result, asserting the
returned tool row keeps provider_call. Failed on list_thread_history
(None), passes on load_context_window.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(reborn-traces): auto-submit traces with PII risk below High

Previously any non-Low residual PII risk was blocked from auto-submission
two ways: the manual-approval eligibility gate held everything != Low, and
the value scorecard halved the score (privacy_gate Medium 0.5) and
subtracted a 0.60-weighted penalty. A minimal tool trace scores ~0.36 at
Low (barely over the 0.35 gate), so any Medium penalty collapsed it to 0 —
nothing below High could ever submit.

Treat below-High residual risk as clean for auto-submission (the
deterministic redactor has already scrubbed detected PII):

- trace_autonomous_eligibility manual-approval gate now holds only High
  (== High, was != Low).
- privacy_gate: Low|Medium => 1.0 (was Medium 0.5); High => 0.0.
- privacy_risk_score: Low|Medium => 0.0 (was Medium 0.5); High => 1.0.

High remains fully blocked: privacy_gate zeros its score and the gate holds
it for manual review. The 0.35 submission gate leaves no headroom for a
partial Medium discount on a minimal trace, so below-High is clean rather
than partially penalized.

TDD: medium_pii_tool_trace_auto_submits_while_high_is_held asserts a
Medium-risk tool trace clears 0.35 and auto-submits while High is held.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(reborn-traces): design for Trace Commons held-trace review

Held traces are currently dropped on the autonomous capture path with no
visibility or authorize path. This plan reuses the existing hold-sidecar
machinery (TraceQueueHold / .held.json / read_trace_queue_holds_for_scope /
ManualReview) and adds: retain held traces, surface a held count+list on
the /traces/credit response, a card/tab UI, and a promote-as-is authorize
endpoint. Four independently-shippable TDD slices.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(reborn-traces): retain manual-review held traces instead of dropping (slice 1)

Autonomous turn-end capture dropped every held trace (logged at debug,
envelope discarded), so PII-gated traces were unrecoverable and invisible.

Slice 1 of the held-review feature retains manual-review holds:

- TraceQueueEligibility::Hold now carries a typed TraceQueueHoldKind
  (ManualReview for the High residual-PII gate; PolicyGate for score /
  tool-allowlist / submission-class gates), replacing reason-string
  classification at the flush call site.
- TraceClientAutonomousCaptureOutcome::Held carries the built envelope and
  its kind so callers can persist it.
- New queue_trace_envelope_as_held_for_scope: queues the envelope plus a
  ManualReview .held.json sidecar under one scope lock; the flush worker
  already skips held sidecars, so it is retained but not submitted.
- capture_turn_trace retains ManualReview holds and still drops PolicyGate
  holds (low-value traces never pollute the review surface).

TDD: held-retain function (RED on missing sidecar -> GREEN), eligibility
kind classification, and caller-level capture tests (an AWS-key message
forces High PII -> retained ManualReview hold; a sub-threshold trace is
dropped, not retained).

Refs docs/plans/2026-06-10-trace-commons-held-review.md

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(webui-v2): surface manual-review held count + list on /traces/credit (slice 2)

Held traces retained by slice 1 were invisible to the UI. Slice 2 surfaces
them on the existing trace-credits response so one fetch powers the whole
card/tab.

- ironclaw_reborn_traces: manual_review_holds_for_scope() returns only
  ManualReview holds (excludes PolicyGate value-gates and transient
  RetryableSubmissionFailure retry holds), via an extracted
  retain_manual_review_holds filter.
- RebornTraceCreditsResponse gains manual_review_hold_count + holds[]
  ({ submission_id, reason }). Sanitized: submission id and the already
  privacy-safe hold reason only, never raw trace content.

TDD: retain_manual_review_holds filter unit test (excludes policy/retry),
disk-level manual_review_holds_for_scope test, and the facade zero-state
test asserts the new fields default empty. webui_v2 handler contract tests
(42) still pass with the propagated fields.

Refs docs/plans/2026-06-10-trace-commons-held-review.md

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(webui-v2): show held-for-review traces on card + Settings tab (slice 3)

Surface the manual-review held count/list from slice 2 in the UI. Both
render only when there are holds, so the common (nothing-held) state is
unchanged.

- Sidebar card: "{count} held for review" line when
  manual_review_hold_count > 0.
- Settings -> Trace Commons tab: a "Held for review" section listing each
  held trace's sanitized reason + submission id from holds[].
- No hook/api change: fetchTraceCredits already returns the raw response,
  so credits.holds / credits.manual_review_hold_count are available.
- Three i18n keys (cardHeld, heldTitle, heldDescription) across all 11
  locales.

The per-trace Authorize action ships with its endpoint in slice 4 (so the
UI never offers a button that 404s). Source-shape assertions extended.

Refs docs/plans/2026-06-10-trace-commons-held-review.md

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(webui-v2): authorize held traces for submission (slice 4)

Complete the held-review feature with a promote-as-is authorize action
across the stack.

ironclaw_reborn_traces:
- TraceContributionEnvelope gains `manual_review_authorized`; an authorized
  envelope submits past every gate in trace_autonomous_eligibility (the flush
  re-evaluates eligibility each pass, so removing the hold sidecar alone is
  not enough to promote).
- authorize_manual_review_hold_for_scope: stamps the envelope (durable
  consent record) BEFORE removing the .held.json sidecar, so a crash between
  the two leaves the trace held (fail closed). Only ManualReview holds are
  authorizable; unknown submissions return Ok(false), not an error.

ironclaw_product_workflow:
- RebornServicesApi::authorize_trace_hold derives scope from the
  authenticated caller (the path submission id is never cross-scope
  authority), validates the id, and returns RebornTraceHoldAuthorizeResponse.

ironclaw_webui_v2:
- POST /api/webchat/v2/traces/holds/{submission_id}/authorize — NoBody,
  mutation rate limit, bearer auth. Descriptor + handler + router + contract
  table (now 46 routes).

Frontend:
- authorizeTraceHold api, an authorize mutation in useTraceCredits that
  invalidates the credits query on success, and a per-hold Authorize button
  on the Settings tab. `authorize`/`authorizing` i18n in all 11 locales.

TDD: authorize promotes a High-PII held envelope past all gates; facade
zero-state; webui_v2 descriptor/handler contracts; composition serve (47);
source-shape assertions. clippy/fmt clean across crates.

Refs docs/plans/2026-06-10-trace-commons-held-review.md

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(traces): loopback dev claim exception + profile_set consent gate

Address the two codex P2 findings from review:

- Preserve loopback claim uploads after onboarding: the loopback-HTTP
  dev invite form stores a loopback claim/ingest endpoint in the
  policy, but the claim/ingest validators required https and rejected
  loopback hosts, so a successful loopback onboarding could never mint
  a claim or submit credits. The validators and the pinned DNS
  resolution now honor the same literal-loopback exception as invite
  parsing (shared is_loopback_host predicate); for loopback hosts the
  pinned resolution additionally requires all resolved addresses to be
  loopback. Non-loopback http, internal hostnames, and private ranges
  stay rejected, and the issuer allowlist still applies.

- Require explicit confirmation before community profile updates:
  trace_commons.profile_set now has the same hard confirmed=true input
  gate as onboarding — it short-circuits with consent_required before
  the enrollment check and any network write, since the capability is
  approval-gate-exempt in local-dev policy. Schema and manifest
  document the field.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(merge): thread attachments field through trace-capture record construction

main added ThreadMessageRecord.attachments (Vec<AttachmentRef>); the
trace-capture reconstruction path and its two test helpers construct
records and must set it. The capture path reconstructs records from a
context window for redaction/scoring and carries no attachment refs of
its own, so Vec::new() is correct.

* fix(traces): adapt v1 autonomous capture to new Held variant shape

The merge brought in slice 1 of the held-trace-review feature, which
changed TraceClientAutonomousCaptureOutcome::Held from
{ submission_id, reason } to { kind, reason, envelope } so manual-review
holds can be retained instead of dropped. The v1 autonomous-capture path
in thread_ops.rs still matched the old shape, breaking the
`--no-default-features --features libsql` build (and default build).

Adapt the v1 path to the new shape and give it the same retain-or-drop
parity as the Reborn capture path
(ironclaw_reborn_composition::trace_capture): ManualReview holds are
retained via queue_held_envelope_for_scope (the on-disk held queue is
shared, so a v1-captured hold surfaces in the v2 review UI); policy/value
gates are dropped as before, just logged.

Behavior mirrors the tested Reborn path
(send_user_message_auto_queues_trace_for_enrolled_scope); the v1
autonomous-capture path is a detached tokio::spawn with no unit-testable
seam, so no focused regression test is added.

[skip-regression-check]

* fix(traces): set manual_review_authorized in reborn-cli test envelope fixture

The merge brought in the held-trace-review manual_review_authorized
field on TraceContributionEnvelope. The reborn-cli trace_queue test
fixture constructs the envelope directly and missed the field, breaking
`cargo clippy --all-features --tests` and `Tests (all-features)` (the
fixture is test-only, so the libsql binary build did not surface it).
Fresh queued envelopes are not yet authorized, so false is correct.

[skip-regression-check]

* test(traces): pass confirmed=true in profile_set parity step

The trace_commons first-party-tools parity test invoked profile_set
without confirmed=true and asserted the NotEnrolled enrollment-gate
result. Commit 6bc776d added the public-attribution consent gate to
dispatch_profile_set, which now short-circuits to consent_required
before the enrollment check when confirmed is unset — so the test's
NotEnrolled assertion failed (the gate output carries no error_code).

Pass confirmed=true so the call clears the consent gate and reaches the
enrollment check, deterministically returning NotEnrolled with no
network (the scope never onboarded). Matches the unit-test pattern
established for the other profile_set tests in the same change.

[skip-regression-check]

* fix(traces): onboarding-security + contribution correctness (coderabbit batch 1)

Addresses 6 coderabbit findings in ironclaw_reborn_traces:

- device_key.rs: re-assert 0o700 on pre-existing key dirs (not just on
  create), so broader perms on an existing device_keys/ or pending/ can't
  leave invite/tenant hashes enumerable.
- device_key.rs: fail closed on load when on-disk public_key/device_key_id
  don't match the loaded private key (tampered/partial files no longer load
  an inconsistent identity that only fails later at remote auth).
- invite.rs: scope the staged pending-key filename by invite ORIGIN, not
  just code, so two issuers reusing one invite code can't share a device key
  (invite_hash stays code-only as the server allowlist subject).
- onboarding/mod.rs: reject ingest_url values with embedded userinfo before
  persisting, so a malicious onboarding response can't smuggle credentials
  into policy.json + outbound requests.
- contribution.rs: preserve mount path prefixes when deriving the
  community-profile endpoint (mirrors trace_submission_status_endpoint);
  a prefixed deployment no longer 404s on profile PUT/DELETE.
- contribution.rs: fail closed in trace_autonomous_eligibility on envelopes
  with no allowed-uses (public_attribution-only) instead of relying on the
  remote to bounce them.

Updated two retry tests that encoded the cross-issuer key-sharing bug now
fixed: they retried against a second mock on a different port; a new
spawn_flaky_mock_issuer keeps the retry on the same origin so it exercises
genuine same-issuer pending-key reuse. Added regression tests for each fix.

* fix(trace-commons): address coderabbit review findings on nearai#4559

- index.html: add type="button" to the Trace Commons settings subtab to
  prevent accidental form submission.
- settings.js + i18n/en.js: route the Trace Commons credits copy through
  I18n.t(...) and register the matching locale keys (matches the existing
  surface pattern; en-only like settings.traceCommons, fallback covers rest).
- factory.rs: drive the malformed SECRETS_MASTER_KEY env case through the
  real caller resolve_local_dev_secret_master_key (via an env-parameterized
  inner) and assert the rejected key is never persisted to the cached file.
- trace_commons_dispatch_e2e.rs: give each test a distinct user/extension
  scope so onboarding state can no longer bleed across tests.
- local_dev_capability_policy.toml: exempt builtin.trace_commons.onboard
  from the REPL approval gate (it has its own confirmed=true consent gate,
  mirroring profile_set).
- docs: fix the onboard prompt-file reference, match the held-trace JSON
  shape to RebornTraceHold (submission_id + reason only), and resolve the
  wire-protocol ownership split (types live locally in onboarding/protocol.rs,
  no shared trace-commons-protocol crate).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>

* fix(traces): tenant-scoping + token leak + read-failure + unbounded scopes (coderabbit batch 2)

Addresses the coupled backend findings:

- Tenant-scope Trace Commons local state across the Reborn paths: new
  trace_scope_key(tenant, user) helper keys policy / device-key / credit /
  profile / capture state by tenant+user, so the same user id in two tenants
  no longer shares state. Applied in host_runtime trace_commons dispatchers,
  product_workflow credits/hold, and composition trace-capture (v1 stays
  user-only — legacy single-tenant). Updated the affected runtime/sink tests
  and added a non-owner attribution assertion.

- Do not return the raw profile token from the model-visible profile_token
  capability: persist it to a 0600 <scope>/profile_token.jwt and return the
  file path + instructions instead, keeping the bearer credential off the LLM
  transcript.

- Stop masking genuine local-state read failures as zero/not-enrolled: the
  status capability and the WebUI credits path now propagate a read/parse
  failure (NotFound is already softened inside read_*_for_scope) so an
  enrolled user with a corrupt policy file is not told they have nothing.

- Bound ObservedTraceScopes: the periodic flush worker now prunes drained
  scopes (new trace_scope_has_pending_queue) after each tick, so the set is
  bounded by actual pending backlog instead of growing one entry per caller
  ever seen.

Note: a v1 caller-level test for the ManualReview hold-retention path is not
included — v1 ingress blocks secrets outright and the outbound leak detector
redacts them, so the High-residual-PII condition that produces a ManualReview
hold cannot be reproduced through process_user_input. The retention logic is
identical to and covered by the Reborn-side
capture_retains_manual_review_hold_for_high_pii_trace.

* test(traces): enroll under tenant-scoped key in webui_v2_serve credits test

trace_credits_reports_enrolled_for_caller_with_enabled_policy wrote the
policy under the bare user id, but the credits route now keys local state
by trace_scope_key(tenant, user). Enroll (and clean up) under the composite
TENANT/user scope so the route sees the enrollment.

* fix(factory): fail closed on explicit-but-unusable SECRETS_MASTER_KEY

An explicitly-set-but-unusable local-dev master key silently fell through
to generating + persisting a fresh key, leaving local-dev secrets
encrypted under an unintended master key the operator never chose:

- resolve_local_dev_secret_master_key used std::env::var(...).ok(), which
  drops VarError::NotUnicode -> treated as absent. Now only NotPresent is
  absent; a non-Unicode value returns InvalidConfig.
- resolve_local_dev_secret_master_key_with_env collapsed a set-but-empty
  (or whitespace-only) value to None via .filter(). Now a set-but-empty
  value returns InvalidConfig instead of generating a key.

Added resolve_local_dev_secret_master_key_rejects_set_but_empty_env_without_persisting
asserting empty/whitespace env values fail closed and persist nothing.
(coderabbit follow-up on nearai#3794)

* fix(factory): reject empty SECRETS_MASTER_KEY before the cached-file read

Follow-up to the prior fix: the empty-env rejection lived in the env
branch, which only runs when no cached key file exists. On a rebuild
where .reborn-local-dev-secrets-master-key already exists, the cached key
was returned first, so an explicitly-set-but-empty SECRETS_MASTER_KEY was
still silently ignored. Hoist the empty/whitespace rejection (and env
normalization) above the cached-file read so it fails closed regardless
of cached state. Added
resolve_local_dev_secret_master_key_rejects_empty_env_even_with_cached_file
asserting the empty env is rejected and the cached key is left unchanged.

* fix(traces): address 14:54 coderabbit re-review (tenant-seed, IO errors, effects, test)

Four outside-diff findings from the re-review:

- runtime.rs: seed ObservedTraceScopes with the runtime owner's
  trace_scope_key(tenant, owner) composite, not the bare owner id, so
  startup pending-queue discovery matches how capture keys state; the
  enrolled-scope test cleanup now removes the composite scope dir too.
- runtime.rs: the trace-queue polling test helper no longer swallows
  read_dir errors via unwrap_or_default() — only NotFound is the expected
  pre-capture fallback; any other IO error panics instead of masking as
  'no queued traces'.
- trace_commons.rs manifests + local_dev grants: onboard (device-key
  material) and profile_token (0600 token file) now declare
  Read/WriteFilesystem effects, and the local-dev grants allow them, so
  the effect model accurately models the local secret-material writes.
- local_dev_authorization test: added local_dev_trace_commons_onboard_skips_approval_gate
  (the onboard exemption was the actual fix; the profile_set-only test
  would pass even if the onboard TOML exemption were dropped).

* fix(factory): validate non-empty SECRETS_MASTER_KEY before the cached-file read

Follow-up: the prior fix rejected an *empty* env value before the cached
read but still validated a non-empty *malformed* value only after it.
So a valid cache + SECRETS_MASTER_KEY=0000... silently ignored the
explicit bad secret config on rebuilds. Move validate_resolved_master_key
into the up-front env normalization so any explicit-but-unusable env key
(empty OR malformed) fails closed regardless of cached state. Added
resolve_local_dev_secret_master_key_rejects_malformed_env_even_with_cached_file.

* fix(traces): address 15:41 coderabbit re-review (credits read-failure + 2 test guards)

- trace_commons.rs dispatch_credits: stop masking genuine records read/parse
  failures as 'no records' (NotFound is already softened inside
  read_local_trace_records_for_scope); report RecordsReadFailed, mirroring
  dispatch_status.
- runtime.rs trace-queue polling helper: fail loud on per-ENTRY read_dir IO
  errors too (map + unwrap_or_else panic) instead of filter_map(e.ok()), so a
  broken entry can't be silently dropped while claiming the queue holds one.
- local_dev_authorization approval-gate test: assert the effects DO require
  approval without the exemption (local_dev_effects_require_approval), so the
  test can't pass via a non-gating default policy if the TOML exemption were
  dropped.

* fix(traces): address Henri review — backend findings (atomic token, error mapping, validation, egress test)

- persist_profile_token now writes atomically (unique 0600 temp + fsync +
  rename) so a reader never observes a half-written or overwritten bearer
  credential under overlapping mints (Henri perf/security Medium).
- dispatch_onboard error mapping: OnboardError::DeviceKey is reported as a
  distinct DeviceKeyError (re-run onboarding) instead of being collapsed into
  PersistError's check-disk-and-permissions guidance (Henri bugs Medium).
- parse_profile_set_input enforces the manifest's declared schema at parse
  time: handle 3-32 ASCII letters/digits/-/_, bio <= 280 bytes (Henri
  conventions Medium). Added schema-limit test.
- Added dispatch_onboard_confirmed_without_host_egress_is_network_denied
  covering the NetworkDenied host-egress-miswiring branch (Henri tests Medium).

* fix(traces): address Henri review — frontend findings (enrolled empty-state + polling)

- v1 credits: TraceCreditResponse now carries `enrolled` (read from the
  standing policy), and settings.js keys the opt-in empty state on
  `!data.enrolled` instead of `!submissions_total` — an enrolled user with
  zero submissions now sees their zero-credit view, not the not-enrolled
  prompt (Henri bugs Medium).
- useTraceCredits: each fetch rebuilds the full server-side credit view, so
  the aggressive 60s poll made an open tab steady O(history) work. Relaxed to
  a 5-min interval + staleTime + no background polling, keeping a focus
  refetch for liveness; mutation invalidation still updates promptly. Added a
  TODO to incrementalize the server-side view (Henri perf Medium).

* perf(traces): memoize server-side credit view by on-disk input signature

Bounds the trace-credits polling cost to O(new submissions) instead of
O(total history). New scoped_credit_view(scope) caches the computed credit
report + manual-review holds keyed by a cheap change signature (submissions
file mtime+len, plus a hash of the held-trace sidecars). On the steady-state
polling case (unchanged history) a request is a couple of stat()s + a clone
rather than reading/parsing the full submissions file and re-aggregating.
On any change the signature differs and it recomputes once. Cache is bounded
(4096 scopes, cleared on overflow).

Wired through the polled WebUI path (local_trace_credits_for_user) and the
model-visible credits capability (dispatch_credits). Added
scoped_credit_view_reflects_record_changes_via_signature covering the
cache-hit path and signature-based invalidation on record changes.

Completes the TODO from the Henri perf-review follow-up (#5).

* fix(traces): gate profile_set behind runtime approval (Henri #1 High)

profile_set publishes a public community profile (an external write to a
public surface). Its `confirmed=true` input is model-controlled, so a
prompt-injected or confused model could supply it. Make the runtime
approval gate the primary, user-controlled consent control:

- Drop `builtin.trace_commons.profile_set` from the local-dev
  approval-gate exemption list (keep `onboard`, which runs its own
  in-turn confirmed=true consent before the network POST).
- Set profile_set's manifest default_permission to Ask (was Allow).
- Split the local-dev authorization test into
  `local_dev_trace_commons_profile_set_requires_approval_gate` (asserts
  Decision::RequireApproval) and
  `local_dev_trace_commons_onboard_skips_approval_gate` (asserts
  Decision::Allow), via a shared `trace_commons_authorize_decision`
  helper that first asserts the effects would gate without an exemption.

Also fix a pre-existing trace_commons harness gap: onboard + profile_token
gained a WriteFilesystem effect (device-key persistence) but the
`trace_commons_tools` harness allow-set was never updated, so those
capabilities were filtered out of the model-visible surface and the
parity/visibility tests failed with driver_unavailable. Grant
WriteFilesystem in the harness allow-set.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(traces): extract onboarding test harness to sibling file (Henri nearai#8)

The onboarding module's ~840-line `#[cfg(test)] mod tests` block (mock
issuer harness, retry/idempotency coverage, URL-validation tests) made
`onboarding/mod.rs` a 1319-line file dominated by test scaffolding. Move
the module body into `onboarding/tests.rs` declared `#[cfg(test)] mod
tests;`, leaving mod.rs focused on production logic (now 480 lines). No
test behavior changes; `use super::*;` still resolves to the onboarding
module.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(webui-v2): update embedded-asset assertion for incrementalized credits poll

The Henri #5 polling fix changed useTraceCredits.js from refetchInterval
60_000 to 300_000 (plus refetchIntervalInBackground: false and
staleTime: 60_000), but the embedded-asset test in assets.rs still
asserted the old 60_000 value and failed in CI. Update the assertion to
lock the new infrequent-poll + paused-while-hidden + focus-refetch shape.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(traces): address CodeRabbit review + stale capability-policy test

CodeRabbit findings on the gating/refactor commits:
- Major: format_profile_token returned the absolute host path of the token
  file (token_file) on the model-visible surface, which violates the
  "never expose absolute paths" guideline. Replace with an opaque
  token_delivery marker; the token is still persisted 0600 for out-of-band
  retrieval by a bearer-auth UI/CLI. Update the message + test accordingly.
- Major (fail-loud): profile_token_error_value and profile_set_error_value
  collapsed "could not read policy" into NotEnrolled, sending enrolled
  users back through onboarding on unreadable/corrupt state. Split into a
  distinct PolicyReadFailed result in both formatters (matches dispatch_status).
- Minor: stale comment claiming profile_set is approval-gate-exempt (it is
  now PermissionMode::Ask and NOT exempt) — corrected.
- Minor: inaccurate harness comments (profile_token writes profile_token.jwt
  not device-key material; yolo auto-approves all Trace Commons Ask-gated
  tools, not just onboard) — corrected.

Also fix bundled_local_dev_capability_policy_parses, which still asserted the
pre-gating policy shape: profile_set as exempt (now onboard exempt /
profile_set NOT exempt), onboard's grant missing the read/write filesystem
effects, and profile_token/profile_set sharing one effect-set assertion even
though profile_token now carries WriteFilesystem and profile_set does not.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* style(traces): collapse single-line use block after Path import removal

rustfmt collapses `use std::{panic, path::PathBuf, sync::Arc}` to one line
once Path was dropped; the prior commit skipped re-running fmt after that
edit, reddening the Formatting CI check.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(traces): consent-gate profile_token + drop fixed-origin profile URL (CodeRabbit)

Two Major CodeRabbit security findings on the profile tools:

- profile_token minted and persisted a bearer credential with no in-turn
  consent gate. PermissionMode::Ask can be auto-approved under local-yolo, so
  a model call could mint a credential without explicit per-conversation
  consent. Add a hard confirmed=true gate (schema + parse + consent_required
  short-circuit) before minting, mirroring dispatch_onboard / dispatch_profile_set.
- format_profile_token and profile_set_success_value hardcoded
  https://tracecommons.ai/profile. The token is scoped to the user's ENROLLED
  issuer (which may be self-hosted or loopback), so steering the user to paste
  a bearer profile-management token at a fixed origin could leak it to the
  wrong host. Drop the fixed profile_url; route through the enrolled profile
  flow / local UI/CLI out of band.

Tests: new dispatch_profile_token_without_confirmed_returns_consent_required_no_mint;
existing without-enrollment test now passes confirmed=true; profile_set success
test asserts no fixed origin; parity step mints with confirmed=true.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(traces): route agent-invoked profile writes through host egress (CodeRabbit #3)

profile_token (upload-claim mint) and profile_set (community-profile PUT/DELETE)
previously made network writes via the ironclaw_reborn_traces crate-local reqwest
client, bypassing the host RuntimeHttpEgress pipeline (private-IP filtering,
redaction, byte accounting) that onboard already uses.

Add a `ContributionHttpSink` port (mirroring `OnboardingHttpSink`): when a sink
is injected, the mint POST and the profile PUT/DELETE run through host egress;
when `None`, the existing hardened crate-local client is used unchanged.
host_runtime supplies `HostEgressContributionSink` (wraps RuntimeHttpEgress,
sanitizes errors via stable_runtime_reason, never leaks URL/token), and
dispatch_profile_token / dispatch_profile_set fail closed with NetworkDenied if
egress is absent (after the enrollment pre-check, so a not-enrolled user still
gets NotEnrolled guidance).

The background trace-upload / status-sync worker and the CLI keep the crate-local
client (pass `None`): that lane is a durable, model-input-free internal task that
sends only already-redacted envelopes to the operator-enrolled endpoint and does
its own SSRF/private-IP validation, so host egress adds complexity without
security benefit. Justification recorded in a comment on `trace_remote_http_client`.

New public surface: ContributionHttpSink/Request/Response/Error/Method,
mint_profile_attribution_token_for_scope_via_sink,
set_community_profile_for_scope_via_sink. Existing public fns keep their
signatures (None path) so CLI/worker/tests are unchanged.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
* fix(reborn): steer routine delivery through outbound targets

* fix(reborn): address outbound delivery review feedback (nearai#4780)

* fix(reborn): fail loud on outbound target registration (nearai#4780)

* fix(slack): harden outbound delivery rendering

* fix(slack): address outbound delivery review feedback

* fix(slack): allow idempotent host remounts
ilblackdragon and others added 28 commits June 15, 2026 16:38
…nearai#4644) (nearai#4871)

* feat(threads): carry image attachment refs on ContextMessage (image-vision step 1)

Foundation for sending image attachments to vision-capable models. The
transcript already holds AttachmentRef { kind: Image, mime_type, storage_key };
this surfaces the image ones structurally on the model-visible ContextMessage
(new ContextImageAttachment { mime_type, storage_key }) instead of only the
textual <attachments> pointer. Cheap metadata — no bytes are read here; the turn
layer will read+encode them later, and only for a vision-capable model. The text
pointer remains the fallback for text-only models. Populated at both context-read
paths (window + by-id) in the in-memory and filesystem stores.

* feat(reborn): translate image attachments into multimodal parts (image-vision step 2)

The model gateway now renders image attachments as real multimodal content for
the provider. HostManagedModelMessage carries encoded image parts
(HostManagedModelImagePart { mime_type, data_base64 }); convert_messages turns a
User message's parts into `ContentPart::ImageUrl` base64 `data:` URLs via
`ChatMessage::user_with_parts` (text rides in `content`, images in
`content_parts` — the provider adapters prepend the text). Text-only user
messages are unchanged. The parts are still empty until step 3 reads+encodes
bytes for vision models; this is the gateway-side translation, unit-tested for
both the image and text-only paths.

* feat(reborn): vision gate + attachment read port (image-vision step 3a)

The model gateway now only attaches image parts for a vision-capable model
(`is_vision_model(provider_model_id)`); a text-only model keeps just the text
(the transcript's `<attachments>` pointer still serves it). And loop_support
gains a `LoopAttachmentReadPort` (read bytes by scope + storage_key) plus a
`with_attachment_read_port` builder: when wired, `resolve_model_messages` reads
each model-visible message's image attachments through the port and
base64-encodes them into `HostManagedModelImagePart`s. Until the port is injected
(step 3b) it stays None, so behavior is unchanged. Read failures are logged and
skipped, never failing the turn. Unit tests cover the gate (vision vs non-vision).

* feat(reborn): wire the attachment read port end-to-end (image-vision step 3b)

Inject the concrete attachment reader so the multimodal path is now live for the
WebChat v2 loop. `ProjectScopedAttachmentReader` reads landed bytes back through
the project workspace filesystem (bounded, re-scoped through the MountView
authority — never a host path) and implements `LoopAttachmentReadPort`. It is
threaded factory → gateway struct → model port (mirroring the
`skill_context_source` plumbing), and `build_reborn_runtime` constructs it from
the local runtime's workspace filesystem and sets it on
`DefaultPlannedRuntimeParts`. When no local runtime is composed it stays `None`
(images remain the textual pointer). With this, an image attached to a
vision-capable model is read, base64-encoded, and sent as a `ContentPart::ImageUrl`.

* docs(threads): image multimodal path is implemented (image-vision step 5)

The attachment_context doc no longer describes the vision path as a separate
future path — it now points at the implemented multimodal flow (model port
reads bytes back, gateway sends ContentPart::ImageUrl) and frames the textual
pointer as the text-only-model fallback.

* test(4644): cover image-vision producer read path + fix reader error taxonomy

Add producer-side integration coverage for the image-vision feature, which
previously had only consumer-side unit tests (convert_messages):

- thread_loop_support_contract: drive the model port with a landed image
  attachment + a stub LoopAttachmentReadPort, asserting the resolved
  HostManagedModelMessage carries the base64-encoded bytes as an image part
  (test-through-the-caller: the read port gates a side effect with the model
  port wrapper between).
- attachment_landing: round-trip (land -> read back), not-found, and oversized
  reader tests.

The oversized test surfaced a bug: read_bytes_bounded returns Ok(None) only
for oversized files (a missing file is Err(NotFound)), so the reader mislabeled
oversized attachments as NotFound and missing ones as a generic Backend error.
Fix the mapping: Ok(None) -> Backend("exceeds limit"), Err(NotFound) ->
NotFound, Err(PermissionDenied) -> Forbidden.

* test(4644): set attachment_read_port in reborn parity harness

The image-vision field added to DefaultPlannedRuntimeParts was not threaded
into the root reborn parity test harness, breaking compilation of every
reborn_*_parity integration target (E0063) and the --all-targets clippy/test
jobs. Set it to None (these parity tests don't exercise the image path).

* fix(attachments): correct vision gate + tidy image-vision read path (review)

Address code-quality review and bot comments on the image-attachment path:

- Vision gate (blocker): is_vision_model missed the current tier-first Claude
  ids (`claude-opus-4-8`, `claude-sonnet-4-6`, Bedrock `anthropic.claude-*`),
  so every current-gen Claude model was mis-classified text-only and silently
  dropped image attachments. Added tier-first patterns + a regression test over
  the real production ids.

- Layering: HostManagedModelImagePart now carries raw bytes, not base64; the
  gateway owns base64/`data:` URL formatting (image_data_url). Drops the base64
  dependency from the neutral ironclaw_loop_support crate. Renamed the producer
  helper encode_image_parts -> read_image_parts to match.

- Documented why the read is not gated on vision capability at the producer:
  the authoritative model id is model_override (resolved in the gateway from its
  routing policy) and can diverge from the run-context route snapshot the port
  holds, so a producer-side gate would risk silently dropping images. The single
  authoritative gate stays in convert_messages; the silent skip is annotated.

- Decomposition: collapsed the 4-site ContextMessage projection (2 stores x 2
  paths) into ContextMessage::from_transcript_message so attachment projection
  lives once and the two stores cannot drift.

- LoopAttachmentReadError now impls std::error::Error.

- Documented attachment_read_port optionality as deliberate (no workspace fs ->
  nothing to read -> degrade to text pointer), not a fail-closed gap.

- Added unit tests for model_image_attachments.

* feat(attachments): vision support across providers + WebUI v2 image thumbnails

Two follow-on gaps from the image-attachment work:

1) Image reading now works on every vision-capable provider. The gateway
   already emitted ContentPart::ImageUrl for vision models, but three adapters
   silently dropped it:
   - anthropic_oauth: emits Anthropic `image` blocks (base64 source).
   - gemini_oauth: emits Gemini `inlineData` parts.
   - bedrock: emits Converse `ImageBlock` (supported formats only; others are
     skipped, keeping the text).
   Added a shared `ImageUrl::decode_data_url()` so every adapter parses the
   gateway's `data:` URL the same way. Each adapter has a unit test. (OpenAI
   Codex Responses API still has no image support upstream — unchanged.)

2) WebUI v2 renders thumbnails for persisted images. The timeline carries only
   attachment refs (no bytes), so a reloaded image showed a file card, not a
   thumbnail. Added:
   - GET /api/webchat/v2/threads/{thread_id}/messages/{message_id}/attachments/
     {attachment_id} — serves landed bytes, scope derived from the
     authenticated caller, storage path resolved server-side, authoritative
     Content-Type + nosniff + short private cache. Keyed by (thread, message,
     attachment) because an attachment id is only unique within its message.
   - RebornServicesApi::read_attachment (default returns NotFound; only
     RebornServices overrides it, so the 10 stub impls are untouched) backed by
     a new InboundAttachmentReader port — the read counterpart of
     InboundAttachmentLander — implemented over the same project workspace mount
     and wired in build_webui_services.
   - Frontend: `<img>` can't send a bearer, so the bubble lazily fetches the
     bytes (authenticated) into a blob URL and revokes it on unmount; the
     timeline projection attaches a `fetch_url` to landed images.
   Contract tests cover the descriptor, the byte response through the real
   router, and the projection URL.

* refactor(reborn): fix trigger-thread attachment scope + dedupe history resolution

Self-review of the attachment-bytes path found a latent bug and a duplication:

- Bug: `read_attachment` resolved the thread via the automation-trigger fallback
  (so a trigger-fired thread is found) but then read the bytes under the
  *caller's* session scope. Trigger threads live under the creator's scope, so
  the reader addressed the wrong project mount and would 404 for trigger-thread
  images — exactly the case the fallback exists to support.

- Duplication: the scope-derivation + history-load + trigger-fallback block was
  copy-pasted between `get_timeline` and `read_attachment` (~25 lines of
  security-sensitive logic that could drift).

Fix: `try_automation_trigger_timeline_fallback` now returns the resolved
`ThreadScope` it already computes alongside the history, and a single
`resolve_thread_history_for_caller` helper owns the primary-load + fallback.
`get_timeline` ignores the scope; `read_attachment` reads bytes under it, so a
trigger-thread thumbnail loads under the creator's mount.

Regression test: `read_attachment_reads_trigger_thread_bytes_under_creator_scope`
asserts the byte read is issued under the trigger creator's scope, not the
caller's session scope (fails on the pre-fix code).

* fix(webui-v2): render attachment thumbnails as data URLs, not blob URLs

The persisted-image thumbnail fetched bytes into a `blob:` object URL, but the
SPA's CSP is `img-src 'self' data:` — so `<img src=blob:…>` was refused
("violates ... img-src 'self' data:"). The optimistic compose-time preview
already uses a `data:` URL (FileReader.readAsDataURL), which is why it rendered
and the persisted one didn't.

Match that convention instead of widening the CSP: `fetchAttachmentDataUrl`
reads the fetched blob into a `data:` URL. This is CSP-compliant and also drops
the object-URL revoke lifecycle (data URLs need none). Same-origin authenticated
fetch is unchanged (allowed by `connect-src 'self'`).

Regression test lives in api.test.mjs (a browser-JS bug, not Rust): it stubs
`URL.createObjectURL` to throw and asserts `fetchAttachmentDataUrl` returns a
`data:` URL — reverting to a blob URL fails the test. [skip-regression-check]
(the hook only scans for Rust #[test]; the test is JS).

* feat(webui-v2): click-to-preview modal for all attachment kinds

Clicking an attachment chip now opens a focused preview modal. Each kind renders
in a CSP-allowed way (classifier `attachmentPreviewMode`):
- image  → inline <img> (data URL; img-src 'self' data:)
- audio/video → inline player (data URL; media-src 'self' data:)
- pdf    → inline <iframe> (blob URL; frame-src 'self' blob:)
- text/JSON/CSV/XML → fetched text in a <pre> (capped, with a truncation note)
- other binary → metadata panel; a Download action is offered in every mode.

CSP: the SPA document shell gains `media-src 'self' data:` and
`frame-src 'self' blob:` (deliberately narrow — locked by the
spa_document_csp_allowlist_is_locked test so they can't widen to */data: frames;
object-src stays 'none').

Plumbing:
- api.js: `fetchAttachmentBlob` is the shared auth-fetch primitive; `blobToDataUrl`
  and the existing `fetchAttachmentDataUrl` build on it. The modal fetches bytes
  once and derives the per-mode representation, revoking the object URL on close.
- history-messages: every landed attachment (not just images) now carries a
  `fetch_url`, so any kind can be previewed/downloaded.
- message-bubble: the chip becomes a button (when it has bytes to show) that
  opens `AttachmentPreviewModal`; non-previewable optimistic rows stay static.

Tests: `attachmentPreviewMode` mode mapping; landed non-image gets a `fetch_url`;
CSP lock-test extended for the new directives.

* docs(threads): correct ContextImageAttachment doc — vision gate is gateway-side

Review (Copilot): the doc claimed the turn layer reads attachment bytes "only
when the active model can actually accept images", but the read is unconditional
when a read port is wired — the vision/non-vision gate lives in the model
gateway, not at read time. Reword to describe the gateway-side gate so the
contract doesn't over-promise. (Companion lib.rs docs were already corrected.)

* fix(attachments): address PR re-review (thumbnail gating, fail-loud, hardening)

Round-2 bot review of the pushed branch:

- message-bubble: gate the thumbnail fetch+render to image kind. `fetch_url` was
  broadened to every landed attachment (for click-to-preview), which made the
  thumbnail fetch a PDF/text and render it as a broken <img>. Non-images keep
  the file icon. (Copilot, CodeRabbit)
- bedrock: normalize MIME to lowercase before format matching; replace the
  `.ok()?` drops in `bedrock_image_block` with logged `// silent-ok` branches so
  a malformed/unsupported inline image is diagnosable, not silently gone.
  (CodeRabbit)
- anthropic_oauth: restrict tool-result coalescing to user blocks that are
  *only* tool_result blocks, so a tool result can't be folded into a multimodal
  (text+image) user prompt. (CodeRabbit)
- reborn_services `read_attachment`: resolve the attachment ref first; if it
  landed (has a storage_key) but no reader is wired, return a sanitized 503
  (composition fault) instead of a 404 that makes real bytes look absent.
  (CodeRabbit — fail loud)
- webui_v2 descriptor: the attachment-bytes route reads workspace-backed bytes,
  so classify it `ProductWorkflow`, not `ProjectionOnly` (fail-closed ingress).
  (CodeRabbit)
- api.js `fetchAttachmentBlob`: reject off-origin URLs before attaching the
  bearer (token sink). (CodeRabbit)
- CSP lock test: assert exact per-directive source lists (media-src/frame-src/
  img-src) instead of substrings, so a widened directive fails. (CodeRabbit)
- tests: trigger-thread regression now records + asserts the storage_key passed
  to the reader (scope + key); reworded a misleading history-messages test
  comment. (CodeRabbit, Copilot)

Skipped: typing `storage_key` as an `AttachmentStorageKey` newtype at the read
port — `storage_key` is `String` across the whole attachment subsystem
(`AttachmentRef`, lander, loop port), and the reader already re-scopes it through
the `MountView`/`ScopedPath` authority so it can't escape the project scope;
a newtype just here would be inconsistent and is a separate cross-cutting
refactor. PR-description "no ironclaw_llm changes" note: will update the PR body.

* chore: remove accidentally-committed runtime attachment + gitignore the dir

`attachments/2026-06-15/26cfd2da-…-sharing-signed-export.png` is a runtime
artifact: the WebChat v2 attachment lander wrote an uploaded image under the
project workspace while `serve` ran from the repo root, and a `git add -A` in
45e4d67 swept it in. It is not source. Remove it and add `/attachments/` to
.gitignore so local-dev uploads can't be committed again.

* fix(webui-v2): fail-fast attachmentUrl + test hygiene; document read_attachment cost

Round-3 review (Copilot, CodeRabbit):

- `attachmentUrl` now throws if threadId/messageId/attachmentId is missing,
  rather than building a `.../undefined/...` path that would later carry the
  bearer via `fetchAttachmentBlob`. history-messages guards all three parts so a
  malformed record yields a plain card (no fetch) instead of throwing mid-
  projection. (Copilot, api.js)
- api.test.mjs: save/restore `URL.createObjectURL` (try/finally) instead of
  deleting it, so the stub can't leak global state across tests. Added a
  `attachmentUrl fails fast` test. (CodeRabbit, outside-diff)
- read_attachment: documented why it loads full thread history (O(messages)) —
  the cost equals the timeline load already incurred when the thread is open and
  each attachment is browser-cached, and a single-message fast path would need a
  new scope-validated "load one message record by id" service method
  (`load_context_messages` only projects image refs). Left as a follow-up rather
  than widening the thread-service contract. (Copilot)

JS regression test in api.test.mjs (browser-JS behavior, not Rust). [skip-regression-check]
…tale pin (nearai#4947)

The /benchmark pre-dispatch suite check looked up
`suites/<name>.toml` at a hardcoded BENCH_PIN SHA, but the dispatched
`bench` job runs `bench-pr-reusable.yml@main` (and is documented to
intentionally track benchmarks main). The pin had drifted, so suites that
exist on benchmarks main were rejected as "Unknown suite" even though the
run would have found them — e.g. `pinchbench26`, `officeqa`,
`terminal-bench-2`.

Point BENCH_PIN at `main` so the validation ref matches the ref we
actually run against; they now stay in lock-step and new suites work the
moment they land on benchmarks main.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…#4950)

* fix(deps): bump wasmtime 44.0.2 -> 44.0.3 to clear RUSTSEC-2026-0182

cargo-deny advisories has been failing repo-wide (incl. main) on
RUSTSEC-2026-0182 — "Leak in WASIp1 fd_renumber implementation"
(wasmtime). The advisory's fix line for the 44.x series is >=44.0.3,
<45.0.0, so this is a patch bump within the existing `^44.0.2`
constraint — no Cargo.toml change, Cargo.lock only. (45.x is available
but requires Rust 1.93.)

Pulls along the matching cranelift 0.131.2 -> 0.131.3 and pulley
44.0.2 -> 44.0.3 patches. `cargo deny check advisories` now reports
`advisories ok`; workspace compiles.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(deps): raise wasmtime manifest floor to 44.0.3 to match lockfile

Review follow-up (gemini): Cargo.lock was bumped to 44.0.3 for
RUSTSEC-2026-0182 but the manifests still declared "44.0.2" with a stale
security-floor comment, leaving the lock/manifest inconsistent and an
accidental downgrade to the vulnerable 44.0.2 possible.

Bump the wasmtime/wasmtime-wasi version floor to 44.0.3 across the root
and all WASM crates, and update the floor comment to document
RUSTSEC-2026-0182 (WASIp1 fd_renumber leak) while retaining the
RUSTSEC-2026-0149 history. Cargo.lock already resolves to 44.0.3, so the
floor bump leaves it unchanged. cargo check --workspace clean;
cargo deny check advisories -> advisories ok.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* chore(deps): add component-model feature to wasmtime in ironclaw_hooks

Review follow-up (coderabbit): ironclaw_hooks declared wasmtime without
the component-model feature that all peer WASM-consuming crates carry.
Align it for workspace feature consistency. Hooks uses only core wasmtime
types (Engine/Module/Instance/Store/Linker/Caller/Config), so this is a
consistency change — Cargo.lock is unchanged (the workspace already
resolves wasmtime with component-model via the peers).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…oop (nearai#4944)

* fix(reborn): surface auth-gate denial to model instead of re-prompt loop

Denying an auth (OAuth/credential) gate in Reborn looped forever: the
Deny decision called TurnCoordinator::cancel_run, ending the run as
terminal Cancelled. The model was never told, the blocking capability
still had no credential, so the work re-submitted and re-blocked on the
same gate — an infinite auth/deny loop.

Deny now RESUMES the parked run (precondition BlockedAuthGate) carrying
a typed disposition instead of cancelling it. The agent loop turns that
into a model-visible capability failure (Authorization, retry Forbidden)
so the model sees "user declined to provide credentials" and continues,
rather than silently re-prompting.

Mechanism (no DB migration — TurnRunRecord persists as a JSON blob):

  AuthInteractionService::resume_denied_auth
    cancels the OAuth flow, then resume_turn(BlockedAuthGate) with
    ResumeTurnRequest.auth_resume_disposition = Some(Denied)
  -> TurnRunRecord / TurnRunState (serde-default field)
  -> LoopRunContext (set in reborn create_host)
  -> planned_driver::resume stamps PendingAuthResume.disposition
  -> CapabilityStage short-circuits the parked auth-resume into a
     model-visible Authorization failure instead of re-dispatching
     (which would re-check the still-missing credential and re-block).

AuthResumeDisposition is a typed enum (Denied { reason } today, room for
Deferred/etc.) so future model-surfacing expands without re-plumbing the
carrier through every layer.

Notes:
- Disposition rides PendingAuthResume so it is consumed once and cannot
  misfire on a different capability's genuine auth need in the same run.
- resume_turn_once overwrites the field each resume; normal credential
  resumes pass None, self-clearing any stale disposition.
- validate_loop_safe_summary bans the substring "authorization:", so the
  short-circuit calls handle_capability_error directly with a clean
  planner summary rather than routing through the Failed-arm summary
  builder.
- Approval-deny (already has durable ApprovalStatus::Denied),
  background-mission Cancelled (no inline wiring), and gate timeout are
  unchanged / out of scope.

Tests: turns serde backward-compat + resume persist/clear; agent_loop
denied-resume continues with cleared slot + Authorization observation;
product_workflow deny-on-parked resumes (not cancels); reborn driver
stamps the disposition and does not re-block.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reborn): partition denied auth-resume by capability; tighten serde test

Address PR review:

- capabilities.rs: the denied auth-resume short-circuit failed EVERY
  visible call. Partition visible_calls by the denied
  pending_auth_resume.capability_id — synthesize the non-retryable
  Authorization failure only for the matching call, clear
  pending_auth_resume, and let the remaining (unrelated) calls dispatch
  normally. Prevents denying one tool from failing unrelated parallel
  calls. Adds multi-call regression coverage.

- ironclaw_turns: rewrite the backward-compat test to assert the
  ResumeTurnRequest struct-level serde contract (missing-field
  deserialize -> None; None -> key absent; Some(Denied) round-trips)
  instead of serializing the Option in isolation.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reborn): make retried auth-deny idempotent (never cancel a resumed run)

PR review (P2): after deny-on-parked resumes the run, a duplicate Deny
(double-click / lost response / retry) found the flow already Canceled
and the run no longer parked, hit the (NotParkedOnGate, Deny) arm, and
called cancel_run — terminating the just-resumed, now-running run.

A Deny only resolves a run parked on this auth gate; NotParkedOnGate
means it was already resolved, so a retry must be an idempotent replay,
never a fresh cancel. Rewrite that arm: when the flow is already
Canceled, read current run state and replay — DenialResumed for a
non-terminal (resumed/running/completed) run, Canceled only if the run
is genuinely terminal — without issuing a new cancel_run. Keep the
StaleAuth guard for a non-Canceled, not-parked flow. Remove the now-dead
cancel_auth helper. Replace the test that asserted the buggy
cancel-on-retry behavior with two idempotency tests (resumed run -> no
cancel; terminal run -> no new cancel).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(reborn): code-review follow-ups — coverage, docs, comments

Straightforward items from multi-agent review (no behavior changes):

- agent_loop: backward-compat serde test now strips PendingAuthResume
  .disposition; document that disposition=Denied skips re-dispatch and
  why the field name is short; comment the unconditional pending_auth_
  resume clear (also consumes disposition when no call matched) and the
  intentionally-empty safe_summary; simplify denied_capability_id
  extraction; add 1-denied + 2-remaining partition test.
- turns: backward-compat test for TurnRunState.auth_resume_disposition
  missing key -> None; turn-coordinator test that resume_turn persists
  the disposition through store -> claim and that a normal resume clears
  it.
- product_workflow: facade test covering DenialResumed -> Resumed;
  deny-on-completed-flow -> StaleAuth test; comment the idempotent-replay
  actor: None.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(reborn): drop unused reason field from AuthResumeDisposition::Denied

Review follow-up (D3): the Denied variant carried a reserved, always-None
`reason: Option<String>`. A raw String reaching the model would bypass the
loop-safe summary validator. Make Denied a unit variant; re-add a properly
typed/validated field if a concrete need arises. Enum stays extensible via
new variants. Wire shape for Denied is now "denied" (brand-new on this
branch, never persisted, so no historical data to preserve).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reborn): guard idempotent deny replay; document deny-after-complete + batch accounting

Review follow-ups:

- D1: the (NotParkedOnGate, Deny) idempotent replay returned DenialResumed
  whenever the flow was Canceled, but a Canceled flow does not prove THIS
  run was deny-resumed by us (it can be canceled by other paths). Only
  replay DenialResumed when the run carries our deny marker
  (TurnRunState.auth_resume_disposition.is_some()); otherwise return
  StaleAuth instead of fabricating a success. Adds a test for the
  canceled-by-another-path case.

- D5: document (no behavior change) that a Deny arriving after the OAuth
  flow already Completed is rejected as StaleAuth and the run proceeds
  with the obtained credential; honoring the late deny is a deliberate
  follow-up non-goal.

- D2: investigated the denied short-circuit's use of a default
  CapabilityBatchTurnSummary. Not a bug — record_result only fires when
  invocation_count > 0, and durable bookkeeping (result_refs, signatures,
  failure kinds) lives on LoopExecutionState, not the batch; the stop
  strategy guards on invocation_count > 0. Documented why default(0) and
  the for_invocation_count reset are safe; no logic change.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(ci): rustfmt + feature-gated TurnRunState construction sites

- cargo fmt across the branch (subagent-written tests were unformatted).
- Add auth_resume_disposition: None to two slack feature-gated TurnRunState
  literals (slack_delivery.rs, slack_serve/e2e_tests.rs) that only compile
  under --all-features, which the default-feature workspace check skipped —
  unblocks the all-features clippy and ironclaw_reborn_cli build.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reborn): review round 2 — terminal-run replay guard, batch policy, cleanups

- M-c: the idempotent deny replay only special-cased Cancelled; a stale
  Deny against a Completed/Failed/RecoveryRequired run returned
  DenialResumed (false "live resume"). Now: Cancelled -> Canceled, any
  other terminal status -> StaleAuth, non-terminal -> DenialResumed
  (via TurnStatus::is_terminal). + test for the get_run_state error path.
- M-d: batch policy / stop_on_first_suspension / summaries /
  CapabilityBatchStarted were computed from the pre-partition
  visible_calls, so a mixed batch dispatched remaining calls under the
  denied call's policy. Compute them after the auth-deny partition from
  the calls that actually reach invoke_capability_batch.
- L-c: single-ownership pass — take pending_auth_resume once instead of
  read-then-post-hoc-clear.
- L-d: trim the oversized auth-deny comment block.
- L-e: drop stale `reason`-slot wording from AuthResumeDisposition doc.
- Nit: remove the PHASE 2a planning label from a test comment.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(reborn): relocate auth_resume_disposition to resume request; drop DenialResumed

Review round 3 (M-b, L-b):

- M-b: auth_resume_disposition was a one-shot resume signal carried on the
  broad, stable LoopRunContext (every host port / hook / prompt / fixture
  carried it; only PlannedDriver::resume read it). Move it onto
  AgentLoopDriverResumeRequest: the turn runner populates it from
  claimed.state, PlannedDriver::resume stamps pending_auth_resume from the
  request. Removes the field + builder from LoopRunContext and the
  create_host wiring. The persisted ResumeTurnRequest/TurnRunRecord/
  TurnRunState fields are unchanged. Behavior identical.

- L-b: delete the DenialResumed response variant — it wrapped the same
  ResumeTurnResponse as Resumed and downstream collapsed to Resumed; the
  denial intent already lives in ResumeTurnRequest.auth_resume_disposition
  (what the executor reads). resume_denied_auth and the idempotent replay
  now return Resumed; removed the dead match arms in workflow.rs and
  reborn_services.rs.

Cancel/Deny behavior unchanged: the facade still maps both
WebUiGateResolution::Denied and ::Cancelled to AuthInteractionDecision::Deny
(both surface a model-visible "user has not provided auth" failure and
continue the loop). A distinct Cancel=stop action is a UI+backend follow-up
(the auth gate has a single action today).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(reborn): harden auth-deny regression tests per review

- T1: assert final_state.result_refs carries both the denied call's
  Authorization failure ref and the remaining call's ref (not only the
  host-appended records) — guards the next-prompt-visibility contract.
- T2: 1-denied + 2-remaining partition test now asserts exact set
  membership {Y, Z} for the dispatched calls (a [Y, Y] regression now
  fails) plus result_refs for all three.
- T3: full TurnPersistenceSnapshot legacy-JSON test — a persisted run
  with auth_resume_disposition absent deserializes to None (guards the
  serde default at the snapshot boundary, not just the bare struct).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* style: rustfmt checkpoint_state_store_contract import block

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(reborn): cover denied-disposition checkpoint round-trip and auth-continuation wrapper

Adds the two regression tests requested in PR nearai#4944 review:
- PendingAuthResume with Some(AuthResumeDisposition::Denied) survives
  checkpoint encode/decode through LoopExecutionState::from_checkpoint_payload.
- dispatch_auth_continuation returns Ok(()) and never calls get_run_state
  or resume_turn for non-turn continuations (SetupOnly / LifecycleActivation
  / ProductActionResume).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
…tion, never-silent feedback, OAuth-only auth (nearai#4946)

* fix(slack): approval/auth UX overhaul — clearer prompts, no silent wedge, OAuth-only auth

Approval gate fixes:
- Re-key the busy/auth hint dedup on the per-message external_event_id (was
  run_id) so every new message re-surfaces the hint instead of going silent
  after the first one.
- Surface WHAT is being approved: look up the ApprovalRequestStore by gate ref
  (the same source the WebUI projection uses) and render tool/reason in the
  Slack prompt.
- Distinguish a stale/already-resolved gate from an active policy denial
  (new ProductRejectionKind::StaleGate) so "deny then approve" no longer reads
  "declined by policy".
- Shorten and correct the prompt copy: What/Why in the body, one accurate
  reply-instruction line in the adapter footer (DM says "here", channel says
  "mention me in this thread"); drop the doubled gate: prefix.

Auth over Slack — OAuth only:
- Live and triggered delivery now allow only link-based OAuth auth
  (authorization_url present). Any credential-entry challenge (PAT / manual
  token) is denied: the run is cancelled (cancel_run, Policy reason) and a
  "set it up in the web app" notice is posted — no credential is solicited
  from the chat surface.
- Triggered automations still render approval and OAuth-auth prompts and record
  delivered gate routes, so an in-thread approve/auth routes back to the
  triggered run (by conversation fingerprint, not thread id) and continues it.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(slack): clearer approval reply copy

The footer overstated where a reply works: "from any DM/channel" implied the
explicit `approve gate:<ref>` form works anywhere, but the bot only acts on
messages it actually receives (this chat). Reword to: reply approve/deny in this
chat (or @-mention in this thread for channels) to respond, and offer the
explicit gate ref only as a disambiguator when several approvals are pending
here — no "from anywhere" claim.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(slack): tolerant approval parse + clearer credential-auth block copy

- parse_approval_resolution: a well-formed `gate:<ref>` now wins even when the
  user pasted the whole instruction line ("approve gate:X or deny gate:X").
  Previously the trailing tokens made it malformed → silently dropped (NoOp) →
  no response. Only a non-gate token with trailing text is still rejected.
- Reword the non-OAuth auth-block message to state the security rationale
  plainly: entering a credential in chat is a risk (stored in the conversation),
  so credential-based connections are web-app only.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(slack): recency-driven gate resolution + never-silent busy hints

Bare approve/deny in a DM bridges to delivered gate routes by
conversation fingerprint. Two defects surfaced:

- Multiple live routes in one DM (top-level DMs share a fingerprint)
  failed closed as Ambiguous, and a list_pending-based pending probe
  false-negatived live gates (the read model's blocked-run lookup is
  scope-fragile — the same reason the reply falls through to the
  fallback) and then pruned the valid route.
- An open gate left unrelated user messages in silence: the busy-thread
  hint was suppressed whenever the conversation binding lookup failed.

Resolution is now recency-ordered and resolve-driven: order candidate
routes most-recent-first and walk them, resolving the newest still-live
gate. Staleness is decided by the authoritative resolve() outcome
(StaleGate/MissingGate), not a separate probe — stale routes are pruned
and skipped (bare lookups only; an exact `gate:<ref>` still surfaces its
own stale error). Same loop applied to the auth path.

Busy-thread hints never go silent and now name what is blocking:
- binding unresolved -> generic busy copy instead of nothing;
- BlockedApproval -> the gate ref AND the tool it would authorize
  (via the same approval_prompt_context_view the prompt uses);
- BlockedAuth -> the gate ref + `auth deny <ref>` / web-app completion
  (there is no auth-approve over Slack by design — providing a
  credential or completing OAuth cannot happen in a chat reply).

Single-active-run is unchanged: an open gate still blocks new turns,
the user just always gets actionable feedback.

Tests: recency + stale-prune contract tests; e2e two-live-routes now
asserts most-recent resolution; busy-hint tests assert the enriched
copy; stale auth-block assertions pinned to the live constant.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(slack): triggered-path parity + clearer gate feedback, plus shared helpers

Make the triggered-automation Slack path behave like regular inbound, and
state what's blocking / what would be approved at every surface.

Behaviour
- Foreign-run guard: a resolution bridged to a triggered run resumes a run
  in the trigger's own scope. The live delivery observer now detects (via a
  ScopeNotFound probe) that the run isn't in this conversation's scope and
  skips cleanly — the triggered delivery loop owns continuation. Fixes the
  spurious "Something went wrong delivering the result here." after each
  approve on a triggered run.
- Triggered approval prompts render through the same shared
  `slack_approval_gate_prompt_view` (What/Why) as the live flow, so the two
  can't drift (they had).
- Busy-thread hint now names the blocking gate ref AND the tool it would
  authorize (BlockedApproval) or the auth gate + `auth deny <ref>`
  (BlockedAuth), via the same approval context view the prompt uses.
- Every triggered Slack message carries a footer naming the routine and the
  surface contract: you can approve/deny or `auth deny` and receive updates
  here; to interact with the run, open the web app. (Triggered Slack is
  output + gate-resolution only — not a conversational channel.)

Refactor (review-driven, modularization)
- Extract shared `cancel_auth_blocked_run` used by both the live and
  triggered auth-deny paths so the cancellation contract can't drift.
- `order_delivered_routes` doc + the shared select comment use
  domain-neutral terms (approval and auth); `dispatch_auth_resolution`
  documents why only MissingAuth (not StaleAuth) falls back on the exact-ref
  path; `resolve_via_delivered_auth_route` matches its approval sibling's
  ack-build shape; new triggered warns carry the module log target.
- Document the triggered Slack surface contract at the single minting site
  (`triggered_notification_for_state`).

Tests
- Foreign-scope resolution ack skips live delivery without a spurious error.
- Triggered approval prompt uses the shared render and carries the footer.

Deferred (noted, not done): unifying the approval/auth resolve walk
(Phase-C `DeliveredRouteResolutionContext`, per existing arch-exempt) and
replacing the foreign-run probe with an ack-level ownership signal.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(slack): split triggered footer — "triggered event", no act-here on results

The single triggered footer claimed "you can reply approve/deny" on every
message, including successful final replies where there is nothing to act on —
confusing. Split into two:

- gate prompts (approval / OAuth auth): "you can respond to this request here —
  to otherwise interact with this run, open the Ironclaw web app."
- updates / final replies (and the non-OAuth auth-unavailable notice): "you
  can't interact with triggered events here — open the Ironclaw web app to
  interact with this run."

Also rename the user-facing wording from "automated routine" to "triggered
event" per product terminology.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(slack): restore 7 auth e2e tests deleted as auto-fixer collateral

The OAuth-only-auth commit (f5c979e) stripped auth_challenges /
slack_auth_prompt_view and the auto-fixer deleted the e2e tests + harness
infra that referenced them — they were never restored, dropping coverage
for behaviour that still exists. A coverage-gap analysis confirmed none of
the seven duplicate existing tests.

Restored (live/inbound path, OAuth-only adapted — no triggered footer):
- DM OAuth auth prompt renders the setup link
- channel auth prompt omits the setup link (private-link stripping)
- DM final reply delivered after auth completes outside Slack (+ cleanup)
- auth deny via @mention in a channel thread cancels the gate, no agent turn
- auth deny in a DM thread cancels the base DM gate, no agent turn
- auth prompt posted exactly once when an AuthResolution ack races the live
  delivery loop (the auth counterpart to the existing approval-fanout test)
- approval -> auth resume completes without a second approval

Restores the co-deleted infra: TurnMode::{BlockAuth, BlockApprovalThenAuth},
RecordingTurnCoordinator::{resume_blocked_run_to_running, complete_blocked_run,
submitted_turn_count}, Harness.auths, the auth-challenge harness builders, the
approval-service approval-then-auth branch, and the auth/thread fixtures.

No production code changed. Assertions match current behaviour (auth deny
posts SLACK_AUTH_CANCELED_MESSAGE; OAuth setup link via FakeAuthChallengeProvider).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reborn): address PR review + fix CI (fmt, StaleGate exhaustiveness)

CI was red on Formatting, Clippy (all-features), and Code Style. Root
causes + review fixes from PR nearai#4946:

- I: add the missing `ProductRejectionKind::StaleGate` arm to the
  openai-compat error mapper (→ 409 Conflict, non-retryable). The
  non-exhaustive match was the all-features clippy failure (E0004).
- Run `cargo fmt --all` over the merge — committed code was never
  formatted (Formatting job diff).

Workflow / approval-resolution review fixes:
- B: fail closed when bare delivered-route recency ties (identical
  recorded_at on the two newest routes → ambiguous).
- C: restore the arch-exempt annotation on `select_delivered_gate_routes`.
- D: normalize a skipped `MissingGate` to `StaleGate` on bare exhaustion
  so the surfaced hint is "no longer pending", not BindingRequired.
- M: propagate delivered-route store read failures as Transient instead
  of collapsing an outage into a route miss; annotate best-effort prune
  with `// silent-ok:`.
- K: scoped approval test double takes run_id from the pending map
  (fail loud) instead of `run_id_hint.unwrap_or_default()`.
- L: add stale-path regression tests (exact-ref StaleAuth / StaleGate
  surface without delivered-route fallback) + openai-compat StaleGate
  mapping test.

Slack delivery review fixes:
- G: scope the foreign-run skip to bridged gate/auth resolution payloads
  only — a normal UserMessage must surface delivery errors, not be
  silently dropped.
- F: use the generic busy copy (not approval-specific) when the
  conversation binding can't be resolved.
- J: drop the duplicated Slack-side StaleGate string; route through the
  shared `user_facing_hint()` / `user_facing_auth_hint()` owners.
- E: document the best-effort approval-context lookup with `// silent-ok:`.
- H: cancel_run e2e fake reports `already_terminal` idempotently.

Slack adapter review fix:
- A: any non-`gate:` approval token is a no-op regardless of trailing
  text, so `approve this` and `approve this please` behave consistently
  (was a hard adapter error on the single-token case).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Co-authored-by: Robert Yan <mstr.raphael@gmail.com>
* fix(reborn): hide active extension activation action

* test(reborn): cover active extension configure modal
…earai#4939)

* fix(reborn): credentials are owner-scoped, not thread-scoped (nearai#4935)

Credentials are owned by tenant/user/agent/project. thread_id, mission_id,
invocation_id, and session-flow provenance are transient invocation context,
never ownership keys. They were leaking into credential identity comparisons in
five places, so a credential consented in one chat thread / OAuth flow was
invisible (or got forked) in the next.

A — OAuth account fork. The bind decision in scoped_update_binding_for_requester
used scope_matches full-equality (incl. the per-flow invocation_id), so a 2nd
OAuth flow could never reuse the existing account and forked a duplicate
UserReusable account every time. The duplicates each report Google's cumulative
include_granted_scopes set, so selection tie-broke by recency and served the
wrong token -> Google 403, "works in one thread but not another". Fixing the
bind decision alone was insufficient: create_flow's validate_scoped_update_binding
and the callback update_bound_oauth_account carried the same invocation leak and
rejected the bound account. All three now resolve at owner granularity via the
shared ironclaw_auth::binding_scope_owns_account helper. Runtime resolution keeps
its provider-scope gate (binding is scope-agnostic on purpose so a reconnect that
grants a new scope updates the existing account instead of forking).

B — runtime resolution thread-binding. account_visible_from_runtime_scope
required exact thread/mission/session for non-UserReusable accounts, so
ExtensionOwned/SharedAdminManaged credentials were thread-bound. Now owner-level
(tenant/user/agent/project); requester authorization is still enforced separately
by the visibility-policy stage / is_authorized_for_requester.

C — gsuite Google resolver. GoogleCredentialResolver never stripped thread_id
before the owner lookup, so even UserReusable Google credentials were thread-bound
at use time. resolve / account_by_id / recovery projection now look up at owner
granularity.

Shared helpers: ResourceScope::credential_owner_scope() (ironclaw_host_api) and
ironclaw_auth::binding_scope_owns_account() are the single source of truth so the
host-runtime resolver, the gsuite resolver, and the OAuth bind/validate/update
paths cannot drift. runtime_account_owner_scope now delegates to the former.

Regression tests: cross-thread resolution for UserReusable + ExtensionOwned on
both resolvers; OAuth reconnect binds to the owner's existing account across a
fresh invocation_id; cross-owner reconnect still does not bind; binding ignores
the provider-scope gate while runtime resolution still enforces it.

Out of scope: manual-token completion still validates through the general-purpose
validate_account_update_target (full scope_matches); it does not affect the OAuth
fork and is left as a follow-up. Pre-existing forked accounts on disk are not
deduped (forward-only).

* refactor(reborn): use canonical binding_scope_owns_account at OAuth bind guard

Replace the inline CredentialAccountOwnerScope::from_scope(&owner_scope).matches()
guard in scoped_update_binding_for_requester with the shared
binding_scope_owns_account helper, and drop the now-unused
CredentialAccountOwnerScope import. Equivalent (matches ignores surface;
owner_scope preserves the original session_id) and keeps the owner-boundary
check routed through the single shared primitive.

* refactor(reborn): tighten the owner-scope ownership boundary

Address structural review feedback so the owner-scoped credential model is
expressed once, in the right layer, instead of being spread across host API
docs, auth helpers, route construction, selector modes, and test boilerplate.

- Layer leak: rename the neutral primitive in ironclaw_host_api to
  ResourceScope::without_thread_and_mission (no credential/OAuth/GSuite policy
  language). The credential-ownership contract now lives in ironclaw_auth as
  AuthProductScope::credential_owner / to_credential_owner, which every runtime
  resolver and the OAuth bind path build on. host_api keeps only neutral
  authority vocabulary, per its guardrails.

- Nullable selection mode -> enum: replace
  configured_accounts_for_requester(scope_gate: Option<(setup, scopes)>) with an
  explicit AccountSelectionPurpose { Runtime { setup, provider_scopes }, Binding }
  so the two materially different selection modes can't collapse into subtle
  None/Some branching.

- Required trait method: select_configured_account_for_binding no longer has a
  CredentialMissing default. An unwired binding path now fails at the type
  level instead of silently no-opping; test doubles return CredentialMissing
  explicitly.

- Test sprawl: add a ConfiguredAccount fixture builder + owner_auth_scope helper
  and apply them across the runtime-credential tests. The copy-heavy ten-field
  NewCredentialAccount literals collapse to intent-revealing builder calls;
  tests.rs drops from ~1100 to 938 lines (below the pre-PR 995) with identical
  coverage.

No behavior change to the owner-scoped resolution itself.

* fix(reborn): enforce exact session_id on OAuth bind/update

binding_scope_owns_account documented that session_id is matched (it is
path-segmenting), but the implementation routed through
CredentialAccountOwnerScope::matches, which wildcards session via is_none_or
when the flow scope's session_id is None. The same gap existed in the binding
account selection (AccountSelectionPurpose::Binding skipped the gate entirely).

Because the bind/update WRITE path is session-segmented on disk
(product_auth_durable keys account records by session_id) and
update_bound_oauth_account reads the account at the flow scope's session path, a
None-session reconnect could wildcard-select a Some(session) account (or
select_latest could pick a different session's duplicate) that the callback can
never update. Require exact session_id equality (including None == None) in both
the helper and the binding selection so a cross-session account is never bound.

Owner-granularity for thread/mission/invocation is unchanged; only session is
tightened, matching the documented path-segmenting contract. Adds regression
tests: a None-session reconnect does not bind a session-scoped account, and a
same-session reconnect still does.

* test(reborn): cross-thread Google bind route test + GSuite helper cleanup

Address review feedback:
- Add extension_google_oauth_start_rebinds_account_authorized_in_a_different_thread:
  a route-level regression proving a Google account authorized in one
  thread/mission is rebound (not forked) when the OAuth reconnect starts from a
  different context. The prior Google route test only varied invocation_id; this
  exercises the user-visible cross-thread regression nearai#4935 fixes end to end.
- Use the canonical AuthProductScope::credential_owner helper in
  GoogleCredentialResolver::account_by_id instead of rebuilding the owner scope
  by hand, so the owner projection has one spelling across the resolver.
- Fix a stale `request.scope` reference in the resolve() comment (the method has
  no `request`); point it at this method's `scope` and access_secret_scope.

* test(reborn): durable cross-invocation reauth regression + record ownership contract

Address review feedback:
- Add filesystem_oauth_reauth_updates_bound_account_across_fresh_invocation:
  drives the durable complete_oauth_callback through a bound reconnect whose flow
  scope differs from the account's creation scope by fresh invocation/thread/
  mission. Existing durable reauth coverage reused one scope for create/claim/
  complete, so it could not catch a regression to full scope_matches in
  validate_scoped_update_binding or update_bound_oauth_account. The new test
  asserts the bound account is updated in place (same id, new access secret, no
  fork); a regression rejects the binding with CrossScopeDenied.
- Record the credential-ownership contract in docs/reborn/contracts/auth-product.md
  (Ownership Boundaries): owner = tenant/user/agent/project; thread/mission/
  invocation are transient provenance; session is path-segmenting for bind/update;
  names the shared helpers (without_thread_and_mission, credential_owner,
  binding_scope_owns_account) and the requester-authorization separation.

* revert: drop accidental run-reborn-webui.sh changes from 63b09ba

The Google OAuth env wiring added to scripts/run-reborn-webui.sh was a local
testing convenience that landed in the PR by mistake. Restore the script to its
state before 63b09ba; the credential/test changes in that commit are unaffected.

* docs: clarify invocation_id is ignored (not stripped) + restore distinct test labels

Address Copilot review:
- binding_scope_owns_account never modifies invocation_id; it is simply not part
  of CredentialAccountOwnerScope's comparison. Reword the credential.rs docstring,
  the auth-product.md contract, and ResourceScope::without_thread_and_mission to
  say invocation is "ignored" / "left unchanged" rather than "stripped"/"dropped",
  so readers don't look for a non-existent strip step.
- The fixture-builder refactor collapsed the distinct account labels that
  resolver_uses_most_recent_account_across_multiple_reusable_logins and
  resolver_resolves_google_capability_labeled_duplicates describe in their
  names/comments. Add a ConfiguredAccount::label setter and restore the distinct
  labels ("personal/work github", "gmail/google-calendar google") so the fixtures
  match the documented capability-labeled-duplicate scenario.

* Address PR review: cross-thread resolve_account test + ambiguous-binding observability

- Add resolve_account_finds_owner_account_authorized_in_a_different_thread
  regression test locking owner-scoped lookup on the account_by_id path
- Log (instead of silently collapsing) the AccountSelectionRequired arm in
  scoped_update_binding_for_requester so an ambiguous reconnect is observable

* Address PR review: owner-granularity on manual-token + surface, fake parity

Three real bugs of the nearai#4935 class that the OAuth path fixed but were missed
elsewhere, plus the consolidation/coverage the review asked for:

- Surface segmentation: binding_scope_owns_account and the binding selection
  pre-filter now require exact AuthSurface equality (surface is path-segmenting
  on disk like session). A bind could otherwise select an account on another
  surface that the callback can never read -> spurious CredentialMissing aborts
  the reconnect.
- Manual-token apply: the bound-reconnect apply path used full scope_matches,
  so a cross-thread manual-token reconnect passed setup but failed at apply
  (re-forking). Added validate_bound_account_update_target (owner granularity)
  and applied it in both the durable path and the fake; the mutation preserves
  the account's own durable scope, mirroring the OAuth callback.
- Fake parity: InMemoryAuthProductServices now routes an exchange with no
  provider account_id but an update_binding to the bound-update path (was
  rejected), and uses owner-granularity in update_bound_callback_account, so
  tests exercise the production reconnect contract.
- Consolidated the generic resolver onto AuthProductScope::credential_owner and
  removed the redundant runtime_account_owner_scope wrapper.

Regression tests: cross-thread manual-token reconnect (contract + durable),
no-account-id OAuth reconnect (fake), cross-surface binding rejection,
binding requester-visibility denial, SharedAdminManaged cross-thread
resolution, and resolve_account cross-thread lookup. Contract doc updated.

* Address auth owner-scope review follow-ups

* fix: mark oauth route tests as test-only

---------

Co-authored-by: Henry Park <henrypark133@gmail.com>
* fix wasm staged credential reuse

* Update crates/ironclaw_host_runtime/tests/runtime_http_egress_contract.rs

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>

---------

Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
* Fix security tool boundary checks

* Address shell wrapper review feedback

* Fix shell clippy lifetime lint

* Fix shell wrapper risk regressions

* test: cover adjacent env split-string payloads

* fix: harden env wrapper risk parsing

---------

Co-authored-by: Codex <codex@openai.com>
* Fix local extension auth owner scope

* Split extension lifecycle auth regression test

* Canonicalize local lifecycle surface context

* Fix local runtime test lifecycle context

* Annotate test helper panics

* Use configured identity for local lifecycle context

* fix(reborn): expose identity facade without webui feature
…4973)

* fix(reborn): add client-only-preserving timeline reload

loadHistory gains a preserveClientOnly option that merges a fresh full
timeline over the current view while keeping client-synthesized err-*
run-failure bubbles, which never persist as timeline records. This lets a
reload triggered on any terminal run status recover tool input/output
previews from the durable timeline without erasing a visible failure
notice.

* fix(reborn): recover tool input/output on every terminal run

Tool input/output cards reach the UI only via the live capability_display_
preview SSE frame or the durable timeline — the projection state carries
only the sparse capability_activity item. When the live preview frame is
dropped (the resumable drain can mark a completed-but-unrecorded preview
NotApplicable past its grace window), the timeline reload is the only
recovery, but it previously fired solely on terminal success.

Settle a run on ANY terminal status (failed, cancelled, recovery_required
included) and reload the timeline so tools that completed before a run
terminated still surface their input/output. The reload uses
preserveClientOnly so the client-side err-* failure bubble survives.

* test(reborn): cover settle-on-every-terminal and client-only reload merge

Rework the useChatEvents harness onto the onRunSettled contract and add
cases asserting success, failure, cancellation, and the typed failed event
each settle the run exactly once (deduped against projection replays), with
success carried through. Add a mergePreservingClientOnly case proving the
timeline wins on shared ids while err-* failure bubbles survive the reload.

Also fix the useHistory source-loader stripper to handle every top-level
export so the previously unparseable test file runs.

* fix(reborn): address review on terminal-run timeline reload

- Write the cache outside setState for every full load so a preserve-client-
  only reload still refreshes the cache when the user switched threads mid-
  fetch (setState bails on a stale thread), avoiding a re-fetch + flicker on
  return; the cache merge runs over the cached messages
- Rename loadHistory's inner options param to loadOptions so it no longer
  shadows the useHistory hook options
- Guard settledRunsRef.current consistently in settleRun so a nullish ref
  can't throw and break the SSE event loop
- Null-safe mergePreservingClientOnly against nullish message elements

Regression coverage lives in the JS suites (useChatEvents.test.mjs,
useHistory.test.mjs); the regression-test check only detects Rust tests, so:
[skip-regression-check]
…8n completion (nearai#4956)

* fix(reborn): don't show NEAR AI as active on a clean inference setup

useLlmProviders fell back to the nearai provider id even when no
selection existed, so the Settings → Inference list promoted NEAR AI
into the ACTIVE group. Resolve activeProviderId to null when nothing
is configured and keep the nearai fallback only for default selection.

* fix(reborn): show a clear error when deleting a running conversation

Deleting a busy/running thread rejects with a 409 "busy" conflict whose
humanized API message is just "Busy", which reads as no feedback. Map
that case to a clear, localized line via a new deleteThreadErrorMessage
helper, with chat.deleteBusy / chat.deleteFailed keys in every locale.

* fix(reborn): make attachment warning banner readable in light theme

The attachment error banner used the un-aliased Tailwind text-red-100
(~#fee2e2), which is near-invisible on the light theme's light-red
banner background. Switch to text-red-200, the design-system class
app.css maps to the theme-aware --v2-danger-text token, matching every
other error banner.

* i18n(automations): localize schedule cadence labels

PR nearai#4920 added new schedule labels (Every minute / Every N minutes /
Hourly at :MM) into the pure scheduleLabel presenter, which renders all
its cadence text in English regardless of locale. Localize the whole
function: sentence templates come from automations.schedule.* keys in
every pack, while time, weekday, and month/date use Intl.DateTimeFormat
for the active locale. Thread the translator + locale from useAutomations
and update the presenter test and the served-bundle guard.

* i18n(automations): localize status pills, summary badges, and delivery panel

Several Automations strings still rendered in English in non-English
locales: run/state status pills (RUNNING/ERROR/…) and date fallbacks
were hardcoded in the presenter, the summary cards showed the raw badge
tone keyword (MUTED/SIGNAL/INFO/DANGER), and the whole delivery-defaults
panel plus the running/failures filter tabs were English placeholders in
every non-en pack.

Key the presenter labels and thread the translator + locale through
normalizeAutomations (dates now use Intl for the active locale); add a
StatCard badgeLabel so the summary chips show a translated tone word;
add the new automations.badge/state/runStatus/lastStatus/date keys and
real translations across all 11 packs; and fill in the delivery/filter
translations that were left as English.

* fix(reborn): keep badge labels on one line for CJK locales

The Badge chip had no whitespace-nowrap, so translated tone labels in
space-free scripts (e.g. Chinese "信号") wrapped character-by-character
and stacked vertically inside the fixed-height pill. Add whitespace-nowrap
and shrink-0 so the label stays on a single line.

* fix(reborn): make attachment error banner dismissible and thread-scoped

The attachment staging error had no close affordance and, because the
composer stays mounted across conversation switches, persisted into every
other thread. Clear the error on thread switch and add a dismiss button.

* fix(reborn): address PR review on automations i18n

- Dedupe the 11 pre-existing duplicate keys per non-en locale pack (an
  early English-placeholder block shadowed by a later translated block),
  keeping the translated occurrence — clears Biome noDuplicateObjectKeys.
- Translate automations.filter.all where it was still English.
- monthDayLabel: use a leap year (2000) as the yearless placeholder so a
  Feb 29 schedule doesn't render as Mar 1.
- isThreadBusyError: match payload.kind === "busy" instead of any 409, so
  non-busy conflicts don't get the running-thread guidance.

* fix(webui): address review feedback on busy and alert styling

---------

Co-authored-by: think-in-universe <46699230+think-in-universe@users.noreply.github.com>
* fix(webui): keep extension cards natural height

* test(webui): harden extension card review coverage

* chore: retrigger regression gate [skip-regression-check]

---------

Co-authored-by: think-in-universe <46699230+think-in-universe@users.noreply.github.com>
* fix(reborn): show friendly google oauth scope failure

* fix(reborn): preserve oauth callback failure status
* fix(webui): avoid token prompt for oauth auth gates

* chore(webui): address oauth auth gate review comments
* feat(reborn): downloadable project files in WebChat v2

Add a generic, path-based project-filesystem read API and surface
agent-produced files (CSV, reports, exports) as download chips in the
WebUI v2 chat.

Backend (reusable for general filesystem navigation):
- ProjectFilesystemReader port + DTOs (list/stat/read) + ProjectFsError
  in ironclaw_product_workflow; facade methods reuse the caller-ownership
  probe and map to the sanitized error taxonomy
- ProjectScopedFilesystemReader over the read-only /workspace mount in
  composition (confines to /workspace, denies sensitive/escape/oversize)
- GET /threads/{id}/files, /files/stat, /files/content webui_v2 routes;
  downloads stream with Content-Disposition: attachment + nosniff
- mime_for_extension helper added to ironclaw_common

Frontend:
- /workspace path references in replies render as download chips
  (filename, size via /files/stat, bearer-authenticated blob fetch)
- shared lib/download.js::saveBlob, deduping the settings-toolbar copy

The agent references files it creates by their /workspace path; the UI
makes those paths downloadable over the generic endpoint. No agent tool
or agent-loop change is needed — download is a presentation concern over
the existing project filesystem.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(reborn): e2e for agent-produced downloadable files

Drive the real reborn agent loop to produce a CSV and a PDF via write_file,
then assert both are listable and downloadable through the v2 project-
filesystem endpoints — the same surface the WebUI download chips call.

- WriteFileGateway: scripted model gateway that registers two write_file
  tool calls (/workspace/report.csv, /workspace/report.pdf) in one round,
  then emits a final reply referencing both paths
- Asserts /files/content serves each with the registry-derived mime
  (text/csv, application/pdf), Content-Disposition: attachment, nosniff,
  and exact bytes; /files lists both; /files/stat reports size; an
  out-of-workspace path is refused
- Parameterize the shared harness with the model gateway and runtime
  policy (the file test uses a minimal-approval LocalYolo policy so an
  in-workspace write auto-proceeds instead of parking on a write gate)

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(reborn): browser e2e for agent-produced file download chips

Playwright scenario driving `ironclaw-reborn serve` on the local-dev-yolo
profile (minimal approvals, so an in-workspace write_file auto-proceeds).
The mock LLM turns the prompt into two write_file tool calls (a CSV and a
PDF) plus a reply referencing their /workspace paths; the SPA renders those
as download chips and clicking one performs the bearer-authenticated blob
fetch and saves the file.

- mock_llm: write_file multi-call pattern + final reply referencing the paths
- chip: data-testid="project-file-chip" + data-file-path for stable selection
  (locked by the static-asset wiring test)
- helpers: SEL_V2 chip selectors
- parameterize the smoke server's config writer with a profile (default
  unchanged); the yolo fixture reuses the startup plumbing

Complements the in-process webui_v2_e2e.rs test (same endpoints, real
agent-produced file). Runs in the CI E2E harness, not under cargo test.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reborn): address PR review on project-file download

Gemini + CodeRabbit findings on nearai#4933:

- download.js: defer URL.revokeObjectURL via setTimeout so the browser's
  async download manager can resolve the blob before the URL is freed
  (Chrome/Firefox/Safari download failure).
- message-bubble.js: render ProjectFileChips for assistant messages only,
  so user-authored /workspace/... text no longer triggers stat/download.
- descriptors.rs: classify the three file routes as ProductWorkflow (they
  enter RebornServicesApi filesystem methods, not projection-stream reads);
  update the locked contract-test expectations to match.
- project_filesystem_reader.rs: re-check realized byte length after
  read_bytes to close a stat/read size-cap race on concurrently-grown
  files; use Path::strip_prefix for workspace confinement; add a
  non-allocating file_name_str helper for mime_for_path.
- handlers.rs: log the response-builder cause before mapping to a
  sanitized 500 instead of dropping it with map_err(|_| ...).
- attachment_format.rs: use direct == in mime_for_extension (input is
  already lowercased) and lock the table extension/alias uniqueness +
  lowercase invariant in the existing table test.
- project-file-paths.js: drop the dead trailing-punctuation strip (the
  regex always ends on an alphanumeric).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reborn): address second-round PR review (Copilot + CodeRabbit)

- project_filesystem_reader.rs: read via read_bytes_bounded so the size
  cap is enforced before oversized content is materialized (closes the
  stat/read TOCTOU more robustly than the prior post-read check); map
  MountNotFound/Unsupported to Unavailable (503 infra) instead of
  InvalidPath (400 caller-blame).
- mock_llm.py: the v2 download-chips canned reply was unreachable — the
  generic multi-tool summary returned first. Add _explicit_canned_response
  and prefer an explicit canned reply over the summary once tool calls have
  run and dedup'd; gmail/single-tool summaries are unaffected.
- webui_v2_e2e.rs: document files_uri as test-only (raw path interpolation;
  callers must avoid URL-special chars).
- docs plan: mark the frontend (task 5) done and use /workspace paths.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(ci): rustfmt merge fallout + bump wasmtime to 44.0.3 for RUSTSEC-2026-0182

- Run rustfmt on the import/route lists hand-resolved during the origin/main
  merge (handlers.rs, router.rs) — fixes the Formatting / Code Style jobs.
- Bump the wasmtime family 44.0.2 → 44.0.3 (patched line >=44.0.3,<45.0.0)
  to clear the cargo-deny vulnerability error for RUSTSEC-2026-0182
  (WASIp1 fd_renumber leak). `cargo deny check` now reports advisories ok.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reborn): address third-round PR review (nearai#4933)

Round 3 of PR nearai#4933 review (human reviewer + later Copilot/CodeRabbit passes).

Bugs:
- project-file-chips: drop dead `|| "Couldn't download…"` fallback — `t()`
  returns the raw key on a miss, so a failed download showed the literal
  `chat.fileDownloadFailed`. Add the key to all 11 locale packs.
- project-file-paths: strip fenced + inline code spans before scanning, so a
  displayed `cat /workspace/.env` no longer renders a one-click download chip.
- project_filesystem_reader: filter sensitive filenames out of `list_dir` via
  `is_sensitive_path_str` — a listing must not enumerate `.env`/`id_rsa` even
  though their bytes stay denied.
- handlers: reject missing/blank `?path=` on stat/download with a field-scoped
  400 instead of forwarding an empty string to the facade.
- handlers: cap the sanitized Content-Disposition filename length so an
  oversized name degrades to a truncated label instead of a 500.

Docs/tests/annotations:
- plan doc: `/project` → `/workspace` to match the implementation.
- mock_llm: suppress the explicit canned reply on a denied/errored tool run so
  a success-style reply can't mask a real failure.
- reborn_services: add `// dispatch-exempt:` annotations on the three read-only
  FS facade methods.
- Unit tests: sanitized_download_filename (injection/truncation/fallback),
  require_project_fs_path, mime_for_extension (canonical/alias/unknown),
  list_dir sensitive-name filter, confine prefix-sibling boundary, oversize
  read, plus three code-span cases for path extraction.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* docs(reborn): mark dropped §3.1–§3.3 as historical in fs-download plan

CodeRabbit follow-up on PR nearai#4933: §3.1–§3.3 were labeled "dropped" in the
design-decision note but still read as active implementation steps. Add a
banner at §3 and tag each of the three headers "(dropped — not implemented)"
so the spec stays authoritative — only §3.4/§3.5 landed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(reborn): unify project-file chips with attachment preview (nearai#4933)

Assistant `/workspace/...` download chips now use the same chip + preview
modal as message attachments, instead of a bespoke download-only button.
Clicking a chip opens `AttachmentPreviewModal` (inline image/pdf/text/audio/
video preview) with the Download action in the footer.

- Extract `AttachmentChip` + `AttachmentThumbnail` from message-bubble.js into
  a shared `attachment-chip.js` (with optional test-hook props); message-bubble
  imports them. The two surfaces can no longer drift.
- project-file-chips.js builds an attachment-shaped descriptor per path
  (`fetch_url` → `/files/content`, `mime_type`/size from `stat`) and renders the
  shared chip + a single shared preview modal.
- Backend: add `mime_type` to `ProjectFsStat` (extension-derived, mirrors the
  download Content-Type) so the WebUI can pick a preview representation before
  fetching bytes; reader populates it via `mime_for_path`.
- api.js: add `projectFileContentUrl` (same-origin relative URL feeding the
  shared `fetchAttachmentBlob`); drop the now-unused `fetchProjectFileBlob`.
- attachment-preview.js: add `data-testid="attachment-download"` hook.
- Tests: reader stat-mime + e2e stat asserts `text/csv`; assets wiring-lock
  updated for the shared components; e2e download scenario now clicks chip →
  preview modal → Download.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* feat(reborn): add inline download icon to attachment chips (nearai#4933)

The shared `AttachmentChip` now renders a trailing one-click download icon
alongside the click-to-preview body: clicking the chip body opens the preview
modal, clicking the icon fetches the bearer-authenticated bytes and saves them
directly (no modal). Applies to both message attachments and project-file
chips since they share the component.

- attachment-chip.js: split the single chip button into a preview-trigger
  button + a separate download-icon button (avoids nested buttons); the icon
  shows only for landed `fetch_url`s and uses `fetchAttachmentBlob` + `saveBlob`.
- project-file-chips.js: pass `downloadTestId="project-file-download"`.
- e2e: CSV asserts the inline download icon saves the bytes directly; PDF still
  covers the preview-modal Download path. New `project_file_download_for`
  selector.
- assets wiring-lock: assert the inline-download hook + fetch/save wiring.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reborn): virtualize host paths in local-dev shell output for download chips

In local-dev-yolo the shell process runs on the host with /workspace aliased
to the real repo dir. The command rewriter maps /workspace -> host path before
exec, so a program that echoes a path it was handed (e.g.
`print("saved /workspace/out.pdf")`) prints the host path. That host path then
flowed back into the model-facing output and the user-visible reply — leaking
host layout and defeating /workspace download-chip detection (the chip
extractor only matches /workspace/... paths).

- Add `rewrite_local_host_output_aliases` (the inverse of the command rewriter):
  map the longest-matching host-path prefix in command output back to its
  virtual alias, with boundary logic mirroring the forward pass so siblings like
  `proj-backup` are left alone. Apply it to the model-facing command-output
  preview in `LocalHostProcessPort::run_command`.
- Fix the default system prompt's Files guidance: reference downloadable files as
  a Markdown link / bare path (not inline code, which the chip extractor strips).

Tests: unit coverage for the rewriter plus caller-level run_command regression
tests (incl. the produced-file-path scenario). The existing workdir-translation
test now asserts the virtualized /workspace path.

* test(reborn): expect virtualized /workspace path in local-dev shell output

Follow-on to the host-path reverse-rewrite: the local-dev-yolo shell test
echoed `$PWD` (the real host workspace dir) and asserted the canonical host
path. The reverse output rewrite now maps that back to the `/workspace` alias
before the result reaches the caller, so the assertion expects
`/workspace/qa-coding-smoke`. This is the end-to-end confirmation, through the
real runtime, that host paths no longer leak into shell output.

* fix(reborn): address PR review — blank list path + stale thumbnail reset

- list_project_files: treat a whitespace-only `?path=` (e.g. `?path=%20%20`)
  as "list the workspace root" instead of forwarding a bogus path the facade
  rejects. Extracted `project_fs_list_path` so it mirrors
  `require_project_fs_path`'s trim-based blank handling, with unit tests.
  (Copilot review.)
- AttachmentThumbnail: fully resync thumbnail state in the effect on every
  `att` change, so a chip reused (React reconciliation) for a different
  attachment no longer shows a stale thumbnail. Lazy state init keeps the
  optimistic-image first render flicker-free. (CodeRabbit review.)

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
# Conflicts:
#	crates/ironclaw_reborn_composition/src/available_extensions.rs
#	crates/ironclaw_webui_v2/src/descriptors.rs
#	crates/ironclaw_webui_v2/src/handlers.rs
#	crates/ironclaw_webui_v2/src/lib.rs
#	crates/ironclaw_webui_v2/src/router.rs
@elliotBraem
elliotBraem merged commit 44f3208 into main Jun 17, 2026
elliotBraem pushed a commit that referenced this pull request Jun 17, 2026
…ing the run (nearai#4954)

* fix(reborn): surface approval-gate denial to model instead of cancelling the run

Approval-gate denial in Reborn cancelled the run (deny_gate /
replay_denied_gate -> cancel_run), so the model never learned the user
declined and the next trigger re-issued the same approval-gated
capability and re-blocked — the same loop class nearai#4944 removed for auth
gates.

Mirror nearai#4944 for approval gates: denial now RESUMES the parked run
carrying a denial disposition; the capability stage converts ONLY the
approval-gated call into a model-visible non-retryable Authorization
failure ("approval gate denied by user", SameCallRetryConstraint::
Forbidden) and the loop continues. Unrelated parallel calls are
unaffected.

Per the maintainability review of the plan, this unifies rather than
duplicates the nearai#4944 plumbing:
- ironclaw_turns: AuthResumeDisposition -> GateResumeDisposition (one
  gate-agnostic enum); ResumeTurnRequest/TurnRunRecord/TurnRunState/
  AgentLoopDriverResumeRequest field auth_resume_disposition ->
  resume_disposition. Serde key pinned to "auth_resume_disposition"
  (rename attr) so persisted run records still deserialize; legacy-key
  round-trip test added.
- ironclaw_agent_loop: PendingApprovalResume gains a disposition field;
  the auth denied short-circuit in CapabilityStage::process is extracted
  into ONE shared short_circuit_denied_resume helper used by both the
  auth and approval paths (no second copy).
- ironclaw_product_workflow: approval deny_gate / replay_denied_gate
  resume instead of cancel; ResolveApprovalInteractionResponse::Denied
  (CancelRunResponse) -> Resumed(ResumeTurnResponse); idempotent replay
  guarded by terminal run status.
- ironclaw_reborn: PlannedDriver::resume stamps the disposition onto the
  pending resume that is set (auth or approval).

Decisions (plan docs/plans/2026-06-15-reborn-approval-deny-continue.md):
both Denied and Cancelled continue (consistent with nearai#4944, no Cancel
variant). The extension_install/extension_search missing-observation gap
is a separate PR; the user-visible extension-install loop is only fully
closed when both land.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reborn): address PR nearai#4954 review — stamp denial on matching gate slot only

Review round 1 fixes:

- planned_driver: the denial disposition was stamped onto BOTH
  pending_auth_resume and pending_approval_resume on a comment-only "one
  slot at a time" invariant. GateStage deliberately preserves a pending
  auth resume when a non-auth gate blocks mid-re-dispatch, so both slots
  can be set at once; stamping both corrupted an unrelated auth resume.
  Now stamps only the pending slot whose gate_ref matches the blocking
  gate (state.last_gate). Adds a regression test asserting the auth slot
  stays None when the approval gate is denied, plus an end-to-end
  resume() drive.
- approval replay: match GateResumeDisposition::Denied explicitly rather
  than is_some(), keeping the gate-agnostic carrier tied to denial.
- tests: real TurnRunRecord struct-level serde test (legacy
  auth_resume_disposition key → resume_disposition) + snapshot-level
  legacy denied-marker test; new deny-path resume-error test asserting
  the record is denied and the run is never cancelled on resume failure.
- arch-exempt annotation on short_circuit_denied_resume's
  too_many_arguments allow (plan nearai#4954); stale comments/typos fixed.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reborn): route denied-gate replay through resume_turn idempotency (auth + approval)

Review finding nearai#7: the denied-gate replay paths derived idempotency from
current run state (TurnStatus::is_terminal() guard) rather than replaying
through resume_turn. After the first Deny resumed the run, a transport
retry with the same idempotency key arriving after the runner completed
returned StaleGate/StaleAuth instead of the original ResumeTurnResponse —
the observable result depended on runner timing.

resume_turn is idempotent by key (memory.rs:665 returns the cached
Result from resume_idempotency before the precondition check). Both the
approval (replay_denied_gate) and auth (resume_denied_auth replay arm)
paths now replay through resume_turn with the same key, deleting the
terminal-guard branching: a retried key replays the original response
regardless of run state; a genuinely stale request with a fresh key
still errors via the precondition. Auth and approval kept symmetric.

FakeTurnCoordinator now models resume idempotency by key so the replay
tests are meaningful; terminal-guard assertions re-framed around
same-key replay vs fresh-key stale, plus an explicit idempotent-replay
test on both services.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reborn): fail closed on ambiguous dual-slot stamp; lock deny-before-resume order

Review round 2 (both Major):

- planned_driver stamp_resume_disposition: the if/else-if silently stamped
  the auth slot if both pending slots matched last_gate. At the denial-
  attribution boundary that could misattribute an approval denial. Now an
  explicit 4-way match fails closed on the ambiguous (true, true) case
  (warn + stamp neither). Test added.
- approval_interaction_contract deny-resume-error test: asserted only
  aggregate call counts, which pass even if call order regressed. Added a
  shared ordered trace across the resolver (deny) and coordinator
  (resume_turn) fakes and assert deny is recorded strictly before resume.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* test(reborn): strengthen replay/checkpoint coverage; downgrade fail-closed log to debug

Review round 3 (straightforward):
- idempotent deny-replay tests (auth + approval) now assert full
  ResumeTurnResponse payload equality, not just run_id.
- stamp_resume_disposition ambiguous-dual-slot diagnostic: warn! -> debug!
  (REPL/TUI logging rule — internal fail-closed diagnostics use debug!).
- executor: assert the first approval BeforeBlock checkpoint carries
  pending_approval_resume.disposition == None before any denial.
- executor: denied-approval short-circuit no-matching-call test (denied X,
  model emits only Y -> X not surfaced, Y dispatches, pending cleared).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* fix(reborn): unify gate Declined resolution; keep WebUI processing on resume

Review round 3 (design + High):

- WebUiGateResolution: the approval card sends `denied`, the auth cards
  send `cancelled`, and both are now treated identically (resume the run
  and surface the decision to the model). Run termination is a separate
  control (the X -> cancelRun route), not a gate resolution. Collapsed
  the two equivalent variants into one `Declined` (serde aliases
  "denied"/"cancelled" keep the wire stable; no JS change). Facade maps
  Declined -> Deny for auth, approval, and the generic fallback.

- #6 WebUI desync (High): useChat.resolveGate kept processing only for
  approved/credential_provided, dropping processing + activeRun on
  denied/cancelled — but those now resume the run. resolveGate now always
  keeps processing/activeRun; the terminal run_status SSE event clears it
  and the X/cancelRun path remains the only stop. Fixes the latent
  auth-cancelled desync from nearai#4944. assets.rs assertion + useChat tests
  updated.

- nearai#7 helper weight: short_circuit_denied_resume no longer returns the
  DeniedResumeOutcome enum / boxes LoopExecutionState / clones the batch.
  It returns ControlFlow<TurnCompletedStep, (state, remaining_calls)>; the
  completed_turn/empty-remaining tail moved to the two call sites. Heavy
  per-denied-call failure synthesis stays shared (one helper).

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

* refactor(reborn): share denied-approval resume between deny_gate and replay

Review (Medium): deny_gate and replay_denied_gate built an identical
ResumeTurnRequest, mapped the same errors, and returned the same Resumed
shape — the only difference was deny_gate's one-off resolver.deny side
effect. Extracted a shared resume_denied(request, run_id) helper; deny_gate
performs the durable denial then delegates to it, and replay_denied_gate
calls it directly. Removes the duplicated request construction / path
handling.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

10 participants