chore: promote staging to staging-promote/733678dd-23996777140 (2026-04-05 13:23 UTC) - #2044
Conversation
* feat(test): add dual-mode live/replay test harness with LLM judge Add a general-purpose test infrastructure for running E2E tests in two modes: - Live mode (IRONCLAW_LIVE_TEST=1): real LLM calls with real tools, records traces to disk for future replay - Replay mode (default): loads saved trace fixtures, deterministic, no API keys The harness uses Config::from_env() in live mode so the test agent mirrors the real binary's behavior (engine_v2, allow_local_tools, approval gates). Includes an LLM judge for semantic verification of non-deterministic output, and saves human-readable session logs alongside trace fixtures for inspection and diffing between live and replay runs. First test case: zizmor security scanner against ironclaw's own workflows. New files: - tests/support/live_harness.rs — LiveTestHarness, builder, LLM judge - tests/e2e_live.rs — zizmor_scan test - tests/fixtures/llm_traces/live/ — recorded trace + session log TestRigBuilder additions: - with_http_interceptor() for injecting RecordingHttpInterceptor - with_config() for real-binary config parity (respects allow_local_tools, engine_v2 from env instead of forcing test defaults) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * test(live): add engine v2 zizmor scan test, add with_engine_v2 to harness Add zizmor_scan_v2 test that exercises the same scenario through engine v2. Documents the current v2 limitation: auto_approve_tools config flag is not honored by EffectBridgeAdapter — it only checks the per-session "always" set, so shell calls pause at the approval gate. Also: - Add with_engine_v2() to LiveTestHarnessBuilder for config override - Refactor v1 test to use shared run_zizmor_scan() helper - V2 test has relaxed assertions matching current v2 behavior Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(test): address PR review feedback - Fix UTF-8 unsafe string truncation in session log (use char_indices to find safe boundary instead of byte-index slicing) - Remove forced auto_approve_tools(true) from LiveTestHarness build_live; let Config::from_env() drive it, with per-test override via new with_auto_approve_tools() builder method - Apply engine_v2 builder override in TestRig's config-override branch so with_engine_v2() is not silently ignored when with_config() is used - Remove unused timeout field and with_timeout() from LiveTestHarnessBuilder - Tighten judge_response parsing to require strict PASS:/FAIL: prefix; anything else is treated as a failure with diagnostic message Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(agent): prevent self-repair notification spam for stuck jobs When a stuck job exceeds max repair attempts, self-repair returns ManualRequired but never transitions the job to a terminal state. detect_stuck_jobs() re-finds it every cycle (~60s), sending a Telegram notification each time — infinite spam. Two-layer fix: 1. repair_stuck_job: transition to Failed before returning ManualRequired, so detect_stuck_jobs stops finding the job 2. Agent loop: HashSet dedup prevents duplicate ManualRequired notifications per job (defense-in-depth if transition fails) Also adds Pending → Failed to the state machine — stuck Pending jobs (dispatched but never started) could not be terminated. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(agent): handle transition failure in ManualRequired and adjust message Log error if Failed transition fails, and adjust the ManualRequired message to accurately reflect whether the job was marked failed or not. Addresses gemini-code-assist feedback on PR #1867. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(agent): flatten nested Result, dedup Failed notifications, fix test state path Adversarial review findings: 1. CRITICAL: update_context returns Result<Result<()>>. Using .is_ok() on the outer only checks job existence, not transition success. Fixed: matches!(result, Ok(Ok(()))). 2. IMPORTANT: RepairResult::Failed arm had same spam potential as ManualRequired. Applied same dedup pattern. 3. IMPORTANT: Test exercised Pending→Failed (bypassing production path). Now transitions through InProgress→Stuck→Failed. 4. Dropped unrelated Cargo.lock version bump. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * docs: update state machine diagram to include Pending -> Failed Adversarial review finding: CLAUDE.md state diagram was the canonical reference but didn't show the new Pending -> Failed transition. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: j-bloggs <j-bloggs@users.noreply.github.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Code reviewNo issues found. Summary: This PR fixes a notification spam bug in self-repair for stuck jobs. The fix is well-designed with two layers:
The state machine is correctly updated to allow |
|
CRITICAL ISSUE FOUND - Memory Leak in Repair Loop: The Impact: Long-running instances will experience unbounded memory growth. After months of operation with millions of failed jobs, the set will consume significant heap memory. Fix options:
Severity: HIGH:100 - Will cause memory issues in production ironclaw/src/agent/agent_loop.rs Lines 531 to 534 in 5083aed |
* fix(bridge): sanitize orphaned tool results in v2 adapter * test(bridge): cover no-tool sanitizer path
* fix: harden WASM channel HTTP SSRF protections * fix: clean up wasm http security helpers for clippy
* Add Telegram local regression test harness * Add local Telegram smoke test runner * test: add high-priority Telegram regression tests Cover 6 previously untested API-level flows using fake axum Telegram servers: photo attachment download, voice attachment download, long message splitting (>4096 chars), Markdown parse error fallback to plain text, sendChatAction typing indicator, and polling mode (getUpdates with offset tracking). Test count: 13 → 19. All use real WASM channel execution with env-var URL override. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * test: add full-process Telegram E2E tests Add 4 end-to-end tests that boot IronClaw, activate the Telegram WASM channel via the setup API, POST webhook updates, and verify the sendMessage round-trip through the mock LLM to a fake Telegram API. Tests cover: - DM round-trip (setup → webhook → LLM → sendMessage) - Edited message handling - Unauthorized user rejection (dm_policy = pairing) - Invalid webhook secret rejection (401) New files: - fake_telegram_api.py: aiohttp server faking the Telegram Bot API - test_telegram_e2e.py: the 4 test scenarios conftest.py changes: - Add fake_telegram_server and telegram_e2e_server fixtures - Extend _wasm_build_symlinks to also cover channels-src/ Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * test: expand Telegram E2E coverage with 8 new regression tests Add 8 new tests covering core functionality gaps and high-priority error resilience scenarios for the Telegram WASM channel: Round 1 (functionality): - Group mention filtering (ignore without @bot, reply with @bot) - Long message chunking (>4096 chars split correctly) - Polling mode roundtrip (getUpdates picks up queued messages) - Markdown fallback (400 parse error triggers plain-text retry) Round 2 (resilience): - Missing webhook secret header (401 rejection) - 429 rate limit resilience (system survives, recovers) - Document download failure (getFile 500, text still processed) - Malformed payload resilience (invalid JSON handled, bot continues) Also extends fake_telegram_api.py with reject_markdown, rate_limit, and fail_downloads simulation flags plus control endpoints, and adds a "long response" canned pattern to mock_llm.py. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix: resolve CI failures in Telegram test suite - Add #[cfg(feature = "integration")] gate to test_bot_mention_detection_case_insensitive and build_telegram_update_value (fixes compilation on default/libsql) - Run cargo fmt on telegram_auth_integration.rs - Fix race condition in fake_telegram_api.py get_updates - Increase rate_limit_count from 5 to 20 for retry resilience - Move helper functions to proper section in test_telegram_e2e.py Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Robert Yan <46699230+think-in-universe@users.noreply.github.com>
…sts (#2064) * fix(staging): repair broken test build and macOS-incompatible SSRF tests Staging tip (f9ed815) was failing `cargo test` for several unrelated reasons. This commit gets the test suite back to green on both Linux CI and macOS. 1. **wrapper.rs:5595** — `PairingStore::new()` was called with zero args in `test_http_request_rejects_private_ip_targets`. The signature changed to `(db, cache)` in #1898 and the other test in the same file (line 5573) was updated to use `PairingStore::new_noop()`, but this one was missed. Switch to `new_noop()` to match. 2. **CLI snapshot** — clap's `render_long_help` for `--auto-approve` now emits an indented blank line between the short and long description (10 spaces, not empty). Update the snapshot to match the new output and refresh `assertion_line` (438 -> 461). 3. **validate_base_url IPv6 bracket bug** — `Url::host_str()` returns IPv6 literals WITH the surrounding brackets (e.g. `[::1]`), but `IpAddr::parse` does not accept brackets — it wants bare `::1`. As a result the IPv6 SSRF defense was effectively dead code: every `[…]` host failed to parse and fell through to the DNS-resolution path. This passed on Linux CI by accident (because `to_socket_addrs` on Linux also fails on bracketed strings), but broke on any host whose resolver returns a public IP for unresolvable lookups (ISP captive portals, ad-injecting DNS providers). Strip the brackets before parsing so the IPv6 detection actually works as intended. 4. **DNS-hijack-tolerant test guards** — two tests (`validate_base_url_rejects_dns_failure`, `test_validate_public_https_url_fails_closed_on_dns_error`) rely on RFC 6761's promise that `.invalid` lookups fail. On networks with DNS hijacking that promise doesn't hold and the lookups succeed (typically resolving to a public ad-server IP). Probe with `ironclaw-dns-hijack-probe.invalid` and skip the test with an eprintln on hijacked-DNS networks. Coverage on CI is unchanged. 5. **ExtensionManager test isolation** — the `extension_manager_with_process_manager_constructs` integration test passes `store: None`, which makes `list()` fall back to file-based `load_mcp_servers()` reading `~/.ironclaw/mcp-servers.json`. Any locally installed MCP server (e.g. notion) leaked into the test and broke the empty assertion. Set `IRONCLAW_BASE_DIR` to a fresh tempdir at the top of the test (before the LazyLock is initialized) to fully isolate. After this commit, `./scripts/dev-setup.sh` runs end-to-end and `cargo test` passes 4709/0 locally on macOS as well as CI. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(staging): address PR review feedback - helpers.rs: simplify IPv6 bracket stripping to `host.trim_matches(...)` per gemini-code-assist suggestion. Functionally equivalent to the chained strip_prefix/strip_suffix/unwrap_or but cleaner; for any host string from `Url::host_str()` (which only ever returns matched brackets), the result is identical. - setup/channels.rs: replace blocking `std::net::ToSocketAddrs` probe with `tokio::time::timeout(2s, tokio::net::lookup_host(...))` per Copilot review. The previous synchronous lookup could block a tokio worker thread inside `#[tokio::test]` and stall the suite on slow or offline DNS. The async resolver with a hard 2-second cap eliminates both risks. - module_init_integration.rs: drop the brittle `is_empty()` assertion and the env-var override entirely, addressing both Copilot's review comment about parallel test ordering and the reviewer's deeper concern about touching process env from an integration test that cannot access the crate-private ENV_MUTEX. The test's actual purpose is to verify that ExtensionManager constructs and `list()` returns Ok — that's exactly what `is_ok()` checks. The empty assertion was always brittle (the test creates empty TOOL/CHANNEL dirs but does not isolate ~/.ironclaw, so any locally installed MCP server leaks in) and trying to "fix" it by mutating IRONCLAW_BASE_DIR from inside an integration test introduces order-dependent behaviour that the user flagged: parallel tests in the same binary can race the LazyLock, and there is no integration-test-visible mutex to serialise env mutations. Removing the assertion is the principled fix. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* test(e2e): expand SSE resilience coverage * fix: address PR review follow-ups * refactor(web): consolidate duplicate chat_events_handler, improve docs - Remove duplicate chat_events_handler from server.rs; wire route to handlers::chat::chat_events_handler instead - Update from_sender doc comment to document boot_id/event-ID reset - Document that WebSocket (subscribe_raw) does not expose event IDs [skip-regression-check] Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * style: fix formatting from merge safety annotations [skip-regression-check] Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address PR #1897 review comments and CI formatting - auth.rs: use safe_truncate() for signature_prefix log (user-supplied string may contain multibyte UTF-8 that panics on byte-index slicing) - scripting.rs: add truncate_for_assert() helper, use it in both resource-limit test assertions (Python stdout may contain multibyte) - app.js: route plan_update through addTrackedEventListener so _lastSseEventId advances on plan_update frames (reconnect dedup) - store_adapter.rs: rewrite slug truncation with explicit ASCII-only byte search and fix formatting (rustfmt moved trailing safety comment onto a new line) [skip-regression-check] Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: ilblackdragon@gmail.com <ilblackdragon@gmail.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…re drift via pre-commit hook (#2065) * feat(i18n): add Korean translation, fix zh-CN drift, cover hardcoded strings Adds Korean (ko) as the third web UI language, brings zh-CN back into parity with en, converts ~80 hardcoded English strings in app.js into i18n keys, and installs a pre-commit hook that prevents future drift. ## Korean web UI - New `src/channels/web/static/i18n/ko.js` — full translation of all 663 keys, mirroring the structure of `en.js`/`zh-CN.js` - New `src/channels/web/server.rs` route `/i18n/ko.js` + handler - New language menu button in `index.html` - Browser auto-detect now special-cases `ko-*` (in addition to `zh-*`) so Korean visitors land on Korean by default - Toast label map in `i18n-app.js` becomes a small lookup table so the next language is a single-line addition ## zh-CN drift fix `zh-CN.js` was missing 9 keys that had been added to `en.js` after the Chinese pack was last touched (`config.telegramOpenBot`, `settings.tools`, and 7 keys under the `tools.*` namespace for the new Tool Permissions tab). Backfilled with Chinese translations so users on the Tools settings panel see proper labels instead of raw key strings. ## Hardcoded strings in app.js `app.js` had ~80 user-facing English string literals that bypassed `I18n.t()` entirely — toasts, confirms, alerts, button labels, meta-item labels for jobs/routines/missions detail panels, the theme dynamic label, dynamic auth states ("Connecting...", "Authenticated"), etc. These were invisible to the language switcher and would always render in English regardless of the user's choice. Replaced every literal with `I18n.t('key', { ...placeholders })` and added the corresponding ~95 new keys to `en.js`, `zh-CN.js`, AND `ko.js` in lockstep so all three packs stay at 663 keys with identical key sets and matching `{name}`-style placeholder tokens. Existing keys were reused where possible (`message.copy`, `approval.approved`, `connection.reconnected`, etc.). ## Pre-commit parity hook New `scripts/check-i18n-parity.sh` (pure POSIX bash, no Node) verifies: 1. No duplicate keys within any single language file 2. Every language has the same key set as `en.js` (the source of truth) 3. Placeholder tokens like `{name}`, `{count}` match across all languages — catches the silent bug where a translator drops an interpolation token Wired into both pre-commit hook install paths: - `scripts/pre-commit-safety.sh` (installed by `dev-setup.sh` as a symlink at `.git/hooks/pre-commit`; symlink is followed via `readlink` so the script location resolves correctly) - `.githooks/pre-commit` (used when devs set `git config core.hooksPath .githooks`) Both block the commit on failure with a clear error message and the `git commit --no-verify` escape hatch. Tested by deliberately removing a key from `ko.js` (caught) and stripping a `{path}` placeholder (caught). ## Korean README New `README.ko.md` — full Korean translation of `README.md`. Follows the layout of `README.ja.md` (6-item single-word ToC to keep anchors clean for non-Latin headings). All code blocks, image paths, and badge URLs preserved verbatim. `한국어` link added to the language switcher in all 5 READMEs (`README.md`, `.zh-CN.md`, `.ru.md`, `.ja.md`, and the new `.ko.md`). ## Verification - `./scripts/check-i18n-parity.sh` — `OK (663 keys × 3 languages)` - `node --check` clean on every modified JS file - Three-way parity: identical sorted key sets across en/zh-CN/ko, zero placeholder mismatches - Hook tested by removing/mutating keys and confirming the commit is blocked Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(i18n): address PR review feedback [skip-regression-check] Addresses 6 review comments on #2065. All changes are in src/channels/web/static/ (per .claude/rules/review-discipline.md exemption) plus a bash helper script — no Rust code is touched. ## scripts/check-i18n-parity.sh - **Portable mktemp** (Copilot): bare `mktemp` works on GNU but BSD/macOS `mktemp` requires an explicit template with at least 6 trailing X's. Wrap in a small `mktemp_file()` helper that always passes a template (`${TMPDIR:-/tmp}/check-i18n-parity.XXXXXX`) so the script runs on every platform. - **Symlink-attack-prone /tmp path** (gemini-code-assist): the placeholder-mismatch buffer was using `/tmp/i18n-ph-mismatch.$$`, which is predictable and vulnerable to symlink races in shared /tmp. Replace with `mktemp_file()` for consistency with the rest of the script. ## src/channels/web/static/app.js - **Hardcoded `'Mode'` label** (gemini): jobs detail meta-grid had `metaItem('Mode', job.job_mode)` — convert to `I18n.t('jobs.mode')` and add the new key to all 3 language packs. - **Hardcoded `'Yes'`/`'No'`** (Copilot): routine detail showed `routine.enabled ? 'Yes' : 'No'` even though the surrounding labels were translated. Reuse the existing `settings.on`/`settings.off` keys ("On"/"Off") which already render in all languages. - **Hardcoded `'N/A'`** (Copilot): mission detail showed `m.next_fire_at ? formatDate(...) : 'N/A'`. Reuse the existing `common.noData` key. Also fixed the same pattern in the TEE popover (`renderTeePopover`) where `'N/A'` was used as a fallback for three different attestation fields, since fixing the pattern across the file is the principled response per the repo's review-discipline rule. ## src/channels/web/static/i18n-app.js - **Hardcoded `LANG_LABELS` map** (gemini): the language-switch toast was reading from a per-call `{ 'en': 'English', 'zh-CN': '简体中文', 'ko': '한국어' }` literal that would grow with every new language and drift from the actual supported set. Move each language's own native name into its own pack under a new `language.name` key: en.js → 'language.name': 'English' zh-CN.js → 'language.name': '简体中文' ko.js → 'language.name': '한국어' Then the toast becomes `I18n.t('language.switch') + ': ' + I18n.t('language.name')` — both halves are read from the language pack that was just switched in, so the entire toast appears in the newly selected language. Adding a future language is now a single key addition with NO changes to i18n-app.js. ## Verification $ ./scripts/check-i18n-parity.sh i18n parity: OK (665 keys × 3 languages) $ cargo test --lib test result: ok. 4241 passed; 0 failed; 3 ignored Three-way parity preserved with the 2 new keys (`jobs.mode` and `language.name`) added to all three language packs in lockstep. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: allow telegram wasm channel name * test: use relative import in wasm setup tests * style: fix wasm setup test import order * test: use noop pairing store in wasm http test * fix(channels): reject reserved names during hot activation
) * fix(safety): add credential patterns and sensitive path blocklist Addresses critical credential leakage found by security testing (~/.ironclaw/tests/SECURITY_REPORT.md, test ce-02). Leak detector (crates/ironclaw_safety/src/leak_detector.rs): - Add OpenRouter API key pattern (sk-or-v1-<hex>) - Add Anthropic OAuth token pattern (sk-ant-oat<NN>-<base64url>) - Add Telegram bot token pattern (word-bounded, 8-12 digit bot ID) - Add Groq API key pattern (gsk_<alphanumeric>) - 12 new tests with synthetic keys (positive, false-positive, integration) File tools (src/tools/builtin/file.rs): - Add sensitive path blocklist to ReadFileTool, WriteFileTool, and ApplyPatchTool (defense-in-depth for all file access vectors) - Blocks: .env (and .env.local/.env.production/etc.), .ssh/, .aws/, .netrc, .pgpass, .npmrc, .pypirc, .docker/config.json, .kube/config, .git-credentials, .gcloud/, .config/gcloud/, .gnupg/, .vault-token, .ironclaw/secrets/ - Allows .env.example, .env.template, .env.sample (safe suffixes) - Case-insensitive; resolves symlinks via canonicalize() before check - 8 new tests covering blocking, safe suffixes, .env variants, case Known gap: shell tool can still `cat ~/.env` — different security domain (denylist-gated in autonomous mode, user-initiated in interactive mode). Tracked for follow-up. Note: new patterns use .unwrap() on Regex::new() matching the established convention of the 16 existing patterns in this file (all use // safety: hardcoded literal). Follow-up to address the existing .unwrap() debt across all patterns. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * Update src/tools/builtin/file.rs Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> * Update crates/ironclaw_safety/src/leak_detector.rs Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> * Update crates/ironclaw_safety/src/leak_detector.rs Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> * Update crates/ironclaw_safety/src/leak_detector.rs Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com> * fix(safety): address review feedback on credential patterns and path blocklist - Move is_sensitive_path to ironclaw_safety crate for shared use - Guard ListDirTool with sensitive path checks (including recursive traversal) - Add missing sensitive paths: ~/.config/gh/hosts.yml, /etc/shadow, ~/.terraform.d/credentials.tfrc.json, ~/.azure/ - Add path traversal regression test and ListDirTool blocking test Addresses review feedback from zmanian, gemini-code-assist, and copilot. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(tools): annotate sensitive dirs as blocked in recursive listing When ListDirTool's recursive traversal encounters a sensitive directory, annotate it with [sensitive - access blocked] so users understand why its contents are suppressed. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(tools): integrate is_sensitive_path into shell tool file-access checking Add defense-in-depth check for shell commands that read sensitive credential files (cat, head, tail, less, cp, etc.). Extracts file path arguments from known file-reading commands and checks them against the shared is_sensitive_path function from ironclaw_safety. This is best-effort — shell-level bypass via aliases, variable expansion, or encoding is still possible. Full mitigation requires filesystem-level sandboxing (seccomp/landlock). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(tools): add output redirection checks, fix segment splitting Address zmanian's re-review findings: - Check > and >> redirection targets against is_sensitive_path (write-path equivalent of read-path protection) - Replace single-char & splitting with proper &&/|| aware parser to avoid fragmenting double operators - Document subshell/command-substitution gap (partially covered by detect_command_injection upstream) - Extract helpers: split_shell_segments, check_segment_file_commands, check_redirect_target, expand_tilde - Add tests for output redirection, chained commands, segment splitting Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(safety): address PR #1675 review feedback and absorb #1713 patterns Absorb #1713's sensitive path patterns into ironclaw_safety crate: - Add shell history files, SSH key types, /etc/gshadow - Add sensitive file extensions (.pem, .key, .p12, .pfx, .jks, .keystore) - Add .dist safe suffix; smart .env matching (excludes .envrc, .environment) - Directory-level blocking for .aws/, .docker/, .kube/ (not just specific files) - Trailing-slash matching so bare directory paths trigger detection Shell tool hardening: - Strip surrounding quotes from tokens before sensitive path check - Add grep, awk, sed to FILE_READ_COMMANDS - Scan ALL redirect operators in a segment, not just the first Other fixes: - Add trailing \b to Telegram bot token regex to prevent over-matching - Update error messages to reference secret_list/secret_create - Strengthen ListDirTool test with tempfile-based .ssh directory Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(safety): close 4 adversarial bypass vectors in leak detector - Fix .env suffix check: use exact remainder matching instead of ends_with, so .env.production.dist is no longer allowed through - Detect process substitution <(...) in redirect checks and scan inner tokens for sensitive paths - Add missing /id_rsa to SENSITIVE_PATH_PATTERNS (other SSH key types were already present) - Check --flag=value tokens for sensitive paths instead of skipping all tokens starting with - Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * Fix 3 items from zmanian re-review on leak detector 1. Document blocking canonicalize() in async context: added comment explaining the trade-off (sub-ms on local FS, could block on NFS) with guidance to make async if needed. 2. Fix overly broad standalone key patterns: moved id_rsa, id_ed25519, id_ecdsa, id_dsa, authorized_keys, known_hosts from substring-based SENSITIVE_PATH_PATTERNS to exact filename matching via SENSITIVE_FILENAMES. This prevents false positives on paths like /project/grid_rsa_data while still blocking /project/test_fixtures/id_rsa. 3. Remove duplicate is_sensitive_path unit tests from file.rs: these belong in sensitive_paths.rs which already has comprehensive coverage. Kept integration-level tests that exercise tool execute() methods. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: j-bloggs <j-bloggs@users.noreply.github.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> Co-authored-by: gemini-code-assist[bot] <176961590+gemini-code-assist[bot]@users.noreply.github.com>
…M install (#2029) * fix(registry): use canonical underscore names in manifests to fix WASM install Manifest `name` fields used hyphens (e.g. "google-calendar") but the internal canonical form uses underscores ("google_calendar"). The release workflow packages .wasm files named after the manifest `name`, so archives contained "google-calendar.wasm". The extension manager canonicalized the name to "google_calendar" before extraction, looked for "google_calendar.wasm", and failed with "tar.gz archive does not contain 'google_calendar.wasm'". Two-part fix: - Update all 9 hyphenated manifest `name` fields and `_bundles.json` refs to use the canonical underscore form. Future releases will package archives with matching filenames. - Add hyphenated-name fallback in both tar.gz extractors so existing v0.22.0 release artifacts (which contain hyphenated filenames) remain installable. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * style: cargo fmt Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address PR review — extract shared helper, rename manifest files, improve errors - Extract `ArchiveFilenames` helper to `naming.rs` to deduplicate alias matching logic between `manager.rs` and `installer.rs` - Rename all 9 manifest JSON files to match their canonical underscore `name` fields (e.g. `google-calendar.json` → `google_calendar.json`) - Improve "not found" error messages to list both canonical and alias filenames that were tried - Update `test_extract_correct_wasm_from_tool_bundle` to use canonical `slack_tool` name matching current production path - Update artifact naming test script for renamed manifests Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* Use info colors for tool call summaries * Revert tool summary background and chevron color changes * fix: color
* fix(acp): propagate follow-up prompt failures as job errors (#1915) The follow-up loop silently swallowed ACP prompt failures — logging and posting a status event but continuing the loop, so the job always reported success: true. Extract the loop into a testable `run_follow_up_loop` function with trait-based dependency injection and return Err on failure so run() reports success: false. Also: downgrade follow-up loop info! logs to debug! (background task), add PartialEq to JobEventPayload, narrow AcpPromptSender to private. * fix(acp): kill child process on protocol failure, race poll against exit Address two review findings from #1981: 1. (gemini-code-assist) poll_prompt() blocked the follow-up loop — child exit wasn't detected during active HTTP polling. Wrap poll_prompt() in tokio::select\! alongside child_exit_rx. 2. (serrrfirat, high severity) Protocol failure could leave the child alive, causing stderr_handle.await to block forever and the job to hang. Add a kill channel: spawn_child_monitor() returns (child_exit_rx, kill_tx); on error, run_acp_session sends the kill signal before awaiting stderr so the pipe closes and cleanup completes. Extract spawn_child_monitor() as a standalone function so both production code and the regression test exercise the same code path. Tests: - follow_up_loop_exits_during_long_poll_when_child_dies (ForeverPromptSource + timeout guard — hangs without the select\! fix) - child_monitor_kills_process_so_stderr_reader_completes (real sleep subprocess + kill signal — hangs without the kill channel) * fix(acp): use non-terminal turn events, add poll retry limits Per-turn ACP results were emitting event_type "result" which maps to AppEvent::JobResult (the terminal signal). This caused job monitors to exit after the first prompt, making all subsequent follow-up outcomes invisible. Changed per-turn events to "turn_result" and emit exactly one terminal "result" from run() after the session ends. Also added retry discrimination for poll_prompt errors: permanent errors (OrchestratorRejected, LlmProxyFailed) fail immediately, transient errors (ConnectionFailed) retry with a cap of 5. --------- Co-authored-by: Rajul Bhatnagar <brajul@amazon.com>
* fix(tools): gate claude_code and acp modes behind enabled flags (#1987) CreateJobTool always exposed claude_code and acp as valid modes in its schema and silently accepted them at runtime, even when CLAUDE_CODE_ENABLED=false / ACP_ENABLED=false. This caused the LLM to sometimes select disabled modes, spawning containers that fail. - Add claude_code_enabled and acp_enabled flags to ContainerJobConfig (following the existing mcp_per_job_enabled pattern) - Expose via ContainerJobManager accessors, query from CreateJobTool through the already-injected job_manager - Dynamically build the mode enum in parameters_schema() — only show enabled modes; omit mode field entirely when only worker is available - Conditionally include agent_name field only when ACP is enabled - Add defense-in-depth guards in execute() rejecting disabled modes with ToolError::InvalidParameters - Add 10 regression tests covering schema gating and runtime rejection * fix(tools,web): address review feedback on mode gating (#2003) - Remove hardcoded "Set mode to claude_code" from CreateJobTool description; mode guidance is already provided dynamically via parameters_schema() - Add check_mode_enabled() guard in jobs_restart_handler to reject disabled modes on job restart via REST API, closing the bypass path - Add ContainerJobManager::is_mode_enabled(mode) to centralize mode validation - Clean up fully-qualified paths in jobs.rs with proper use imports - 5 regression tests (description, restart rejection, is_mode_enabled) * fix: harden mode gating with defense-in-depth and synchronous persistence - Add ModeDisabled variant to OrchestratorError and guard inside ContainerJobManager::create_job() so disabled modes are rejected even if callers forget to validate - Make job mode persistence synchronous instead of fire-and-forget to prevent silent mode loss on transient DB errors (restarts would silently downgrade to worker mode) - Refactor parameters_schema() to build serde_json::Map directly, removing an unreachable if-let guard and an .expect() call Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Rajul Bhatnagar <brajul@amazon.com> Co-authored-by: serrrfirat <f@nuff.tech> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(web): emit Done status after response to fix SSE ordering (#2079) Move the terminal "Done" status out of thread_ops and emit it only after the gateway successfully responds via a new respond_then_done() helper in agent_loop. This guarantees the browser receives the assistant message before the turn-closing event, preventing the web UI from appearing stuck. Adds a regression test asserting the response event is captured before the Done status in the ordered event log. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(web): add frontend safety net for lost SSE response events (#2079) Track whether a `response` SSE event was received for the current turn. When "Done" arrives without a preceding response, schedule a loadHistory() call after 1500ms so the user sees the answer even if the response event was lost to broadcast lag or a brief disconnect. This is the second prong of the fix described in #2079 — the backend ordering fix alone prevents the race, but this fallback handles residual edge cases (proxy buffering, SSE reconnection gaps). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address review findings — Done on all paths, fix frontend timer leaks Backend: - Send Done status when BeforeOutbound hook blocks the response, so the client still knows the turn is complete. - Send Done status for empty/suppressed responses (e.g. approval handled via send_status) to match pre-refactor behavior. Frontend: - Set _turnResponseReceived on stream_chunk events so streaming responses don't trigger a spurious loadHistory() when Done arrives. - Clear _doneWithoutResponseTimer on sendMessage() to prevent stale timers from a previous turn firing during the new one. - Clear turn-tracking state on switchThread() to prevent cross-thread contamination of the timer and flag. - Clear turn-tracking state on SSE reconnect (eventSource.onopen) to prevent stale timers from before the disconnect. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address PR review — always emit Done, extract helper, add test - respond_then_done now emits Done regardless of respond outcome so the client always knows the turn ended, even on delivery failure - Extract send_done() helper to deduplicate the inline Done+warn blocks in the hook-blocked and empty-response paths - Add done_emitted_for_empty_response test covering the empty-response branch ordering invariant - Lift 1500ms magic number to DONE_WITHOUT_RESPONSE_TIMEOUT_MS constant - Add comment explaining _turnResponseReceived single-thread tracking Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(agent): suppress Done while awaiting approval; introduce HandleOutcome Distinguish "no response, turn complete" from "no response, turn paused" in handle_message's return type so the run loop can decide whether to emit the terminal Done status. The previous code lumped both into Ok(Some("")), causing v1 NeedApproval to incorrectly emit Done after ApprovalNeeded — which then tripped the new web UI safety net and triggered a spurious loadHistory() under the live approval prompt. - New HandleOutcome enum with Shutdown / Respond / NoResponse / Pending - SubmissionResult::NeedApproval now maps to HandleOutcome::Pending - Bridge handlers wrapped via HandleOutcome::from_legacy (their approval flows return non-empty descriptive text, so they never need Pending) - Regression test no_done_emitted_while_awaiting_approval drives a v1 Always-approval probe and asserts no Done is captured - Repaired pre-existing done_emitted_for_empty_response test, which asserted the wrong invariant: the dispatcher substitutes empty LLM responses with a fallback message, so a truly empty response never reaches the run loop. Renamed and updated to assert the ordering. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> Co-authored-by: ilblackdragon@gmail.com <ilblackdragon@gmail.com>
* feat(slack): implement on_broadcast and fix message tool channel hints Implement the on_broadcast callback for the Slack WASM channel, enabling proactive message delivery to Slack channels/users via the message tool. Previously this was a stub returning "not implemented". The implementation: - Uses the user_id parameter as the broadcast target (channel ID or user ID) - Strips leading # from targets for convenience - Warns when target doesn't look like a Slack ID (C/U/D/G prefix) - Posts via chat.postMessage with host-injected Bearer token - Tracks active threads for broadcast replies (consistent with on_respond) Also fixes the message tool's channel parameter description to list 'slack' alongside 'slack-relay' and clarifies that Slack targets must be IDs. [skip-regression-check] Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(slack): extract shared post helper, harden broadcast validation Address review findings on the slack broadcast implementation: - Extract `post_slack_message()` shared by `on_respond` and `on_broadcast`, eliminating ~40 lines of duplicated HTTP-call-then-parse logic. - Log `track_active_thread` errors at Warn level instead of silently swallowing them with `let _ =` (restores observability lost in original). - Make non-ID broadcast targets a hard error instead of a soft warning — consistent with the message tool schema that says "must be an ID, not a name". - Fix empty-target error message to not assume a name was provided. - `resolve_broadcast_target` now returns `&str` (avoids allocation). - Add 2 tests covering the resolve+validate pipeline end-to-end. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(slack): track broadcast message ts so replies are recognized Address Gemini review: broadcast messages now track the Slack-returned timestamp as an active thread (falling back from response.thread_id to the posted message ts). This ensures that if a user replies to a broadcast, the agent recognizes the reply as an active thread. Also fix stale doc comment on looks_like_slack_id (was missing W prefix). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(staging): repair 4 categories of CI test failures
1. Telegram token test flaky race: add env mutex guard so concurrent
tests that override IRONCLAW_TEST_TELEGRAM_API_BASE_URL don't
pollute the unguarded read in the colon-preservation test.
2. SSE/connection E2E tests: #sse-status element was removed from HTML
and replaced with #sse-dot colored indicator. Update tests to check
the dot's CSS class instead of text content. Add SSE-ready wait to
the page fixture so chat tests don't race against connection setup.
3. Tool approval E2E tests: API unified legacy pending_approval and
engine v2 gates into a single pending_gate response field. Update
all E2E test helpers to use the correct field name.
4. WASM tar.gz extraction bug: canonicalized extension names use
underscores (web_search) but release archives use hyphens
(web-search.wasm). Accept both filename forms when extracting.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: address review feedback — stronger SSE signal, consistent naming
- SSE wait: use sseHasConnectedBefore JS flag (set in onopen) instead
of checking #sse-dot CSS class, which defaults to connected state
before SSE actually connects
- Rename _wait_for_pending_approval → _wait_for_pending_gate and
_wait_for_no_pending_approval → _wait_for_no_pending_gate
- Update all docstrings/error messages to say pending_gate
- Deduplicate name.replace('_', '-') in tar.gz extraction and include
both accepted filenames in the error message
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: address second round of review feedback
- Add comment documenting invariant: canonical names use underscores,
archives may use hyphens, reverse is not supported
- Fix quote wrapping in single-name error case
- Remove dead SEL["sse_status"] selector from helpers.py
- Simplify SSE wait: use window.sseHasConnectedBefore === true
(fails fast on rename instead of silent 10s timeout)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: use global binding for sseHasConnectedBefore, not window property
sseHasConnectedBefore is declared with let at global scope, which
does not create a window property. window.sseHasConnectedBefore
would always be undefined. Use typeof guard + direct reference instead.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(security): cap tar.gz entry pre-allocation to MAX_ENTRY_SIZE
The tar header's declared size is attacker-controlled. Without capping,
Vec::with_capacity could attempt a huge allocation and OOM before the
read_to_end take() limit kicks in.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* style: cargo fmt
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: make test_sse_status_shows_connected non-redundant, document flag reset
- test_sse_status_shows_connected now checks #sse-dot CSS class (visual
indicator) instead of re-checking sseHasConnectedBefore which the
page fixture already guarantees
- Add comment to test_sse_reconnect_after_disconnect explaining why
sseHasConnectedBefore is reset and that the history-reload path is
covered by test_sse_reconnect_preserves_chat_history
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* Improve channel onboarding and Telegram pairing flow * fix: remove dead restart_required code, fix review findings, and stabilize polling E2E test - Remove restart_required/needs_restart dead code from 6 files (no real extension uses it; all channels hot-activate at runtime) - Remove dead extensions.configuredRestart i18n key from all 3 locales - Fix pairing test asserting wrong upsert semantics (test expected idempotent behavior but impl always rotates codes) - Fix pairing test using expired code for approval (req.code -> req_again.code) - Fix missing i18n fallback for auth.extensionTokenPlaceholder - Validate setup_url scheme (https?://) before assigning to <a>.href - Replace hardcoded English "Approve"/"Pairing code is required" with i18n keys - Demote misleading "bot is open to all users" Telegram log from Warn to Debug - Move polling E2E test to run first (polling loop dies during refresh_active_channel) - Add poll_interval_ms config field to Telegram WASM channel - Fix conversations.rs compilation (missing ? operator) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address remaining review comments (i18n regressions) - Remove hardcoded pairing_instructions() function from server.rs; use onboarding metadata from ExtensionManager instead (fixes i18n regression where pairing instructions were always English) - Restore i18n calls for stepper labels in renderWasmChannelStepper (was using hardcoded English strings instead of missions.step* keys) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: restart polling loop on channel refresh, address Copilot review - Add WasmChannel::ensure_polling() that stops any stale polling task and starts a fresh one from the on_start config - Call ensure_polling() in refresh_active_channel after re-running on_start, fixing the root cause of the dead polling loop in E2E tests - Move polling test back to its original position (no longer order-dependent) - Fix requires_pairing in channel_onboarding_for_state to use channel_requires_pairing() instead of legacy owner_id-only check - Add rel='noopener noreferrer' to all setup_url target=_blank links Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: collapse nested if-let to satisfy clippy collapsible_if Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: return promise from approvePairing, stop polling unconditionally in ensure_polling - Add missing `return` before apiFetch in approvePairing() so callers can await/chain the result - Move poll_shutdown_tx.take() before the enabled check in ensure_polling() so switching from polling to webhook stops the old polling task Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: remove trailing commas in JSON test fixtures after restart_required removal Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(web): intercept approval text input ("yes"/"no"/"always") in chat
When a tool requires approval in the web UI, typing "yes", "no", or
"always" in the chat input now resolves the approval card directly
instead of sending a regular message. This prevents duplicate approval
prompts and "No pending approval" errors that occurred when text went
through the backend message pipeline.
The frontend intercepts approval keywords in sendMessage() and routes
them through sendApprovalAction() — the same code path as clicking
the Approve/Deny/Always buttons on the card.
[skip-regression-check]
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: find most recent unresolved approval card for text interception
Address review feedback: instead of checking the last card and then
separately checking if it's resolved, find the most recent unresolved
card directly. Handles the edge case where the last card is resolved
(during 1.5s removal animation) but an earlier one isn't.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* test: add E2E test for skip-resolved-card behavior
Addresses review feedback: adds a test where two approval cards are
visible, the newer one is resolved via button click, then typing "yes"
correctly targets the older unresolved card instead of falling through.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…0555 chore: promote staging to staging-promote/5083aed4-24002418644 (2026-04-06 06:33 UTC)
3b55836
into
staging-promote/733678dd-23996777140
…4002418644 chore: promote staging to staging-promote/c551c30b-23996777140 (2026-04-05 13:23 UTC)
Auto-promotion from staging CI
Batch range:
a55aff980a4e235590c3af57ded2542512e2f9f6..5083aed462fa5659472dc92e41209d46a39d0f95Promotion branch:
staging-promote/5083aed4-24002418644Base:
staging-promote/733678dd-23996777140Triggered by: Staging CI batch at 2026-04-05 13:23 UTC
Commits in this batch (27):
Current commits in this promotion (2)
Current base:
staging-promote/733678dd-23996777140Current head:
staging-promote/5083aed4-24002418644Current range:
origin/staging-promote/733678dd-23996777140..origin/staging-promote/5083aed4-24002418644Auto-updated by staging promotion metadata workflow
Waiting for gates:
Auto-created by staging-ci workflow