feat(gateway): add attachment flows, v2 skill install coverage, and e2e stabilization - #2385
Conversation
There was a problem hiding this comment.
Code Review
This pull request implements a robust file attachment system that supports various document types, persists them to project-local storage, and indexes them for the engine. It also enhances skill management by supporting multi-file bundles from ZIPs and GitHub, and introduces a unified slash-command autocomplete. Feedback from the review suggests refining the slash command regex to prevent accidental matches with file paths and ensuring consistent MIME type handling and XML decoding between the frontend and backend. Specifically, support for docx and xlsx should be added to the frontend inference logic, and MIME mapping should be unified to prevent inconsistencies in file naming.
serrrfirat
left a comment
There was a problem hiding this comment.
Security and code quality review -- 12 findings (2 Critical, 4 High, 6 Medium). Inline comments below.
There was a problem hiding this comment.
Pull request overview
Adds end-to-end coverage and implementation support for gateway file attachments, v2 skill install/routing behavior, and stabilizes the E2E/browser harness across multi-tenant vs single-tenant gateway behavior.
Changes:
- Introduces gateway “inline attachment” flows (frontend UI + API plumbing) and extends agent/engine handling to preserve attachment context and persist project-local copies.
- Enhances v2 skill install lifecycle (ZIP bundle installs, GitHub repo installs, metadata propagation) and improves v2 skill activation/routing semantics.
- Hardens E2E stability (process teardown, mock resets, auth/approval flows, Slack/Telegram fixtures) and expands scenario coverage (attachments, slash autocomplete, OAuth matrices).
Reviewed changes
Copilot reviewed 58 out of 59 changed files in this pull request and generated 6 comments.
Show a summary per file
| File | Description |
|---|---|
| tests/engine_v2_skill_codeact.rs | Updates v2 skill metadata test fixture with new fields. |
| tests/e2e/scenarios/test_widget_customization.py | Adds isolated single-tenant gateway fixture and adjusts customization assertions for multi-tenant behavior. |
| tests/e2e/scenarios/test_v2_kernel_auth_preflight.py | Improves subprocess teardown and pins mock API URL per test. |
| tests/e2e/scenarios/test_v2_kernel_auth_gateway_flow.py | Improves subprocess teardown and pins mock API URL per test. |
| tests/e2e/scenarios/test_v2_engine_oauth_google.py | Improves teardown and adds autouse mock URL pinning for OAuth tests. |
| tests/e2e/scenarios/test_v2_engine_error_handling.py | Improves subprocess teardown to reduce flakiness. |
| tests/e2e/scenarios/test_v2_engine_auth_cancel.py | Improves teardown and pins mock API URL per test. |
| tests/e2e/scenarios/test_v2_engine_approval_flow.py | Improves subprocess teardown to reduce flakiness. |
| tests/e2e/scenarios/test_v2_auth_oauth_matrix.py | Stabilizes OAuth matrix tests and improves extension name handling and REPL behavior. |
| tests/e2e/scenarios/test_tool_approval.py | Uses a writable-chat helper and strengthens thread scoping checks. |
| tests/e2e/scenarios/test_telegram_e2e.py | Stabilizes Telegram tests (update IDs, richer diagnostics, isolated server for polling). |
| tests/e2e/scenarios/test_sse_reconnect.py | Stabilizes refresh/history expectations by persisting history before reload. |
| tests/e2e/scenarios/test_slack_e2e.py | Completes Slack DM pairing during activation and stabilizes URL verification test. |
| tests/e2e/scenarios/test_skill_oauth_flow.py | Adjusts auth-required trigger and broadens auth indicator matching. |
| tests/e2e/scenarios/test_routine_event_batch.py | Uses writable-chat helper to avoid read-only thread flake. |
| tests/e2e/scenarios/test_ownership_model.py | Adds guard to ensure chat input is writable before sending. |
| tests/e2e/scenarios/test_owner_scope.py | Adds guard to ensure chat input is writable before sending. |
| tests/e2e/scenarios/test_chat.py | Adds attachment E2E coverage and slash autocomplete coverage. |
| tests/e2e/scenarios/test_auth_no_duplicate_response.py | Improves teardown and pins mock API URL per test. |
| tests/e2e/mock_llm.py | Extends mock LLM to support skill install flows, richer tool summaries, and attachment-aware triggers. |
| tests/e2e/helpers.py | Adds selectors + ensure_writable_chat_input helper used across tests. |
| tests/e2e/conftest.py | Adds mock reset fixtures, improves subprocess teardown, and introduces isolated Telegram server fixture. |
| tests/e2e/CLAUDE.md | Updates scenario documentation for new chat/attachment coverage. |
| tests/e2e_attachments.rs | Adds v2 channel attachment persistence coverage and a CWD guard for tests. |
| src/tools/builtin/skill_tools.rs | Adds skill bundle ZIP install support, GitHub repo installs, and dependency install support for bundles. |
| src/llm/transcription/mod.rs | Updates attachment struct initialization for new field. |
| src/llm/rig_adapter.rs | Adds ImageDetail mapping and tests for detail defaults. |
| src/document_extraction/mod.rs | Updates attachment struct initialization for new field. |
| src/channels/web/ws.rs | Adds websocket support for non-image attachments and tests forwarding. |
| src/channels/web/types.rs | Adds AttachmentData to send message requests and extends SkillInfo response fields. |
| src/channels/web/server.rs | Adds attachment decoding, MIME→ext mapping, and request logging updates. |
| src/channels/web/handlers/skills.rs | Enriches skills API with usage/setup hints and install metadata exposure. |
| src/channels/web/CLAUDE.md | Updates web API docs to reflect attachment uploads and body limit intent. |
| src/channels/wasm/wrapper.rs | Propagates new local_path attachment field through WASM channel plumbing. |
| src/channels/wasm/host.rs | Adds local_path to WASM host attachment struct and tests. |
| src/channels/tui.rs | Updates attachment struct initialization for new field. |
| src/channels/http.rs | Updates attachment struct initialization for new field. |
| src/channels/channel.rs | Adds local_path to core IncomingAttachment. |
| src/bridge/skill_migration.rs | Syncs v1 skills into v2 store with new metadata fields and adds a public sync helper. |
| src/bridge/router.rs | Persists incoming attachments to project-local storage and indexes them into v2 notes; improves v2 auth pending handling. |
| src/bridge/effect_adapter.rs | Syncs installed skills into v2 docs and improves latent extension readiness behavior. |
| src/bridge/auth_manager.rs | Adds execution-oriented readiness path that can explicitly activate latent providers. |
| src/agent/thread_ops.rs | Runs safety/policy/secret scanning against attachment-augmented content and allows attachment-only sends. |
| src/agent/mod.rs | Re-exports attachment augmentation and bridge pending sentinel. |
| src/agent/attachments.rs | Adds project-path metadata to rendered attachment markup and tests for rendering. |
| src/agent/agent_loop.rs | Adds a bridge “pending” sentinel to correctly represent paused states. |
| crates/ironclaw_skills/src/v2.rs | Extends v2 skill metadata with bundle path and source URL. |
| crates/ironclaw_skills/src/registry.rs | Adds bundle install support (extra files + install metadata) and improves delete semantics. |
| crates/ironclaw_gateway/static/style.css | Adds styling for attachment previews and message attachment rendering. |
| crates/ironclaw_gateway/static/index.html | Expands file input accept list and updates attach button copy/labels. |
| crates/ironclaw_gateway/static/i18n/zh-CN.js | Adds new attachment-related i18n strings. |
| crates/ironclaw_gateway/static/i18n/ko.js | Adds new attachment-related i18n strings. |
| crates/ironclaw_gateway/static/i18n/en.js | Adds new attachment-related i18n strings. |
| crates/ironclaw_gateway/static/app.js | Implements attachment staging/validation/rendering, per-user history parsing, approval gate hydration, and slash skill autocomplete. |
| crates/ironclaw_engine/src/runtime/mission.rs | Updates skill metadata construction for new fields. |
| crates/ironclaw_engine/src/memory/skill_tracker.rs | Updates skill metadata construction for new fields. |
| crates/ironclaw_engine/orchestrator/default.py | Improves v2 skill selection/scoring and explicit /<skill> routing behavior. |
| .gitignore | Ignores .ironclaw/ in repo working copies. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 58 out of 59 changed files in this pull request and generated 4 comments.
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
…2e stabilization (nearai#2385) * feat(gateway): add attachment flows and slash-skill coverage * feat(v2): persist project attachments across channels * feat(skills): install GitHub skill bundles * feat(v2): cover live skill install and setup flow * test(e2e): stabilize gateway and auth coverage * test(e2e): stabilize post-merge warnings and browser flows * fix(review): address follow-up PR feedback * fix(review): address remaining attachment and skill install comments * Address remaining attachment review comments * fix(ci): allowlist ws.rs → server::inline_attachments_to_incoming ws.rs was already allowlisted for the attachment shim symbols (`images_to_attachments`, the rate limiter types, etc.) so the new unified entrypoint added by this branch (combining images and generic attachments before validation) follows the same pattern. The entry will be removed together with the rest of the ws.rs server:: block once the attachment helpers migrate into platform/. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(e2e): attachment persistence path and Slack activate signature Two e2e-surfacing regressions after merging staging: 1. `persist_project_attachments` was writing to `<base_dir>/projects/.ironclaw/attachments/...` because PR nearai#2385's reviewer-requested switch from `std::env::current_dir()` to an explicit `project_root` kept the `.ironclaw/` prefix baked into `PROJECT_ATTACHMENT_DIR` while rooting at `ironclaw_base_dir()/projects`. Point `resolve_project_root()` at the parent of the base dir so `<parent>/.ironclaw/attachments/<owner>/<project>/...` matches the prompt's `project_path` and the user's expectation when base dir is `~/.ironclaw`. Updates the corresponding assertion in test_v2_engine_auth_flow.py to resolve paths against the fixture's home tempdir instead of the repo root. 2. `activate_slack()` grew a required `http_url` arg during the skill-install branch work but the `active_slack` fixture in test_slack_e2e.py still passed the old three-arg shape. That tripped every Slack scenario at setup (TypeError). Thread `http_url` from `slack_e2e_server` through the fixture. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(engine-v2): auth-prompt surfacing, bundle_path injection, attachment-only inputs - Orchestrator formatter now writes `Installed bundle path on disk:` into each skill block so the skill body sees the bundle location it needs to reference (e.g. running `pip install -r <bundle>/requirements.txt`). Previously the bundle_path metadata field was populated but never surfaced into the prompt, so skills that rely on filesystem paths silently no-op'd. - The router no longer rejects messages whose text body is empty when the payload carries attachments. Safety validation's empty-input guard is a v1 input-sanity check; a pure-attachment follow-up (image upload with no caption) is a legitimate submission in the v2 gateway contract and previously tripped "Input cannot be empty". - The engine auth-flow e2e tests now detect gate-paused state via `HistoryResponse.pending_gate` (and `resume_kind.Authentication`) rather than scanning the turn response text for "paste your token". Auth instructions live in the `onboarding_state` SSE event, not in the chat response (see `test_auth_no_duplicate_response.py`); the old string-matching assertion was checking the wrong surface. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(e2e): switch approval/auth-prompt probes to pending_gate Approval and auth prompts are surfaced through HistoryResponse.pending_gate and the onboarding_state/gate_required SSE events, not as text in turns[-1].response — the duplicate-response regression guard in test_auth_no_duplicate_response.py explicitly forbids them from appearing in the chat transcript. Update the helpers in test_v2_engine_approval_flow.py, test_v2_engine_auth_cancel.py, and test_v2_kernel_auth_preflight.py to poll pending_gate instead of scanning turn text for "requires approval" or "paste your token". Unblocks 5 approval, 1 auth-cancel, and 3 preflight tests. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(e2e): google-oauth _wait_for_auth_prompt / _wait_for_response use pending_gate Bring the Google Drive / skill-OAuth regression file in line with the rest of the v2 e2e helpers: poll `HistoryResponse.pending_gate` for auth/approval prompts, and accept a pending_gate as a valid terminal state for `_wait_for_response` (an auth-retry chain that hits another gate is still progress, not a hang). Unblocks the oauth-cancel, invalid-token-paste, and api-key-then-api-call scenarios; the lingering token-refresh scenario still exposes a real v2 auto-refresh regression (the engine prompts the user instead of issuing a refresh against the stored refresh_token). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(e2e): relax a few stale v2-surface assertions - `test_skill_oauth_flow::test_auth_required_sse_event` was pinned to the old `onboarding_state/auth_required` SSE payload. The v2 gate pipeline delivers credential gates as `gate_required` (resume_kind `Authentication`) or, when preflight falls through to approval first, `approval_needed`. Accept any of those three, and treat a `thinking` "Running <tool>" status as evidence the tool call fired when no standalone `tool_started` event is emitted. - `test_message_persistence` helpers asserted HTTP 200 on `/api/chat/send`, but the gateway now returns 202 ACCEPTED (fire-and-forget). Accept both. - `test_project_detail` flipped the wrong global (`engineV2`) instead of `engineV2Enabled`, leaving the `data-v2-only` Projects tab hidden so the click timed out. Set the real flag. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(review): address attachment index note correctness Two Copilot review findings on the attachment persistence path: - `attachment_index_note` in `src/bridge/router.rs` used the raw user-supplied filename in the markdown `# Uploaded attachment:` header and in the memory-doc `title` field. A filename with newlines / backticks / control characters would corrupt the agent-visible transcript and break searchable titles. Route the filename through a new `sanitize_filename_for_display` that strips control chars, collapses newlines/tabs to spaces, swaps backticks for apostrophes, truncates at 256 chars, and falls back to `"attachment"` when the sanitized result is empty. - `persist_project_attachments` cleared `attachment.data` before calling `attachment_index_note`, so the `size_bytes.unwrap_or( data.len() as u64)` fallback reported `0` bytes whenever the channel hadn't pre-populated `size_bytes`. Swap the order — build the index note while the buffer is still populated, then drop the bytes. Also adjust `src/agent/attachments.rs::format_attachment` for the Image arm: when `data` has been cleared but `local_path` is set (the engine-v2 persist-then-clear flow), the "visual content not available in this conversation" message is misleading — the image is available, just on disk. Surface a dedicated prompt that tells the agent to reference the project file path instead of trying to load bytes from memory. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(e2e): cancel_during_auth asserts pending_gate clears, not chat text The test polled \`turns[-1].response\` for "cancel" but the cancel flow never writes an assistant row to the chat-history DB: resolve_gate returns \`BridgeOutcome::Respond("Cancelled.")\` which broadcasts via SSE and calls \`stop_thread\` on the engine thread, neither of which goes through the DB persistence path that populates turn responses. Switch the test to verify the user-visible signal the gateway actually emits — \`history.pending_gate\` disappears after "cancel" resolves the gate. Matches the approach used in \`test_v2_engine_approval_flow.py\`'s deny-flow tests. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(e2e): pairing approve test tolerates ExtensionName boundary reject Staging's new `features/pairing/` slice (ironclaw#2599 stage 4b) validates the `{channel}` URL segment through `ExtensionName::new` at the handler boundary: a path-traversal / control-character / whitespace-containing segment (like `evil.Ignore all`) now returns 400 instead of silently routing to a pairing-store miss. The regression test used to assert the older 200+JSON shape. Relax it to accept either 200 (generic `Invalid or expired pairing code.`) or 400 (boundary validation); the real invariant the test exists to protect — the raw injection-shaped channel string must not echo back into the response — is still asserted. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(review): preserve image bytes through LLM call + document drive mock pin Two review findings: - `src/bridge/router.rs::persist_project_attachments` was clearing `attachment.data` after writing the file to disk. The very next step in `handle_with_engine_inner` is `augment_with_attachments`, which only emits a multimodal `image_parts` entry when `att.data` is non-empty — so every engine-v2 image upload was silently dropped from the LLM request even though the file landed on disk. The `persisted_attachments` Vec is local to the dispatch and is dropped as soon as the engine call returns, so the "storage hygiene" comment the clear used to justify was a no-op. Stop clearing; let RAII free the bytes. Updates `src/agent/attachments.rs`'s Image-arm prompt to reflect the refined invariant (`data.is_empty()` now implies a downstream caller or channel stripped the buffer, not the normal persist path). - `tests/e2e/scenarios/test_v2_engine_oauth_google.py::_pin_mock_drive_api_url` posts to `/__mock/set_github_api_url`. The wire name is historical — the Drive suite reused the knob — but the fixture name made the intent hard to follow. Adds a docstring that calls out the shared `_github_api_url` in `mock_llm.py` and explains why the endpoint rename would cascade into every other test that uses it. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(review): address remaining Copilot feedback on PR 2385 - audio attachments: include `mime` (and size) attribute in `<attachment>` XML for parity with image/document so the frontend can render MIME and size in attachment cards - /api/skills list/search: parallelize per-skill filesystem I/O (`read_install_metadata`, `try_exists`, `metadata`) via `futures::future::join_all` instead of awaiting serially — keeps the handler O(n) in wall time for large skill sets - history parseUserMessageContent: only strip the trailing `<attachments>…</attachments>` block when at least one `<attachment>` tag is parsed from inside it, otherwise leave the raw text intact so user messages that legitimately end with that markup are preserved - sync_v1_skill_to_store: look up existing shared skill doc via `list_skills_global()` instead of `list_shared_memory_docs(project_id)` so shared skills installed under one project are updated in place when re-synced from another project (prevents duplicate shared docs across per-user projects) and preserve the original `project_id` on in-place update Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(tests): close the staging test backlog — rust suite green, e2e 14→4 A pass over staging turned up 12 rust test failures and 14 playwright e2e failures + 1 fixture error. Most were wiring/invariant drift or stale test expectations around engine v2. This patch cleans up the ones with clear root causes. Rust (12 → 0): - `tools::builtin::skill_tools` (8 tests): ripped out hand-rolled ZIP byte blobs that were missing the EOCD record since the extractor switched to `zip::ZipArchive::new` in nearai#2385. Tests now build through `zip::ZipWriter`, matching the production path. Drops the obsolete nested-path assertion whose assumption conflicts with intentional GitHub-archive root stripping. - `extensions::manager::test_telegram_token_colon_preserved_in_validation_url`: `src/pairing/approval.rs::propagate_approval_restores_runtime_state_when_on_start_fails` was mutating the `IRONCLAW_TEST_TELEGRAM_API_BASE_URL` runtime-env overlay without holding `ENV_MUTEX`. Now acquires `lock_env()` so concurrent readers see a stable value. - `bridge::router::handle_with_engine_persists_attachment_files_and_indexes_them`: two distinct `ENGINE_STATE_TEST_LOCK` statics (one in `test_support`, one in the sibling `tests` module) meant cross-module tests raced on the shared `ENGINE_STATE` `OnceLock`. Replaced the private duplicate with `use super::test_support::ENGINE_STATE_TEST_LOCK`. - `e2e_attachments::engine_v2_channel_attachments_persist_for_telegram_and_whatsapp`: attachment persistence resolves paths through the cached `bootstrap::ironclaw_base_dir()`, not the test's tempdir CWD. Added `bridge::override_engine_project_root_for_test` and wired the test to use it. - `telegram_auth_integration::test_group_message_emits_chat_type_metadata`: local fix — rebuild `channels-src/telegram` so the WASM picks up the April-17 `chat_type` emit from nearai#2513. CI rebuilds the module per run, so no binary committed here. Playwright (14 failed + 1 error → 4 failed + 1 error): - `test_chat.py::test_gateway_attachment_flow_renders_thread_and_reaches_llm` and the unextractable variant: a legacy change listener on `#image-file-input` fired before the unified `handleAttachmentFiles` path, cleared `e.target.value`, and left the FileList empty by the time the unified handler ran. Removed the duplicate wiring in `crates/ironclaw_gateway/static/js/surfaces/chat.js`. - `test_chat.py::test_slash_autocomplete_shows_commands_and_skills`: `SLASH_COMMANDS` never merged installed skills. Added `refreshSlashSkillEntries()` that fetches `/api/skills` on menu open and re-runs the filter once the skills land. - `test_pending_user_messages.py::test_pending_message_survives_sse_reconnect`: the SSE open handler only reloads history when `disconnectMs > SSE_RELOAD_THRESHOLD_MS`; the test's instant reconnect skipped that. Ages `_sseDisconnectedAt` past threshold. - `test_pending_user_messages.py::test_welcome_card_hidden_when_pending`: `_create_new_thread` returned `currentThreadId` before the new-thread API round-trip set it, so callers got the pre-click id and keyed `_pendingUserMessages` on the wrong thread. Now waits for the id to change. - `TestV2EngineSkillInstallFlow` (7 → 2 failures): - Skill card template didn't render `usage_hint`, `has_requirements`, `has_scripts`, or `install_source_url`. Extended `renderSkillCard` in `surfaces/skills.js`. - The deny message `"Do not execute it; choose an alternative approach"` accidentally matched `user_signals_execution_intent`'s EXEC_PHRASES ("execute it"), re-arming `require_action_attempt` on resume and nudging the LLM into another tool call. Rephrased to `"Do not retry; choose a different approach"` in `src/bridge/router.rs`. Partial progress (still failing, needs deeper engine-v2 work): - `test_v2_engine_oauth_google::test_oauth_token_refresh_on_expiry`: added an `oauth:` block to the test's `google_drive` skill (which registers a refresh config via `credential_spec_to_oauth_refresh`) and aligned `GOOGLE_OAUTH_CLIENT_ID` with the mock proxy's expected `hosted-google-client-id`. Thread still hits the auth gate instead of refreshing — the pre-flight path isn't reaching `oauth_refresh_for_secret("google_drive_token")`; needs instrumentation on the engine-v2 gate pipeline. Net: rust suite green, playwright 4 failures left (2 skill-install approval-flow edge cases, 1 OAuth refresh, 1 REPL auth that flakes only under full-suite load) + 1 restart-fixture health-check timeout that flakes under 20-min suite pressure. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(tests): close the final 4 e2e failures + add copy-button coverage Follow-up to the earlier staging test pass. Drives the remaining playwright failures to green and adds the missing test for the per-message Copy button. New coverage: - `test_chat.py::test_message_copy_button_writes_raw_text`: clicking the per-message Copy button writes the raw text (user turn) or the raw markdown (assistant, via `data-raw`) to navigator.clipboard and flashes the button label to "Copied!" then back to "Copy". The existing `test_copy_from_chat_forces_plain_text` only covered the Cmd+C selection handler, so a regression to the button path was invisible. Fixes: - `TestV2EngineSkillInstallFlow::test_implicit_skill_activation_works_immediately_after_install`: the pika skill manifest uses the legacy `metadata.openclaw.requires` shape without a top-level `activation:` block, so `score_skill` scored 0 for every prompt and the skill never activated unless the user typed `/pikastream-video-meeting`. "Please use pikastream-video-meeting to prepare this call" should activate just like the slash form. `score_skill` now treats the skill name (and the hyphen/underscore-normalized form) as an implicit keyword, gated at ≥4 chars so short generic names don't false-match. `test_installed_skill_does_not_overfire_on_unrelated_prompt` still passes — a grocery-list prompt doesn't accidentally trigger pika. - `TestV2EngineSkillInstallFlow::test_duplicate_install_is_idempotent_and_keeps_single_card`: the test was waiting for an approval card on the second install, but `SkillInstallTool::requires_approval` short-circuits to `ApprovalRequirement::Never` when the skill is already loaded — asking the user to approve a guaranteed no-op is pure friction, and the test was asserting against that intentional behavior. Rewrote the test to skip the approval step and assert on the terminal message's idempotent "already installed / no install needed" wording, which matches the actual production output. - Mock LLM: the pattern branch in `match_tool_call` was re-emitting a matching tool call on every LLM round because "last user content" doesn't change across turns, so the engine looped until it hit the multi-result summary path. Added a guard that falls through to the text-response path when the matching tool_name is already present in `recent_tool_results` — mirroring real LLM behavior. - `test_v2_engine_oauth_google::test_oauth_token_refresh_on_expiry`: two compounding issues blocked the refresh path. (1) The mock `/oauth/refresh` handler validates `client_id == "hosted-google- client-id"`, but the fixture env set `test-google-client-id`. (2) Proxy URL points at `http://127.0.0.1:<port>` (the mock LLM) and the production SSRF guard blocks loopback by default; mock E2E tests opt in via `IRONCLAW_OAUTH_PROXY_ALLOW_LOOPBACK=1`. Also added an `oauth:` block to the test's `google_drive` skill so `credential_spec_to_oauth_refresh` registers a refresh config under `google_drive_token`. Finally, the refresh path needs a stored refresh token — paste-based auth (the earlier tests' fallback when no google-drive WASM binary is available) only persists the access token, so the test now skips in that configuration rather than asserting on a refresh that can't happen, matching the pattern already used by `test_oauth_redirect_flow`. Remaining after this PR: `test_repl_http_auth_prompt_accepts_token_and_retries` passes in isolation but flakes under full-suite load (the PTY REPL sibling test is already `@pytest.mark.skip` for the same reason); and `test_always_approve_survives_restart` which times out the `/api/health` probe under full-suite pressure. Both are PTY / fixture-startup concurrency issues, not product regressions. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(tests): close the last 2 e2e failures — full suite green (401 passed, 0 failed) Root-causes the two tests left open after the previous commit. Both were real bugs/config drift masquerading as flakiness. - `test_repl_http_auth_prompt_accepts_token_and_retries`: `CLI_MODE` defaults to `tui` (the ratatui full-screen UI), which reads stdin keystroke-by-keystroke and renders into a framebuffer. The PTY-driven tests in this file send whole lines via `os.write(master_fd, b"prompt\n")` and match for specific text in the raw stream — under the default TUI that line-based send never reaches the agent, so the auth card never fires and `_read_repl_until` times out with only cursor-position escape sequences captured. Pinning `CLI_MODE=repl` on the fixture routes the test back onto the plain line-based REPL surface it's written against. Confirmed passing 5/5 in isolation and under full-suite load. - `test_always_approve_survives_restart`: the fixture's ironclaw subprocess was dying at startup with `Channel webhook_server failed to start: Failed to bind to 127.0.0.1:8080: Address already in use (os error 98)` — the fixture picked a free `GATEWAY_PORT` but left `HTTP_HOST`/ `HTTP_PORT` unset, so the HTTP channel tried to claim the default port 8080 and collided with every other e2e server (and anything else on 8080). Every `/api/health` probe was hitting a dead process, which showed up as a 60 s timeout instead of a bind error because the subprocess's stderr was never drained — a full 64 KiB pipe buffer made the child block on its next write before it could even log the bind failure. Fix: - allocate a second free TCP port for `HTTP_PORT` (mirrors the sibling `v2_approval_server` fixture); - wire `stdout`/`stderr` through background drain tasks so `RUST_LOG=ironclaw=debug` output can't back-pressure the child into a startup hang; - surface the last 32 KiB of stderr in the timeout error so future regressions (panic, bind conflict) show up in the failure message instead of being silently swallowed. Full-suite e2e: 401 passed, 8 skipped, 0 failed, 0 errored (17:31). Rust unit + integration tests still green, clippy clean, fmt clean. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(tests): address PR nearai#2744 review + reconcile with staging ## Review feedback - **Slash-skill cache spam (3× — Gemini + 2× Copilot):** the previous `refreshSlashSkillEntries()` re-fetched `/api/skills` on every keystroke in `filterSlashCommands`; the in-flight guard only suppressed concurrent duplicates. Added a 30 s TTL and an `invalidateSlashSkillCache()` hook that the install/remove flows in `surfaces/skills.js` call so the menu picks up install/remove changes immediately instead of waiting for the TTL. - **Wrong module path in comment (Copilot):** `src/bridge/router.rs` comment referenced `llm::reasoning::user_signals_execution_intent` but `reasoning` isn't `pub` — the helper is re-exported as `crate::llm::user_signals_execution_intent`. Updated the comment to use the canonical path and cross-reference the defining file. - **Misleading `#[tokio::test]` justification (Copilot):** prior comment said "single-threaded tokio and cannot deadlock" without pinning the runtime flavor. `#[tokio::test]` *does* default to the current-thread runtime in this crate, but spelling it out is safer against future defaults drifting. Pinned `#[tokio::test(flavor = "current_thread")]` explicitly and reworded the comment to name the runtime kind. - **Drain tasks cancelled but not awaited (Copilot):** the restart fixture in `test_v2_engine_approval_flow.py` cancelled the stdout/stderr drainers on `stop()` without awaiting them, causing "Task was destroyed but it is pending!" warnings and, on stop→start cycles, zombie readers. Now cancels *and* `asyncio.gather (..., return_exceptions=True)` awaits them. ## Merge reconciliation with `origin/staging` Staging merge introduced: - A strict MIME allowlist on `/api/chat/send` attachments (nearai#2332). `test_gateway_attachment_unextractable_file_uses_placeholder` previously relied on `application/octet-stream` reaching `document_extraction` and triggering the "[Failed to extract …]" placeholder; the new gateway-side allowlist rejects that MIME outright at the HTTP layer, so the test never exercised the fallback path. Updated the test to upload a corrupt PDF (`%PDF-1.4` magic + garbage body) which passes MIME + header checks but fails extraction — the exact scenario the placeholder was designed for. - A conflict in `src/pairing/approval.rs` where staging added `#[ignore]` to the propagate-approval test (needs a pre-built telegram WASM binary) and this branch added `#[allow(clippy::await_holding_lock)]`. Merged both, plus pinned the explicit `current_thread` runtime flavor per review. ## Pre-existing failures left alone `test_portfolio.py::test_portfolio_chat_keyword_triggers_skill` and `test_portfolio_wallet_address_triggers_skill` both fail identically against plain `origin/staging` (verified via `git stash` + checkout of the staging versions of the test file and `crates/ironclaw_engine/orchestrator/default.py`). Root cause is unrelated to this PR — appears to be the mock LLM's portfolio response text tripping the engine's tool-intent nudge path before reaching the canned response the test asserts on. Out of scope here. ## Verification - `cargo fmt` - `cargo clippy --all --benches --tests --examples --all-features` — zero warnings - `cargo test --lib` — 5329 passed, 7 ignored, 0 failed - `pytest scenarios/test_chat.py scenarios/test_v2_engine_approval_flow.py scenarios/test_v2_engine_auth_flow.py::TestV2EngineSkillInstallFlow scenarios/test_v2_auth_oauth_matrix.py scenarios/test_pending_user_messages.py` — **58 passed, 1 skipped, 0 failed** Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(fmt): collapse override_engine_project_root call onto single line rustfmt on staging collapses this call; my earlier `cargo fmt` ran before the `project_root.clone()` edit landed so the local check missed it. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * Address PR nearai#2744 review: startup-timeout leak + attachment test isolation 1. `test_v2_engine_approval_flow.py` — `start()` re-raised `TimeoutError` from `wait_for_ready` without tearing the subprocess down. Because `await start()` runs before the fixture's `try/finally`, a startup timeout would leak the child process and its bound ports into the rest of the test run. Snapshot the stderr tail before teardown, `await stop()` (which kills the proc and cancels/awaits the drain tasks), then re-raise with the captured tail. 2. `tests/e2e_attachments.rs` — the `engine_v2_project_root()` helper derived from `bootstrap::ironclaw_base_dir()` is a process-global `LazyLock` that resolves to `$HOME/.ironclaw` on dev machines and CI runners. Passing its parent as the engine's project_root meant this test was writing real attachment files into `~/.ironclaw/attachments` every time it ran. Allocate a per-test `tempfile::TempDir` instead and point `override_engine_project_root_for_test` at it — now writes are fully contained. The `engine_v2_attachment_root_lock` mutex stays (still required to serialize mutations of the process-global engine state across concurrent tests). Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Summary
.ironclaw/attachments/...storage and verify attachment handling across gateway, Telegram, and WhatsApp paths/<skill>routing behaves correctlystagingby fixing warning noise, flaky reconnect/history assumptions, writable-thread handling, approval-flow isolation, Telegram stale-update handling, and initial auth-page bootstrap racesTest plan
cargo clippy --all --benches --tests --examples --no-default-features --features libsql -- -D warningscargo fmt --all -- --checktests/e2e/.venv/bin/pytest tests/e2e/scenarios/ -q -s --maxfail=1 -W error::DeprecationWarning -W error::pytest.PytestUnraisableExceptionWarning