Skip to content

engine-v2: make available_actions callable-only for blocked providers - #2868

Merged
henrypark133 merged 20 commits into
stagingfrom
v2-engine-callable-only-cleanup
Apr 25, 2026
Merged

henrypark133 merged 20 commits into
stagingfrom
v2-engine-callable-only-cleanup

Conversation

@henrypark133

@henrypark133 henrypark133 commented Apr 22, 2026 •

Copy link
Copy Markdown
Collaborator

Context

This PR now includes the follow-up work from #2869, #2876, and #2889, so it should be read against the engine-v2 epic comment in #2767:
#2767 (comment)

The main alignment is:

  • Task 4: available_actions() callable-only cleanup
  • Task 5: prompt / loop alignment
  • Task 6: canonical callable-action discovery metadata
  • cleanup of the earlier Section 7 prototype so we do not keep a hidden inspect/promote path around in the runtime
  • the later enablement follow-up that makes blocked integrations capability-first and routes model-visible enablement through tool_activate(name=...)

Old Experience

Before this stack:

  • available_actions() mixed together the wrong provider states in engine v2. Some blocked provider actions (needs_auth, needs_setup, inactive, not-installed, routed-only) could still show up as normal callable tools even though they were not directly invokable on that turn, while installed-but-not-yet-activated latent WASM provider actions could disappear entirely and fail to trigger auth-on-first-call.
  • The CodeAct system prompt duplicated callable tool inventory in prose. That meant the same tool information could appear both in provider schemas and in prompt text, and resume / compaction paths could keep stale or duplicated metadata instead of rebuilding one canonical view.
  • Action discovery was split across multiple partially-connected paths. Registry-backed tools and engine-native mission actions did not share one discovery envelope, tool_info was not consistently reading the same callable snapshot the executor was using, and hyphen/underscore aliases were easy to handle inconsistently at different call sites.
  • Blocked integrations did not have one clean model contract. Some paths still assumed first-call recovery on blocked provider actions; others treated blocked state as background metadata; tool_install, tool_auth, and latent synthetic actions all leaked overlapping lifecycle detail into the surfaced experience.
  • Some deferred-tool scaffolding still implied a tool_info -> promotion flow. That was leftover prototype behavior, not the contract we want to preserve.

New Experience

After this stack:

  • The provider-facing callable surface is the callable-now surface. Ready actions remain callable; blocked provider/channel integrations are no longer normal model-facing callable tools.
  • The prompt now has three distinct model-visible layers:
    • normal callable actions/tool inventory
    • Capabilities for contextual/background runtime information
    • Activatable Integrations for blocked managed integrations
  • tool_activate(name=...) is the single model-facing enablement path in engine v2 and CodeAct. The harness owns whether that means install, auth, activation, or a surfaced manual-setup failure.
  • Capability entries for blocked integrations now include compact action previews so the model can see what enablement unlocks without dumping large tool families into the default callable surface.
  • tool_info now reads the current callable snapshot / inventory, resolves canonical names and hyphen/underscore aliases consistently, fails cleanly for names that are not callable in the current execution context, and does not gate or promote a new tool into the next LLM step's callable set.
  • Newly enabled tools become visible on the next top-level turn / orchestrator iteration, not mid-CodeAct step.

What Changed

  • add ActionDiscoveryMetadata, ActionDiscoverySummary, ActionInventory, and ActionDef::matches_name() so callable action metadata is canonical and alias-safe
  • rework bridge projection so ActionProjector only surfaces callable-now actions and CapabilityProjector owns blocked provider/channel background plus action previews
  • add engine-native mission action discovery metadata in src/bridge/engine_actions.rs
  • thread action snapshots through orchestrator and scripting paths so policy checks, execution, and tool_info all read the same callable set
  • replace the old prompt-side callable tool list with separate Capabilities and Activatable Integrations sections, and refresh engine-owned system prompts across resume / compaction without duplicating appended sections
  • make tool_activate the normal surfaced enablement operation, while preserving install approval semantics when enablement would auto-install a not-yet-installed integration
  • remove the short-lived extension-list cache in the effect bridge so immediate post-auth prompt state stays fresh on the next turn
  • pin isolated E2E mock-LLM fixtures explicitly to the mock openai_compatible provider instead of relying on env-vs-DB precedence
  • replace the stale v2 auth preflight surface test with a dedicated browser E2E for the new tool_activate / Activatable Integrations contract
  • update the relevant CLAUDE docs and inline type/tool comments so the repo documentation matches the new callable vs contextual vs activatable split

Key Coverage Added In This Stack

  • available_actions_omit_installed_needs_auth_provider_action
  • available_actions_omit_installed_needs_setup_provider_action
  • available_actions_omit_installed_inactive_provider_action
  • available_actions_omit_latent_inactive_provider_actions
  • available_capabilities_include_latent_provider_background
  • available_capabilities_include_latent_provider_activation_entry
  • tool_activate_requires_approval_before_auto_installing_integration
  • skill_prompt_context_survives_pause_and_resume
  • skill_prompt_context_survives_compaction_and_resume
  • resume_refreshes_checkpointed_system_prompt_metadata
  • tool_info_does_not_gate_callable_tool_into_next_llm_callable_set
  • execute_code_propagates_snapshot_to_tool_execution_context
  • test_v2_tool_activate_surface.py

Testing

All of the following were run on this branch:

cargo fmt --all
cargo clippy --all-targets -- -D warnings
cargo test effect_adapter:: --lib

tests/e2e/.venv/bin/pytest \
  tests/e2e/scenarios/test_v2_auth_oauth_matrix.py::test_settings_first_gmail_auth_then_chat_runs \
  tests/e2e/scenarios/test_v2_auth_oauth_matrix.py::test_mcp_oauth_roundtrip_via_browser \
  tests/e2e/scenarios/test_telegram_hot_activation.py::test_telegram_auth_required_shows_configure_modal_and_can_cancel \
  tests/e2e/scenarios/test_telegram_hot_activation.py::test_telegram_hot_activation_transitions_installed_to_pairing \
  tests/e2e/scenarios/test_v2_tool_activate_surface.py -v

Result: all passed.

https://github.com/nearai/ironclaw/actions/runs/24864084782

Copilot AI review requested due to automatic review settings April 22, 2026 20:12
@github-actions github-actions Bot added size: XL 500+ changed lines risk: low Changes to docs, tests, or low-risk modules contributor: core 20+ merged PRs labels Apr 22, 2026

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Aligns engine-v2 available_actions() with the “callable-only” surface policy by filtering out extension-backed provider actions when the provider is installed-but-blocked (needs auth/setup/inactive), and adds coverage to prevent regressions.

Changes:

  • Removed the special-case bypass that kept NeedsAuth provider tools in available_actions within ActionProjector.
  • Added ActionProjector tests covering omission for NeedsAuth, NeedsSetup, Inactive, and routed-only channels.
  • Added EffectBridgeAdapter-level tests using an installed-provider fixture to assert blocked installed provider actions are omitted from available_actions.

Reviewed changes

Copilot reviewed 2 out of 2 changed files in this pull request and generated 1 comment.

File Description
src/bridge/effect_adapter.rs Adds adapter-level fixture/tests for omitting installed-but-blocked provider actions from available_actions.
src/bridge/action_projector.rs Removes NeedsAuth bypass and expands projector-level tests for blocked/routed providers.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread src/bridge/effect_adapter.rs

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request updates the ActionProjector to omit tools that are inactive or require authentication or setup from the available_actions list. It also introduces comprehensive test helpers and integration tests in action_projector.rs and effect_adapter.rs to verify these filtering rules. A review comment identifies a resource leak in a test helper where std::mem::forget prevents temporary directory cleanup, suggesting a refactor to return the TempDir object instead.

Comment thread src/bridge/effect_adapter.rs Outdated
Comment thread src/bridge/action_projector.rs Outdated
Comment thread src/bridge/effect_adapter.rs
* fix(engine): align prompt metadata refresh with resume state

* fix(engine): finish prompt refresh compaction coverage (#2869)

* fix(engine): preserve prompt refresh on resume (#2869)

* Add engine v2 action discovery metadata (#2876)

* Add engine v2 action discovery metadata

* fix(engine): address action discovery review (#2876)

* fix(engine): address follow-up review comments (#2876)

* fix(engine): satisfy clippy in orchestrator lookup

* fix(engine): propagate action snapshots in executor paths (#2876)

* fix(bridge): restrict tool_info to callable actions (#2876)

* [codex] Finish engine v2 deferred action inventory cleanup (#2889)

* Add deferred action inventory groundwork

* fix(engine): address deferred action inventory follow-up

* fix(engine): address deferred inventory review feedback

* test: fix fmt and clippy failures
Copilot AI review requested due to automatic review settings April 23, 2026 21:55
@github-actions github-actions Bot added scope: channel/web Web gateway channel scope: tool Tool infrastructure scope: tool/builtin Built-in tools scope: tool/wasm WASM tool sandbox scope: llm LLM integration scope: extensions Extension management risk: medium Business logic, config, or moderate-risk modules and removed risk: low Changes to docs, tests, or low-risk modules labels Apr 23, 2026

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Copilot reviewed 39 out of 39 changed files in this pull request and generated 3 comments.


💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread src/bridge/action_projector.rs Outdated
Comment thread crates/ironclaw_engine/src/executor/prompt.rs
Comment thread crates/ironclaw_engine/src/executor/loop_engine.rs Outdated

@henrypark133 henrypark133 left a comment

Copy link
Copy Markdown
Collaborator Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Review: callable-only surface and prompt/inventory alignment look good

I did a full pass on the stacked engine-v2 changes here and did not find any verified blocker-level issues.

What looks good:

  • The callable-surface contract is much cleaner now: blocked provider actions move out of available_actions, while background capability state stays model-visible through the canonical capabilities section.
  • The snapshot plumbing is materially better. tool_info, structured execution, scripting, and orchestrator paths now read the same callable snapshot instead of drifting on alias handling or live registry state.
  • The prompt refresh work is directionally right. Resume/compaction now rebuild the engine-owned system prompt while preserving appended step-zero context like prior knowledge and active skills.
  • The new caller-level regression coverage is the right shape for this area, especially around resume/compaction behavior and tool_info not mutating the next-step callable set.

Low-priority notes:

  • refresh_system_prompt() still fetches available_action_inventory() even though the current prompt builder ignores that inventory. That looks like avoidable start/resume work unless a deferred inventory section is about to land.
  • project_tool_action() currently emits discovery metadata for every callable tool even when it only repeats the callable name and has no summary/schema override. Tightening that would trim some per-turn payload size.

Verification:

  • cargo test --test engine_v2_gate_integration tool_info_does_not_gate_callable_tool_into_next_llm_callable_set -- --nocapture
  • cargo test --test engine_v2_skill_codeact skill_prompt_context_survives_pause_and_resume -- --nocapture
  • cargo test --test engine_v2_skill_codeact skill_prompt_context_survives_compaction_and_resume -- --nocapture

Copilot AI review requested due to automatic review settings April 23, 2026 22:53
@henrypark133
henrypark133 deleted the v2-engine-callable-only-cleanup branch April 25, 2026 02:39
nickpismenkov added a commit that referenced this pull request Apr 28, 2026
Same pipe-deadlock fix as scripts/live_canary/common.py f59981d,
applied to tests/e2e/scenarios/test_v2_auth_oauth_matrix.py's
_start_auth_matrix_server. The auth-matrix fixture spawns ironclaw
with stdout=PIPE + stderr=PIPE and never drains them, so under
sustained log volume the kernel pipe buffer fills, ironclaw blocks
on its next stdout write, and any test that relies on subsequent
gateway responses (auth gate emission, SSE events, chat replies)
hangs until pytest-timeout fires.

This fix doesn't make the auth-full lane's failing test pass — the
real bug is engine-v2 silently dropping `auth_required` SSE events
for unauthenticated extensions (introduced by #2868). But it makes
the failure mode debuggable: gateway log is captured to
/tmp/ironclaw-auth-matrix-gateway.log (overridable via
IRONCLAW_AUTH_MATRIX_LOG env), and RUST_LOG passes through from the
test runner so we can crank up verbosity without rebuilding.

Without this change, the failing test's log was empty after the
extension-install line; with this change you see the engine-v2
trace summary that surfaces the actual NotCallable-without-auth-gate
bug. That diagnostic visibility is the value here.

- _drain_stream_to_file: asyncio drainer mirroring common.py's sync
  threading version
- _start_auth_matrix_server: drain stdout/stderr to log_path
- _shutdown_auth_matrix_server: cancel drain_tasks for clean exit
- env: RUST_LOG forwarding so debug runs work

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
nickpismenkov added a commit that referenced this pull request Apr 28, 2026
…+ GH issues

Three additions to scripts/live-canary/notify_slack.py to make the
6h Slack report actionable instead of just informational:

1) **Per-lane rich failure block** — Haiku now extracts four
   structured fields when status==fail: test_name, error, root_cause,
   fix. The Slack section renders them in the issue-friendly shape
   the reviewer asked for:

       ❌ auth-full (mock) — 11/13 passed, 1 failed in 213s
         Test: `test_wasm_tool_first_chat_auth_attempt_emits_auth_url`
         Error: SSE stream closed; auth_required event never arrived
         Root Cause: bridge gate not wired for installed-but-unauthed
                     extensions (#2868 fallout)
         Fix: route Extension::NeedsAuth through effect_adapter.rs

   For passing/skipped lanes the existing single-line `> reason` is
   preserved so the green-path Slack output is unchanged.

2) **Cross-lane "Summary by Category" block** — second Haiku pass
   over all failed-lane summaries that groups them by shared root
   cause (e.g. "WASM tool dispatch regression — Auth Full, Auth
   Smoke, Auth Live Seeded"). Only fires when there are 2+
   failures (single-failure runs are already obvious from the
   per-lane block). Rendered as a Slack mrkdwn bulleted list since
   Block Kit doesn't support real tables.

3) **Auto-opened GitHub issues** — opt-in via CANARY_CREATE_ISSUES=1
   env var (gated to scheduled runs only in live-canary.yml so
   workflow_dispatch debugging doesn't flood the tracker). For each
   failed lane:
   - Search for an OPEN issue with title `[canary] <lane>: <test>`.
   - If found: comment "another occurrence on <run_url>".
   - If not found: open a new issue with the rich body + labels
     `canary-failure` + `lane:<lane>`.

   Strategy chosen to avoid issue spam while still surfacing
   recurring failures. Uses GITHUB_TOKEN + the repo's existing
   `permissions: issues: write` block — no new secrets.

All three additions degrade silently — Haiku failure stamps
.notable but doesn't block the post; categorization failure produces
an "_(unavailable)_" placeholder; issue-creation errors are logged
to stderr only. The notifier still exits 0 in every failure path so
a flaky webhook can't fail the canary run.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
nickpismenkov added a commit that referenced this pull request Apr 28, 2026
* fix(oauth): remove pending flow on provider-error callback

The /oauth/callback handler's ?error= branch (RFC 6749 §4.1.2.1
provider-side failures — user cancels consent, scope denied, etc.)
returned the error page immediately without removing the flow from
ext_mgr.pending_oauth_flows(). The ghost entry then lingered until
the 5-minute expiry sweep, and any subsequent auth dance for the
same (extension, user) pair had to dedupe against it.

Mirror the happy-path cleanup: decode the state param, remove the
keyed flow, then return the error page.

Surfaced during live-canary auth-full repro: after
test_wasm_tool_oauth_provider_error_leaves_extension_unauthed ran,
the stale flow sat in the shared auth_matrix_server fixture.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(e2e): widen auth OAuth matrix timeouts for CI load

Four tests in live-canary auth-full were failing in CI with
`Page.wait_for_function: Timeout 60000ms exceeded`,
`ClientConnectionError('Connection closed')`, and
`Timed out waiting for OAuth refresh request` — all inside 60/20s
deadlines that are tuned for a dev laptop and don't leave margin
for ubuntu-latest's 2-vCPU runner under full suite load.

Raise the per-call deadlines so the inner budgets fit comfortably
inside pyproject.toml's 120s per-test cap:

  _wait_for_refresh_request default: 20.0s -> 60.0s
  _wait_for_auth_event call site:      60   -> 90
  _wait_for_auth_prompt call site:     60   -> 90
  send_chat_and_wait_for_terminal_message call sites: 60000 -> 90000
  _wait_for_mock_google_tokens call site: 60.0 -> 90.0
  _wait_for_response_contains (gmail) call site: 60.0 -> 90.0

Strictly widening; no passing test is slowed, no semantics change.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(canary): Haiku-powered Slack report job

Replace the team's raw Slack subscription (firehose of workflow
notifications) with one curated per-run summary:

  Canary: 9 passed, 1 failed of 10 lanes
  :x: auth-full (mock) — 12/13 passed, 1 failed in 350s
  > test_wasm_tool_first_chat_auth_attempt_emits_auth_url timed
  >   out waiting for auth_required SSE event on the fresh thread
  tools: shell, http_request, gmail (~6 calls)
  ...
  commit `abc1234` • <github run link>

New `canary-report` job (needs: every lane, if: always) downloads
all lane artifacts, parses junit + summary + log tail per lane, and
asks claude-haiku-4-5 to return a compact JSON per lane
({status, reason, tool_calls_total, tools_used, notable}). That's
aggregated into a single Slack block message and posted via
incoming webhook.

Safety shape:
- Script exits 0 even on Haiku/Slack failure so the notifier never
  masks the underlying canary signal.
- Missing ANTHROPIC_API_KEY falls back to raw junit-only phrasing.
- Slack POST failure falls back to plain-text "X/Y lanes failed"
  with the GH run URL so the channel still hears something.
- No new Python deps — pure stdlib (urllib.request, xml.etree).
- 20 KB log-tail cap per lane to keep Haiku token usage bounded.

Secrets:
- ANTHROPIC_API_KEY (already present, used by provider-matrix)
- SLACK_WEBHOOK_URL (new — create an incoming webhook in Slack
  and add as repo secret; notifier prints to stdout otherwise)

Testing:
- Trigger manually via Actions -> "Live Canary" -> "Run workflow"
  with any single lane; canary-report runs after regardless of
  which lanes executed.
- Run locally with --dry-run to preview the Slack payload.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(canary): post_json error handling + robust Haiku JSON extraction

Address gemini-code-assist review on scripts/live-canary/notify_slack.py:

1. `post_json` unreachable error branch: `urllib.request.urlopen`
   raises `urllib.error.HTTPError` for 4xx/5xx before reaching the
   `if resp.status >= 300` check, so the error body was never
   surfaced. Wrap in try/except and read the body from the
   HTTPError instance — that's where Anthropic's "invalid API key"
   / "rate limited" detail lives.

2. Haiku JSON extraction was fragile: `startswith("```")` assumed
   the response had no prose preamble and only handled one fence
   shape. Replace with `re.search(r"\{.*\}", text, re.DOTALL)` so
   we pick the outermost JSON object regardless of any wrapper
   markdown or leading/trailing text. Greedy + DOTALL is correct
   for the single top-level object our schema requires.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(e2e): raise pytest timeout + bump multi-user chat wait to 180s

The CI run on feat/canary-report surfaced that 90s was still not
enough for test_mcp_same_server_multi_user_via_browser on
ubuntu-latest — it timed out at the inner Playwright
wait_for_function deadline with "Timeout 90000ms exceeded" after
118s of total test time.

The test opens two browser contexts + two SSE streams and drives a
full chat turn per user in sequence. Under 2-vCPU contention the
compound pipeline genuinely takes over 90s.

- tests/e2e/pyproject.toml: timeout 120 -> 240 (pytest-level cap)
- test_v2_auth_oauth_matrix.py: send_chat_and_wait_for_terminal_message
  call sites 90000 -> 180000 (two owner/member turns, each budgeted
  for one runner-slow turn)

180s < 240s, so the inner deadline fires first with the useful
Playwright traceback instead of the generic pytest SIGTERM.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(e2e): fix pytest-timeout CLI override + widen Mode-C deadlines

The previous commit (c3c9bbab) raised tests/e2e/pyproject.toml's
timeout from 120 to 240, but the auth canary runs the suite via
scripts/auth_canary/run_canary.py which hardcodes
`--timeout=120` on the pytest command line. The CLI flag wins
over pyproject's ini_options, so the 240 bump was invisible to
the auth lanes. That's why auth-smoke on the canary `all` run
still failed with "Timeout (>120.0s) from pytest-timeout" even
after our 180s inner widening — the outer CLI cap was firing at
120s first.

Fix the override and widen the two remaining Mode-C deadlines
that blew in the same run:

  scripts/auth_canary/run_canary.py: --timeout=120 -> 240
  _wait_for_refresh_request default: 60.0 -> 120.0
    (test_wasm_tool_oauth_refresh_on_demand and
     test_mcp_oauth_refresh_on_demand both use the default)
  test_settings_first_gmail_auth_then_chat_runs call sites:
    _wait_for_mock_google_tokens 90.0 -> 120.0
    _wait_for_response_contains 90.0 -> 120.0

All remain comfortably under the new 240s pytest-level cap so a
real hang still fails fast with a useful traceback.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(e2e): opt-in text-match predicate for multi-user browser test

Ship the structural fix that was overdue. Repeated budget bumps on
send_chat_and_wait_for_terminal_message weren't holding under
ubuntu-latest "all"-mode parallelism — 120s, 180s both exceeded on
test_mcp_same_server_multi_user_via_browser. The underlying race is
in the JS predicate: it waits for the assistant bubble AND the
data-streaming attribute cleared AND the chat input re-enabled.
Under 2-vCPU contention an SSE reconnect can drop the final
attribute-clearing delta, and the compound predicate never flips
even though the response text arrived long ago.

Add an opt-in `expected_text_contains` parameter. When supplied,
the predicate succeeds the moment the expected substring appears in
the new assistant message — regardless of data-streaming or input
state. Callers that already assert on specific response text (the
existing MCP / gmail tests) can now short-circuit the race without
compromising correctness: the test's own content assertions remain
the gate.

Default behavior unchanged for the ~30 existing call sites across
test_chat.py, test_sse_reconnect.py, test_tool_approval.py,
test_portfolio.py, test_message_persistence.py, test_agent_loop_recovery.py,
test_pending_user_messages.py, test_widget_customization.py.

Applied to the two multi-user call sites with
expected_text_contains="Mock MCP search result" — that's exactly
what the test's next two assertions verify.

Local run of the flaky test alone: 40s, green.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* ci(canary): move auth-smoke to self-hosted runner

Multi-user browser test (test_mcp_same_server_multi_user_via_browser)
consistently exceeds the Playwright budget on GH ubuntu-latest under
the 2-vCPU parallelism pressure of an "all" canary run — a single
compound chat turn burns >180s, with each budget bump we apply it
ratchets the flake, not the fix.

Pilot move onto the [self-hosted, ironclaw-live] runner that
private-oauth already uses. Same runner label means no new
infrastructure required; if the self-hosted box has Python 3.12 and
Playwright browsers installed (or can provision them via the existing
setup-python + scripts/live-canary/run.sh's `PLAYWRIGHT_INSTALL=with-deps`
flow), this is a zero-code-change canary fix.

If the pilot works, auth-full is the next candidate. If the runner
queues become a bottleneck, we'd scale to multiple workers under
the same label rather than revert to ubuntu-latest.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* ci(canary): revert auth-smoke to ubuntu-latest + widen budgets to 300s/360s

Railway self-hosted runner ('railway-private-oauth' on a small Docker
container) turned out to be no faster than GH ubuntu-latest for the
multi-user browser flow — both take ~194–196s for
test_mcp_same_server_multi_user_via_browser. The runner container is
evidently provisioned at a similar vCPU allocation, so the move
bought nothing.

Revert to ubuntu-latest (parallel canary shape preserved; avoids
serialising auth lanes behind private-oauth on the single
self-hosted worker) and widen deadlines for the last CI-load hop:

  test_v2_auth_oauth_matrix.py multi-user call sites:
    Playwright wait_for_function 180000 -> 300000 ms
  scripts/auth_canary/run_canary.py:
    --timeout=240 -> 360 (outer pytest cap)
  tests/e2e/pyproject.toml:
    timeout = 240 -> 360

300s inner fits inside the new 360s outer with 60s margin. Local
run of the same test alone completes in ~40s, so we have plenty
of headroom against real hangs still surfacing fast with a
useful traceback.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* disable report

* scripts(auth-canary): add Google storage-state bootstrap helper

The auth-browser-consent lane drives Google's real OAuth consent UI in
Playwright, but Google's risk engine routinely interrupts the flow with
a "Verify it's you" challenge that handle_google_popup cannot solve, so
the test stalls on the password screen.

Bypass: log in once interactively in Playwright Chromium, save cookies
+ localStorage to a storage_state.json, point AUTH_BROWSER_GOOGLE_-
STORAGE_STATE_PATH at it. Subsequent canary runs spawn contexts with
that state preloaded, so the popup arrives at consent with no login or
challenge in the way.

- scripts/auth_live_canary/bootstrap_google_storage_state.py: new
  one-shot interactive helper that writes
  ~/.ironclaw/auth-canary/google_storage_state.json by default
- scripts/auth_live_canary/README.md: document the bypass under
  "Browser-consent Google challenge bypass"

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* canary(auth-browser-consent): fix Google account-picker + chat drift

The auth-browser-consent google case was failing on two distinct
issues, the first masking the second:

1) Account picker. When AUTH_BROWSER_GOOGLE_STORAGE_STATE_PATH is set
   (the recommended path — username/password automation gets blocked
   by Google's risk engine), Google's OAuth popup lands on a "Choose
   an account" picker before the consent screen. handle_google_popup
   only knew how to fill email + password and click Continue/Allow,
   so the popup sat on the picker until complete_provider_auth's
   120s callback wait timed out. Added a picker-detection step that
   tries selectors in order — username text, [data-identifier], and
   a generic "any visible @-bearing text not equal to 'Use another
   account'" XPath — and clicks the first hit, with debug logging
   so future regressions surface in the run output.

2) Tool-name and response-text drift. After the OAuth fix unblocked
   the rest of the probe, browser_chat still failed because:
   - case.expected_tool_name was "gmail", but the gateway records
     the tool call under its WASM module name "gmail_tool"
   - case.expected_text was "Gmail" (case-sensitive), but real LLM
     responses to "check gmail unread" against an empty inbox vary
     ("Your inbox is clear...", "Inbox is empty", etc.) and rarely
     emit literal "Gmail"
   Updated BROWSER_CASES["google"] to expected_tool_name="gmail_tool"
   and expected_text="inbox", and made the browser_chat assertion's
   text comparison case-insensitive so the canary doesn't depend on
   exact wording.

After both fixes the auth-browser-consent google lane runs green:
  ✓ browser_oauth   (popup -> /oauth/callback)
  ✓ browser_chat    (assistant references inbox)
  ✓ responses_api   (real Gmail tool call)

Not addressed here: BROWSER_CASES["github"] likely has the same
expected_tool_name drift ("github" vs probably "github_tool"); needs
verification with real GitHub OAuth creds before changing.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* canary(auth-browser-consent): robust account-picker fallback + browser channel

Two follow-ups discovered during local debugging of the auth-browser-
consent google lane:

1) Account-picker fallback was matching hidden <style> blocks. The XPath
   `//*[contains(text(), '@') ...]` matched any element whose text
   contains `@`, which includes <style> tags carrying CSS at-rules
   (@font-face, @media). Replaced the XPath with role-based locators
   (get_by_role link/button) filtered by an email regex — only
   interactive elements match, no false positives from style blocks.
   Verified locally that the fallback now clicks the right account row
   even when AUTH_BROWSER_GOOGLE_USERNAME is unset.

2) Bootstrap script: Google's anti-automation blocks Playwright's
   default Chromium (Chrome for Testing) at sign-in with "This browser
   or app may not be secure". Added a --browser flag with a default of
   firefox (Marionette is less aggressively fingerprinted than CDP),
   plus chrome (system Google Chrome) and chromium (override) options.
   For accounts where Google blocks even those — typically brand-new
   Gmails or accounts with high risk scores — the fallback path is to
   launch Chrome manually with --remote-debugging-port and connect via
   playwright.chromium.connect_over_cdp; documented in the README.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* canary(auth-live-canary): include observed extension state in timeout error

When `wait_for_extension_state` times out the bare error
"Timed out waiting for extension state: gmail" is unhelpful for
diagnosing CI failures, since CI artifacts don't capture IronClaw's
gateway logs — there's no way to tell whether the extension never
appeared, appeared but never authenticated, or authenticated but
never activated.

Track the last-observed extension on each poll and surface
authenticated/active in the timeout message. After this change a
failed run says e.g.
"Timed out waiting for extension state: gmail (expected
authenticated=True, active=True; last observed: authenticated=False,
active=False)", which immediately separates token-exchange failures
from activation-state-machine bugs.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* canary(auth-live-canary): widen chat-wait deadlines 120s -> 300s

The auth-browser-consent google probe completed OAuth + extension
activation successfully on CI but timed out at the next step
(send_chat_and_wait_for_terminal_message), with the agent stuck on
"Thinking (step 1)" for the full 120s budget. Local runs on the
same code path complete the chat in ~36s, but ubuntu-latest 2-vCPU
runners under cold-start load (gateway restart, mock LLM bootstrap,
WASM tool first-invocation) need substantially more headroom.

300s matches the precedent set by `d8765714 ci(canary): revert
auth-smoke to ubuntu-latest + widen budgets to 300s/360s` for the
auth-smoke lane on the same runner class.

Both call sites widened — the seeded Responses-API probe at line 221
and the browser_oauth probe at line 800.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* canary(common): drain gateway/mock_llm stdout pipes (was deadlocking CI)

scripts/live_canary/common.py spawns the IronClaw gateway and the
mock LLM with stdout=PIPE + stderr=STDOUT, reads one line of mock_llm
output to discover its bound port, then never reads from either pipe
again. On Linux the kernel pipe buffer caps at 64 KiB; once a
sustained chat request fills it with `RUST_LOG=info` output, the
child blocks on its next stdout write and the request handler
freezes mid-response.

That's why every auth-browser-consent CI run got stuck on
"Thinking (step 1)..." for the full chat-wait budget while the same
test passes locally — macOS pipe buffers are larger and the test
completes before the buffer fills.

Fix: spawn a daemon thread per subprocess that drains the pipe to a
log file under the run's output_dir. Two wins:

- Pipes never fill, child never blocks.
- gateway.log and mock_llm.log become CI artifacts, so the next
  failure that doesn't have a clear runner-side error message is
  immediately debuggable from IronClaw's own logs.

Verified locally that the lane still passes after the change and
both log files are produced. Locally each is < 10 KiB; CI runs may
be larger but well under any artifact size limit.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* canary: pin LLM backend via settings API + add LLM_API_KEY (root cause of CI freeze)

The auth-browser-consent google lane has been freezing on CI at
"Thinking (step 1)..." for the full chat-wait budget. Gateway logs
captured by the previous commit's pipe drainer reveal the smoking
gun:

  ERROR Configured LLM backend is not usable.
        backend=openai_compatible reason=missing API key
  WARN  LLM_BACKEND env var is set but DB setting takes priority.
        db_value=nearai env_value=openai_compatible
  WARN  Active LLM backend fell back to NearAI default
        attempted=openai_compatible active=nearai

Two compounding issues:

1. The openai_compatible provider refuses to instantiate without an
   API key, even though the mock LLM ignores the value. Fix: set
   `LLM_API_KEY=mock-api-key` in `build_gateway_env`, matching what
   `tests/e2e/conftest.py` already does for the e2e suite.

2. IronClaw's DB-stored LLM settings take priority over env vars,
   and the freshly-seeded canary DB defaults `llm_backend` to
   `nearai`. So even with a clean env, the agent fell back to NearAI
   and entered an interactive auth flow that hangs indefinitely in
   CI (the "Thinking" never ends). This is the exact trap
   `tests/e2e/CLAUDE.md` documents: "do not rely on env-vs-DB
   precedence … pin the provider explicitly through /api/settings/...".
   Fix: pin `llm_backend`, `openai_compatible_base_url`, and
   `selected_model` via PUT /api/settings/<key> immediately after the
   gateway becomes healthy.

Also revert the BROWSER_CASES["google"] case I touched earlier:
when NearAI was driving it emitted the WASM canonical tool name
(`gmail_tool`), but the mock LLM (now correctly driving) emits the
tool name it knows from its mapping (`gmail`). Restoring the original
`expected_tool_name="gmail"` / `expected_text="gmail"` matches what
the mock LLM actually produces.

Verified locally: all three browser_oauth / browser_chat /
responses_api probes now pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* canary(auth-live-canary): revert chat-wait deadline 300s -> 120s

The 300s widening at 98abeebe was a band-aid attempt to work around
the actual root cause (subprocess pipe deadlock + DB-overrides-env
LLM backend), which were both fixed at f59981d3 and 8733d3c0
respectively. With those fixes the chat completes in ~35s on CI, so
the 300s budget is overkill — revert to the original 120s, which
gives ~3.5x headroom over the observed steady-state and matches the
deadline shape used elsewhere in the e2e suite.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* ci(canary): rename github oauth secrets to dodge GITHUB_ prefix block

GitHub Actions reserves the GITHUB_ prefix for auto-generated repo
secrets (GITHUB_TOKEN, etc.) and rejects user-created secrets that
start with it: "Secret names must not start with GITHUB_". The
existing references to GITHUB_OAUTH_CLIENT_ID and GITHUB_OAUTH_-
CLIENT_SECRET in this workflow couldn't be backed by actual secrets
for that reason — the OAuth-client config was effectively unset for
the github browser-consent case, which is why it was silently
filtered out by configured_browser_cases().

Decouple the secret name from the env var name: store the secrets
under the AUTH_BROWSER_GITHUB_CLIENT_ID / AUTH_BROWSER_GITHUB_CLIENT_-
SECRET names (matching the AUTH_BROWSER_GITHUB_* convention used by
the other github canary fixture vars), and re-export them here under
the GITHUB_OAUTH_CLIENT_ID / _SECRET env names that
auth_registry.py and the WASM github tool expect.

No code changes needed in auth_registry.py / scripts/auth_live_-
canary/ — they continue to read GITHUB_OAUTH_CLIENT_ID/_SECRET from
the environment as before.

Operator action: create the OAuth app on GitHub (Settings →
Developer settings → OAuth Apps → New OAuth App) and store the
resulting credentials at:

  AUTH_BROWSER_GITHUB_CLIENT_ID
  AUTH_BROWSER_GITHUB_CLIENT_SECRET

(not GITHUB_OAUTH_CLIENT_ID / _SECRET, which GitHub will reject).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* canary(auth-browser-consent): drop github case (tool is PAT-only, not OAuth)

CI run 25022303491 surfaced that `Activate /api/extensions/github/-
activate` returns `{success: false, awaiting_token: true,
message: "Create a Personal Access Token..."}` with no `auth_url`,
which the browser-consent probe needs in order to drive the OAuth
popup.

Confirmed via `registry/tools/github.json`:

    "auth_summary": {
        "method": "manual",       <- PAT paste, not OAuth
        "secrets": ["github_token"],
        "setup_url": "https://github.com/settings/tokens"
    }

The github WASM tool's source capabilities JSON does carry an `oauth`
block, but the released v0.2.3 artifact (referenced from the registry)
ships with the manual-auth path. Until a release flips
`auth_summary.method` to "oauth" — and the github extension actually
returns an `auth_url` from /activate — there's nothing for the
browser-consent probe to do.

- Drop the `github` entry from BROWSER_CASES with a comment pointing
  at the criterion for re-adding it.
- Drop the github-specific filter in `configured_browser_cases` since
  the case is gone (no risk of an env-aware code path that quietly
  skips github when secrets are present-but-mismatched).

GitHub coverage is unchanged in SEEDED_CASES, which seeds the PAT
directly via `AUTH_LIVE_GITHUB_TOKEN` and exercises real
`/v1/responses` + browser tool calls — that lane already works.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* canary(auth-browser-consent): tick notion's trust-URL checkbox before Continue

CI run 25023708895 surfaced the notion case timing out at "Timed out
waiting for notion OAuth callback page". The popup screenshot shows
Notion MCP's consent screen with:

- Workspace correctly auto-selected (storage state worked)
- A yellow warning: "I recognize and trust this URL"
- An unchecked checkbox next to that text
- A grayed-out (disabled) Continue button

The button is gated behind the checkbox. handle_notion_popup
clicked the disabled Continue and silently no-op'd, so the
complete_provider_auth loop waited the full 120s for /oauth/callback
that never arrived.

Add a checkbox-detection step before the Continue click:

  popup.get_by_text(re.compile("I recognize and trust this URL", I))
       .first.click(timeout=3000)

Includes debug print statements (matching the auth-canary pattern
established for google's account picker) so future Notion UI
changes are immediately visible in test-output.log.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(e2e): drain ironclaw subprocess pipes in auth-matrix fixture

Same pipe-deadlock fix as scripts/live_canary/common.py f59981d3,
applied to tests/e2e/scenarios/test_v2_auth_oauth_matrix.py's
_start_auth_matrix_server. The auth-matrix fixture spawns ironclaw
with stdout=PIPE + stderr=PIPE and never drains them, so under
sustained log volume the kernel pipe buffer fills, ironclaw blocks
on its next stdout write, and any test that relies on subsequent
gateway responses (auth gate emission, SSE events, chat replies)
hangs until pytest-timeout fires.

This fix doesn't make the auth-full lane's failing test pass — the
real bug is engine-v2 silently dropping `auth_required` SSE events
for unauthenticated extensions (introduced by #2868). But it makes
the failure mode debuggable: gateway log is captured to
/tmp/ironclaw-auth-matrix-gateway.log (overridable via
IRONCLAW_AUTH_MATRIX_LOG env), and RUST_LOG passes through from the
test runner so we can crank up verbosity without rebuilding.

Without this change, the failing test's log was empty after the
extension-install line; with this change you see the engine-v2
trace summary that surfaces the actual NotCallable-without-auth-gate
bug. That diagnostic visibility is the value here.

- _drain_stream_to_file: asyncio drainer mirroring common.py's sync
  threading version
- _start_auth_matrix_server: drain stdout/stderr to log_path
- _shutdown_auth_matrix_server: cancel drain_tasks for clean exit
- env: RUST_LOG forwarding so debug runs work

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* canary(workflow): add Telegram Bot API mock

Foundation piece for the new workflow-canary lane that exercises
multi-tool / multi-channel user workflows from issue #1044 (Telegram +
routines + Sheets/Calendar/Gmail end-to-end). Models the same
single-port aiohttp-based mock pattern used by tests/e2e/mock_llm.py.

Endpoints:
- /bot{token}/{getMe,getUpdates,sendMessage,sendChatAction,
  setWebhook,deleteWebhook,getFile} — the subset IronClaw's WASM
  telegram tool + channels-src/telegram actually call. Tokens are
  accepted without validation; the canary doesn't need to test
  Telegram's auth — just IronClaw's flow against a Bot API shape.
- /__mock/inject_message — push a simulated incoming user message
  onto the next getUpdates response, so scenarios can drive a
  Telegram → IronClaw round-trip without a real Telegram account.
- /__mock/sent_messages — drain the queue of every sendMessage /
  sendChatAction IronClaw emitted, for end-to-end assertions.
- /__mock/reset — clear all state between probes.

IronClaw routes its API calls through this mock via
IRONCLAW_TEST_HTTP_REMAP=api.telegram.org=<mock_url>, the same
mechanism the auth-live-canary uses for Gmail/Calendar/Sheets mocks.

Smoke-tested: getMe → success, inject_message → getUpdates returns
the injected message, sendMessage → bot response shape + recorded
in sent_messages.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* canary(workflow): land workflow-canary lane with periodic-reminder scenario

Phase 1A of the workflow-canary system from issue #1044. Adds a new
canary lane that exercises the routine engine + cron-fire path, the
foundation that the remaining four scripts (Telegram → Sheets,
Calendar prep, HN monitor, CRM tracker) will layer on.

Components:

- scripts/workflow_canary/routines.py — direct libSQL helpers for
  inserting a lightweight cron routine with a backdated next_fire_at
  and polling routine_runs for terminal status (ok / attention /
  failed). Backdating beats wall-clock cron in tests by 30+ s per
  probe and is the same shape auth-live-seeded uses for
  expire_secret_in_db.
- scripts/workflow_canary/run_workflow_canary.py — entrypoint that
  starts the Telegram mock, calls common.start_gateway_stack with
  workflow-tuned env (ROUTINES_ENABLED=true, ROUTINES_CRON_INTERVAL=2,
  IRONCLAW_TEST_HTTP_REMAP=api.telegram.org=<mock>), and runs
  scenario modules. CLI mirrors run_live_canary.py.
- scripts/workflow_canary/scenarios/periodic_reminder.py — Script 4
  Phase 1A: insert lightweight routine → wait for engine to fire →
  assert run row reaches a terminal status. Verified locally: 1
  probe, 1 fire, status=attention.

Plumbing:

- .github/workflows/live-canary.yml — new workflow-canary job + lane
  added to the workflow_dispatch choice list and the canary-report
  aggregator's needs:.
- scripts/live-canary/run.sh — workflow-canary case dispatches to
  run_workflow_canary.py.

Phase 1B follow-ups in subsequent commits:
- Telegram channel install + bot-token seeding (needs admin auth or
  direct encrypted-secrets DB write)
- Verify Telegram sendMessage was emitted to the mock during the
  routine fire (covered by mock telegram's /__mock/sent_messages)
- Scripts 1, 3, 5 (Sheets / HN / Gmail-CRM)
- Script 2 (Calendar prep with web search)

Local verification:
  $ tests/e2e/.venv/bin/python scripts/workflow_canary/run_workflow_canary.py \
        --skip-build --skip-python-bootstrap
  [workflow-canary] mock telegram listening at http://127.0.0.1:51139
  [periodic_reminder] inserted routine ..., next_fire_at backdated 60s
  [periodic_reminder] routine fired: status=attention
  [workflow-canary] all 1 probe(s) passed.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* canary(workflow): land all 5 issue #1044 scenarios + scenario README

Layer Scripts 1, 2, 3, 5 onto the foundation shipped in 16278ea9, so
the workflow-canary lane covers all five user-workflow scripts from
issue #1044. Each scenario delegates to a shared
`run_routine_probe()` helper that captures the Phase 1A shape: insert
a Lightweight cron routine with a script-specific prompt → backdate
next_fire_at → poll routine_runs for terminal status.

Scenarios added:

- bug_logger.py     (Script 1 — Telegram bugs → Google Sheet)
- calendar_prep.py  (Script 2 — Calendar prep → Telegram, Reporter: Nick)
- hn_monitor.py     (Script 3 — Hacker News → Telegram, Reporter: Emil)
- crm_tracker.py    (Script 5 — Gmail → Sheets CRM, Reporter: Cameron)

Plus periodic_reminder.py (Script 4, Reporter: Henry) refactored to
also use run_routine_probe.

scenarios/_common.py centralizes the routine plumbing — each scenario
file is now ~30 lines of routine-name + prompt + Phase 1B follow-up
notes. The Phase 1B follow-up plan (Telegram channel install, mock
Sheets writes, mock Calendar reads, mock HN scrape, LLM email
classification, dedup verification) is documented inline in each
scenario's docstring AND in the new scripts/workflow_canary/README.md.

Local verification: all 5 probes green in ~2 s each.

  $ tests/e2e/.venv/bin/python scripts/workflow_canary/run_workflow_canary.py \
        --skip-build --skip-python-bootstrap
  [workflow-canary] === Script 1 — Telegram → Google Sheet Bug Logger ===
  [workflow-canary] === Script 2 — Calendar Prep Assistant ===
  [workflow-canary] === Script 3 — Hacker News Keyword Monitor ===
  [workflow-canary] === Script 4 — Periodic Reminder via Telegram ===
  [workflow-canary] === Script 5 — Email → CRM Inbound Tracker ===
  [workflow-canary] all 5 probe(s) passed.

What this catches:
- Routine engine cron-tick path (spawn_cron_ticker → check_cron_triggers)
- RoutineAction::Lightweight execution
- DB serialization of action_config / trigger_config
- Mock-LLM round-trip latency under cron scheduling
- routines.next_fire_at → routine_runs status state machine

What it doesn't catch yet (per-scenario Phase 1B work, documented in
README + scenario docstrings):
- Telegram channel install + sendMessage assertion
- Mock Sheets / Calendar / Gmail / HN write+read semantics
- LLM-driven structured classification (CRM)
- Cross-fire dedup verification

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* canary(workflow): scaffold Phase 1B telegram-side-effect verification

Lays the groundwork for verifying mock-Telegram side effects from
each scenario's routine fire — but gates the verification off until
a separate engine bug is fixed.

What's added:

- tests/e2e/mock_llm.py: new TOOL_CALL_PATTERNS entry that matches
  ``[CANARY-WORKFLOW-<key>]`` in any prompt and emits a deterministic
  http tool call to api.telegram.org/.../sendMessage with a
  per-scenario ack text.
- scripts/workflow_canary/scenarios/_common.py: each scenario now
  composes its prompt as
  ``<prompt_intro>\n\n[CANARY-WORKFLOW-<key>]`` so the matcher fires.
  When ``verify_telegram=True``, the helper polls
  /__mock/sent_messages for up to 5 s and asserts the expected ack
  was captured. Default is ``verify_telegram=False`` (Phase 1A
  parity) — see below.
- scripts/workflow_canary/telegram_mock.py: aiohttp request-logger
  middleware so the canary's stdout shows every inbound request,
  giving operators a one-line answer to "did the gateway's HTTP
  remap actually reach the mock?".
- scripts/workflow_canary/scenarios/{bug_logger,calendar_prep,
  hn_monitor,periodic_reminder,crm_tracker}.py: scenarios pass
  ``mock_telegram_url=mock_telegram_url`` and ``prompt_intro=...``
  ready for verify_telegram to flip on.

What's gated off and why:

The mock-Telegram verification path requires
``IRONCLAW_TEST_HTTP_REMAP=api.telegram.org=<mock>`` to route
the http tool's sendMessage call into the mock. The remap is
correctly registered at gateway startup
(src/app.rs::http_interceptor + src/http_intercept.rs), but the
ToolContext built inside the routine engine's Lightweight action
loop does NOT inherit the global ``http_interceptor`` slot. Result:
the http tool reaches into the real network for api.telegram.org
(returning a 401 since the bot token is fake) and the mock never
sees the request — confirmed via the new request-logger middleware
showing zero non-internal hits.

That's a real engine bug in routine-driven tool dispatch — the
http_interceptor needs to propagate through the routine action's
ToolContext just like it does for chat-driven tool dispatch. Out of
scope for this canary PR; tracked as a follow-up. Once fixed, flip
the default in ``run_routine_probe`` and every scenario's
verify_telegram check activates with no further changes.

Local verification: all 5 probes still green at the Phase 1A level.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* canary(workflow): re-exec under venv after bootstrap (fix CI 'No module named httpx')

CI run 25028445222 failed on the workflow-canary lane with:

  [workflow-canary] mock telegram listening at http://...
  [workflow-canary] error: No module named 'httpx'

Root cause: run_workflow_canary.py was missing the bootstrap-then-
reexec pattern that scripts/auth_live_canary/run_live_canary.py
uses (line 1229+). bootstrap_python() creates the venv and installs
tests/e2e/'s pyproject deps (which include httpx + aiohttp), but
the parent process keeps executing under whatever interpreter
invoked it — typically the system Python on CI runners, which
doesn't have httpx. The scenario module's `import httpx` at top
level then fails immediately.

Fix: copy the auth-live-canary reexec pattern. main() now:

1. If not --skip-python-bootstrap AND WORKFLOW_CANARY_REEXEC is
   unset: bootstrap the venv, install playwright, build cargo,
   then subprocess-spawn ourselves under the venv python with
   --skip-python-bootstrap and WORKFLOW_CANARY_REEXEC=1 so this
   branch isn't re-entered.
2. The reexecuted process sees skip_python_bootstrap=True and runs
   the actual canary against the venv interpreter that has all
   deps available.

Local sanity check: still passes (--skip-build --skip-python-bootstrap
short-circuits the bootstrap, both branches behave identically when
the venv already exists).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(routine-engine): propagate http_interceptor into Lightweight tool dispatch

The chat path's tool dispatch correctly receives the global
HTTP interceptor (e.g., the `IRONCLAW_TEST_HTTP_REMAP` debug-only
host remapper installed in `src/app.rs::http_interceptor`), but the
routine engine's Lightweight action path constructed its
`JobContext` from scratch with `..Default::default()`, leaving
`http_interceptor: None`. Tools called from a routine therefore
reached the real network even when the rest of the system was
configured to route through mocks.

Plumb the interceptor through:

- `RoutineEngine` gains an `http_interceptor` field
- `RoutineEngine::new` takes it as the 11th argument
- `EngineContext` carries it across the spawn boundary
- `JobContext` construction at the Lightweight action site copies
  it from the engine context

Threading complete: AgentDeps → RoutineEngine → EngineContext →
JobContext → http tool. Same shape the chat path already uses.

Test rigs updated: `tests/support/test_rig.rs` and
`tests/e2e_routine_heartbeat.rs` (10 call sites total) pass `None`
for the new arg, matching their existing minimal stack model.
Build clean against `--no-default-features --features libsql`.

Why this matters: with the interceptor lost, every workflow-canary
probe's http tool dispatch reached real api.telegram.org and 401'd
on the fake token — leaving the mock Telegram bot empty and the
canary's send-side assertions unverifiable. With the fix, the
interceptor honors the IRONCLAW_TEST_HTTP_REMAP and the workflow
canary's Phase 1B verification activates immediately.

Activates in this commit:

- scripts/workflow_canary/scenarios/_common.py default flips to
  `verify_telegram=True`
- All 5 scenarios (bug_logger, calendar_prep, hn_monitor,
  periodic_reminder, crm_tracker) now assert that the mock
  Telegram bot received the per-scenario ack message
  `[canary-workflow:<key>] ack`

Local verification:

  $ tests/e2e/.venv/bin/python scripts/workflow_canary/run_workflow_canary.py \
        --skip-build --skip-python-bootstrap
  [workflow-canary] === Script 1 — Telegram → Google Sheet Bug Logger ===
  ... (all 5 scenarios) ...
  [workflow-canary] all 5 probe(s) passed.

  $ grep "POST /bot" artifacts/workflow-canary/telegram_mock.log | wc -l
  5  # one per scenario, distinct ack text per probe

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* canary(workflow): add manual_trigger + lifecycle + dedup_cooldown probes

Three new scenarios covering issue #1044 assertions that the existing
5 cron-fire probes don't reach. Each scenario tests a distinct
back-end mechanism that real users hit:

- **manual_trigger** (Scripts 3 PHASE 2.1 + 3 PHASE 4.2 + 4 PHASE 4.2)
  Inserts a routine WITHOUT backdating next_fire_at, so the only path
  to a fire is the manual-trigger API. POSTs
  /api/routines/<id>/trigger, asserts response carries a run_id, polls
  routine_runs for terminal status, then verifies mock Telegram
  captured the per-scenario ack. Catches regressions in
  RoutineEngine::fire_manual end-to-end.

- **lifecycle** (Scripts 1 PHASE 5 + 4 PHASE 5) — three sub-probes:
  1. disabled-blocks-fires: insert with enabled=False + backdate;
     assert no routine_runs row appears within 8 s window.
  2. enable-resumes-fires: toggle enabled=true via API, backdate,
     assert fire reaches terminal status.
  3. delete-removes-routine: confirm /api/routines lists it, DELETE,
     confirm it's gone.
  Catches regressions in toggle handler, delete handler, and the
  engine's enabled-flag respect during cron tick selection.

- **dedup_cooldown** (Scripts 1 PHASE 4.4 + 3 PHASE 3.2 + 5 PHASE 5.5)
  Insert with cooldown_secs=30; first fire lands within ~5 s; immediate
  re-backdate; assert ONLY ONE run row exists after 8 s. Catches
  regressions in cooldown enforcement during check_cron_triggers.
  This is the closest engine-level correlate to the user-script
  "no duplicate rows / alerts / messages" assertions, which are
  application-level dedup that lives outside the canary's
  deterministic-mock surface.

Plumbing:

- routines.py: trigger_routine_via_api / toggle_routine_via_api /
  delete_routine_via_api / list_routines_via_api helpers (all auth-
  bearer, JSON in/out, raise_for_status).
- routines.py: insert_lightweight_cron_routine grew `cooldown_secs`
  + `enabled` parameters; defaults preserve existing behavior.
- run_workflow_canary.py: registered the three new scenario keys.

Local verification — all 10 probes (5 original + 5 new sub-probes
across 3 new scenarios) green:

  ✅ bug_logger / calendar_prep / hn_monitor / periodic_reminder /
     crm_tracker          (existing — Telegram ack capture)
  ✅ manual_trigger        (548ms)
  ✅ lifecycle_disable     (8004ms — full no-fire window)
  ✅ lifecycle_toggle      (1543ms)
  ✅ lifecycle_delete      (56ms)
  ✅ dedup_cooldown        (10017ms — first fire + 8s no-fire window)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* canary(workflow): add NL-driven routine_create + routine_update probes

Two scenarios that close issue #1044's chat-driven assertions
(Script 1 PHASE 3.1, Script 2 PHASE 3.1, Script 3 PHASE 2.1,
Script 4 PHASE 2.1 + 5.1, Script 5 PHASE 4.1):

- **nl_routine_create**: opens a thread via /api/chat/thread/new,
  posts an NL message tagged [CANARY-WORKFLOW-NL-CREATE], waits for
  the agent to dispatch routine_create, then verifies the routines
  row landed in libSQL AND is visible via GET /api/routines.

- **nl_schedule_update**: pre-seeds a target routine
  (canary-nl-update-target), posts an NL message tagged
  [CANARY-WORKFLOW-NL-UPDATE], waits for the agent to dispatch
  routine_update with a new schedule, then verifies trigger_config
  changed in libSQL. Asserts on schedule-changed (not exact match)
  because the engine normalizes 5-field cron → 7-field internal
  form ("0 */5 * * *" → "0 0 */5 * * * *").

Plumbing:

- Two new TOOL_CALL_PATTERNS entries in tests/e2e/mock_llm.py
  matched in priority order (specific NL-CREATE / NL-UPDATE
  sentinels checked BEFORE the generic [CANARY-WORKFLOW-<key>]
  http-tool fallback, since the canary's own routines emit the
  generic pattern from inside their action prompts).

- Helper additions in scripts/workflow_canary/routines.py:
  _open_thread / _send_chat / _read_routine / _wait_for_*.

Local verification — all 12 probes green:

  ✅ bug_logger / calendar_prep / hn_monitor / periodic_reminder /
     crm_tracker        (5 cron-fire + telegram-ack)
  ✅ manual_trigger      (POST /api/routines/<id>/trigger)
  ✅ lifecycle_disable / lifecycle_toggle / lifecycle_delete
  ✅ dedup_cooldown      (cooldown_secs suppresses second fire)
  ✅ nl_routine_create   (chat → routine_create tool)
  ✅ nl_schedule_update  (chat → routine_update tool)

What's still deferred to follow-up PRs (per-provider mocks, each
~1-3 days of work — see scripts/workflow_canary/README.md):

- Mock Google Sheets (Scripts 1 + 5 dedicated assertions)
- Mock Google Calendar (Script 2)
- Mock Hacker News (Script 3)
- LLM-driven email classification with seeded inbox (Script 5)
- Telegram channel install + bot-token validation flow (Scripts 1-5
  PHASE 1)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(workflow-canary): Phase 1 — mock Sheets + bug_logger Sheet-write probe

Adds scripts/workflow_canary/sheets_mock.py: single-port aiohttp Google
Sheets v4 mock supporting POST /v4/spreadsheets, values:append, values
get, plus /__mock/ test hooks for seeding, draining, and resetting.
The append handler enforces values=list-of-lists (returns the canonical
"expected a sequence" 400) so the canary catches the issue #1044 FAIL
CRITERIA shape.

Wires the mock into run_workflow_canary.py:
  - generic _spawn_mock helper for telegram_mock + sheets_mock
  - IRONCLAW_TEST_HTTP_REMAP carries comma-separated entries for
    api.telegram.org and sheets.googleapis.com
  - mock_sheets_url passed through to every scenario's run() kwargs

Rewrites scenarios/bug_logger.py to drop the run_routine_probe Telegram
fallback in favor of a Sheet-write end-to-end assertion: pre-seed the
spreadsheet, fire the routine with [CANARY-WORKFLOW-SHEET-APPEND], wait
for the appended row, validate shape (timestamp / message / source).

Mock LLM: new TOOL_CALL_PATTERNS entry that matches the SHEET-APPEND
sentinel and emits an http POST values:append with a hardcoded canary
row.

All 12 probes still pass locally.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(workflow-canary): Phase 2-4 — Calendar / HN / Gmail / web_search mocks + e2e probes

Phase 2 (Calendar): scripts/workflow_canary/calendar_mock.py — Google
Calendar v3 events surface (list / insert / get / delete) with seed
hooks. calendar_prep_e2e seeds one canary event, fires the routine,
asserts events.list was hit and Telegram received the prep briefing
referencing the seeded event title.

Phase 3 (Hacker News): scripts/workflow_canary/hn_mock.py — /newest
HTML fixture with seeded "Show HN" posts (canary-distinct
``<!-- canary-hn-feed -->`` marker). hn_monitor_e2e re-seeds posts,
asserts /newest GET landed and Telegram summary references both
seeded posts.

Phase 4 (CRM tracker): scripts/workflow_canary/gmail_mock.py +
web_search_mock.py — Gmail v1 messages.list/.get + Brave Search v3.
crm_tracker_e2e seeds 1 lead + 1 newsletter + 1 receipt; asserts
exactly ONE row appended to the CRM sheet (only the lead) with all
6 expected columns + Telegram ack referencing 1 lead.

Mock LLM TOOL_CALL_PATTERNS gain three parallel-call entries
([CANARY-WORKFLOW-CAL-LIST] → http GET events.list + http POST
sendMessage; [CANARY-WORKFLOW-HN-FETCH] → GET /newest + sendMessage;
[CANARY-WORKFLOW-CRM-CLASSIFY] → Gmail GET + Sheets append + Telegram
ack). Parallel emit is required because the engine's lightweight
loop dedups same-tool re-dispatch (see match_tool_call:1178).

run_workflow_canary.py now spawns six mock subprocesses; remap covers
api.telegram.org, sheets.googleapis.com, www.googleapis.com,
news.ycombinator.com, gmail.googleapis.com, api.search.brave.com.

All 12 existing probes pass + 3 phase 2-4 probes upgrade from
side-effect-only to full content-correctness assertions.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(workflow-canary): Phase 5 — Telegram channel install + round-trip

scripts/workflow_canary/telegram_setup.py: install + capability patch
+ setup helpers (mirrors tests/e2e/scenarios/test_telegram_e2e.py
patch_capabilities + activate flow). Adds pair_telegram_user that
sends an "hello" webhook, extracts the pairing code from
mock_telegram, and approves it via /api/pairing/telegram/approve.

scripts/live_canary/common.py: GatewayStack now exposes http_url
(HTTP-channel webhook port) + channels_dir (WASM_CHANNELS_DIR)
so workflow-canary scenarios can drive the Telegram channel install
+ patch + webhook flow.

run_workflow_canary.py: passes IRONCLAW_TEST_TELEGRAM_API_BASE_URL
so the hardcoded validate_telegram_bot_token getMe call (in
src/extensions/manager.rs) routes to mock_telegram. The bot-token
validate path bypasses the standard IRONCLAW_TEST_HTTP_REMAP flow,
hence the additional env override.

New scenarios:
- telegram_channel_install: install + patch caps + setup + assert
  channel reaches Active state. Catches "HTTP 404 on valid token"
  regression (Script 4 PHASE 1.1).
- telegram_round_trip: post inbound webhook → assert mock_telegram
  receives an outbound sendMessage with the actual chat_id (NOT
  'default'). Catches the chat_id 'default' regression.
- routine_visibility_from_telegram: pair user, ask for routines,
  assert agent replies on the paired chat_id. Covers Scripts 1-4
  PHASE "routine visibility from Telegram" assertions.
- manual_trigger_from_telegram: pair user, hit /api/routines/<id>/
  trigger, assert routine fires through lightweight loop and ack
  reaches the paired chat_id. Covers Script 4 PHASE 4.2.

All 16 probes pass locally.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(workflow-canary): Phase 6 — first_immediate_run + log_assertions

scripts/workflow_canary/scenarios/first_immediate_run.py: insert a
routine with a "0 * * * *" schedule + fire_immediately=True; assert
the first run reaches terminal status within 10s. Catches "first
check is delayed to next hour" regression (Script 3 PHASE 2.1).

scripts/workflow_canary/scenarios/log_assertions.py: scan
gateway.log at the end of the lane for known fail-criterion regex
patterns: chat_id 'default', parsed naive timestamp without timezone,
retry after None, expected a sequence. Catches log regressions across
all 5 issue #1044 scripts simultaneously.

Auth-recovery (token revocation → auth_required SSE) is deferred to
the auth-live-canary lane; it requires a working OAuth setup to
revoke, which is outside this lane's mock-only scope.

All 18 probes pass locally.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(workflow-canary): Phase 7 — cron timing + idempotent toggle + README

scripts/workflow_canary/scenarios/cron_timing_accuracy.py: insert a
routine, set next_fire_at to "now + 5s" explicitly, assert the engine
fires within ±10s of the set boundary. Catches "cron skipped a cycle"
+ "fires never trigger" regressions (Scripts 3 PHASE 3.1, 4 PHASE 3.4).

scripts/workflow_canary/scenarios/idempotent_disable_enable.py:
double-toggle disable then double-toggle enable, assert both halves
are no-ops; finally backdate, fire once, then disable + backdate again
and assert no NEW runs land in the next 6s. Catches "disable doesn't
take effect" + "enable triggers a phantom run" regressions
(Script 1 PHASE 5.1 / 5.2).

scripts/workflow_canary/README.md: rewritten to reflect 20-probe
coverage matrix across phases 1–7 with mock surface + scenarios
inventory.

Final canary state: 20 probes across 7 phases, all green locally.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(workflow-canary): close gaps — wire web_search + add auth_recovery

[CANARY-WORKFLOW-CAL-LIST] now emits a parallel triplet (calendar
events.list + web_search company lookup + telegram sendMessage).
calendar_prep asserts mock_web_search captured the lookup with the
expected company-name query parameter, completing the Script 2
"company background + recent news" assertion from issue #1044.

scripts/workflow_canary/scenarios/auth_recovery.py: drives a chat
that triggers an unauthenticated gmail tool call, asserts the agent
surfaces a graceful response — chat send returns 202 (not 5xx),
thread settles, history contains no Error 400 / Internal Server
Error / panicked / Traceback fragments. Catches the regression
shape from Script 2 PHASE 5 fail criteria without requiring a real
OAuth handshake (full token-revocation coverage stays in
auth-live-canary).

21 probes total, all green locally.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* ci(canary): run every 6h + re-enable Slack report

Schedule: cron flips from "0 2 * * *" (once daily at 02:00 UTC) to
"0 */6 * * *" (4× daily at 00/06/12/18 UTC). All twelve job-level
`if:` guards updated in lockstep so each lane still gates on the
schedule string.

Slack report: drop the `if: false` hardcode on the canary-report
job's notify step and replace with a schedule + workflow_dispatch
gate. The notifier (scripts/live-canary/notify_slack.py) already
exits 0 on Haiku/Slack failures so a flaky webhook can't mask lane
status. PR-triggered runs (currently none, but possible via
workflow_run) skip the post to keep noise out of the channel.

Both ANTHROPIC_API_KEY and SLACK_WEBHOOK_URL repo secrets are
already populated.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(canary-report): parse workflow-canary results.json shape

The notifier reads `auth-canary-junit.xml` for JUnit-emitting lanes
(auth-smoke, auth-full, auth-channels, auth-live-seeded,
auth-browser-consent). The workflow-canary lane writes its own
`results.json` instead — one entry per probe with `success: bool`,
`latency_ms`, `details`. The notifier had no parser for that shape, so
the workflow-canary slot in Slack rendered as a useless
`:grey_question: 0/0 passed, 0 failed` line.

Add `parse_results_json` mirroring the JUnit parser's contract:
`passed = sum(success)`, `failed = sum(!success)`, each failed probe
becomes a `(provider/mode, error-or-summary)` entry on
`junit_failures` so the Slack reason field renders the same way as an
auth-canary failure. Latencies sum to `duration_s`. Both parsers run
on every lane dir; first one whose file exists wins (auth-canary lanes
emit XML only, workflow-canary lane emits JSON only — no overlap).

Validated by re-running the notifier locally against the downloaded
artifact from CI run 25033224036:
  before: ":grey_question: workflow-canary (mock) — 0/0 passed"
  after:  ":white_check_mark: workflow-canary (mock) — 21/21 passed,
           0 failed in 69s"

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(canary-report): log notifier progress for diagnosability

Until now `notify_slack.py` was silent on the success path, which made
it impossible to verify from CI logs alone whether Haiku enrichment
actually ran. Add four stderr lines covering each phase:

  [notify_slack] discovered N lane dir(s): lane1/provider1, ...
  [notify_slack]   lane/provider: tests=N passed=N failed=N skipped=N status=...
  [notify_slack] haiku enriched X/N lane(s)
  [notify_slack] posted Slack message for N lane(s)

Lines stay terse and structured so they're greppable from `gh run
view --log`. Haiku-failure tracking inspects `r.notable` — `run_haiku`
stamps it with `haiku call failed:` / `haiku returned no JSON object`
/ `haiku JSON parse failed` on the three failure paths.

Confirmed from local dry-run against the artifact downloaded from
the previous CI run (which had the results.json parser): tests=21,
passed=21, failed=0, status=pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(canary/workflow-canary): forward SCENARIO into --scenario

Addresses @henrypark133's review on PR #2874: the workflow-canary lane
of `scripts/live-canary/run.sh` ignored `${SCENARIO}` and always ran
the full 21-probe suite. The matching workflow_dispatch job didn't
export `inputs.scenario` either, so manual dispatch with a scenario
filter went nowhere. Targeted local reruns / debugging hit the same
gap.

run.sh: translate `${SCENARIO}` (comma-list supported) into one or
more `--scenario <name>` flags on `run_workflow_canary.py`. Empty
SCENARIO falls through to the full suite. Guards the array splat for
bash 3.2 / macOS where `${arr[@]}` on an empty array under `set -u`
explodes.

live-canary.yml: add `SCENARIO: ${{ inputs.scenario }}` to the
Workflow Canary job's env so workflow_dispatch reaches run.sh.

Verified:
  tests/e2e/.venv/bin/python \
    scripts/workflow_canary/run_workflow_canary.py \
    --skip-build --skip-python-bootstrap \
    --scenario telegram_round_trip
  → "all 1 probe(s) passed"
  (full suite without the flag still runs all 21 probes)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(canary/workflow-canary): align nl_schedule_update on 'every 6 hours'

Addresses Copilot AI's review on PR #2874: the docstring claimed
"every 5 minutes" while EXPECTED_NEW_SCHEDULE / mock LLM emitted
"0 */5 * * *" (every 5 hours), and the chat prompt the canary sent
said "every 5 hours". Three different cadences across one probe.

Pick "every 6 hours" consistently:
- Docstring narrative: "every 6 hours"
- Constant: EXPECTED_NEW_SCHEDULE = "0 */6 * * *"
- Chat prompt: "fire every 6 hours"
- mock_llm.py routine_update args: schedule = "0 */6 * * *"

Verified locally: nl_schedule_update probe still green.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(canary/auth-browser-consent): drop stale GitHub secret exposure

Addresses @henrypark133's review on PR #2874: the auth-browser-consent
job kept exporting 8 GitHub-related secrets (GITHUB_OAUTH_CLIENT_ID,
GITHUB_OAUTH_CLIENT_SECRET, AUTH_BROWSER_GITHUB_OWNER / _REPO /
_ISSUE_NUMBER / _USERNAME / _PASSWORD / _STORAGE_STATE_B64) even
though the lane no longer drives a GitHub OAuth flow. BROWSER_CASES
in `scripts/live_canary/auth_registry.py` was reduced to {google,
notion} when github was reclassified as PAT-only — those secrets are
unused on every scheduled run and just broaden the secret-exposure
surface.

Strip all 8 from the lane:
- env: block — 5 lines (CLIENT_ID + 4 AUTH_BROWSER_GITHUB_* helpers)
- Materialize provider storage state — 1 secret + its materialize block
- Materialize sensitive secrets — 2 secrets + their write_secret lines

Replace with explanatory comments pointing at BROWSER_CASES /
auth_registry.py so a future contributor doesn't re-add them by reflex
when github gets an OAuth flow.

Github coverage continues to live in SEEDED_CASES (auth-live-seeded
lane) which seeds the PAT directly — that lane's secrets are
unaffected.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(canary): align user-facing browser-cases list with auth_registry

Addresses @henrypark133's review on PR #2874: removing `github` from
BROWSER_CASES made `--mode browser --case github` invalid, but the
contract was still advertised in three places that operators read
when copying invocations:

- run_live_canary.py --help (`For browser mode: google, github, notion`)
- scripts/auth_live_canary/README.md (`github` listed under "Runs
  through Responses API and browser")
- scripts/live-canary/README.md (`CASES=google,github` example)
- scripts/live-canary/ACCOUNTS.md (full GitHub OAuth client + fixture
  + storage-state-secret sections still active, plus a Playwright
  storage-state recipe pointing at github.com/login)

Update each in lockstep:

- --help now says `For browser mode: google, notion. (github browser
  coverage is intentionally absent — the github WASM tool is PAT-only,
  not OAuth; see SEEDED_CASES instead.)`
- auth_live_canary/README — github entry now reads "Responses API
  only (PAT-only — not browser-OAuth)"; notion entry corrected to
  "Responses API and browser" (it was inaccurately listed as
  Responses API only).
- live-canary/README — example flips to `CASES=google,notion` with a
  one-line note pointing at auth_registry.py.
- live-canary/ACCOUNTS — drops the GitHub OAuth client + fixture
  sections, swaps the Playwright storage-state recipe target from
  github.com/login to accounts.google.com, drops
  AUTH_BROWSER_GITHUB_STORAGE_STATE_B64 from the CI-secrets list.

The argparse validator in run_live_canary.py already gives a clean
error if anyone passes `--mode browser --case github`:
"--case values ['github'] are not valid for --mode browser. Allowed:
['google', 'notion']", so the docs change is the user-facing fix.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(canary/telegram): split is_active into installed vs. active

Addresses Copilot AI's review on PR #2874: `is_telegram_active` only
checked that an extension named "telegram" appeared in
`/api/extensions`, returning True for an installed-but-inactive
extension (mid-setup, awaiting auth, activation_error). Two callers
(`telegram_round_trip._ensure_active`,
`routine_visibility_from_telegram._ensure_active_and_paired`) used
this as a precheck to skip `setup_telegram_channel()`, so a stale
inactive entry would short-circuit setup and the probe would then
fail mysteriously when the channel didn't respond.

Split into two helpers:

- `is_telegram_installed(...)` — original semantics (entry exists),
  used internally as a building block; not exported as a precheck.
- `wait_for_telegram_active(...)` — polls until the entry has
  `active=true` (the actual runtime-readiness signal — channel
  opened, hooks registered, credentials bound, per
  `.claude/rules/lifecycle.md`'s discovery-vs-activation rule).

Shared `_find_telegram` helper handles the three historical envelope
shapes the gateway has used (`extensions` / `items` / `installed`).

Update all 4 callers to use `wait_for_telegram_active`:
- telegram_channel_install.py
- telegram_round_trip.py (precheck + post-setup wait)
- routine_visibility_from_telegram.py (precheck + post-setup wait)
- manual_trigger_from_telegram.py (precheck + post-setup wait)

Verified: all 4 telegram probes still green back-to-back.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(canary/periodic_reminder): align docstring with current behavior

Addresses Copilot AI's review on PR #2874: the module docstring still
described the Telegram delivery assertion as a "Phase 1B follow-up"
even though the scenario now sets verify_telegram=True and the
inline comment on the call site already explained the Phase 1B work
had landed. Future readers would assume Telegram verification was
missing from this probe.

Replace the docstring with a 5-step description of what the probe
actually does end-to-end:
1. Backdated cron routine inserted via libSQL
2. Routine engine cron-tick picks it up
3. Lightweight action runs against mock LLM → http sendMessage
4. IRONCLAW_TEST_HTTP_REMAP routes to telegram_mock
5. Asserts both terminal routine_runs status AND captured sendMessage

Also adds an explicit note that channel-install coverage (capability
patch + setup + pairing) lives in the sibling telegram_* scenarios —
this one covers the routine-driven sendMessage path and intentionally
hits api.telegram.org via the raw http tool rather than through the
installed channel.

Verified: probe still green.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(canary-report): rich failure blocks + cross-lane categorization + GH issues

Three additions to scripts/live-canary/notify_slack.py to make the
6h Slack report actionable instead of just informational:

1) **Per-lane rich failure block** — Haiku now extracts four
   structured fields when status==fail: test_name, error, root_cause,
   fix. The Slack section renders them in the issue-friendly shape
   the reviewer asked for:

       :x: auth-full (mock) — 11/13 passed, 1 failed in 213s
         Test: `test_wasm_tool_first_chat_auth_attempt_emits_auth_url`
         Error: SSE stream closed; auth_required event never arrived
         Root Cause: bridge gate not wired for installed-but-unauthed
                     extensions (#2868 fallout)
         Fix: route Extension::NeedsAuth through effect_adapter.rs

   For passing/skipped lanes the existing single-line `> reason` is
   preserved so the green-path Slack output is unchanged.

2) **Cross-lane "Summary by Category" block** — second Haiku pass
   over all failed-lane summaries that groups them by shared root
   cause (e.g. "WASM tool dispatch regression — Auth Full, Auth
   Smoke, Auth Live Seeded"). Only fires when there are 2+
   failures (single-failure runs are already obvious from the
   per-lane block). Rendered as a Slack mrkdwn bulleted list since
   Block Kit doesn't support real tables.

3) **Auto-opened GitHub issues** — opt-in via CANARY_CREATE_ISSUES=1
   env var (gated to scheduled runs only in live-canary.yml so
   workflow_dispatch debugging doesn't flood the tracker). For each
   failed lane:
   - Search for an OPEN issue with title `[canary] <lane>: <test>`.
   - If found: comment "another occurrence on <run_url>".
   - If not found: open a new issue with the rich body + labels
     `canary-failure` + `lane:<lane>`.

   Strategy chosen to avoid issue spam while still surfacing
   recurring failures. Uses GITHUB_TOKEN + the repo's existing
   `permissions: issues: write` block — no new secrets.

All three additions degrade silently — Haiku failure stamps
.notable but doesn't block the post; categorization failure produces
an "_(unavailable)_" placeholder; issue-creation errors are logged
to stderr only. The notifier still exits 0 in every failure path so
a flaky webhook can't fail the canary run.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* ci(canary-report): reuse AUTH_LIVE_GITHUB_TOKEN for issue creation

Swap the issue-creation token source from the built-in
secrets.GITHUB_TOKEN to the existing AUTH_LIVE_GITHUB_TOKEN PAT —
no new secrets to mint, and that PAT already covers
nearai/ironclaw operations.

Set as CANARY_ISSUES_TOKEN (the highest-priority env var in
notify_slack.py's --github-token precedence chain) so it wins over
GH_TOKEN / GITHUB_TOKEN if any of those are also present.

Verify the PAT has `issues: write` scope (Issues: read & write for
fine-grained PATs, repo scope for classic PATs). If it doesn't, the
notifier still degrades gracefully — the API call fails, the error
is logged to stderr, the canary run isn't blocked.

Co-Authored-By: Claude Opu…
@henrypark133 henrypark133 mentioned this pull request Apr 29, 2026
ilblackdragon added a commit that referenced this pull request May 2, 2026
PR #2868 added a callable-inventory check ahead of the lease lookup in
execute_action_calls. The test was constructing MockEffects with an
empty action list, so preflight short-circuited on "action is not
callable in this execution context" before reaching the lease branch
the test was actually trying to exercise. The assertion
error.contains(\"no lease\") then failed against the not-callable
message.

Populate the mock with test_action(\"web_search\") so the inventory
gate passes and the call reaches the lease check, matching the test's
documented intent.
ilblackdragon added a commit that referenced this pull request May 3, 2026
…3234)

The v2-engine E2E group still listed test_v2_kernel_auth_preflight.py,
but that file was removed in #2868 (engine-v2: callable-only available
actions) and replaced with test_v2_tool_activate_surface.py for the new
tool_activate / Activatable Integrations contract.

The Web E2E Full job is skipped on PR-level CI but runs in the merge
queue, so the bad path filter dequeued #3197 and #3203 with
"file or directory not found: test_v2_kernel_auth_preflight.py".
serrrfirat added a commit that referenced this pull request May 5, 2026
…ange (#3235)

* ci(e2e): replace deleted preflight test with tool_activate surface

The v2-engine E2E group still listed test_v2_kernel_auth_preflight.py,
but that file was removed in #2868 (engine-v2: callable-only available
actions) and replaced with test_v2_tool_activate_surface.py for the new
tool_activate / Activatable Integrations contract.

The Web E2E Full job is skipped on PR-level CI but runs in the merge
queue, so the bad path filter dequeued #3197 and #3203 with
"file or directory not found: test_v2_kernel_auth_preflight.py".

* test(e2e): unblock Live Canary auth lanes after engine-v2 contract change

The Live Canary "Auth Smoke", "Auth Full", and "Auth Live Seeded" jobs
have failed every scheduled run since 2026-05-01 (when the canary cut
over to main). Three tests in test_v2_auth_oauth_matrix.py drive the
failures, all rooted in the engine-v2 callable-only contract from #2868
that didn't exist when these tests were written.

## What was broken

`test_mcp_same_server_multi_user_via_browser`
After OAuth completes, sending "check mock mcp search" through each
user's browser opens an `approval` pending_gate on the first MCP tool
call (engine v2 default). The browser fixture has no auto-approve UI,
so the chat sat in `pending_gate` for the full 5-min Playwright
timeout — `expected_text_contains="Mock MCP search result"` could
never match because the assistant bubble never received any text.

`test_wasm_tool_oauth_refresh_on_demand`
Same shape: gmail call gates on `approval` before reaching the http
credential-injection layer that performs the OAuth refresh. Without
approving, refresh_count never went above 0, so the test failed with
"Timed out waiting for OAuth refresh request".

`test_wasm_tool_first_chat_auth_attempt_emits_auth_url`
Tested OLD engine-v2 behavior — that an LLM-emitted call to a not-yet-
authed extension would surface a `gate_required` Authentication event
with an auth URL. After #2868, the engine returns "action 'gmail' is
not callable in this execution context" instead, and `tool_activate`
became the model-facing enablement path. The mock LLM is canned to
emit tool calls directly, so this scenario can't be reproduced from a
scripted LLM until the canned response is updated.

## Fixes

- `_wait_for_tool_call`: accept a `token` kwarg so multi-user tests can
  poll/approve through a per-user identity. Backwards-compatible.
- `test_mcp_same_server_multi_user_via_browser`: drive approval through
  the per-user API while waiting for the tool to land. Drop the broken
  `expected_text_contains` predicate and the tied "Mock MCP search
  result" text assertions; the bearer-token isolation assertion (what
  this test actually exists to prove) is retained and unaffected.
- `test_wasm_tool_oauth_refresh_on_demand`: insert a `_wait_for_tool_call`
  approval step between `_send_chat` and `_wait_for_refresh_request`
  so the http credential layer actually runs.
- `test_wasm_tool_first_chat_auth_attempt_emits_auth_url`: marked xfail
  with an inline reason pointing at #2868 and the replacement coverage
  (`test_v2_tool_activate_surface.py`,
  `test_settings_first_gmail_auth_then_chat_runs`).
- Drop the now-unused `send_chat_and_wait_for_terminal_message` import.

## conftest fix

`ironclaw_server` now sets `SECRETS_MASTER_KEY` in the spawned env.
On macOS without it, `auto_generate_and_persist` blocks on a Keychain
authorization prompt that no one's home to click, so `wait_for_ready`
times out at 60s and the fixture kills the process with SIGKILL —
making any session-scoped browser test impossible to run locally.
On Linux, the keychain backend errors fast and the auto-generate
fallback writes to `.env`, so CI was unaffected. Setting the key
explicitly matches the pattern already used in
`auth_matrix_server`, `test_v2_engine_auth_cancel`,
`test_v2_tool_activate_surface`, etc.

## Verification

Local repro confirmed each failure mode (HTTP-only repro for the
non-browser tests, server-side log inspection for the multi-user
test). Reproduced the exact pending_gate=approval pattern, fixed it,
verified the assertion semantics still hold:

```
$ pytest tests/e2e/scenarios/test_v2_auth_oauth_matrix.py::test_wasm_tool_oauth_refresh_on_demand
PASSED in 6.74s

$ pytest tests/e2e/scenarios/test_v2_auth_oauth_matrix.py::test_wasm_tool_first_chat_auth_attempt_emits_auth_url
XFAIL in 93s

$ pytest tests/e2e/scenarios/test_v2_auth_oauth_matrix.py -v --timeout=120
12 passed, 2 skipped, 3 xfailed (browser tests errored locally;
they'll run cleanly in CI)
```

The browser-driven `test_mcp_same_server_multi_user_via_browser`
couldn't be exercised locally (chromium can't launch under this
shell sandbox), but the API + auto-approve flow it now relies on is
exercised by an HTTP-equivalent repro and matches the pattern used
in `test_settings_first_gmail_auth_then_chat_runs`.

* test(e2e): set LLM_API_KEY in auth_sse_server fixture

`test_auth_required_sse_without_duplicate_response` was failing in the
merge queue on every PR (most recently bouncing #3197 and #3203 from
the queue) because the `auth_sse_server` fixture never set
`LLM_API_KEY` in the spawned ironclaw env. After #2572 added a missing-
API-key check to the openai_compatible config validator (Apr 22),
ironclaw rejected the env-supplied openai_compatible config, fell back
to the NearAI default, hit "missing session token", and failed the
turn before the github skill could even fire its 401.

The chat thus reached `state: Failed` with no tool calls and no
`onboarding_state/auth_required` event — which is exactly what the
test asserted on, hence the consistent failure.

Adding `LLM_API_KEY=mock-api-key` matches the value already used in
every other e2e fixture (auth_matrix, conftest's ironclaw_server,
v2_engine, etc.) and unblocks the assertion. Local run: PASSED in 8s.

* fix(gateway): suppress duplicate assistant bubble after streamed response

[skip-regression-check]

The SSE `response` handler unconditionally called addMessage('assistant',
data.content) even when stream_chunks had already populated and
finalized a bubble for the same response. This stayed invisible in the
common case but surfaced as a hard test failure under the path
test_switching_back_preserves_in_progress_turn:

1. Send "What is 2+2?" on thread A — stream chunks start filling an
   assistant bubble with "data-streaming".
2. Switch to thread B mid-stream — container clears (history reload).
3. Switch back to thread A — history rehydration shows the in-progress
   turn with no response yet, so 0 assistant bubbles in DOM.
4. Stream chunks continue to fire for A — appendToLastAssistant creates
   a new bubble and accumulates the response into it.
5. response event fires — flushes any remaining buffer, removes the
   data-streaming flag (good) — then addMessage('assistant', content)
   creates a SECOND identical bubble.

Result: locator(".message.assistant").filter(has_text="4") matches
two elements, Playwright strict mode rejects the wait_for, the test
fails. Outside the test, two identical bubbles render to the user.

Fix: only call addMessage in the response handler when there was no
in-flight streaming bubble. If one existed, the streamed content is
already correct (chunks accumulate `data.content` verbatim) and the
data-streaming flag has just been cleared. Non-streaming responses
(no chunks fired) still take the addMessage branch.

Regression coverage: tests/e2e/scenarios/test_message_persistence.py::
test_switching_back_preserves_in_progress_turn already reproduces this
exact scenario and was failing in the merge queue. With this fix it
passes; skip-regression-check used because the existing E2E test is
the regression test, and the gateway doesn't have a JS unit test
harness for SSE handler state.

* fix(gateway): dedupe history-rendered SSE responses

* test(e2e): set mock LLM API key in standalone fixtures

* fix(e2e): make v2 approval tests deterministic

* test(e2e): stabilize duplicate skill install assertion

* test(e2e): assert duplicate install stays ungated

* fix(skills): skip approval for disk-installed duplicates

* fix(v2): honor no-op skill installs without approval

* test(e2e): wait for pending send marker to clear

---------

Co-authored-by: Firat Sertgoz <f@nuff.tech>
pull Bot pushed a commit to Stars1233/ironclaw that referenced this pull request May 7, 2026
…arai#3157)

* fix(engine): inline gate await for Tier 0 + Tier 1 Approval gates

CodeAct scripts that hit a tool requiring approval surfaced as
`RuntimeError: execution paused by gate 'approval'` inside the script
instead of pausing for the user. The async tool-resolve path
converted `EngineError::GatePaused` into a Python exception; the
sync preflight path returned `need_approval` to the orchestrator,
which on resume re-ran the LLM step and re-executed any
non-idempotent earlier tool calls in the same script.

Replace both with a host-supplied `GateController` that pauses the
live execution in place. The Monty VM (Tier 1) and the Tier 0 batch
loop both stay alive across the user's approval; on `Approved` the
gated action re-executes (lease re-consumed, auto-approve installed
before delivery so subsequent gates short-circuit); on `Denied`
the script raises a typed `RuntimeError("user denied tool 'X': ...")`
that the script can catch.

Auth and External resume kinds keep the legacy thread re-entry path
- their resolution installs new state (credentials, callback
payloads) that only takes effect on the next run-through.

Boot sweep invalidates `Approval`-kind `PendingGate` rows from a
prior process so a stranded gate after restart fails fast instead
of taking the legacy path and re-executing earlier mutations.

Tests: 3 new regression tests in scripting.rs (approve / deny /
no-controller fallback) and 4 in gate_controller.rs covering the
resolution registry's one-shot, dropped-receiver, and unknown-
request semantics. Pre-existing failures `call_id_preserved_when_no_lease`
and `stop_thread_works` reproduce on staging without these changes.

Design: docs/plans/2026-05-01-codeact-inline-gate-await.md.

* test(engine): live regression for inline gate await with CodeAct

Two integration tests in engine_v2_gate_integration.rs that exercise
the inline gate-await flow end-to-end through `ThreadManager` →
`ExecutionLoop` → orchestrator → CodeAct → `EffectExecutor` →
`GateController`:

1. `codeact_inline_gate_await_resumes_user_reproducer` reproduces the
   exact reported bug shape: a CodeAct script issuing
   `await github_tool(action="search_issues_pull_requests", ...)` for
   "what are p1 bugs in nearai/ironclaw filed in last 7 days". The
   github_tool returns `EngineError::GatePaused` mid-execution; the
   test's `OneShotApprovingGateController` marks the effects mock
   approved and returns `Approved`; the engine retries inline; the
   tool succeeds; the script's `FINAL("Found 0 P1 bugs ...")` reaches
   the user. Asserts the controller saw exactly one pause request,
   github_tool was called twice, and the response is the script's
   FINAL — not the pre-fix `RuntimeError: execution paused by gate`.

2. `codeact_inline_gate_await_denial_does_not_retry` covers the deny
   path: controller returns `Denied { reason: "not now" }`. Asserts
   github_tool was called exactly once (no retry on denial), the
   typed `user denied tool 'github_tool': not now` message appears in
   the failure events, and the pre-fix `execution paused by gate`
   string does NOT appear anywhere.

Also restored a fmt-only line shape in `gate_controller.rs` from
`cargo fmt`.

* test(engine): use realistic mock issues in inline gate-await fixture

The pre-fix fixture returned `{"items": []}` which made the script's
FINAL emit "Found 0 P1 bugs in nearai/ironclaw" — misleading, since
the repo actually has open P1 issues (e.g. nearai#2818, nearai#2997). Update the
mock to return two such items and tighten the assertion to verify
the exact count flowed from tool result through CodeAct to FINAL().

* refactor(engine): require gate_controller, bound retry, drop V1 fallback

Removes the `Option<Arc<dyn GateController>>` foot-gun: the field's
`None` arm in the executors silently re-emitted the original
`"execution paused by gate 'approval'"` RuntimeError, which is exactly
the bug this PR exists to fix. With the field required and a named
`CancellingGateController` as the explicit drop-in for non-pausing
paths (post-resolution replay, mission protected writes, tests),
forgetting to wire a controller is a compile error.

Other follow-ups in the same change to keep them on one commit:

- `MAX_INLINE_GATE_RETRIES = 3` shared by `scripting::drive_inline_gate`
  (Tier 1 async output) and `structured::execute_with_inline_gate_retry`
  (Tier 0 mid-execution). A misbehaving tool that keeps gating after
  each approval surfaces a clean error instead of pinning a CPU.
- `denial_reason_for_resolution` helper centralizes the
  `GateResolution -> reason` mapping so denial messages can't drift
  between Tier 0 and Tier 1.
- `invalidate_stranded_approval_gates_evicts_only_approval_kind` unit
  test covers the boot sweep with a mixed Approval/Auth/External
  population.
- Existing `codeact_gate_without_controller_falls_back_to_runtime_error`
  test rewritten as `codeact_default_controller_cancels_approval_gates`
  to assert the inverted invariant: the legacy bug message must NEVER
  appear, even with the inert default controller.
- Design doc updated to reflect as-shipped shape (required field, the
  bounded-retry constant, denial helper, `max_duration` stays at 30 s).

* fix(engine): bound inline pause, propagate one-shot approval, race fix

Addresses review on PR nearai#3157 (serrrfirat + Copilot + gemini-code-assist).

Four blocking correctness fixes:

1. **Bounded BridgeGateController::pause.** The await on the resolution
   oneshot now races against `pending.expires_at`. Without this, a user
   ignoring the prompt past expiry would strand the engine: the DB row
   expires, the oneshot stays open, the VM keeps running. On expiry we
   discard the pending row, drop the registry entry, and return
   `Cancelled` so the VM unwinds cleanly.

2. **One-shot approval threaded through retry.** Add
   `call_approval_granted: bool` on `ThreadExecutionContext` (default
   false). Inline retry paths (`drive_inline_gate`,
   `execute_with_inline_gate_retry`,
   `execute_single_action_with_inline_retry`, scripting sync preflight)
   set it to true on the retry call so `EffectBridgeAdapter::execute_action`
   forwards it as `approval_already_granted=true` to the host's tool
   approval check. Mirrors the legacy `execute_resolved_pending_action`
   contract; without this, tools with `ApprovalRequirement::Always`
   gated again on every retry until the bound tripped, and `always=false`
   AskEachTime gates re-prompted immediately after approval.

3. **Per-execution context registration race fixed.**
   `set_execution_context` was called AFTER `handle_user_message().await`
   returned the thread_id — but the engine task is already running,
   so a fast tool gate could reach `pause()` before the entry existed
   and get `Cancelled`. Now the bridge calls
   `set_pre_execution_context` (per-user) BEFORE
   `handle_user_message`, then promotes to (user, thread)-keyed once
   thread_id is known. `pause()` falls back to the per-user entry on
   miss.

4. **Tier 0 parallel-batch coverage.** New
   `execute_single_action_with_inline_retry` wraps `execute_single_action`
   with the same bounded retry shape used by Tier 1. Both the
   single-runnable and multi-runnable branches of
   `handle_execute_actions_parallel` go through it, so simultaneous
   gates in a parallel batch no longer fall through to the legacy
   re-entry path.

Smaller fixes:

- PROJECTION lint annotations on the two `broadcast_for_user` sites
  added by this PR (gate_controller emit_gate_prompt; resolve_gate
  inline-await fast-path resolution event).

Test coverage:

- `GatingThenOkEffects` now records `context.call_approval_granted`
  per call; `codeact_gate_inline_await_approved_delivers_result`
  asserts the retry observes `true`. Locks in the one-shot approval
  propagation contract.

* test(engine): update gate integration tests for inline-await semantics

CI was failing on 5 tests in engine_v2_gate_integration.rs that asserted
the legacy `ThreadOutcome::GatePaused` unwind for `Approval` gates. With
inline-await + the required `CancellingGateController` (default), Approval
gates resolve inline and the thread completes (or fails) instead of
pausing — these tests were exercising the pre-PR flow that this PR
replaces.

- `gate_paused_transitions_thread_to_waiting` → renamed to
  `approval_gate_resolves_inline_via_controller`. Wires
  `AutoApprovingGateController`, asserts thread completes after inline
  approval, ApprovalRequested + ActionExecuted both recorded.
- `gate_paused_thread_resumes_to_completion` → renamed to
  `approval_denied_inline_completes_thread_with_failed_action`.
  Asserts the default `CancellingGateController` cancels the gate and
  the thread completes with a failed action (no stranded pending gate).
- `approval_chains_directly_into_auth_for_install_flow` rewritten:
  Approval handled inline by controller, Auth gate (still legacy path)
  bubbles up as ThreadOutcome::GatePaused, legacy auth-resume drives
  completion.
- `approval_resolution_executes_pending_call_directly` and
  `gate_resume_with_execution_obligation` reduced to no-op stubs with
  inline rationale documenting where the post-PR equivalent coverage
  lives. Removing entirely would erase the breadcrumb in git log.

Also threads the `ApprovalRequested` event through
`execute_single_action_with_inline_retry`: the wrapper now returns
`Vec<EventKind>` per call so the caller can append both the approval
prompt and the post-retry outcome to the thread event log. Without
this, observers saw only the final ActionExecuted/Failed event with
no record that a gate had fired.

Drops a few dead test helpers (`ApprovalTool`, `make_caps_with_approval_tool`,
unused imports) that the deleted assertions no longer reference.

Quality gate: fmt clean, clippy zero warnings, 513 engine lib + 443
bridge + 29 gate integration tests pass (2 pre-existing engine-lib
staging failures unrelated to this PR).

* test(engine): update skill_codeact integration tests for inline-await

Two more tests in engine_v2_skill_codeact.rs were asserting the
legacy `ThreadOutcome::GatePaused` → `resume_thread` flow for
`Approval` gates. With PR nearai#3157 the engine catches Approval gates
inline and the thread runs to completion in a single `join_thread`.

- `skill_prompt_context_survives_pause_and_resume`: wires
  `AutoApprovingHttpController`, asserts thread completes after inline
  approval, drops the now-impossible `resume_thread` step.
- `skill_prompt_context_survives_compaction_and_resume`: same; also
  drops the assertion on the persisted transcript at the *pause point*
  (no longer externally observable). The load-bearing post-compaction
  LLM-call assertions remain — they're the actual contract this test
  was protecting.

Adds an `AutoApprovingHttpController` test helper local to this file
that approves gates by marking the underlying `PausingHttpMockEffects`
approved before returning `Approved`.

* fix(engine): bridge cleanup on dropped sender + ActionFailed on preflight denial

Addresses Copilot review on PR nearai#3157.

- `BridgeGateController::pause`: when the resolution oneshot's sender
  is dropped (process shutdown / registry cleared), discard the
  `PendingGate` row before returning `Cancelled`. Without this, a
  stranded prompt remained visible in the UI and `pending_gates.insert`
  rejected duplicates for the same `(user, thread)` so a follow-up
  gate could not register. Same cleanup as the expiry branch already
  performed.

- `scripting.rs` sync-preflight denial path: emit `EventKind::ActionFailed`
  before resuming Monty with `RuntimeError`, so the thread event log
  is consistent with the other denial paths (`drive_inline_gate`,
  `structured.rs`). Auditing why a tool didn't run is now possible
  from the events alone.

- Delete the two empty `#[tokio::test]` stubs left in
  `engine_v2_gate_integration.rs`
  (`approval_resolution_executes_pending_call_directly_via_resolved_pending_action`,
  `gate_resume_with_execution_obligation`). The post-PR equivalent
  coverage is in `codeact_inline_gate_await_*` (this file) and the
  scripting unit tests; the rationale is preserved in `git log`
  (commit 87fe4fc) without an empty test slot misleading coverage
  signals.

* fix(engine): finish inline-await migration; remove legacy gate-paused for Tier 0 policy gates

Address PR nearai#3157 review:

- router.rs: clear pre_execution slot on handle_user_message error so a
  failed engine spawn doesn't leak a stale (user, conversation) entry
  that would mis-route the next gate prompt.

- router.rs:await_thread_outcome: on the 5-min request deadline, return
  BridgeOutcome::Pending instead of join_thread() — joining would block
  for up to the gate's 30-min expires_at when the parked task is in
  pause(). The PendingGate row stays live for resolution.

- gate_controller.rs: re-key pre_execution by (user_id, conversation_id)
  instead of user_id alone. Two concurrent conversations / browser
  tabs for the same user no longer clobber each other's slot. Plumbs
  conversation_id through ThreadExecutionContext and GatePauseRequest
  so pause() can match a gate to its originating conversation.

- gate_controller.rs: serialize concurrent inline gates per
  (user, thread) via a per-key tokio Mutex held across the
  PendingGateStore::insert + select-await window. A parallel batch
  where two tools both gate now queues the second behind the first
  rather than silently surfacing it as Cancelled on (user, thread)
  uniqueness collision.

- orchestrator.rs (Tier 1 / CodeAct): remove the legacy gate_paused JSON
  sentinel + thread re-entry path for Approval gates. Both
  __execute_action__ and __execute_actions_parallel__ now pause inline
  on PolicyDecision::RequireApproval (mirroring structured.rs) and
  route tool-raised gates through the existing
  execute_single_action_with_inline_retry wrapper. Authentication and
  External resume kinds keep the legacy re-entry path because their
  resolution installs new state that only takes effect on the next
  thread run-through.

Test fixtures across bridge/effect_adapter, action_projector, and the
gate/sandbox integration suites get the new conversation_id field
defaulted to None.

Refs: comments 3173757791, 3176248954, 3176248976, 3176249001, 3176249032

* fix(engine): defer post-Pending context cleanup; route parallel JoinSet through inline-retry

Address PR nearai#3157 review on commit d211bfc:

- router.rs: when await_thread_outcome returns BridgeOutcome::Pending and
  the engine task is still running (typically parked in
  BridgeGateController::pause), defer clear_execution_context to a
  spawned watcher task that polls is_running until the thread completes
  and then clears the (user, thread) context + gate-locks entry. Without
  this, the unconditional clear after a 5-min request timeout stranded
  the parked thread: the eventual gate resolution would call pause()
  for any subsequent gate with no registered context and surface as
  silent Cancelled. Watcher caps at 60 minutes (well past the 30-min
  PendingGate expiry) as a defensive safety bound.

- structured.rs: route the multi-runnable JoinSet branch in
  execute_action_calls through execute_with_inline_gate_retry, matching
  the single-runnable fast path. Without this, an Approval gate raised
  mid-execution in a parallel batch with >1 runnable tool call
  bubbled out as Err(GatePaused) and went through the legacy
  gate-paused / re-entry path, re-introducing the double-execution bug
  for already-completed sibling calls in the same batch. The signature
  of execute_action_calls switches leases from &LeaseManager to
  &Arc<LeaseManager> so the JoinSet tasks can clone an owned handle;
  all current callers are tests already constructing
  Arc::new(LeaseManager::new()), so this is a no-op call-site change.

Refs: comments 3176333657, 3176333685

* test(engine): fix call_id_preserved_when_no_lease MockEffects inventory

PR nearai#2868 added a callable-inventory check ahead of the lease lookup in
execute_action_calls. The test was constructing MockEffects with an
empty action list, so preflight short-circuited on "action is not
callable in this execution context" before reaching the lease branch
the test was actually trying to exercise. The assertion
error.contains(\"no lease\") then failed against the not-callable
message.

Populate the mock with test_action(\"web_search\") so the inventory
gate passes and the call reaches the lease check, matching the test's
documented intent.

* fix(engine): use gate-provided params in inline-await pause; drop Thread clone in parallel branches

- orchestrator.rs::execute_single_action_with_inline_retry now sources
  GatePauseRequest.parameters from the gate-paused payload (which
  reflects safety-layer transformations/redactions) rather than the
  original caller params, matching structured.rs::execute_with_inline_gate_retry.
- Both inline-retry helpers now take ThreadId + user_id instead of
  &Thread, so the parallel JoinSet branches no longer clone the full
  Thread (with message/event transcripts) per spawned task. The
  per-task ThreadExecutionContext clone is what the helpers actually
  need; in orchestrator.rs we build a base ctx once and override
  current_call_id per task instead of re-running thread_execution_context.

Addresses Copilot PR nearai#3157 review comments 3176594866, 3176594886, 3176594898.

* fix(engine): address serrrfirat review — audit event + stop-during-wait

Three latest serrrfirat review threads on PR nearai#3157:

* Medium: structured inline approval drops ApprovalRequested audit
  event (REAL — fixed). `execute_with_inline_gate_retry` previously
  swallowed `Err(GatePaused)` and returned only the post-retry
  outcome, so `classify_exec_result` never saw the gate and the
  `ApprovalRequested` event was lost. The orchestrator (Tier 1) path
  emits both events. Mirror that shape here:

  - `execute_with_inline_gate_retry` now returns
    `(Result<ActionResult, EngineError>, Vec<EventKind>)`. The Vec
    carries one `ApprovalRequested` per retry iteration that gated.
  - Slot type widens to `(ActionResult, EventKind, Vec<EventKind>)`
    so the merge phase flushes pre-terminal events before the
    terminal event in original-call order.
  - Two regression tests pin the contract (approved → Executed,
    denied → Failed); both assert ApprovalRequested precedes the
    terminal event.

* High: stop/cancel does not unblock thread parked in inline gate
  await (REAL — fixed). `BridgeGateController::pause()` selected
  only on the resolution oneshot + 30-min expiry; `stop_thread()`
  sent `ThreadSignal::Stop` but the parked engine task wasn't
  polling the signal channel.

  - Adds `GateController::cancel_thread(thread_id)` to the
    engine-side trait (default no-op).
  - `ThreadManager::stop_thread()` calls `cancel_thread()` BEFORE
    sending `ThreadSignal::Stop` so any parked `pause()` future
    wakes promptly with `GateResolution::Cancelled`.
  - `BridgeGateController` tracks in-flight pauses in
    `active_pauses: HashMap<ThreadId, HashSet<Uuid>>` and walks
    that set on cancel, delivering `Cancelled` via the existing
    `GateResolutions::try_deliver` channel and discarding pending
    rows via `PendingGateStore::discard_for_thread`.
  - `pause()` always untrack on exit (idempotent — `cancel_thread`
    may have already removed the entry).
  - Two new bridge-tier unit tests: parked pause wakes within 2s
    of cancel_thread; cancel_thread on a thread with no parked
    pause is a no-op.

* Medium: live inline gate waits are uncapped per user/process
  (PARTIAL — push back, follow-up). The reviewer's own comment
  notes this is a design-doc follow-up, not a blocker. The current
  bound is implicit: one pending gate per (user, thread) × the
  existing thread-creation budget × the 30-min expiry, which is
  enough to ship the inline-await substrate. Adding a per-user
  semaphore correctly requires designing the cap UX, fairness, and
  rejection error shape — out of scope for this PR. Add a TODO
  inside `pause()` pointing at the design doc and the
  follow-up issue, with the reasoning written down so the next
  contributor doesn't have to re-derive it.

Pre-existing on this branch (NOT introduced by these changes):
`runtime::manager::tests::stop_thread_works` flakes on the branch;
the original PR commit message acknowledges "stop_thread_works
reproduce on staging without these changes." Confirmed by stashing
this commit and running the test: still fails. Out of scope here.

Test totals after this commit:
  - executor::structured: 19 passing (incl. 2 new)
  - bridge::gate_controller: 6 passing (incl. 2 new)
  - cargo fmt + clippy --lib clean

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(engine): typed DenialOutcome + plug exhaustion-path lease leak

Three review-driven fixes to the inline gate-await path.

1. CancellingGateController and bridge expiry/shutdown surfaced as
   `RuntimeError("user denied tool 'X': cancelled")` — misleading
   because the user never saw a prompt. Replace `Option<String>`
   helper with a typed `DenialOutcome { DeniedByUser, Unavailable }`
   so cancelled/expired/no-handler gates render as "approval for
   tool 'X' unavailable: …" while real user denials keep the
   "user denied" framing. Updated all six call sites in
   scripting.rs / structured.rs / orchestrator.rs through the
   typed surface so wording can't drift.

2. Both `execute_with_inline_gate_retry` and the orchestrator's
   `execute_single_action_with_inline_retry` consumed a fresh
   lease use on the final approved iteration, then exited the
   loop without ever calling execute_action. The exhaustion-path
   error didn't trip the caller's refund check, so a misbehaving
   tool that gates after every approval would slowly drain
   `max_uses`. Refund the unused lease before returning.

3. Drop the dead `parameters: serde_json::Value` field on
   `PendingFuture::Tool` and the matching `_parameters` arg on
   `resolve_tool_future`. The gate's own parameter snapshot is the
   source of truth on retry; threading the original through made
   the signature read like there was a second source.

Quality gate: cargo fmt, cargo clippy --all --benches --tests
--examples --all-features (clean), cargo test -p ironclaw_engine
--lib (520 pass), cargo test --test engine_v2_gate_integration
(27 pass), cargo test --lib bridge:: (456 pass).

---------

Co-authored-by: Nikolay Pismenkov <nickpismenkov@gmail.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This was referenced May 7, 2026
pull Bot pushed a commit to soitun/ironclaw that referenced this pull request May 13, 2026
…arai#3589)

Both xfails in tests/e2e/scenarios/test_v2_auth_oauth_matrix.py are
stale and pass against current code:

- test_wasm_tool_first_chat_auth_attempt_emits_auth_url
  Marked xfail in nearai#3235 because the engine-v2 callable-only contract
  (nearai#2868) stopped emitting an auth gate on direct LLM-driven tool
  calls. PR nearai#3157 (auth-preflight + inline-await) restored the
  behavior the test asserts: when the LLM emits a direct call to a
  not-yet-authed extension, the bridge raises an Authentication gate
  with auth_url populated (src/bridge/effect_adapter.rs:1356-1392).
  Marker removed; test passes.

- test_settings_first_custom_mcp_auth_then_chat_runs
  The xfail reason claimed post-auth tool-output propagation was
  broken. Real cause: engine-v2 gates the first MCP tool call on
  `approval` and the browser fixture has no auto-approve UI, so the
  chat sat in pending_gate forever. Same shape as the bugs fixed in
  nearai#3235 for test_wasm_tool_oauth_refresh_on_demand and
  test_mcp_same_server_multi_user_via_browser. Inserted
  _wait_for_tool_call between _send_chat and _wait_for_response_contains
  to drive approval through the API; test passes.

Verified locally: both tests pass back-to-back in 27s on a fresh
auth_matrix_server.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
ilblackdragon added a commit that referenced this pull request May 13, 2026
… + auto-approve footgun (#3533) (#3559)

* fix(extensions): restore chat-driven tool_install + fix double-invoke + auto-approve footgun (#3533)

"Connect my telegram" was giving the user two options and not actually
installing anything because three layered issues had accumulated since
engine v2:

1. **`tool_install` was hidden from the agent** (#2868). The unified
   `tool_activate` it was meant to be subsumed by was later removed in
   #3166, but the hidden-from-callable-surface gate stayed. Restored
   by dropping `hidden_from_model_callable_surface` from
   `bridge::action_projector`. User consent is mediated by the tool's
   own `ApprovalRequirement::UnlessAutoApproved` and the seeded
   `AskEachTime` permission.

2. **Two competing Telegram registry entries** (`telegram` channel and
   `telegram_mtproto` tool) both surfaced in the agent prompt's
   `Activatable Integrations` section. The LLM correctly enumerated
   them as "Option 1" and "Option 2" instead of installing the
   canonical bot channel. Added a `hidden: bool` field to
   `ExtensionManifest` / `RegistryEntry`, set `telegram_mtproto` to
   `hidden: true`, and filter hidden entries out of the
   "available-but-not-installed" appendix in `ExtensionManager::list`.
   Hidden entries remain installable by explicit name.

3. **Updated the agent prompt** so `Activatable Integrations` instructs
   the model to call `tool_install(name="<name>")` directly rather than
   describing manual UI steps.

Fixes the double-`tool_install` invocation that surfaced once the agent
could install from chat:

- **`InlineGate` discarded cached output.** The bridge raised an
  Authentication gate after `tool_install` succeeded, and the
  inline-await retry re-executed the action (re-downloading the WASM
  bundle) instead of returning the already-computed output. Added
  `resume_output: Option<serde_json::Value>` to `InlineGate`; on
  approval, return the cached output if present. Mirror fix in the
  orchestrator's `execute_single_action_with_inline_retry` (reading
  `result_json["resume_output"]`) and the structured-batch retry path.
- **`effect_adapter::auth_gate_from_extension_result`** now passes
  `Some(output_value.clone())` as the gate's `resume_output` so the
  retry has cached state to short-circuit on.
- **OAuth callback double-fired.** `oauth_callback_handler` now skips
  the `ExternalCallback` re-entry when the inline-await path already
  woke a parked waiter — eliminates the "thread already running" race.
- **`resolve_inline_gates_for_credential`** now also discards matching
  Authentication rows from `pending_gates` so the row doesn't linger
  in `HistoryResponse.pending_gate` after inline resolution.

Fixes the auto-approve footgun:

- **`ToolPermissionSnapshot::resolve_permission`** now collapses DB
  values that match the seeded default to `explicit = None`. Before
  this, the boot-time `seed_tool_permissions` write of `tool_install ->
  AskEachTime` was indistinguishable from a user-explicit override,
  causing `effect_adapter::enforce_tool_permission`'s `is_explicit_ask`
  check to refuse `AGENT_AUTO_APPROVE_TOOLS=true`. Real overrides
  (`AlwaysAllow`, `Disabled`) still surface as `Some(...)`.

Tests
- Unit: 4980/4980 pass (host) + 525/525 pass (engine).
- Unit: new tests in `bridge::tool_permissions::tests` lock seeded-vs-
  explicit collapse; new test in `bridge::action_projector::tests`
  asserts `tool_install` is callable; new manifest hidden-flag tests
  in `registry::manifest::tests` and `extensions::manager::tests`.
- E2E: removed `@pytest.mark.xfail` on
  `test_chat_first_gmail_installs_prompts_and_retries` (now passes
  end-to-end via the chat-driven install path). Added
  `test_chat_install_approval_then_auth_card` driving the
  explicit-approval variant with a single Approve click (no Always
  workaround needed) — wired into the `auth-full` canary lane.
- Mock LLM: extended the gmail-install-then-retry pattern to recognize
  both the legacy "Extension not installed:" and the post-#3533 "is
  not callable in this execution context" error strings, and to retry
  `gmail(action="list_messages")` after a successful `tool_install`.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(permissions): address #3559 review (permission bypass, lease accounting, hidden search filter)

Five fixes from the #3559 review (4× Copilot doc nits + 3× serrrfirat
security/correctness findings):

1. **Permission bypass (High).** Pre-#3559's `resolve_permission`
   collapsed any DB row whose value matched the seeded default to
   `explicit = None`, so a user who deliberately set `tool_install =
   AskEachTime` had their explicit choice silently dropped and
   `AGENT_AUTO_APPROVE_TOOLS=true` bypassed the gate. Provenance is now
   handled at write time: `seed_tool_permissions` is gone and a
   one-shot, sentinel-gated migration (`cleanup_ghost_seeded_tool_permissions`)
   deletes existing ghost-seeded rows at startup. With no ghost rows,
   the resolver treats every DB row as user-explicit and honors it.

2. **Lease/event accounting on `resume_output` replay (Medium).**
   Inline-gate handlers in `structured.rs`, `scripting.rs`
   (`resolve_tool_future` + `drive_inline_gate` retry loop), and
   `orchestrator.rs` refunded the lease use the action just consumed,
   then returned the cached `resume_output` on approval without
   re-consuming — netting successful side-effecting actions to zero
   lease uses. Skip the refund when the gate carries cached output.

3. **Hidden registry filter on `tool_search` (Medium).**
   `RegistryCatalog::search` did not filter `hidden: true` entries,
   so `telegram_mtproto` could resurface through the search path and
   reintroduce the "two Telegram options" outcome that #3533 fixes
   for the default-list path. Added the filter and a regression test.

4-7. Copilot doc nits: outdated `_set_tool_permission` docstring;
   misleading "bridge-side auto-install implemented" comment in
   `mock_llm.py`; `tool_install` described as "non-agent surface" in
   `src/bridge/CLAUDE.md` while a paragraph below says the model
   calls it directly; dangling `issue #3533 / PR —` placeholders
   in both CLAUDE.md docs.

Regression tests:
- `bridge::tool_permissions::user_explicit_value_matching_seeded_default_stays_explicit` —
  the original Copilot/serrrfirat bug case.
- `app::cleanup_ghost_seeded_tool_permissions_removes_seed_matching_rows` —
  idempotent migration + sentinel.
- `extensions::registry::test_search_skips_hidden_entries` — hidden
  entries excluded from search but still installable by exact name.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(#3559): caller-level regression coverage for review findings 1 & 2

Two follow-up regression tests for the #3559 security review, plus a
real bug surfaced by the first one.

1. `executor::structured::resume_output_replay_consumes_exactly_one_lease_use`
   exercises the post-execution Authentication gate inline-retry path
   with `max_uses=1` and asserts:
   - Cached output is returned as a successful `ActionResult`.
   - Exactly one `ActionExecuted` event is emitted.
   - The lease budget is exhausted after one execution (refund-skip
     keeps the consumption from being undone).

   Writing this test surfaced a real bug: the structured cached-output
   branch pushed `ActionExecuted` into `emitted_events`, and the
   caller's `classify_exec_result` emitted ANOTHER terminal
   `ActionExecuted` for the same Ok result — double-emit for one
   action. Tier 1 (`scripting::drive_inline_gate`) and Tier 1 alt
   (`orchestrator::execute_action_with_inline_gate`) emit themselves
   because their callers don't run an Ok-branch classifier; structured
   was the outlier. Dropped the redundant push; the classifier emits
   the single canonical event.

2. `bridge::effect_adapter::explicit_ask_each_time_for_seeded_default_tool_still_gates`
   drives `execute_action` end-to-end (the side-effecting caller) with
   a tool whose `name()` matches a seeded-`AskEachTime` baseline
   (`tool_install`) and an explicit `AskEachTime` user override. The
   resolver collapse-to-implicit bug would have shown up here — not
   just in the helper-level test that already exists in
   `bridge::tool_permissions::tests`. Per `.claude/rules/testing.md`
   "Test Through the Caller, Not Just the Helper".

   Added `SeededAskEachTimeTestTool` as a `tool_install`-named test
   fixture with `requires_approval: UnlessAutoApproved` to mirror the
   real tool's contract.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
serrrfirat added a commit that referenced this pull request May 26, 2026
* fix(e2e): restore auth and approval coverage (#3430)

* test(e2e): avoid REPL auth retry race (#3437)

* feat: add pairing_approve tool for Slack binding via chat (#3396)

* feat: add pairing_approve tool for Slack binding via chat

Users can now paste their Slack pairing code in the IronClaw chat and
the LLM will call pairing_approve to bind their accounts. No need to
use the API directly.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: require approval before pairing + fix formatting

Address review comment: pairing_approve now requires UnlessAutoApproved
approval before executing, preventing accidental account binding.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* test: add regression test for pairing_approve tool

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: address review — Always approval, lock channel, add to protected list

1. Changed ApprovalRequirement to Always (not bypassable by auto-approve)
2. Locked channel to slack-relay constant (removed generic channel param)
3. Added pairing_approve to PROTECTED_TOOL_NAMES
4. Added tests: always-approval, protected-name, channel constant

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(web): isolate cross-tenant SSE/WS status events and thread access (#3390)

* fix(web): isolate cross-tenant SSE/WS status events and thread access

Plug a multi-tenant leak where unscoped `sse.broadcast(...)` calls
from `GatewayChannel::send_status`, sandbox `JobEvent` dispatch, the
WASM/Slack OAuth completion handlers, and any producer that lost
`metadata.user_id` along the way fan out to every connected
subscriber — exposing another tenant's tool calls, tool output,
onboarding state, and job lifecycle to anyone with an open SSE/WS.

Changes
- Extract `dispatch_status_event(sse, multi_tenant_mode, user_id, ev)`
  from `Channel::send_status`. In multi-tenant mode an unscoped event
  is dropped (with a WARN naming the producer to fix); single-tenant
  keeps the global broadcast since there is one subscriber population.
- `IncomingMessage::new` now defaults `metadata` to `{"user_id": ...}`,
  and `with_metadata` preserves the key so downstream `send_status`
  consumers always have an owner to scope by.
- WASM/Slack OAuth completion broadcasts route through
  `broadcast_for_user(&owner_id, ...)`. Sandbox `JobEvent` dispatch
  in `main.rs` respects `multi_tenant_mode` for the empty-`user_id`
  fallback.
- New pre-commit check #10 (`MULTITENANT`) flags unscoped
  `sse.broadcast(...)` lines without a `// multi-tenant-safe: <reason>`
  marker or a transport-only exemption. Marker regex accepts the marker
  anywhere in a `//` comment so compound annotations on a single line
  work.

Tests
- `src/channels/web/tests/status_event_isolation.rs` — 5 unit tests
  covering both modes and the per-variant drop invariant.
- `src/channels/web/platform/sse.rs` — 2 quadrant tests for the
  `subscribe_raw` filter (scoped/unscoped × matching/mismatched).
- `tests/thread_isolation_integration.rs` — 9 HTTP-level checks that
  Bob cannot reach Alice's chat history (paginated and not), threads
  list, engine v2 detail/steps/events, or Responses GET, plus an
  unauthenticated-rejection guard.
- 8 new self-test cases for the `MULTITENANT` script check, including
  the compound projection-exempt + multi-tenant-safe annotation case.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(web): pin cross-tenant boundaries on jobs, files, routines

Audit of the protected route surface for the same bug shape #3390 fixed
(handler that takes a user-controlled id and reads without an ownership
predicate) found that the implementations were correct but four
boundaries had no integration test. Lock them in before they regress.

- Sandbox job persisted-events history (`/api/jobs/{id}/events`):
  Bob → 404 on Alice's job; Alice → 200 on her own.
- Sandbox job workspace listing (`/api/jobs/{id}/files/list`):
  Bob → 404 on Alice's job.
- Sandbox job file read (`/api/jobs/{id}/files/read`): Bob → 404 on
  Alice's job; Alice → 200 on her own; Alice → 403/404 on
  `?path=../outside.txt` (path-traversal pin against the
  `canonicalize() + starts_with(base_canonical)` guard).
- Routine run history (`/api/routines/{id}/runs`): Bob → 404; Alice
  → 200 with at least one seeded run.

The OAuth-state and NEAR-nonce stores were also flagged in the audit
but neither is a real cross-tenant bug: both are pre-auth, single-use,
and the token IS the secret. Documenting here so a future audit
doesn't re-flag them.

Project-static (`/projects/{id}/...`) is left for a follow-up — it
relies on `ironclaw_base_dir()` which is a process-wide `LazyLock`,
making per-test override fragile in the integration runner.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(web): address PR #3390 review — forge-resistant metadata, OAuth toast routing

Addresses six review comments from gemini-code-assist, copilot, and
serrrfirat on PR #3390. False-positives and the perf nit on
`with_metadata` are explained in the reply thread, not changed in code.

- (HIGH, serrrfirat) `IncomingMessage::with_metadata` now ALWAYS sets
  `metadata.user_id` from `self.user_id`, dropping any caller-supplied
  value. A WASM channel emitting `{"user_id":"victim"}` via
  `apply_emitted_metadata` can no longer reroute downstream
  `ToolStarted` / `ToolResult` SSE events into another tenant's
  stream. New unit tests pin the forgery-resistance invariant.
- (MEDIUM, copilot + serrrfirat) Slack relay OAuth callback now
  broadcasts the completion toast to the resolved `oauth_user`
  (the IronClaw user who initiated the flow) rather than
  `state.owner_id`. In multi-tenant deployments those differ and the
  previous routing delivered the toast to the wrong browser tab.
  Extracted the lookup into `resolve_relay_oauth_user`; two unit
  tests cover the secret-present and secret-missing cases.
- (MEDIUM, gemini) `dispatch_status_event` treats empty-string
  `user_id` the same as `None` so producers that lost the field
  along the way fail-closed instead of falling through to a global
  broadcast in multi-tenant mode.
- (MEDIUM, gemini) `main.rs` sandbox JobEvent dispatch now reuses
  `dispatch_status_event` instead of duplicating the drop / WARN /
  broadcast policy. `dispatch_status_event` is bumped from
  `pub(crate)` to `pub` so the binary crate can call it.
- (LOW, copilot) Pre-commit `MULTITENANT` self-tests gain three
  cases (`state.sse.broadcast(`, `gw_state.sse.broadcast(`,
  annotated receiver-prefixed) to lock the existing boundary regex
  behaviour against future tightening.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(channels): preserve i64 metadata.user_id from Telegram in with_metadata

PR #3390's forge-resistance fix made `IncomingMessage::with_metadata`
*always* overwrite `metadata.user_id` with `self.user_id` as a String.
That broke the Telegram WASM channel: it persists Telegram's chat user
ID as `metadata.user_id: i64` and re-deserializes it into
`TelegramMessageMetadata { user_id: i64, ... }` in `on_respond` /
`on_status`. After the fix, `respond` blew up with
`invalid type: string "999", expected i64 at line 1 column 87`,
failing 3 Telegram integration tests in CI.

Narrow the carve-out: overwrite only when the existing `user_id` is a
String (or missing). Non-string values are channel-private and the
SSE routing layer reads via `as_str()` — non-strings already fail
closed in multi-tenant mode, so the forge threat (WASM emits
`{"user_id":"victim"}` as a string) is still mitigated, while
Telegram's i64 use case survives.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(web): redact WARN payload + harden dotdot traversal test (PR #3390)

Two follow-up fixes from the second pass of review on #3390:

1. `dispatch_status_event`'s WARN log used `?event`, which on
   `AppEvent::Response` / `Thinking` / `ToolResult` carries
   user-authored content into operator logs in a multi-tenant
   deployment. Replace with `event_kind = event.event_type()`
   (the wire-stable variant name) — enough to identify the
   misbehaving producer without leaking tenant data. Picked up via
   Copilot's review on `src/channels/web/mod.rs:666`.

2. `alice_job_file_read_rejects_dotdot_traversal` planted
   `outside.txt` under `outer.path()` (the `start_server_with_db`
   fixture's tempdir holding `test.db`) but probed
   `?path=../outside.txt` relative to `alice_proj` — a separate
   `tempfile::tempdir()` rooted at the OS temp directory. The two
   paths were unrelated, so the test could pass even if `..`
   traversal was permitted (probe just hit empty space). Build the
   directory tree by hand instead: `parent/alice_proj/` with the
   planted file at `parent/outside.txt`, so the probe deterministically
   resolves to the planted bytes. Add a body-content assertion that
   fails loudly if those bytes leak. Picked up via Copilot's review on
   `tests/cross_tenant_resource_isolation.rs:355`.

Plus a `cargo fmt` fix for `src/channels/channel.rs:1202` that was
breaking the Formatting CI check.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(web): address PR #3390 follow-ups — multi-tenant fallback WARN, exhaustive variant pin

- `resolve_relay_oauth_user`: take `multi_tenant_mode`; emit WARN when the
  `relay:{ext}:oauth_user` secret is missing in multi-tenant mode so the
  unrecoverable-initiator case surfaces in operator logs. Single-tenant
  fallback stays silent (owner == only user).
- `dispatch_status_event`: doc note clarifying the function is `pub` only
  for the sandbox JobEvent rx loop in `main.rs`; not part of a stable
  public API.
- `_compile_time_appevent_variant_check`: exhaustive-match helper paired
  with `unscoped_drop_holds_for_every_status_variant_in_multi_tenant`.
  Adding a new `AppEvent` variant now fails the test build, prompting an
  update to both the helper and the runtime leak-candidate list.
- `tests/thread_isolation_integration.rs`: honest scope note on the
  engine-v2 thread tests — they pin handler shape (404/empty for
  unknown id), not the cross-tenant ownership branch. Cross-tenant
  engine-v2 coverage requires an `ENGINE_STATE` test fixture; tracked
  as a follow-up in the comment block.

Tests: 8 unit (5 status_event_isolation + 3 resolve_relay_oauth_user)
and 9 thread_isolation_integration pass; clippy clean; pre-commit
safety scripts pass (regression suite 27 cases).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(web): unconditionally consume relay:{ext}:oauth_user secret in OAuth callback

The previous cleanup site lived inside the `if let Some(pairing_store)`
branch of the result block, which was unreachable on three failure
paths:

1. `pairing_store` is None (no identity pairing wired up)
2. The result block `?`-short-circuits before reaching the `if let`
   (e.g. `set_setting` fails, `activate_stored_relay` fails, an inner
   `relay_config()` / `list_connections()` errors)
3. The `if let` body itself errors before reaching the delete (e.g.
   `list_connections` returns no matching team)

Leaving the secret behind lets a subsequent OAuth callback for the
same extension read a stale initiating user and misroute the
completion toast — Copilot review on PR #3390 (comment id 3211833864).

Move the delete to right after `resolve_relay_oauth_user` returns,
where it always runs once the value has been captured, regardless of
downstream failure mode. Updated the inner comment to document that
the secret is already gone by the time the pairing branch reads
`oauth_user`.

Regression test: `test_relay_oauth_callback_consumes_oauth_user_secret_on_failure_path`
seeds the secret, fires the callback against a fixture with no real
relay backend (so the result block deterministically errors), and
asserts the secret is gone afterward. Pre-fix this would have left
the secret behind on the failure path.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(auth): tighten Telegram pairing UX and OAuth-failure recovery (#3317, #3319, #3320) (#3381)

* fix(auth): tighten Telegram pairing UX and OAuth-failure recovery (#3317, #3319, #3320)

Three Bug Bash P1 issues from the same user journey: setup → use → fail.
The unifying root cause was per-channel auth tested in isolation; cross-channel
flows (Telegram → Gmail OAuth → resume) had no coverage and three small leaks
combined into a stuck conversation.

#3317 — Telegram pairing reply now names every IronClaw surface explicitly
(web settings, agent chat, terminal). The agent submission parser learns
`approve <channel> <code>`, dispatched through a new bridge handler that
mirrors `POST /api/pairing/{channel}/approve`.

#3319 — OAuth callback failures now log a category + correlation ID so a
user-reported "I saw 400" maps to one log line. Adds the
`OauthCallbackFailure` enum and `oauth_failure_correlation_id` helper.

#3320 — Two cleanup gaps fixed: (a) `/clear` now drains
`pending_oauth_flows` for the user (otherwise stale flows linger 5min and
mask new auth attempts); (b) OAuth provider-error and exchange-failure paths
now auto-cancel the engine pending auth gate via `clear_engine_pending_auth`,
so the conversation isn't blocked waiting for a resume that will never arrive.

Tests: 5 new submission-parser tests, 2 new bridge-handler tests, and one
new OAuth callback test verifying the pending-flow drain on provider error.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(auth): cross-channel pairing claim coverage + canary lane (#3317)

Adds the structural coverage that was missing when #3317 shipped:

- E2E (`tests/e2e/scenarios/test_telegram_pairing_chat_claim.py`):
  three scenarios that drive the full Telegram pairing flow through
  the gateway. Asserts the bot reply names every IronClaw surface
  (web Settings, agent chat, terminal CLI), drives `approve telegram
  CODE` through `/api/chat/send` and verifies the paired user
  exchanges messages without re-prompting, and confirms invalid
  codes get a clear rejection instead of an LLM-improvised reply.

- Rust integration (`tests/telegram_pairing_chat_claim_integration.rs`):
  drives `Submission::PairingClaim` through a real `Agent` →
  `bridge::handle_pairing_claim` → `PairingStore::approve` chain
  using `TestRig` with engine v2 enabled. Covers the happy path
  (mints a code, claims it via chat, asserts `Pairing approved`)
  and the invalid-code rejection. The unit tests in `bridge/router`
  cover only the no-extension-manager and invalid-channel branches —
  this test exercises the wiring between submission parser, agent
  loop dispatch, and bridge handler that #3317 specifically broke.

- Canary (`scripts/live_canary/auth_registry.py`): adds the two
  user-visible scenarios to `AUTH_CHANNEL_TESTS` so the auth-channels
  lane (scheduled every 6h) catches the regression class in CI
  before any real user encounters it.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs: align pairing-claim and oauth-correlation comments with code (#3381)

Three Copilot review comments on PR #3381 flagged docstring/code drift in
already-merged PR #3317/#3319/#3320 changes. No behavior change — only
the doc strings move:

- `Submission::PairingClaim.code` and the inline `approve <channel> <code>`
  parser comment claimed the user's casing was preserved, but the parser
  builds the code from `lower` and the regression tests already lock in
  the lowercased shape (`code == "abc12345"`). Update both comments to
  describe the actual normalize-then-store contract.

- `oauth_failure_correlation_id` claimed the correlation appeared in the
  user-facing error subtitle, but the failure path renders
  `landing_html(label, false)` whose subtitle is fixed and never receives
  the correlation. Mark the helper as logs-only and note that plumbing
  the ID through the HTML is a follow-up.

[skip-regression-check] doc-only, behavior already covered by existing
pairing-claim parser tests in submission.rs.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore(telegram): bump channel registry to 0.2.11

The pairing-reply wording was updated in channels-src/telegram/, which
the version-check CI requires be matched by a registry version bump.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(telegram): fix import path in pairing chat claim e2e test

The scenarios/ folder is a Python package (has __init__.py), so a
flat `from test_telegram_e2e import …` fails with
ModuleNotFoundError during pytest collection. Switch to a relative
import that matches the package layout, and drop the unused
OWNER_USER_ID symbol while we're here.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(auth): address Copilot review on /clear OAuth drain + correlation doc

Two follow-ups on PR #3381's Copilot pass:

1. `Agent::process_clear` (engine v1 path) now drains in-flight OAuth
   flows for the clearing user, mirroring the engine-v2 cleanup added in
   `bridge::router::clear_engine_conversation`. Without this, `/clear`
   was a clean slate on v2 but v1 left ghost flows in
   `extension_manager.pending_oauth_flows()` until the 5-minute
   `OAUTH_FLOW_EXPIRY` ticked over — same regression class #3320 fixed
   on v2.

2. `oauth_failure_correlation_id`'s docstring previously said "redacted
   state fingerprint", but callers seed it with the raw `state` query
   value (or `flow.extension_name` for post-resolution failures).
   Updated the doc to describe the actual behaviour: an arbitrary seed
   that is hashed before any hex output, with a pointer to
   `redact_oauth_state_for_logs` for the log-safe fingerprint.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(e2e): make Telegram pairing chat-claim suite actually run

The scenario landed in PR #3317 was orphaned — never wired into any CI
lane and could not pass even when run by hand. Three structural
issues, all fixed here:

1. `install_telegram` now overlays the locally-built WASM (and matching
   capabilities file) on top of the registry-downloaded artifact when
   present. The pairing-reply wording lives inside the WASM binary, so
   without this overlay the test was asserting source-tree text against
   the previous release's bytes. The overlay is best-effort: when the
   local WASM is absent (CI groups that don't build the channel), the
   test that depends on it skips with a clear message and the canary
   lane in `scripts/live_canary/auth_registry.py` still covers the
   wording end-to-end against the deployed binary.

2. `Submission::PairingClaim` is handled out-of-band by the bridge
   layer; the response is delivered via `WebChannel::respond` →
   `AppEvent::Response` over SSE only — no `Turn` is persisted, so
   polling `/api/chat/history` could never see it. Refactored
   `test_chat_surface_approves_pairing_code` and
   `test_chat_surface_rejects_invalid_pairing_code` onto a
   `_send_and_collect_response` helper that opens the SSE stream first
   (so the broadcast doesn't fan out to zero subscribers) and matches
   on the `response` event for the test thread.

3. Wired the file into `e2e.yml`'s `extensions` group so the suite
   actually runs on every PR.

Verified locally: all three scenarios pass, plus the existing 22
Telegram e2e tests still green with the install-overlay change.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(bridge): bound and sanitize invalid-channel echo in pairing claim

Address Copilot review on PR #3381: `handle_pairing_claim`'s
invalid-channel branch was rendering the raw `channel` token (and the
underlying `IdentityError`, which itself echoes the offending input)
back to the user. Both routes are unbounded and could carry control
characters or markup since `channel` comes from chat input — a
hostile prompt could blow up the SSE / Telegram / TUI reply or smuggle
backticks/escape sequences through.

Cap the echo at 32 ASCII-alphanumeric (or `-`/`_`) characters and
replace the verbatim error with a fixed category description, so the
reply size and shape are bounded by what we render explicitly. Add a
regression test that drives a 200-char hostile blob (control chars +
backticks) through the handler and asserts the rendered reply stays
clean and short.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(bridge): include hyphens in invalid-channel error copy

Address Copilot review on PR #3381: the invalid-channel reply
listed "lowercase letters, digits, or underscores" as the valid
character set, but `ExtensionName::new` (and `web::features::pairing::
parse_channel`) intentionally accept hyphens too — they're folded to
underscores during canonicalization. A user typing `slack-relay` would
otherwise get an "invalid name" reply listing rules that contradict
the actual validator.

Updated the message to include hyphens with `telegram` and
`slack-relay` as concrete examples, and tightened the regression test
to assert against the user-controlled preview region between the
delimiter backticks rather than a global backtick count (which was
fragile to copy that includes example slugs in backticks).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(auth): address PR #3381 review on credential-scoped gate cleanup and Telegram surface promise

Three reviewer findings, one commit:

- OAuth provider-error and exchange-failure paths used
  `clear_engine_pending_auth(user, None)`, which discards every
  Authentication gate for the user. A failed Gmail callback could
  silently wipe an unrelated Slack/MCP gate waiting on a different
  thread. New `clear_engine_pending_auth_for_credential(user, credential)`
  helper in bridge::router scopes cleanup to the failed flow.
  Provider-error path tracks `removed_secret_name` alongside
  `removed_user_id` so the scoped variant is callable.
- Expired-flow branch in the OAuth callback handler had two bugs: it
  never cleared the engine pending auth gate (so the conversation sat
  blocked forever, same #3320 class the provider-error fix addresses),
  and the broader `clear_auth_mode` it called would re-discard via
  the unscoped helper anyway. Now calls the credential-scoped helper
  and the legacy-v1-only `clear_session_auth_mode_for_thread`.
- Telegram pairing reply advertised `approve telegram CODE` as
  usable "in any IronClaw chat (TUI / web / Telegram)", but an
  unpaired Telegram DM is intercepted by the allowlist gate before
  the agent parser sees the command — the user would just get
  another pairing reply. Reply now lists only the surfaces that
  actually work (web / TUI / CLI) and a comment explains why.

Regression coverage:

- `clear_engine_pending_auth_for_credential_only_clears_matching_credential`
  locks in helper scoping (Gmail/Slack two-gate scenario).
- `oauth_callback_expired_flow_clears_credential_scoped_engine_gate`
  drives the full callback through axum oneshot with engine state
  seeded; asserts the matching gate clears and the unrelated gate
  survives.
- E2E `test_telegram_dm_approve_command_is_intercepted_by_allowlist_gate`
  exercises the Telegram webhook path (not /api/chat/send) to lock in
  the channel-layer interception, wired into the auth canary lane.
- Existing E2E pairing-reply test gains an assertion that
  "TUI / web / Telegram" is *not* in the reply.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(oauth): mirror failure-cleanup contract on provider-error and reconcile stale comment

Copilot review on PR #3381 caught two real issues in the credential-scoped cleanup
landed in d45cf1bd8:

- Provider-error branch (`?error=access_denied`) returned the error page
  without broadcasting `OnboardingState::Failed` or clearing the legacy v1
  session `pending_auth`. The exchange-failure and expiry branches do both.
  Net effect: the auth card stayed spinning and the next user message was
  intercepted as a token. Now mirrors the other failure paths — keep the
  full `flow`, emit Failed SSE, clear v1 session, clear credential-scoped
  engine gate, then return the error page.
- Post-exchange comment said "failed callbacks should leave the gate
  visible for retry" — that was the pre-#3320 contract. Rewrote it to
  describe the new shape: each failure mode clears its own gate at the
  failure site; this section only handles legacy-v1 session cleanup that
  runs regardless of outcome. Also explains why we use
  `clear_session_auth_mode_for_thread` here instead of `clear_auth_mode`
  (the latter would re-clear the engine gate on the *success* path and
  break the `ExternalCallback` resume).

Regression: `test_oauth_callback_provider_error_broadcasts_onboarding_failed`
in the oauth tests module — drives an `?error=access_denied` callback with
a flow whose `sse_manager` is attached, asserts the receiver gets
`OnboardingState::Failed` with the provider's `error_description` as the
message body.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(common): describe paths and platform helpers in crate description (#3498)

Align the package `description` and the lib.rs crate-level doc with
the modules now exposed from `ironclaw_common` (paths, platform,
env_helpers, attachment), which #3387 lifted out of `src/`. The
previous wording predates that extraction and only mentioned "types
and utilities".

This is also the release-plumbing trigger for v0.28.1: release-plz
proposes a leaf bump on source-path changes, and once `ironclaw_common`
crosses to a new patch, the root `ironclaw` bump can be added on top
of the release-plz branch (same approach as commit 9e69f22d2 for
v0.28.0). See PR #3372 for the equivalent v0.28.0 trigger.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore: release

* chore(release): bump ironclaw to 0.28.1

cargo-semver-checks did not detect API-breaking changes in the
ironclaw_common 0.4.1 -> 0.4.2 leaf bump, so release-plz did not
cascade a bump into the root ironclaw package. Add the root version
bump and CHANGELOG entry manually so this release-plz PR produces
an ironclaw-v0.28.1 tag and triggers cargo-dist.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(llm): hide provider-specific auth, model fetch, and embeddings config behind facades (#3416)

* refactor(llm): hide provider-specific auth, model fetch, and embeddings config behind facades

External callers were reaching into provider-specific modules of
`ironclaw_llm` (`gemini_oauth::CredentialManager`,
`github_copilot_auth::*`, `OpenAiCodexSessionManager`,
`codex_auth::*`, `BedrockConfig`, etc.). Closes those leaks behind a
small set of verb-based public surfaces while keeping per-provider
behaviour inside the LLM crate.

Changes:

1. Extract `oauth_helpers.rs` into a new `ironclaw_oauth` crate. The
   loopback OAuth callback listener (port 9876, landing pages,
   `OAUTH_CALLBACK_HOST` rules) is shared by every IronClaw OAuth flow
   (NEAR AI session login, WASM tool auth, MCP) and never depended on
   `ironclaw_llm`. `src/auth/oauth.rs` now `pub use ironclaw_oauth::*`
   directly. `ironclaw_llm` no longer depends on `ironclaw_oauth` —
   the helper had zero internal callers.

2. Add `ironclaw_llm::auth` facade (`start_login`, `validate_token`,
   `default_headers`, `load_persisted_credentials`,
   `default_credentials_path`) with backend-agnostic types
   (`AuthPrompt`, `LoginRequest`, `AuthOutcome`, `PersistedCredentials`,
   `OpenAiCodexLoginOptions`, `AuthBackend`, `CredentialSource`).
   Privatize `gemini_oauth`, `github_copilot_auth`, `openai_codex_session`,
   `codex_auth` (`pub(crate) mod`). Migrate the wizard, the
   `ironclaw login --openai-codex` CLI subcommand, and the LLM config
   loader to the facade. Wizard introduces a single `WizardAuthPrompt`
   that handles device-code prompts + browser launch for all backends.

3. Add `ironclaw_llm::models::fetch_models_for(provider_id, &opts)`
   facade. Privatize `fetch_anthropic_models`, `fetch_openai_models`,
   `fetch_ollama_models`, `fetch_openai_compatible_models`,
   `is_openai_chat_model`, `openai_model_priority`, `sort_openai_models`.
   Wizard's per-backend match collapses to one call. Move classifier
   unit tests into `crates/ironclaw_llm/src/models.rs`; rewrite the
   two wizard fallback tests through the public API.

4. Decouple embeddings from `ironclaw_llm::BedrockConfig`. New
   `crate::workspace::BedrockEmbeddingSetup { region, profile }` carries
   only what `BedrockEmbeddings` actually needs. `EmbeddingsConfig::create_provider`
   and `BedrockEmbeddings::new` take the new type; callers translate from
   `LlmConfig.bedrock` at the boundary (`src/app.rs`, `src/cli/mod.rs`).

5. Add `ironclaw_llm::testing::nearai_test_config(model)` helper for
   tests that need a minimal `LlmConfig` shape (no retries, no caching,
   NEAR AI backend). Replaces two duplicated 30-line struct literals
   in the gateway settings hot-reload tests.

Boundary cleanup is behaviour-preserving: 4,932 main-binary unit tests,
729 ironclaw_llm unit tests, 4 ironclaw_oauth tests, 3 architecture
boundary tests all pass; `cargo clippy --all --benches --tests
--examples --all-features` is clean.

Three `pub` methods on `gemini_oauth::CredentialManager` /
`GeminiOauthProvider` (`get_valid_access_token`, `last_response_meta`,
`count_tokens`) and the `GeminiResponseMeta` struct are now reachable
only crate-internally and have no callers; marked `#[allow(dead_code)]`
with a comment rather than deleted to keep this PR purely a boundary
move (delete in a follow-up if no caller emerges).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(llm): promote dedicated backends into the registry; absorb config validation, defaults, and per-provider overrides into ironclaw_llm

Continues the LLM boundary cleanup from 0addf3ac2. After that commit
provider-specific auth, model fetch, and embeddings config lived behind
facades inside `ironclaw_llm`, but four backend-specific knowledge
sources still leaked out:

  1. Validation rules and default values for the dedicated-config
     backends (Bedrock cross-region prefixes, OpenAI Codex endpoints
     and client_id, Gemini OAuth credentials path defaults) lived
     inline in `src/config/llm.rs::resolve`.
  2. The dispatcher in `create_llm_provider` matched on backend strings
     ("nearai", "bedrock", ...) instead of a typed protocol value. The
     same booleans (`is_nearai`, `is_bedrock`, `is_gemini_oauth`,
     `is_openai_codex`) recurred across `src/config/llm.rs`,
     `src/app.rs`, `src/cli/models.rs`, and the wizard.
  3. The setup wizard had per-backend specialization in
     `step_inference_provider` and `run_provider_setup` (manual menu
     pushes for nearai/bedrock/codex/gemini_oauth, four dedicated
     `setup_*` entry points dispatched on string compares).
  4. `Settings` carried named `bedrock_region`, `bedrock_cross_region`,
     `bedrock_profile` columns even though no other dedicated backend
     had named columns and adding a new one would mean schema churn.

Layers A-D address each in turn:

* Layer A — `BedrockConfig::build`, `OpenAiCodexConfig::build`, and
  `GeminiOauthConfig::build` own validation + defaults inside the
  crate. `LlmConfigError` (`MissingRequired` / `InvalidValue`) carries
  the failures across the boundary, with a `From` impl into the
  binary's `ConfigError`. `src/config/llm.rs` calls the builders;
  named-string defaults are gone from the binary. The orphaned
  `tests/gemini_oauth_regression.rs` husk is deleted.

* Layer B — `ProviderProtocol` gains four new variants
  (`Bedrock`, `OpenAiCodex`, `GeminiOauth`, `NearAi`) plus a
  `has_dedicated_config()` predicate. The four dedicated-config
  backends (with all aliases) become first-class registry entries in
  `providers.json`, so `is_known()` / `model_env_var()` / the wizard /
  the gateway handler iterate the registry uniformly. The
  `is_nearai`/`is_bedrock`/`is_gemini_oauth`/`is_openai_codex` boolean
  spaghetti collapses to protocol comparisons. `OpenAiCodex` and
  `NearAi` carry explicit `#[serde(rename = "openai_codex" / "nearai",
  alias = ...)]` so the wire-stable adapter strings the gateway and
  frontend already use keep working. `LlmConfig::active_model_name()`
  is now consumed by `cli/doctor.rs` instead of an inlined partial
  dispatch.

* Layer C — `SetupHint` gains four credential-collection variants
  (`AwsCredentials`, `OAuthDeviceCode`, `FileBasedCredentials`,
  `SessionToken`). The wizard's `step_inference_provider` builds its
  menu from a single `registry.selectable()` iteration with generic
  env-detection (declared `api_key_env`, plus an Anthropic-specific
  OAuth fallback). `run_provider_setup` dispatches on the SetupHint
  variant; the remaining `def.id == "..."` checks live inside the
  `ApiKey` arm only because Anthropic and GitHub Copilot present a
  hybrid choice (API key OR OAuth) the simple `ApiKey` hint doesn't
  capture. The synthetic bedrock + nearai entries in
  `handlers/llm.rs::build_llm_providers` are deleted; a single
  registry-driven loop covers both. ADAPTER_LABELS in
  `static/js/surfaces/config.js` gains entries for the new protocols.

* Layer D — `LlmBuiltinOverride` gains a generic
  `extras: HashMap<String, String>` bag with `extra(key)` /
  `set_extra(key, value)` accessors. The bedrock resolver and wizard
  read/write through this bag; `Settings::migrate_legacy_provider_fields()`
  drains the named `bedrock_*` columns into `extras` on
  `Settings::load_from()` so existing `settings.json` files migrate
  losslessly. The named columns are kept (deprecated, marked with
  `#[serde(skip_serializing_if = "Option::is_none")]`) for one
  release; tracked for deletion in #3443.
  `strip_admin_only_llm_keys` and `llm_setting_requires_reload` now
  match dotted-path subkeys under `llm_builtin_overrides.*` so a
  write to e.g. `llm_builtin_overrides.bedrock.extras.region`
  triggers the right gating + chain reload.

Boundary cleanup is behaviour-preserving: 4,933 main-binary unit tests,
739 ironclaw_llm unit tests pass; `cargo clippy --all --benches --tests
--examples --all-features` is clean. New regression tests:
`crates/ironclaw_llm/src/config.rs` (6 builder tests),
`crates/ironclaw_llm/src/registry.rs::dedicated_config_backends_are_in_registry_and_selectable`,
and `src/setup/wizard.rs::legacy_bedrock_fields_migrate_into_extras_on_load`.

Three follow-ups tracked in #3443: delete the deprecated `bedrock_*`
named columns, move `BedrockEmbeddings` out of `src/workspace/` into
the LLM crate (last cargo-feature leak), and drive
`LlmConfig::active_model_name()` off `ProviderProtocol` instead of
backend strings.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test: add bug-bash regression-snapshot harness

Bug-bash fixtures pin specific open bugs to a deterministic snapshot.
When a bug is fixed, the snapshot diff is the reviewable proof; when
someone reintroduces the bug, the snapshot drifts and CI blocks the
merge.

This commit lands the harness plus the first recorded fixture for
issue #2541 (agent must call a tool, not answer from training data):

  tests/e2e_bug_bash_snapshots.rs
    `snapshot_summarization_uses_tools` replays the fixture, captures
    `ReplayOutcome`, and asserts the YAML snapshot. Gated on
    `feature = "libsql"`, same as other replay-snapshot tests.

  tests/fixtures/llm_traces/bug_bash/summarization_uses_tools.json
    Two-step recorded LLM trace (tool_call -> text) keyed off the
    user prompt via `request_hint.last_user_message_contains`.

  tests/fixtures/llm_traces/bug_bash/README.md
    Coverage map for #2540-#2546 (one recorded, six TODO) plus the
    `IRONCLAW_RECORD_TRACE` recording workflow.

  tests/snapshots/replay__bug_bash_summarization_uses_tools.snap
    Insta YAML snapshot pinning `tool_calls: [echo]`, 2 LLM calls,
    and the event-kind histogram. Drift = regression.

* fix(settings): preserve pre-existing extras during legacy bedrock migration

`migrate_legacy_provider_fields` claimed to be idempotent and to drain
named `bedrock_*` columns into `llm_builtin_overrides["bedrock"].extras`
once on load. The previous implementation drained correctly but used
`HashMap::insert` unconditionally, which means a settings file
carrying BOTH a legacy `bedrock_region` column AND an already-populated
`extras["region"]` (manual hand-edit, or a future writer emitting both
shapes during a transition) would silently downgrade to the legacy
value.

Guard each `set_extra` call with `entry.extra(key).is_none()` so the
new-shape value always wins. Clarify the docstring to state this
explicitly.

Add three regression tests in `settings::tests`:

- `legacy_bedrock_migration_round_trips_through_save` — legacy JSON ->
  load_from -> serialize -> reload, asserts the deprecated columns are
  not re-emitted and extras survive the round trip.
- `legacy_bedrock_migration_preserves_existing_extras` — file with both
  shapes; asserts the pre-existing extras value is kept and absent
  extras are still backfilled from legacy fields.
- `legacy_bedrock_migration_is_idempotent_in_memory` — calling the
  migration twice on the same Settings is a no-op (compares serialized
  shape, since LlmBuiltinOverride does not derive PartialEq).

* fix(pr-3416): address PR review — migration on DB/TOML, admin-key gate, codex login, credential_kind/has_credentials

Addresses comments from gemini-code-assist, Copilot, and serrrfirat on PR #3416.

## Bugs

**Legacy bedrock fields not migrated on DB/TOML loads** (serrrfirat, High).
`Settings::load_from` (JSON) ran `migrate_legacy_provider_fields`, but
`from_db_map` and `load_toml` did not. Existing operators with
`bedrock_*` settings persisted in the DB or `config.toml` would silently
lose their AWS region/profile/cross-region after upgrade because the
resolver now reads only from `llm_builtin_overrides["bedrock"].extras`.
Both loaders now call the migration; added round-trip tests for each.

**Admin-only key write gate had narrower matching than read gate**
(Copilot, High). `strip_admin_only_llm_keys` matches both exact keys
and dotted subpaths under admin-only roots; `is_admin_only_setting_key`
in the web settings handler used `.contains(&key)` only. A non-admin
could write `llm_builtin_overrides.bedrock.extras.region` directly,
bypassing the gate. Promoted `is_admin_only_llm_key` to `pub(crate)`,
made the web write-side gate call it, added regression tests covering
dotted subpaths.

**`ironclaw login --openai-codex` dropped TOML/DB config** (Copilot,
High). The pre-refactor code resolved `Config::from_env` and used
`config.llm.openai_codex` so endpoint / client-id / session-path
overrides committed via TOML or DB stuck. The post-refactor code only
read env vars via `OpenAiCodexLoginOptions::from_env`. Added
`OpenAiCodexLoginOptions::from_resolved_config(&OpenAiCodexConfig)`;
the login command now prefers the resolved config when present and
falls back to env-only when `Config::from_env` itself fails (fresh
machine, no DB).

**Dedicated-auth backends marked configured without credentials**
(serrrfirat, Medium). `nearai` / `gemini_oauth` / `openai_codex` ship
`api_key_required: false` because they don't authenticate via a bearer
API key. The frontend `isProviderConfigured` treated that as "no
credentials needed" and rendered the Use button on a fresh install,
where clicking could trigger an interactive device-code OAuth from
inside a settings request.

Added `credential_kind` (wire-stable snake_case discriminator matching
`SetupHint::kind()`, e.g. `session_token`, `o_auth_device_code`,
`file_based_credentials`, `aws_credentials`) and `has_credentials`
(backend-authoritative; checks AWS env vars for Bedrock, codex session
file existence, file-based credential path expansion + existence) to
the web LLM providers payload. Frontend `isProviderConfigured` /
`providerMissingReason` now gate non-api-key kinds on `has_credentials`.

## Nits

**`fetch_models_for` doc overclaimed "Always returns something"**
(Copilot). The generic openai-compatible branch returns `vec![]` when
`base_url` is empty. Updated the docstring to call this out so callers
know to handle the empty case.

**`AuthError::Other` used for "validation not applicable"** (Gemini
bot). Added a dedicated `AuthError::TokenValidationNotSupported { backend }`
variant; `validate_token` now returns it for Gemini / OpenAiCodex
instead of stringly-formatted `Other`.

**Bug-bash regression-harness URLs pointed at `near/ironclaw`**
(Copilot, x2). The canonical tracker is `nearai/ironclaw`. Rewrote
all seven URLs in `tests/fixtures/llm_traces/bug_bash/README.md` and
the one in `tests/e2e_bug_bash_snapshots.rs`.

## Declined

The Gemini bot's MalformedConfig suggestion at
`crates/ironclaw_llm/src/models.rs:46` was not adopted: the call site is
the openai-compatible model-listing path, not a security-sensitive
request. The fetcher early-returns `vec![]` on empty `base_url` — no
URL parsing happens — and the docstring tightening above covers the
observable surprise. Promoting it to a typed error would change the
public-facing `fetch_models_for` signature for no behavioural gain.

## Tests

- `cargo fmt --check` clean
- `cargo clippy --all --benches --tests --examples --all-features` zero warnings
- `cargo test --lib` 4,941 / 4,941 pass
- `cargo test --features libsql --test e2e_bug_bash_snapshots` 1 / 1 pass
- New regression tests:
  - `settings::tests::legacy_bedrock_fields_migrate_into_extras_on_db_load`
  - `settings::tests::legacy_bedrock_fields_migrate_into_extras_on_toml_load`
  - `channels::web::features::settings::tests::test_admin_only_setting_keys_cover_dotted_subpaths`
  - `channels::web::handlers::llm::tests::test_llm_providers_expose_credential_kind_and_has_credentials`
  - `channels::web::handlers::llm::tests::test_nearai_has_credentials_true_when_session_token_loaded`

* fix(pr-3416): tighten Bedrock/Codex has_credentials probes; collapse set_extra into one .into()

- `backend_has_credentials` for AWS now requires `AWS_PROFILE` OR
  (`AWS_ACCESS_KEY_ID` AND `AWS_SECRET_ACCESS_KEY`). The lone
  `AWS_ACCESS_KEY_ID` / `AWS_SESSION_TOKEN` arms previously flipped
  has_credentials true even though the AWS SDK can't sign without the
  secret key, so the UI was rendering Bedrock as configured on hosts
  that would fail at first call.
- `backend_has_credentials` for OpenAI Codex now honours
  `OPENAI_CODEX_SESSION_PATH` via `read_env` before falling back to
  the default session path under `~/.ironclaw/`. Users with a custom
  session location were seeing "not configured" despite a valid login.
- New regression tests `test_bedrock_partial_aws_env_reports_not_configured`
  and `test_openai_codex_honours_session_path_env` drive the
  `build_llm_providers` call site (not just the helper) so both gaps
  stay closed.
- Tidied `LlmBuiltinOverride::set_extra` to convert the key once and
  reuse it across the remove/insert branches; behaviour identical.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore(providers): default nearai model to "auto"

Switch the nearai registry entry's `default_model` from
`claude-sonnet-4-5` to `auto`, NEAR AI's server-side routing alias.
New installs without `NEARAI_MODEL` set now get auto-routed instead
of being pinned to a specific Anthropic model.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Make Skills E2E lifecycle deterministic (#3309)

* test(e2e): make skills lifecycle deterministic

* test(e2e): address skills review comments (#3309)

* test(e2e): unxfail two auth-matrix tests now that contracts match (#3589)

Both xfails in tests/e2e/scenarios/test_v2_auth_oauth_matrix.py are
stale and pass against current code:

- test_wasm_tool_first_chat_auth_attempt_emits_auth_url
  Marked xfail in #3235 because the engine-v2 callable-only contract
  (#2868) stopped emitting an auth gate on direct LLM-driven tool
  calls. PR #3157 (auth-preflight + inline-await) restored the
  behavior the test asserts: when the LLM emits a direct call to a
  not-yet-authed extension, the bridge raises an Authentication gate
  with auth_url populated (src/bridge/effect_adapter.rs:1356-1392).
  Marker removed; test passes.

- test_settings_first_custom_mcp_auth_then_chat_runs
  The xfail reason claimed post-auth tool-output propagation was
  broken. Real cause: engine-v2 gates the first MCP tool call on
  `approval` and the browser fixture has no auto-approve UI, so the
  chat sat in pending_gate forever. Same shape as the bugs fixed in
  #3235 for test_wasm_tool_oauth_refresh_on_demand and
  test_mcp_same_server_multi_user_via_browser. Inserted
  _wait_for_tool_call between _send_chat and _wait_for_response_contains
  to drive approval through the API; test passes.

Verified locally: both tests pass back-to-back in 27s on a fresh
auth_matrix_server.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(extensions): restore chat-driven tool_install + fix double-invoke + auto-approve footgun (#3533) (#3559)

* fix(extensions): restore chat-driven tool_install + fix double-invoke + auto-approve footgun (#3533)

"Connect my telegram" was giving the user two options and not actually
installing anything because three layered issues had accumulated since
engine v2:

1. **`tool_install` was hidden from the agent** (#2868). The unified
   `tool_activate` it was meant to be subsumed by was later removed in
   #3166, but the hidden-from-callable-surface gate stayed. Restored
   by dropping `hidden_from_model_callable_surface` from
   `bridge::action_projector`. User consent is mediated by the tool's
   own `ApprovalRequirement::UnlessAutoApproved` and the seeded
   `AskEachTime` permission.

2. **Two competing Telegram registry entries** (`telegram` channel and
   `telegram_mtproto` tool) both surfaced in the agent prompt's
   `Activatable Integrations` section. The LLM correctly enumerated
   them as "Option 1" and "Option 2" instead of installing the
   canonical bot channel. Added a `hidden: bool` field to
   `ExtensionManifest` / `RegistryEntry`, set `telegram_mtproto` to
   `hidden: true`, and filter hidden entries out of the
   "available-but-not-installed" appendix in `ExtensionManager::list`.
   Hidden entries remain installable by explicit name.

3. **Updated the agent prompt** so `Activatable Integrations` instructs
   the model to call `tool_install(name="<name>")` directly rather than
   describing manual UI steps.

Fixes the double-`tool_install` invocation that surfaced once the agent
could install from chat:

- **`InlineGate` discarded cached output.** The bridge raised an
  Authentication gate after `tool_install` succeeded, and the
  inline-await retry re-executed the action (re-downloading the WASM
  bundle) instead of returning the already-computed output. Added
  `resume_output: Option<serde_json::Value>` to `InlineGate`; on
  approval, return the cached output if present. Mirror fix in the
  orchestrator's `execute_single_action_with_inline_retry` (reading
  `result_json["resume_output"]`) and the structured-batch retry path.
- **`effect_adapter::auth_gate_from_extension_result`** now passes
  `Some(output_value.clone())` as the gate's `resume_output` so the
  retry has cached state to short-circuit on.
- **OAuth callback double-fired.** `oauth_callback_handler` now skips
  the `ExternalCallback` re-entry when the inline-await path already
  woke a parked waiter — eliminates the "thread already running" race.
- **`resolve_inline_gates_for_credential`** now also discards matching
  Authentication rows from `pending_gates` so the row doesn't linger
  in `HistoryResponse.pending_gate` after inline resolution.

Fixes the auto-approve footgun:

- **`ToolPermissionSnapshot::resolve_permission`** now collapses DB
  values that match the seeded default to `explicit = None`. Before
  this, the boot-time `seed_tool_permissions` write of `tool_install ->
  AskEachTime` was indistinguishable from a user-explicit override,
  causing `effect_adapter::enforce_tool_permission`'s `is_explicit_ask`
  check to refuse `AGENT_AUTO_APPROVE_TOOLS=true`. Real overrides
  (`AlwaysAllow`, `Disabled`) still surface as `Some(...)`.

Tests
- Unit: 4980/4980 pass (host) + 525/525 pass (engine).
- Unit: new tests in `bridge::tool_permissions::tests` lock seeded-vs-
  explicit collapse; new test in `bridge::action_projector::tests`
  asserts `tool_install` is callable; new manifest hidden-flag tests
  in `registry::manifest::tests` and `extensions::manager::tests`.
- E2E: removed `@pytest.mark.xfail` on
  `test_chat_first_gmail_installs_prompts_and_retries` (now passes
  end-to-end via the chat-driven install path). Added
  `test_chat_install_approval_then_auth_card` driving the
  explicit-approval variant with a single Approve click (no Always
  workaround needed) — wired into the `auth-full` canary lane.
- Mock LLM: extended the gmail-install-then-retry pattern to recognize
  both the legacy "Extension not installed:" and the post-#3533 "is
  not callable in this execution context" error strings, and to retry
  `gmail(action="list_messages")` after a successful `tool_install`.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(permissions): address #3559 review (permission bypass, lease accounting, hidden search filter)

Five fixes from the #3559 review (4× Copilot doc nits + 3× serrrfirat
security/correctness findings):

1. **Permission bypass (High).** Pre-#3559's `resolve_permission`
   collapsed any DB row whose value matched the seeded default to
   `explicit = None`, so a user who deliberately set `tool_install =
   AskEachTime` had their explicit choice silently dropped and
   `AGENT_AUTO_APPROVE_TOOLS=true` bypassed the gate. Provenance is now
   handled at write time: `seed_tool_permissions` is gone and a
   one-shot, sentinel-gated migration (`cleanup_ghost_seeded_tool_permissions`)
   deletes existing ghost-seeded rows at startup. With no ghost rows,
   the resolver treats every DB row as user-explicit and honors it.

2. **Lease/event accounting on `resume_output` replay (Medium).**
   Inline-gate handlers in `structured.rs`, `scripting.rs`
   (`resolve_tool_future` + `drive_inline_gate` retry loop), and
   `orchestrator.rs` refunded the lease use the action just consumed,
   then returned the cached `resume_output` on approval without
   re-consuming — netting successful side-effecting actions to zero
   lease uses. Skip the refund when the gate carries cached output.

3. **Hidden registry filter on `tool_search` (Medium).**
   `RegistryCatalog::search` did not filter `hidden: true` entries,
   so `telegram_mtproto` could resurface through the search path and
   reintroduce the "two Telegram options" outcome that #3533 fixes
   for the default-list path. Added the filter and a regression test.

4-7. Copilot doc nits: outdated `_set_tool_permission` docstring;
   misleading "bridge-side auto-install implemented" comment in
   `mock_llm.py`; `tool_install` described as "non-agent surface" in
   `src/bridge/CLAUDE.md` while a paragraph below says the model
   calls it directly; dangling `issue #3533 / PR —` placeholders
   in both CLAUDE.md docs.

Regression tests:
- `bridge::tool_permissions::user_explicit_value_matching_seeded_default_stays_explicit` —
  the original Copilot/serrrfirat bug case.
- `app::cleanup_ghost_seeded_tool_permissions_removes_seed_matching_rows` —
  idempotent migration + sentinel.
- `extensions::registry::test_search_skips_hidden_entries` — hidden
  entries excluded from search but still installable by exact name.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(#3559): caller-level regression coverage for review findings 1 & 2

Two follow-up regression tests for the #3559 security review, plus a
real bug surfaced by the first one.

1. `executor::structured::resume_output_replay_consumes_exactly_one_lease_use`
   exercises the post-execution Authentication gate inline-retry path
   with `max_uses=1` and asserts:
   - Cached output is returned as a successful `ActionResult`.
   - Exactly one `ActionExecuted` event is emitted.
   - The lease budget is exhausted after one execution (refund-skip
     keeps the consumption from being undone).

   Writing this test surfaced a real bug: the structured cached-output
   branch pushed `ActionExecuted` into `emitted_events`, and the
   caller's `classify_exec_result` emitted ANOTHER terminal
   `ActionExecuted` for the same Ok result — double-emit for one
   action. Tier 1 (`scripting::drive_inline_gate`) and Tier 1 alt
   (`orchestrator::execute_action_with_inline_gate`) emit themselves
   because their callers don't run an Ok-branch classifier; structured
   was the outlier. Dropped the redundant push; the classifier emits
   the single canonical event.

2. `bridge::effect_adapter::explicit_ask_each_time_for_seeded_default_tool_still_gates`
   drives `execute_action` end-to-end (the side-effecting caller) with
   a tool whose `name()` matches a seeded-`AskEachTime` baseline
   (`tool_install`) and an explicit `AskEachTime` user override. The
   resolver collapse-to-implicit bug would have shown up here — not
   just in the helper-level test that already exists in
   `bridge::tool_permissions::tests`. Per `.claude/rules/testing.md`
   "Test Through the Caller, Not Just the Helper".

   Added `SeededAskEachTimeTestTool` as a `tool_install`-named test
   fixture with `requires_approval: UnlessAutoApproved` to mirror the
   real tool's contract.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore: release

* feat(engine): IRONCLAW_DISABLE_CODEACT flag to disable v2 CodeAct (#3665)

* add flag to disable codeact on engine v2

* fmt

* fix(engine): keep compact actions reachable when CodeAct is disabled

With IRONCLAW_DISABLE_CODEACT=true the structured-tools prompt told
the model to use the provider's tool_calls interface for every action,
but the bridge filtered the provider tool list down to
emits_full_schema_tool(). Most tools default to CompactToolInfo
(mission_create, gmail_send, notion_search, ...), so they appeared in
the prompt as "available" while being absent from the provider tool
list — i.e. unreachable. Addresses serrrfirat's review on PR #3665.

Fix coordinates both halves of the surface:

- src/bridge/llm_adapter.rs: in disabled-CodeAct mode, drop the
  emits_full_schema_tool() filter and emit every action into the
  provider tool list with its full schema.
- crates/ironclaw_engine/src/executor/prompt.rs: in disabled-CodeAct
  mode, skip the "## Enabled Tools" section. The compact-form listing
  with the tool_info(detail="schema") instruction is meaningless when
  the provider already sends full schemas, and would just duplicate
  the surface. "## Activatable Integrations" stays — the model still
  needs to know what tool_install can target.

Test seam: build_codeact_system_prompt_inner now takes disable_codeact
as an explicit parameter, called once at the public entry points. This
lets prompt tests exercise both branches without process-global env
mutation.

Tests:
- executor::prompt::tests::disabled_codeact_omits_enabled_tools_section_and_keeps_activatable
- bridge::llm_adapter::tests::complete_emits_compact_actions_when_codeact_disabled
- existing complete_with_tools_only_emits_full_schema_provider_tools
  now serialized via lock_env() so env mutation in the new test
  can't leak across parallel runs.

cargo test -p ironclaw_engine --lib: 527 passed
cargo test --lib bridge::: 469 passed
cargo clippy -p ironclaw_engine --all-targets -- -D warnings: clean
cargo clippy --lib --tests -- -D warnings: clean
cargo fmt --check: clean

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Emil Bogomolov <emil.bogomolov@near.ai>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Fix markdown_to_mrkdwn to avoid converting emphasis inside generated <… (#3532)

* agent: Fix markdown_to_mrkdwn to avoid converting emphasis inside g…

* agent: Fix rustfmt/clippy CI failure by removing extra blank line b…

* agent: slack: fix markdown_to_mrkdwn replacement order to satisfy p…

* slack: protect generated links and sanitize sentinels in markdown_to_mrkdwn

Two issues raised on PR #3532 review:

1. Emphasis inside generated `<url|text>` was still rewritten because the
   global `**`/`~~` → `*`/`~` substitution ran after link materialization.
   Push the generated link span into the same protected arena used for
   Slack-native `<...>` constructs so subsequent global replacements can't
   reach inside it. Matches the PR's stated goal.

2. Untrusted input containing the private-use sentinel chars
   (U+E000 / U+E001) could forge a protected-span reference and pull in
   another span's content. Strip those chars from input up front.

Adds regression tests for both. Bumps registry/channels/slack.json
0.3.2 → 0.3.3 to satisfy the channel-source version-bump check.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* slack: escape link labels, drop pipe/gt URLs, expand nested sentinels

Addresses two follow-up review concerns on PR #3532:

Copilot review: `<url|text>` was built by string concatenation, so
`|` or `>` inside the URL would corrupt the entity, and `<` / `>`
inside the label would open/close a Slack span and break the link.
The label now escapes `<` → `&lt;` and `>` → `&gt;` (Slack's documented
literal-character form); a URL containing `<`, `>`, or `|` falls back
to leaving the original markdown form intact (those chars are not
valid URL characters per RFC 3986 anyway).

Latent nested-sentinel bug introduced by the previous fix: a markdown
link whose label contained a Slack-native `<...>` span (e.g.
`[<@U1> hi](url)`) ended up with the inner sentinel buried inside the
arena entry for the outer link span. The final restore pass advances
past the outer sentinel without rescanning what it just emitted, so
the raw U+E000/U+E001 characters would leak into the output. URL and
label are now pre-expanded before the link span is pushed.

While here, factor the duplicated restore loop into
`expand_protected_spans`, reused by both the pre-link expansion and
the final restore, and lift the sentinel constants to file scope.

Adds three regression tests covering label-bracket escaping, the URL
pipe/gt fallback, and the nested-span case. Bumps
registry/channels/slack.json 0.3.3 → 0.3.4.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(gateway): add logs download button (#3588)

* feat(web): support externally-provided tools in Responses API (#3122)

* feat(web): support externally-provided tools in Responses API

Lets callers of `/v1/responses` (and `/api/v1/responses`) declare their
own `function`-typed tools and feed back results via
`function_call_output` items, matching the OpenAI Responses wire shape.

Since IronClaw's engine has no per-request tool surface, integration
happens at the prompt level: the catalog is rendered as
`<external-tools>` in the user message and the agent signals a call by
ending its response with a fenced ```` ```tool_call ```` block. When
that fence is recognised, the reply is split into a leading `Message`
plus a `function_call` `ResponseOutputItem`.

Validation rejects unsupported tool types (`web_search`, `file_search`,
`code_interpreter`) and tools missing `name` with 400, with two new
integration tests covering both paths.

* refactor(responses-api): switch external tools to engine v2 native path

Replace the prompt-level fence protocol from PR #3122 with engine v2
native tool calls: caller-supplied `tools[]` are surfaced as real
LLM-callable actions, the engine pauses with `ResumeKind::External`
when one is invoked, and the bridge router projects the pause to a
new `AppEvent::ExternalToolCall` carrying the OpenAI-shaped
`function_call` wire fields.

The integration is small because v2 already has the right primitives:

- `ResumeKind::External { callback_id }` and
  `GateResolution::ExternalCallback { payload }` already existed for
  OAuth-style callbacks.
- `agent_loop.rs:1480` already routes Responses API messages to
  `handle_with_engine` when `ENGINE_V2=true`, so no v2 migration of
  the endpoint itself is needed.
- `EffectBridgeAdapter::execute_action` is the single chokepoint
  where caller tools can be detected before they reach the dispatch
  pipeline.

Changes:

- New `src/bridge/external_tools.rs` (`ExternalToolCatalog`) — per-thread
  registry of caller-supplied `ActionDef`s, plus the `ext_tool:`
  callback-id helpers used to disambiguate external-tool pauses from
  OAuth/pairing pauses (which also use `ResumeKind::External`).
- `EffectBridgeAdapter` consults the catalog: any name in it is
  short-circuited to a `GatePaused { resume_kind: External {
  callback_id: ext_tool:<call_id> } }` before any registry dispatch,
  and `available_action_inventory` merges the catalog into the
  LLM-visible action surface (internal beats external on collision).
- `Submission::ExternalCallback` gains an optional `payload` field;
  `bridge::handle_external_callback` plumbs it into
  `GateResolution::ExternalCallback { payload }`. Fallback predicate
  `gate_resume_is_external` lets non-auth External pauses (i.e.
  caller-tool resumes) resolve through the same handler.
- New `AppEvent::ExternalToolCall` projected by `notify_pending_gate`
  when a paused gate carries an `ext_tool:` callback id; OAuth/
  pairing flows keep flowing through the existing `GateRequired`
  channel.
- `responses_api.rs` is gutted of the prompt rendering and fence
  parsing (`render_external_tools_preamble`, `extract_trailing_tool_call`,
  `parse_external_tool_call`, `ParsedToolCall`, `external_tool_names`
  accumulator field, and the `TOOL_CALL_FENCE` constants). The handler
  now: rejects `tools[]` when `ENGINE_V2=false`, registers caller
  tools in the catalog under the resolved thread id, detects resume
  requests (`previous_response_id` + `function_call_output` items in
  `input`) and submits them as `Submission::ExternalCallback` with
  the outputs as the resolution payload, and surfaces
  `AppEvent::ExternalToolCall` as a `function_call` `ResponseOutputItem`
  in both streaming (`output_item.added`+`done`) and non-streaming.
- All existing OAuth/pairing `ExternalCallback` constructors updated
  to pass `payload: None` (no behaviour change).
- Fence-protocol unit tests removed; replaced with coverage for the
  new `responses_tools_to_action_defs` converter and the accumulator's
  `ExternalToolCall` arm.

Existing 9 integration tests in `tests/responses_api_path_prefix.rs`
still pass.

Note for reviewers:
- The accumulator-side text response no longer tries to split the
  reply on a fenced `tool_call` block. The wire shape that callers
  receive for caller-tool invocations is purely event-driven now.
- Internal vs external collision is handled silently by the dedup in
  `available_action_inventory` (internal wins). A request-time
  rejection for shadowing names is a follow-up — the current behavior
  is safe (the LLM only sees the internal version) but could surprise
  a caller who expects their tool to run.

* test(responses-api): cover ENGINE_V2-off and resume-without-pending-gate

Two new integration tests for behaviours added by the engine-native
external-tool refactor:

- `external_tools_rejected_when_engine_v2_disabled`: a request with
  caller-supplied `tools[]` while `ENGINE_V2` is off must 400 with a
  message naming the flag, not silently fall through.
- `resume_without_pending_gate_returns_400`: a request with
  `function_call_output` items and a `previous_response_id` that
  doesn't correspond to a live external-tool gate must 400, not start
  a fresh turn against the (unrelated) thread.

Both tests drive the full router (`start_test_server` + bearer auth)
per `.claude/rules/testing.md` "Test Through the Caller".

* test(responses-api): integration tests + drop unsafe env mutation

Three groups of changes:

1. **Drop unsafe env-var mutation in tests.** `responses_api.rs` no
   longer reads `ENGINE_V2` directly: it keys off the presence of the
   live `ExternalToolCatalog` (initialized by `init_engine`) as the
   "engine v2 is up" signal. The path-prefix test that exercises the
   no-engine branch no longer needs `unsafe { std::env::remove_var }`
   — the absence of `init_engine` in `TestGatewayBuilder` is what
   makes the catalog absent, which is what makes the request reject.

2.…
theredspoon pushed a commit to theredspoon/ironclaw that referenced this pull request Jun 21, 2026
…nearai#2868)

* engine-v2: make available_actions callable-only for blocked providers

* fix(engine): address review fixture tempdir leak (nearai#2868)

* engine-v2: refresh canonical prompt metadata on resume (nearai#2869)

* fix(engine): align prompt metadata refresh with resume state

* fix(engine): finish prompt refresh compaction coverage (nearai#2869)

* fix(engine): preserve prompt refresh on resume (nearai#2869)

* Add engine v2 action discovery metadata (nearai#2876)

* Add engine v2 action discovery metadata

* fix(engine): address action discovery review (nearai#2876)

* fix(engine): address follow-up review comments (nearai#2876)

* fix(engine): satisfy clippy in orchestrator lookup

* fix(engine): propagate action snapshots in executor paths (nearai#2876)

* fix(bridge): restrict tool_info to callable actions (nearai#2876)

* [codex] Finish engine v2 deferred action inventory cleanup (nearai#2889)

* Add deferred action inventory groundwork

* fix(engine): address deferred action inventory follow-up

* fix(engine): address deferred inventory review feedback

* test: fix fmt and clippy failures

* engine-v2: trim unused callable discovery payload

* tests: restore env vars in review-fix cases

* engine-v2: populate callable snapshots consistently

* Unify v2 integration enablement on tool_activate

* engine-v2: tighten tool_info inventory and approvals

* llm: normalize tool_info hint syntax

* engine-v2: tighten tool_activate install approval lookup

* tests: align gmail settings-first flow with approval contract

* engine-v2: fix remaining tool surface review issues

* engine-v2: restore auto-approve defaults

* fix(engine): align v2 tool permissions with defaults

* fix(engine): close v2 callable snapshot gaps

* fix(bridge): label latent-only providers accurately
theredspoon pushed a commit to theredspoon/ironclaw that referenced this pull request Jun 21, 2026
* fix(oauth): remove pending flow on provider-error callback

The /oauth/callback handler's ?error= branch (RFC 6749 §4.1.2.1
provider-side failures — user cancels consent, scope denied, etc.)
returned the error page immediately without removing the flow from
ext_mgr.pending_oauth_flows(). The ghost entry then lingered until
the 5-minute expiry sweep, and any subsequent auth dance for the
same (extension, user) pair had to dedupe against it.

Mirror the happy-path cleanup: decode the state param, remove the
keyed flow, then return the error page.

Surfaced during live-canary auth-full repro: after
test_wasm_tool_oauth_provider_error_leaves_extension_unauthed ran,
the stale flow sat in the shared auth_matrix_server fixture.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(e2e): widen auth OAuth matrix timeouts for CI load

Four tests in live-canary auth-full were failing in CI with
`Page.wait_for_function: Timeout 60000ms exceeded`,
`ClientConnectionError('Connection closed')`, and
`Timed out waiting for OAuth refresh request` — all inside 60/20s
deadlines that are tuned for a dev laptop and don't leave margin
for ubuntu-latest's 2-vCPU runner under full suite load.

Raise the per-call deadlines so the inner budgets fit comfortably
inside pyproject.toml's 120s per-test cap:

  _wait_for_refresh_request default: 20.0s -> 60.0s
  _wait_for_auth_event call site:      60   -> 90
  _wait_for_auth_prompt call site:     60   -> 90
  send_chat_and_wait_for_terminal_message call sites: 60000 -> 90000
  _wait_for_mock_google_tokens call site: 60.0 -> 90.0
  _wait_for_response_contains (gmail) call site: 60.0 -> 90.0

Strictly widening; no passing test is slowed, no semantics change.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(canary): Haiku-powered Slack report job

Replace the team's raw Slack subscription (firehose of workflow
notifications) with one curated per-run summary:

  Canary: 9 passed, 1 failed of 10 lanes
  :x: auth-full (mock) — 12/13 passed, 1 failed in 350s
  > test_wasm_tool_first_chat_auth_attempt_emits_auth_url timed
  >   out waiting for auth_required SSE event on the fresh thread
  tools: shell, http_request, gmail (~6 calls)
  ...
  commit `abc1234` • <github run link>

New `canary-report` job (needs: every lane, if: always) downloads
all lane artifacts, parses junit + summary + log tail per lane, and
asks claude-haiku-4-5 to return a compact JSON per lane
({status, reason, tool_calls_total, tools_used, notable}). That's
aggregated into a single Slack block message and posted via
incoming webhook.

Safety shape:
- Script exits 0 even on Haiku/Slack failure so the notifier never
  masks the underlying canary signal.
- Missing ANTHROPIC_API_KEY falls back to raw junit-only phrasing.
- Slack POST failure falls back to plain-text "X/Y lanes failed"
  with the GH run URL so the channel still hears something.
- No new Python deps — pure stdlib (urllib.request, xml.etree).
- 20 KB log-tail cap per lane to keep Haiku token usage bounded.

Secrets:
- ANTHROPIC_API_KEY (already present, used by provider-matrix)
- SLACK_WEBHOOK_URL (new — create an incoming webhook in Slack
  and add as repo secret; notifier prints to stdout otherwise)

Testing:
- Trigger manually via Actions -> "Live Canary" -> "Run workflow"
  with any single lane; canary-report runs after regardless of
  which lanes executed.
- Run locally with --dry-run to preview the Slack payload.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(canary): post_json error handling + robust Haiku JSON extraction

Address gemini-code-assist review on scripts/live-canary/notify_slack.py:

1. `post_json` unreachable error branch: `urllib.request.urlopen`
   raises `urllib.error.HTTPError` for 4xx/5xx before reaching the
   `if resp.status >= 300` check, so the error body was never
   surfaced. Wrap in try/except and read the body from the
   HTTPError instance — that's where Anthropic's "invalid API key"
   / "rate limited" detail lives.

2. Haiku JSON extraction was fragile: `startswith("```")` assumed
   the response had no prose preamble and only handled one fence
   shape. Replace with `re.search(r"\{.*\}", text, re.DOTALL)` so
   we pick the outermost JSON object regardless of any wrapper
   markdown or leading/trailing text. Greedy + DOTALL is correct
   for the single top-level object our schema requires.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(e2e): raise pytest timeout + bump multi-user chat wait to 180s

The CI run on feat/canary-report surfaced that 90s was still not
enough for test_mcp_same_server_multi_user_via_browser on
ubuntu-latest — it timed out at the inner Playwright
wait_for_function deadline with "Timeout 90000ms exceeded" after
118s of total test time.

The test opens two browser contexts + two SSE streams and drives a
full chat turn per user in sequence. Under 2-vCPU contention the
compound pipeline genuinely takes over 90s.

- tests/e2e/pyproject.toml: timeout 120 -> 240 (pytest-level cap)
- test_v2_auth_oauth_matrix.py: send_chat_and_wait_for_terminal_message
  call sites 90000 -> 180000 (two owner/member turns, each budgeted
  for one runner-slow turn)

180s < 240s, so the inner deadline fires first with the useful
Playwright traceback instead of the generic pytest SIGTERM.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(e2e): fix pytest-timeout CLI override + widen Mode-C deadlines

The previous commit (c3c9bbab) raised tests/e2e/pyproject.toml's
timeout from 120 to 240, but the auth canary runs the suite via
scripts/auth_canary/run_canary.py which hardcodes
`--timeout=120` on the pytest command line. The CLI flag wins
over pyproject's ini_options, so the 240 bump was invisible to
the auth lanes. That's why auth-smoke on the canary `all` run
still failed with "Timeout (>120.0s) from pytest-timeout" even
after our 180s inner widening — the outer CLI cap was firing at
120s first.

Fix the override and widen the two remaining Mode-C deadlines
that blew in the same run:

  scripts/auth_canary/run_canary.py: --timeout=120 -> 240
  _wait_for_refresh_request default: 60.0 -> 120.0
    (test_wasm_tool_oauth_refresh_on_demand and
     test_mcp_oauth_refresh_on_demand both use the default)
  test_settings_first_gmail_auth_then_chat_runs call sites:
    _wait_for_mock_google_tokens 90.0 -> 120.0
    _wait_for_response_contains 90.0 -> 120.0

All remain comfortably under the new 240s pytest-level cap so a
real hang still fails fast with a useful traceback.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(e2e): opt-in text-match predicate for multi-user browser test

Ship the structural fix that was overdue. Repeated budget bumps on
send_chat_and_wait_for_terminal_message weren't holding under
ubuntu-latest "all"-mode parallelism — 120s, 180s both exceeded on
test_mcp_same_server_multi_user_via_browser. The underlying race is
in the JS predicate: it waits for the assistant bubble AND the
data-streaming attribute cleared AND the chat input re-enabled.
Under 2-vCPU contention an SSE reconnect can drop the final
attribute-clearing delta, and the compound predicate never flips
even though the response text arrived long ago.

Add an opt-in `expected_text_contains` parameter. When supplied,
the predicate succeeds the moment the expected substring appears in
the new assistant message — regardless of data-streaming or input
state. Callers that already assert on specific response text (the
existing MCP / gmail tests) can now short-circuit the race without
compromising correctness: the test's own content assertions remain
the gate.

Default behavior unchanged for the ~30 existing call sites across
test_chat.py, test_sse_reconnect.py, test_tool_approval.py,
test_portfolio.py, test_message_persistence.py, test_agent_loop_recovery.py,
test_pending_user_messages.py, test_widget_customization.py.

Applied to the two multi-user call sites with
expected_text_contains="Mock MCP search result" — that's exactly
what the test's next two assertions verify.

Local run of the flaky test alone: 40s, green.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* ci(canary): move auth-smoke to self-hosted runner

Multi-user browser test (test_mcp_same_server_multi_user_via_browser)
consistently exceeds the Playwright budget on GH ubuntu-latest under
the 2-vCPU parallelism pressure of an "all" canary run — a single
compound chat turn burns >180s, with each budget bump we apply it
ratchets the flake, not the fix.

Pilot move onto the [self-hosted, ironclaw-live] runner that
private-oauth already uses. Same runner label means no new
infrastructure required; if the self-hosted box has Python 3.12 and
Playwright browsers installed (or can provision them via the existing
setup-python + scripts/live-canary/run.sh's `PLAYWRIGHT_INSTALL=with-deps`
flow), this is a zero-code-change canary fix.

If the pilot works, auth-full is the next candidate. If the runner
queues become a bottleneck, we'd scale to multiple workers under
the same label rather than revert to ubuntu-latest.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* ci(canary): revert auth-smoke to ubuntu-latest + widen budgets to 300s/360s

Railway self-hosted runner ('railway-private-oauth' on a small Docker
container) turned out to be no faster than GH ubuntu-latest for the
multi-user browser flow — both take ~194–196s for
test_mcp_same_server_multi_user_via_browser. The runner container is
evidently provisioned at a similar vCPU allocation, so the move
bought nothing.

Revert to ubuntu-latest (parallel canary shape preserved; avoids
serialising auth lanes behind private-oauth on the single
self-hosted worker) and widen deadlines for the last CI-load hop:

  test_v2_auth_oauth_matrix.py multi-user call sites:
    Playwright wait_for_function 180000 -> 300000 ms
  scripts/auth_canary/run_canary.py:
    --timeout=240 -> 360 (outer pytest cap)
  tests/e2e/pyproject.toml:
    timeout = 240 -> 360

300s inner fits inside the new 360s outer with 60s margin. Local
run of the same test alone completes in ~40s, so we have plenty
of headroom against real hangs still surfacing fast with a
useful traceback.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* disable report

* scripts(auth-canary): add Google storage-state bootstrap helper

The auth-browser-consent lane drives Google's real OAuth consent UI in
Playwright, but Google's risk engine routinely interrupts the flow with
a "Verify it's you" challenge that handle_google_popup cannot solve, so
the test stalls on the password screen.

Bypass: log in once interactively in Playwright Chromium, save cookies
+ localStorage to a storage_state.json, point AUTH_BROWSER_GOOGLE_-
STORAGE_STATE_PATH at it. Subsequent canary runs spawn contexts with
that state preloaded, so the popup arrives at consent with no login or
challenge in the way.

- scripts/auth_live_canary/bootstrap_google_storage_state.py: new
  one-shot interactive helper that writes
  ~/.ironclaw/auth-canary/google_storage_state.json by default
- scripts/auth_live_canary/README.md: document the bypass under
  "Browser-consent Google challenge bypass"

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* canary(auth-browser-consent): fix Google account-picker + chat drift

The auth-browser-consent google case was failing on two distinct
issues, the first masking the second:

1) Account picker. When AUTH_BROWSER_GOOGLE_STORAGE_STATE_PATH is set
   (the recommended path — username/password automation gets blocked
   by Google's risk engine), Google's OAuth popup lands on a "Choose
   an account" picker before the consent screen. handle_google_popup
   only knew how to fill email + password and click Continue/Allow,
   so the popup sat on the picker until complete_provider_auth's
   120s callback wait timed out. Added a picker-detection step that
   tries selectors in order — username text, [data-identifier], and
   a generic "any visible @-bearing text not equal to 'Use another
   account'" XPath — and clicks the first hit, with debug logging
   so future regressions surface in the run output.

2) Tool-name and response-text drift. After the OAuth fix unblocked
   the rest of the probe, browser_chat still failed because:
   - case.expected_tool_name was "gmail", but the gateway records
     the tool call under its WASM module name "gmail_tool"
   - case.expected_text was "Gmail" (case-sensitive), but real LLM
     responses to "check gmail unread" against an empty inbox vary
     ("Your inbox is clear...", "Inbox is empty", etc.) and rarely
     emit literal "Gmail"
   Updated BROWSER_CASES["google"] to expected_tool_name="gmail_tool"
   and expected_text="inbox", and made the browser_chat assertion's
   text comparison case-insensitive so the canary doesn't depend on
   exact wording.

After both fixes the auth-browser-consent google lane runs green:
  ✓ browser_oauth   (popup -> /oauth/callback)
  ✓ browser_chat    (assistant references inbox)
  ✓ responses_api   (real Gmail tool call)

Not addressed here: BROWSER_CASES["github"] likely has the same
expected_tool_name drift ("github" vs probably "github_tool"); needs
verification with real GitHub OAuth creds before changing.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* canary(auth-browser-consent): robust account-picker fallback + browser channel

Two follow-ups discovered during local debugging of the auth-browser-
consent google lane:

1) Account-picker fallback was matching hidden <style> blocks. The XPath
   `//*[contains(text(), '@') ...]` matched any element whose text
   contains `@`, which includes <style> tags carrying CSS at-rules
   (@font-face, @media). Replaced the XPath with role-based locators
   (get_by_role link/button) filtered by an email regex — only
   interactive elements match, no false positives from style blocks.
   Verified locally that the fallback now clicks the right account row
   even when AUTH_BROWSER_GOOGLE_USERNAME is unset.

2) Bootstrap script: Google's anti-automation blocks Playwright's
   default Chromium (Chrome for Testing) at sign-in with "This browser
   or app may not be secure". Added a --browser flag with a default of
   firefox (Marionette is less aggressively fingerprinted than CDP),
   plus chrome (system Google Chrome) and chromium (override) options.
   For accounts where Google blocks even those — typically brand-new
   Gmails or accounts with high risk scores — the fallback path is to
   launch Chrome manually with --remote-debugging-port and connect via
   playwright.chromium.connect_over_cdp; documented in the README.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* canary(auth-live-canary): include observed extension state in timeout error

When `wait_for_extension_state` times out the bare error
"Timed out waiting for extension state: gmail" is unhelpful for
diagnosing CI failures, since CI artifacts don't capture IronClaw's
gateway logs — there's no way to tell whether the extension never
appeared, appeared but never authenticated, or authenticated but
never activated.

Track the last-observed extension on each poll and surface
authenticated/active in the timeout message. After this change a
failed run says e.g.
"Timed out waiting for extension state: gmail (expected
authenticated=True, active=True; last observed: authenticated=False,
active=False)", which immediately separates token-exchange failures
from activation-state-machine bugs.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* canary(auth-live-canary): widen chat-wait deadlines 120s -> 300s

The auth-browser-consent google probe completed OAuth + extension
activation successfully on CI but timed out at the next step
(send_chat_and_wait_for_terminal_message), with the agent stuck on
"Thinking (step 1)" for the full 120s budget. Local runs on the
same code path complete the chat in ~36s, but ubuntu-latest 2-vCPU
runners under cold-start load (gateway restart, mock LLM bootstrap,
WASM tool first-invocation) need substantially more headroom.

300s matches the precedent set by `d8765714 ci(canary): revert
auth-smoke to ubuntu-latest + widen budgets to 300s/360s` for the
auth-smoke lane on the same runner class.

Both call sites widened — the seeded Responses-API probe at line 221
and the browser_oauth probe at line 800.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* canary(common): drain gateway/mock_llm stdout pipes (was deadlocking CI)

scripts/live_canary/common.py spawns the IronClaw gateway and the
mock LLM with stdout=PIPE + stderr=STDOUT, reads one line of mock_llm
output to discover its bound port, then never reads from either pipe
again. On Linux the kernel pipe buffer caps at 64 KiB; once a
sustained chat request fills it with `RUST_LOG=info` output, the
child blocks on its next stdout write and the request handler
freezes mid-response.

That's why every auth-browser-consent CI run got stuck on
"Thinking (step 1)..." for the full chat-wait budget while the same
test passes locally — macOS pipe buffers are larger and the test
completes before the buffer fills.

Fix: spawn a daemon thread per subprocess that drains the pipe to a
log file under the run's output_dir. Two wins:

- Pipes never fill, child never blocks.
- gateway.log and mock_llm.log become CI artifacts, so the next
  failure that doesn't have a clear runner-side error message is
  immediately debuggable from IronClaw's own logs.

Verified locally that the lane still passes after the change and
both log files are produced. Locally each is < 10 KiB; CI runs may
be larger but well under any artifact size limit.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* canary: pin LLM backend via settings API + add LLM_API_KEY (root cause of CI freeze)

The auth-browser-consent google lane has been freezing on CI at
"Thinking (step 1)..." for the full chat-wait budget. Gateway logs
captured by the previous commit's pipe drainer reveal the smoking
gun:

  ERROR Configured LLM backend is not usable.
        backend=openai_compatible reason=missing API key
  WARN  LLM_BACKEND env var is set but DB setting takes priority.
        db_value=nearai env_value=openai_compatible
  WARN  Active LLM backend fell back to NearAI default
        attempted=openai_compatible active=nearai

Two compounding issues:

1. The openai_compatible provider refuses to instantiate without an
   API key, even though the mock LLM ignores the value. Fix: set
   `LLM_API_KEY=mock-api-key` in `build_gateway_env`, matching what
   `tests/e2e/conftest.py` already does for the e2e suite.

2. IronClaw's DB-stored LLM settings take priority over env vars,
   and the freshly-seeded canary DB defaults `llm_backend` to
   `nearai`. So even with a clean env, the agent fell back to NearAI
   and entered an interactive auth flow that hangs indefinitely in
   CI (the "Thinking" never ends). This is the exact trap
   `tests/e2e/CLAUDE.md` documents: "do not rely on env-vs-DB
   precedence … pin the provider explicitly through /api/settings/...".
   Fix: pin `llm_backend`, `openai_compatible_base_url`, and
   `selected_model` via PUT /api/settings/<key> immediately after the
   gateway becomes healthy.

Also revert the BROWSER_CASES["google"] case I touched earlier:
when NearAI was driving it emitted the WASM canonical tool name
(`gmail_tool`), but the mock LLM (now correctly driving) emits the
tool name it knows from its mapping (`gmail`). Restoring the original
`expected_tool_name="gmail"` / `expected_text="gmail"` matches what
the mock LLM actually produces.

Verified locally: all three browser_oauth / browser_chat /
responses_api probes now pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* canary(auth-live-canary): revert chat-wait deadline 300s -> 120s

The 300s widening at 98abeebe was a band-aid attempt to work around
the actual root cause (subprocess pipe deadlock + DB-overrides-env
LLM backend), which were both fixed at f59981d3 and 8733d3c0
respectively. With those fixes the chat completes in ~35s on CI, so
the 300s budget is overkill — revert to the original 120s, which
gives ~3.5x headroom over the observed steady-state and matches the
deadline shape used elsewhere in the e2e suite.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* ci(canary): rename github oauth secrets to dodge GITHUB_ prefix block

GitHub Actions reserves the GITHUB_ prefix for auto-generated repo
secrets (GITHUB_TOKEN, etc.) and rejects user-created secrets that
start with it: "Secret names must not start with GITHUB_". The
existing references to GITHUB_OAUTH_CLIENT_ID and GITHUB_OAUTH_-
CLIENT_SECRET in this workflow couldn't be backed by actual secrets
for that reason — the OAuth-client config was effectively unset for
the github browser-consent case, which is why it was silently
filtered out by configured_browser_cases().

Decouple the secret name from the env var name: store the secrets
under the AUTH_BROWSER_GITHUB_CLIENT_ID / AUTH_BROWSER_GITHUB_CLIENT_-
SECRET names (matching the AUTH_BROWSER_GITHUB_* convention used by
the other github canary fixture vars), and re-export them here under
the GITHUB_OAUTH_CLIENT_ID / _SECRET env names that
auth_registry.py and the WASM github tool expect.

No code changes needed in auth_registry.py / scripts/auth_live_-
canary/ — they continue to read GITHUB_OAUTH_CLIENT_ID/_SECRET from
the environment as before.

Operator action: create the OAuth app on GitHub (Settings →
Developer settings → OAuth Apps → New OAuth App) and store the
resulting credentials at:

  AUTH_BROWSER_GITHUB_CLIENT_ID
  AUTH_BROWSER_GITHUB_CLIENT_SECRET

(not GITHUB_OAUTH_CLIENT_ID / _SECRET, which GitHub will reject).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* canary(auth-browser-consent): drop github case (tool is PAT-only, not OAuth)

CI run 25022303491 surfaced that `Activate /api/extensions/github/-
activate` returns `{success: false, awaiting_token: true,
message: "Create a Personal Access Token..."}` with no `auth_url`,
which the browser-consent probe needs in order to drive the OAuth
popup.

Confirmed via `registry/tools/github.json`:

    "auth_summary": {
        "method": "manual",       <- PAT paste, not OAuth
        "secrets": ["github_token"],
        "setup_url": "https://github.com/settings/tokens"
    }

The github WASM tool's source capabilities JSON does carry an `oauth`
block, but the released v0.2.3 artifact (referenced from the registry)
ships with the manual-auth path. Until a release flips
`auth_summary.method` to "oauth" — and the github extension actually
returns an `auth_url` from /activate — there's nothing for the
browser-consent probe to do.

- Drop the `github` entry from BROWSER_CASES with a comment pointing
  at the criterion for re-adding it.
- Drop the github-specific filter in `configured_browser_cases` since
  the case is gone (no risk of an env-aware code path that quietly
  skips github when secrets are present-but-mismatched).

GitHub coverage is unchanged in SEEDED_CASES, which seeds the PAT
directly via `AUTH_LIVE_GITHUB_TOKEN` and exercises real
`/v1/responses` + browser tool calls — that lane already works.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* canary(auth-browser-consent): tick notion's trust-URL checkbox before Continue

CI run 25023708895 surfaced the notion case timing out at "Timed out
waiting for notion OAuth callback page". The popup screenshot shows
Notion MCP's consent screen with:

- Workspace correctly auto-selected (storage state worked)
- A yellow warning: "I recognize and trust this URL"
- An unchecked checkbox next to that text
- A grayed-out (disabled) Continue button

The button is gated behind the checkbox. handle_notion_popup
clicked the disabled Continue and silently no-op'd, so the
complete_provider_auth loop waited the full 120s for /oauth/callback
that never arrived.

Add a checkbox-detection step before the Continue click:

  popup.get_by_text(re.compile("I recognize and trust this URL", I))
       .first.click(timeout=3000)

Includes debug print statements (matching the auth-canary pattern
established for google's account picker) so future Notion UI
changes are immediately visible in test-output.log.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(e2e): drain ironclaw subprocess pipes in auth-matrix fixture

Same pipe-deadlock fix as scripts/live_canary/common.py f59981d3,
applied to tests/e2e/scenarios/test_v2_auth_oauth_matrix.py's
_start_auth_matrix_server. The auth-matrix fixture spawns ironclaw
with stdout=PIPE + stderr=PIPE and never drains them, so under
sustained log volume the kernel pipe buffer fills, ironclaw blocks
on its next stdout write, and any test that relies on subsequent
gateway responses (auth gate emission, SSE events, chat replies)
hangs until pytest-timeout fires.

This fix doesn't make the auth-full lane's failing test pass — the
real bug is engine-v2 silently dropping `auth_required` SSE events
for unauthenticated extensions (introduced by #2868). But it makes
the failure mode debuggable: gateway log is captured to
/tmp/ironclaw-auth-matrix-gateway.log (overridable via
IRONCLAW_AUTH_MATRIX_LOG env), and RUST_LOG passes through from the
test runner so we can crank up verbosity without rebuilding.

Without this change, the failing test's log was empty after the
extension-install line; with this change you see the engine-v2
trace summary that surfaces the actual NotCallable-without-auth-gate
bug. That diagnostic visibility is the value here.

- _drain_stream_to_file: asyncio drainer mirroring common.py's sync
  threading version
- _start_auth_matrix_server: drain stdout/stderr to log_path
- _shutdown_auth_matrix_server: cancel drain_tasks for clean exit
- env: RUST_LOG forwarding so debug runs work

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* canary(workflow): add Telegram Bot API mock

Foundation piece for the new workflow-canary lane that exercises
multi-tool / multi-channel user workflows from issue #1044 (Telegram +
routines + Sheets/Calendar/Gmail end-to-end). Models the same
single-port aiohttp-based mock pattern used by tests/e2e/mock_llm.py.

Endpoints:
- /bot{token}/{getMe,getUpdates,sendMessage,sendChatAction,
  setWebhook,deleteWebhook,getFile} — the subset IronClaw's WASM
  telegram tool + channels-src/telegram actually call. Tokens are
  accepted without validation; the canary doesn't need to test
  Telegram's auth — just IronClaw's flow against a Bot API shape.
- /__mock/inject_message — push a simulated incoming user message
  onto the next getUpdates response, so scenarios can drive a
  Telegram → IronClaw round-trip without a real Telegram account.
- /__mock/sent_messages — drain the queue of every sendMessage /
  sendChatAction IronClaw emitted, for end-to-end assertions.
- /__mock/reset — clear all state between probes.

IronClaw routes its API calls through this mock via
IRONCLAW_TEST_HTTP_REMAP=api.telegram.org=<mock_url>, the same
mechanism the auth-live-canary uses for Gmail/Calendar/Sheets mocks.

Smoke-tested: getMe → success, inject_message → getUpdates returns
the injected message, sendMessage → bot response shape + recorded
in sent_messages.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* canary(workflow): land workflow-canary lane with periodic-reminder scenario

Phase 1A of the workflow-canary system from issue #1044. Adds a new
canary lane that exercises the routine engine + cron-fire path, the
foundation that the remaining four scripts (Telegram → Sheets,
Calendar prep, HN monitor, CRM tracker) will layer on.

Components:

- scripts/workflow_canary/routines.py — direct libSQL helpers for
  inserting a lightweight cron routine with a backdated next_fire_at
  and polling routine_runs for terminal status (ok / attention /
  failed). Backdating beats wall-clock cron in tests by 30+ s per
  probe and is the same shape auth-live-seeded uses for
  expire_secret_in_db.
- scripts/workflow_canary/run_workflow_canary.py — entrypoint that
  starts the Telegram mock, calls common.start_gateway_stack with
  workflow-tuned env (ROUTINES_ENABLED=true, ROUTINES_CRON_INTERVAL=2,
  IRONCLAW_TEST_HTTP_REMAP=api.telegram.org=<mock>), and runs
  scenario modules. CLI mirrors run_live_canary.py.
- scripts/workflow_canary/scenarios/periodic_reminder.py — Script 4
  Phase 1A: insert lightweight routine → wait for engine to fire →
  assert run row reaches a terminal status. Verified locally: 1
  probe, 1 fire, status=attention.

Plumbing:

- .github/workflows/live-canary.yml — new workflow-canary job + lane
  added to the workflow_dispatch choice list and the canary-report
  aggregator's needs:.
- scripts/live-canary/run.sh — workflow-canary case dispatches to
  run_workflow_canary.py.

Phase 1B follow-ups in subsequent commits:
- Telegram channel install + bot-token seeding (needs admin auth or
  direct encrypted-secrets DB write)
- Verify Telegram sendMessage was emitted to the mock during the
  routine fire (covered by mock telegram's /__mock/sent_messages)
- Scripts 1, 3, 5 (Sheets / HN / Gmail-CRM)
- Script 2 (Calendar prep with web search)

Local verification:
  $ tests/e2e/.venv/bin/python scripts/workflow_canary/run_workflow_canary.py \
        --skip-build --skip-python-bootstrap
  [workflow-canary] mock telegram listening at http://127.0.0.1:51139
  [periodic_reminder] inserted routine ..., next_fire_at backdated 60s
  [periodic_reminder] routine fired: status=attention
  [workflow-canary] all 1 probe(s) passed.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* canary(workflow): land all 5 issue #1044 scenarios + scenario README

Layer Scripts 1, 2, 3, 5 onto the foundation shipped in 16278ea9, so
the workflow-canary lane covers all five user-workflow scripts from
issue #1044. Each scenario delegates to a shared
`run_routine_probe()` helper that captures the Phase 1A shape: insert
a Lightweight cron routine with a script-specific prompt → backdate
next_fire_at → poll routine_runs for terminal status.

Scenarios added:

- bug_logger.py     (Script 1 — Telegram bugs → Google Sheet)
- calendar_prep.py  (Script 2 — Calendar prep → Telegram, Reporter: Nick)
- hn_monitor.py     (Script 3 — Hacker News → Telegram, Reporter: Emil)
- crm_tracker.py    (Script 5 — Gmail → Sheets CRM, Reporter: Cameron)

Plus periodic_reminder.py (Script 4, Reporter: Henry) refactored to
also use run_routine_probe.

scenarios/_common.py centralizes the routine plumbing — each scenario
file is now ~30 lines of routine-name + prompt + Phase 1B follow-up
notes. The Phase 1B follow-up plan (Telegram channel install, mock
Sheets writes, mock Calendar reads, mock HN scrape, LLM email
classification, dedup verification) is documented inline in each
scenario's docstring AND in the new scripts/workflow_canary/README.md.

Local verification: all 5 probes green in ~2 s each.

  $ tests/e2e/.venv/bin/python scripts/workflow_canary/run_workflow_canary.py \
        --skip-build --skip-python-bootstrap
  [workflow-canary] === Script 1 — Telegram → Google Sheet Bug Logger ===
  [workflow-canary] === Script 2 — Calendar Prep Assistant ===
  [workflow-canary] === Script 3 — Hacker News Keyword Monitor ===
  [workflow-canary] === Script 4 — Periodic Reminder via Telegram ===
  [workflow-canary] === Script 5 — Email → CRM Inbound Tracker ===
  [workflow-canary] all 5 probe(s) passed.

What this catches:
- Routine engine cron-tick path (spawn_cron_ticker → check_cron_triggers)
- RoutineAction::Lightweight execution
- DB serialization of action_config / trigger_config
- Mock-LLM round-trip latency under cron scheduling
- routines.next_fire_at → routine_runs status state machine

What it doesn't catch yet (per-scenario Phase 1B work, documented in
README + scenario docstrings):
- Telegram channel install + sendMessage assertion
- Mock Sheets / Calendar / Gmail / HN write+read semantics
- LLM-driven structured classification (CRM)
- Cross-fire dedup verification

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* canary(workflow): scaffold Phase 1B telegram-side-effect verification

Lays the groundwork for verifying mock-Telegram side effects from
each scenario's routine fire — but gates the verification off until
a separate engine bug is fixed.

What's added:

- tests/e2e/mock_llm.py: new TOOL_CALL_PATTERNS entry that matches
  ``[CANARY-WORKFLOW-<key>]`` in any prompt and emits a deterministic
  http tool call to api.telegram.org/.../sendMessage with a
  per-scenario ack text.
- scripts/workflow_canary/scenarios/_common.py: each scenario now
  composes its prompt as
  ``<prompt_intro>\n\n[CANARY-WORKFLOW-<key>]`` so the matcher fires.
  When ``verify_telegram=True``, the helper polls
  /__mock/sent_messages for up to 5 s and asserts the expected ack
  was captured. Default is ``verify_telegram=False`` (Phase 1A
  parity) — see below.
- scripts/workflow_canary/telegram_mock.py: aiohttp request-logger
  middleware so the canary's stdout shows every inbound request,
  giving operators a one-line answer to "did the gateway's HTTP
  remap actually reach the mock?".
- scripts/workflow_canary/scenarios/{bug_logger,calendar_prep,
  hn_monitor,periodic_reminder,crm_tracker}.py: scenarios pass
  ``mock_telegram_url=mock_telegram_url`` and ``prompt_intro=...``
  ready for verify_telegram to flip on.

What's gated off and why:

The mock-Telegram verification path requires
``IRONCLAW_TEST_HTTP_REMAP=api.telegram.org=<mock>`` to route
the http tool's sendMessage call into the mock. The remap is
correctly registered at gateway startup
(src/app.rs::http_interceptor + src/http_intercept.rs), but the
ToolContext built inside the routine engine's Lightweight action
loop does NOT inherit the global ``http_interceptor`` slot. Result:
the http tool reaches into the real network for api.telegram.org
(returning a 401 since the bot token is fake) and the mock never
sees the request — confirmed via the new request-logger middleware
showing zero non-internal hits.

That's a real engine bug in routine-driven tool dispatch — the
http_interceptor needs to propagate through the routine action's
ToolContext just like it does for chat-driven tool dispatch. Out of
scope for this canary PR; tracked as a follow-up. Once fixed, flip
the default in ``run_routine_probe`` and every scenario's
verify_telegram check activates with no further changes.

Local verification: all 5 probes still green at the Phase 1A level.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* canary(workflow): re-exec under venv after bootstrap (fix CI 'No module named httpx')

CI run 25028445222 failed on the workflow-canary lane with:

  [workflow-canary] mock telegram listening at http://...
  [workflow-canary] error: No module named 'httpx'

Root cause: run_workflow_canary.py was missing the bootstrap-then-
reexec pattern that scripts/auth_live_canary/run_live_canary.py
uses (line 1229+). bootstrap_python() creates the venv and installs
tests/e2e/'s pyproject deps (which include httpx + aiohttp), but
the parent process keeps executing under whatever interpreter
invoked it — typically the system Python on CI runners, which
doesn't have httpx. The scenario module's `import httpx` at top
level then fails immediately.

Fix: copy the auth-live-canary reexec pattern. main() now:

1. If not --skip-python-bootstrap AND WORKFLOW_CANARY_REEXEC is
   unset: bootstrap the venv, install playwright, build cargo,
   then subprocess-spawn ourselves under the venv python with
   --skip-python-bootstrap and WORKFLOW_CANARY_REEXEC=1 so this
   branch isn't re-entered.
2. The reexecuted process sees skip_python_bootstrap=True and runs
   the actual canary against the venv interpreter that has all
   deps available.

Local sanity check: still passes (--skip-build --skip-python-bootstrap
short-circuits the bootstrap, both branches behave identically when
the venv already exists).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(routine-engine): propagate http_interceptor into Lightweight tool dispatch

The chat path's tool dispatch correctly receives the global
HTTP interceptor (e.g., the `IRONCLAW_TEST_HTTP_REMAP` debug-only
host remapper installed in `src/app.rs::http_interceptor`), but the
routine engine's Lightweight action path constructed its
`JobContext` from scratch with `..Default::default()`, leaving
`http_interceptor: None`. Tools called from a routine therefore
reached the real network even when the rest of the system was
configured to route through mocks.

Plumb the interceptor through:

- `RoutineEngine` gains an `http_interceptor` field
- `RoutineEngine::new` takes it as the 11th argument
- `EngineContext` carries it across the spawn boundary
- `JobContext` construction at the Lightweight action site copies
  it from the engine context

Threading complete: AgentDeps → RoutineEngine → EngineContext →
JobContext → http tool. Same shape the chat path already uses.

Test rigs updated: `tests/support/test_rig.rs` and
`tests/e2e_routine_heartbeat.rs` (10 call sites total) pass `None`
for the new arg, matching their existing minimal stack model.
Build clean against `--no-default-features --features libsql`.

Why this matters: with the interceptor lost, every workflow-canary
probe's http tool dispatch reached real api.telegram.org and 401'd
on the fake token — leaving the mock Telegram bot empty and the
canary's send-side assertions unverifiable. With the fix, the
interceptor honors the IRONCLAW_TEST_HTTP_REMAP and the workflow
canary's Phase 1B verification activates immediately.

Activates in this commit:

- scripts/workflow_canary/scenarios/_common.py default flips to
  `verify_telegram=True`
- All 5 scenarios (bug_logger, calendar_prep, hn_monitor,
  periodic_reminder, crm_tracker) now assert that the mock
  Telegram bot received the per-scenario ack message
  `[canary-workflow:<key>] ack`

Local verification:

  $ tests/e2e/.venv/bin/python scripts/workflow_canary/run_workflow_canary.py \
        --skip-build --skip-python-bootstrap
  [workflow-canary] === Script 1 — Telegram → Google Sheet Bug Logger ===
  ... (all 5 scenarios) ...
  [workflow-canary] all 5 probe(s) passed.

  $ grep "POST /bot" artifacts/workflow-canary/telegram_mock.log | wc -l
  5  # one per scenario, distinct ack text per probe

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* canary(workflow): add manual_trigger + lifecycle + dedup_cooldown probes

Three new scenarios covering issue #1044 assertions that the existing
5 cron-fire probes don't reach. Each scenario tests a distinct
back-end mechanism that real users hit:

- **manual_trigger** (Scripts 3 PHASE 2.1 + 3 PHASE 4.2 + 4 PHASE 4.2)
  Inserts a routine WITHOUT backdating next_fire_at, so the only path
  to a fire is the manual-trigger API. POSTs
  /api/routines/<id>/trigger, asserts response carries a run_id, polls
  routine_runs for terminal status, then verifies mock Telegram
  captured the per-scenario ack. Catches regressions in
  RoutineEngine::fire_manual end-to-end.

- **lifecycle** (Scripts 1 PHASE 5 + 4 PHASE 5) — three sub-probes:
  1. disabled-blocks-fires: insert with enabled=False + backdate;
     assert no routine_runs row appears within 8 s window.
  2. enable-resumes-fires: toggle enabled=true via API, backdate,
     assert fire reaches terminal status.
  3. delete-removes-routine: confirm /api/routines lists it, DELETE,
     confirm it's gone.
  Catches regressions in toggle handler, delete handler, and the
  engine's enabled-flag respect during cron tick selection.

- **dedup_cooldown** (Scripts 1 PHASE 4.4 + 3 PHASE 3.2 + 5 PHASE 5.5)
  Insert with cooldown_secs=30; first fire lands within ~5 s; immediate
  re-backdate; assert ONLY ONE run row exists after 8 s. Catches
  regressions in cooldown enforcement during check_cron_triggers.
  This is the closest engine-level correlate to the user-script
  "no duplicate rows / alerts / messages" assertions, which are
  application-level dedup that lives outside the canary's
  deterministic-mock surface.

Plumbing:

- routines.py: trigger_routine_via_api / toggle_routine_via_api /
  delete_routine_via_api / list_routines_via_api helpers (all auth-
  bearer, JSON in/out, raise_for_status).
- routines.py: insert_lightweight_cron_routine grew `cooldown_secs`
  + `enabled` parameters; defaults preserve existing behavior.
- run_workflow_canary.py: registered the three new scenario keys.

Local verification — all 10 probes (5 original + 5 new sub-probes
across 3 new scenarios) green:

  ✅ bug_logger / calendar_prep / hn_monitor / periodic_reminder /
     crm_tracker          (existing — Telegram ack capture)
  ✅ manual_trigger        (548ms)
  ✅ lifecycle_disable     (8004ms — full no-fire window)
  ✅ lifecycle_toggle      (1543ms)
  ✅ lifecycle_delete      (56ms)
  ✅ dedup_cooldown        (10017ms — first fire + 8s no-fire window)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* canary(workflow): add NL-driven routine_create + routine_update probes

Two scenarios that close issue #1044's chat-driven assertions
(Script 1 PHASE 3.1, Script 2 PHASE 3.1, Script 3 PHASE 2.1,
Script 4 PHASE 2.1 + 5.1, Script 5 PHASE 4.1):

- **nl_routine_create**: opens a thread via /api/chat/thread/new,
  posts an NL message tagged [CANARY-WORKFLOW-NL-CREATE], waits for
  the agent to dispatch routine_create, then verifies the routines
  row landed in libSQL AND is visible via GET /api/routines.

- **nl_schedule_update**: pre-seeds a target routine
  (canary-nl-update-target), posts an NL message tagged
  [CANARY-WORKFLOW-NL-UPDATE], waits for the agent to dispatch
  routine_update with a new schedule, then verifies trigger_config
  changed in libSQL. Asserts on schedule-changed (not exact match)
  because the engine normalizes 5-field cron → 7-field internal
  form ("0 */5 * * *" → "0 0 */5 * * * *").

Plumbing:

- Two new TOOL_CALL_PATTERNS entries in tests/e2e/mock_llm.py
  matched in priority order (specific NL-CREATE / NL-UPDATE
  sentinels checked BEFORE the generic [CANARY-WORKFLOW-<key>]
  http-tool fallback, since the canary's own routines emit the
  generic pattern from inside their action prompts).

- Helper additions in scripts/workflow_canary/routines.py:
  _open_thread / _send_chat / _read_routine / _wait_for_*.

Local verification — all 12 probes green:

  ✅ bug_logger / calendar_prep / hn_monitor / periodic_reminder /
     crm_tracker        (5 cron-fire + telegram-ack)
  ✅ manual_trigger      (POST /api/routines/<id>/trigger)
  ✅ lifecycle_disable / lifecycle_toggle / lifecycle_delete
  ✅ dedup_cooldown      (cooldown_secs suppresses second fire)
  ✅ nl_routine_create   (chat → routine_create tool)
  ✅ nl_schedule_update  (chat → routine_update tool)

What's still deferred to follow-up PRs (per-provider mocks, each
~1-3 days of work — see scripts/workflow_canary/README.md):

- Mock Google Sheets (Scripts 1 + 5 dedicated assertions)
- Mock Google Calendar (Script 2)
- Mock Hacker News (Script 3)
- LLM-driven email classification with seeded inbox (Script 5)
- Telegram channel install + bot-token validation flow (Scripts 1-5
  PHASE 1)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(workflow-canary): Phase 1 — mock Sheets + bug_logger Sheet-write probe

Adds scripts/workflow_canary/sheets_mock.py: single-port aiohttp Google
Sheets v4 mock supporting POST /v4/spreadsheets, values:append, values
get, plus /__mock/ test hooks for seeding, draining, and resetting.
The append handler enforces values=list-of-lists (returns the canonical
"expected a sequence" 400) so the canary catches the issue #1044 FAIL
CRITERIA shape.

Wires the mock into run_workflow_canary.py:
  - generic _spawn_mock helper for telegram_mock + sheets_mock
  - IRONCLAW_TEST_HTTP_REMAP carries comma-separated entries for
    api.telegram.org and sheets.googleapis.com
  - mock_sheets_url passed through to every scenario's run() kwargs

Rewrites scenarios/bug_logger.py to drop the run_routine_probe Telegram
fallback in favor of a Sheet-write end-to-end assertion: pre-seed the
spreadsheet, fire the routine with [CANARY-WORKFLOW-SHEET-APPEND], wait
for the appended row, validate shape (timestamp / message / source).

Mock LLM: new TOOL_CALL_PATTERNS entry that matches the SHEET-APPEND
sentinel and emits an http POST values:append with a hardcoded canary
row.

All 12 probes still pass locally.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(workflow-canary): Phase 2-4 — Calendar / HN / Gmail / web_search mocks + e2e probes

Phase 2 (Calendar): scripts/workflow_canary/calendar_mock.py — Google
Calendar v3 events surface (list / insert / get / delete) with seed
hooks. calendar_prep_e2e seeds one canary event, fires the routine,
asserts events.list was hit and Telegram received the prep briefing
referencing the seeded event title.

Phase 3 (Hacker News): scripts/workflow_canary/hn_mock.py — /newest
HTML fixture with seeded "Show HN" posts (canary-distinct
``<!-- canary-hn-feed -->`` marker). hn_monitor_e2e re-seeds posts,
asserts /newest GET landed and Telegram summary references both
seeded posts.

Phase 4 (CRM tracker): scripts/workflow_canary/gmail_mock.py +
web_search_mock.py — Gmail v1 messages.list/.get + Brave Search v3.
crm_tracker_e2e seeds 1 lead + 1 newsletter + 1 receipt; asserts
exactly ONE row appended to the CRM sheet (only the lead) with all
6 expected columns + Telegram ack referencing 1 lead.

Mock LLM TOOL_CALL_PATTERNS gain three parallel-call entries
([CANARY-WORKFLOW-CAL-LIST] → http GET events.list + http POST
sendMessage; [CANARY-WORKFLOW-HN-FETCH] → GET /newest + sendMessage;
[CANARY-WORKFLOW-CRM-CLASSIFY] → Gmail GET + Sheets append + Telegram
ack). Parallel emit is required because the engine's lightweight
loop dedups same-tool re-dispatch (see match_tool_call:1178).

run_workflow_canary.py now spawns six mock subprocesses; remap covers
api.telegram.org, sheets.googleapis.com, www.googleapis.com,
news.ycombinator.com, gmail.googleapis.com, api.search.brave.com.

All 12 existing probes pass + 3 phase 2-4 probes upgrade from
side-effect-only to full content-correctness assertions.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(workflow-canary): Phase 5 — Telegram channel install + round-trip

scripts/workflow_canary/telegram_setup.py: install + capability patch
+ setup helpers (mirrors tests/e2e/scenarios/test_telegram_e2e.py
patch_capabilities + activate flow). Adds pair_telegram_user that
sends an "hello" webhook, extracts the pairing code from
mock_telegram, and approves it via /api/pairing/telegram/approve.

scripts/live_canary/common.py: GatewayStack now exposes http_url
(HTTP-channel webhook port) + channels_dir (WASM_CHANNELS_DIR)
so workflow-canary scenarios can drive the Telegram channel install
+ patch + webhook flow.

run_workflow_canary.py: passes IRONCLAW_TEST_TELEGRAM_API_BASE_URL
so the hardcoded validate_telegram_bot_token getMe call (in
src/extensions/manager.rs) routes to mock_telegram. The bot-token
validate path bypasses the standard IRONCLAW_TEST_HTTP_REMAP flow,
hence the additional env override.

New scenarios:
- telegram_channel_install: install + patch caps + setup + assert
  channel reaches Active state. Catches "HTTP 404 on valid token"
  regression (Script 4 PHASE 1.1).
- telegram_round_trip: post inbound webhook → assert mock_telegram
  receives an outbound sendMessage with the actual chat_id (NOT
  'default'). Catches the chat_id 'default' regression.
- routine_visibility_from_telegram: pair user, ask for routines,
  assert agent replies on the paired chat_id. Covers Scripts 1-4
  PHASE "routine visibility from Telegram" assertions.
- manual_trigger_from_telegram: pair user, hit /api/routines/<id>/
  trigger, assert routine fires through lightweight loop and ack
  reaches the paired chat_id. Covers Script 4 PHASE 4.2.

All 16 probes pass locally.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(workflow-canary): Phase 6 — first_immediate_run + log_assertions

scripts/workflow_canary/scenarios/first_immediate_run.py: insert a
routine with a "0 * * * *" schedule + fire_immediately=True; assert
the first run reaches terminal status within 10s. Catches "first
check is delayed to next hour" regression (Script 3 PHASE 2.1).

scripts/workflow_canary/scenarios/log_assertions.py: scan
gateway.log at the end of the lane for known fail-criterion regex
patterns: chat_id 'default', parsed naive timestamp without timezone,
retry after None, expected a sequence. Catches log regressions across
all 5 issue #1044 scripts simultaneously.

Auth-recovery (token revocation → auth_required SSE) is deferred to
the auth-live-canary lane; it requires a working OAuth setup to
revoke, which is outside this lane's mock-only scope.

All 18 probes pass locally.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(workflow-canary): Phase 7 — cron timing + idempotent toggle + README

scripts/workflow_canary/scenarios/cron_timing_accuracy.py: insert a
routine, set next_fire_at to "now + 5s" explicitly, assert the engine
fires within ±10s of the set boundary. Catches "cron skipped a cycle"
+ "fires never trigger" regressions (Scripts 3 PHASE 3.1, 4 PHASE 3.4).

scripts/workflow_canary/scenarios/idempotent_disable_enable.py:
double-toggle disable then double-toggle enable, assert both halves
are no-ops; finally backdate, fire once, then disable + backdate again
and assert no NEW runs land in the next 6s. Catches "disable doesn't
take effect" + "enable triggers a phantom run" regressions
(Script 1 PHASE 5.1 / 5.2).

scripts/workflow_canary/README.md: rewritten to reflect 20-probe
coverage matrix across phases 1–7 with mock surface + scenarios
inventory.

Final canary state: 20 probes across 7 phases, all green locally.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(workflow-canary): close gaps — wire web_search + add auth_recovery

[CANARY-WORKFLOW-CAL-LIST] now emits a parallel triplet (calendar
events.list + web_search company lookup + telegram sendMessage).
calendar_prep asserts mock_web_search captured the lookup with the
expected company-name query parameter, completing the Script 2
"company background + recent news" assertion from issue #1044.

scripts/workflow_canary/scenarios/auth_recovery.py: drives a chat
that triggers an unauthenticated gmail tool call, asserts the agent
surfaces a graceful response — chat send returns 202 (not 5xx),
thread settles, history contains no Error 400 / Internal Server
Error / panicked / Traceback fragments. Catches the regression
shape from Script 2 PHASE 5 fail criteria without requiring a real
OAuth handshake (full token-revocation coverage stays in
auth-live-canary).

21 probes total, all green locally.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* ci(canary): run every 6h + re-enable Slack report

Schedule: cron flips from "0 2 * * *" (once daily at 02:00 UTC) to
"0 */6 * * *" (4× daily at 00/06/12/18 UTC). All twelve job-level
`if:` guards updated in lockstep so each lane still gates on the
schedule string.

Slack report: drop the `if: false` hardcode on the canary-report
job's notify step and replace with a schedule + workflow_dispatch
gate. The notifier (scripts/live-canary/notify_slack.py) already
exits 0 on Haiku/Slack failures so a flaky webhook can't mask lane
status. PR-triggered runs (currently none, but possible via
workflow_run) skip the post to keep noise out of the channel.

Both ANTHROPIC_API_KEY and SLACK_WEBHOOK_URL repo secrets are
already populated.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(canary-report): parse workflow-canary results.json shape

The notifier reads `auth-canary-junit.xml` for JUnit-emitting lanes
(auth-smoke, auth-full, auth-channels, auth-live-seeded,
auth-browser-consent). The workflow-canary lane writes its own
`results.json` instead — one entry per probe with `success: bool`,
`latency_ms`, `details`. The notifier had no parser for that shape, so
the workflow-canary slot in Slack rendered as a useless
`:grey_question: 0/0 passed, 0 failed` line.

Add `parse_results_json` mirroring the JUnit parser's contract:
`passed = sum(success)`, `failed = sum(!success)`, each failed probe
becomes a `(provider/mode, error-or-summary)` entry on
`junit_failures` so the Slack reason field renders the same way as an
auth-canary failure. Latencies sum to `duration_s`. Both parsers run
on every lane dir; first one whose file exists wins (auth-canary lanes
emit XML only, workflow-canary lane emits JSON only — no overlap).

Validated by re-running the notifier locally against the downloaded
artifact from CI run 25033224036:
  before: ":grey_question: workflow-canary (mock) — 0/0 passed"
  after:  ":white_check_mark: workflow-canary (mock) — 21/21 passed,
           0 failed in 69s"

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(canary-report): log notifier progress for diagnosability

Until now `notify_slack.py` was silent on the success path, which made
it impossible to verify from CI logs alone whether Haiku enrichment
actually ran. Add four stderr lines covering each phase:

  [notify_slack] discovered N lane dir(s): lane1/provider1, ...
  [notify_slack]   lane/provider: tests=N passed=N failed=N skipped=N status=...
  [notify_slack] haiku enriched X/N lane(s)
  [notify_slack] posted Slack message for N lane(s)

Lines stay terse and structured so they're greppable from `gh run
view --log`. Haiku-failure tracking inspects `r.notable` — `run_haiku`
stamps it with `haiku call failed:` / `haiku returned no JSON object`
/ `haiku JSON parse failed` on the three failure paths.

Confirmed from local dry-run against the artifact downloaded from
the previous CI run (which had the results.json parser): tests=21,
passed=21, failed=0, status=pass.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(canary/workflow-canary): forward SCENARIO into --scenario

Addresses @henrypark133's review on PR #2874: the workflow-canary lane
of `scripts/live-canary/run.sh` ignored `${SCENARIO}` and always ran
the full 21-probe suite. The matching workflow_dispatch job didn't
export `inputs.scenario` either, so manual dispatch with a scenario
filter went nowhere. Targeted local reruns / debugging hit the same
gap.

run.sh: translate `${SCENARIO}` (comma-list supported) into one or
more `--scenario <name>` flags on `run_workflow_canary.py`. Empty
SCENARIO falls through to the full suite. Guards the array splat for
bash 3.2 / macOS where `${arr[@]}` on an empty array under `set -u`
explodes.

live-canary.yml: add `SCENARIO: ${{ inputs.scenario }}` to the
Workflow Canary job's env so workflow_dispatch reaches run.sh.

Verified:
  tests/e2e/.venv/bin/python \
    scripts/workflow_canary/run_workflow_canary.py \
    --skip-build --skip-python-bootstrap \
    --scenario telegram_round_trip
  → "all 1 probe(s) passed"
  (full suite without the flag still runs all 21 probes)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(canary/workflow-canary): align nl_schedule_update on 'every 6 hours'

Addresses Copilot AI's review on PR #2874: the docstring claimed
"every 5 minutes" while EXPECTED_NEW_SCHEDULE / mock LLM emitted
"0 */5 * * *" (every 5 hours), and the chat prompt the canary sent
said "every 5 hours". Three different cadences across one probe.

Pick "every 6 hours" consistently:
- Docstring narrative: "every 6 hours"
- Constant: EXPECTED_NEW_SCHEDULE = "0 */6 * * *"
- Chat prompt: "fire every 6 hours"
- mock_llm.py routine_update args: schedule = "0 */6 * * *"

Verified locally: nl_schedule_update probe still green.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(canary/auth-browser-consent): drop stale GitHub secret exposure

Addresses @henrypark133's review on PR #2874: the auth-browser-consent
job kept exporting 8 GitHub-related secrets (GITHUB_OAUTH_CLIENT_ID,
GITHUB_OAUTH_CLIENT_SECRET, AUTH_BROWSER_GITHUB_OWNER / _REPO /
_ISSUE_NUMBER / _USERNAME / _PASSWORD / _STORAGE_STATE_B64) even
though the lane no longer drives a GitHub OAuth flow. BROWSER_CASES
in `scripts/live_canary/auth_registry.py` was reduced to {google,
notion} when github was reclassified as PAT-only — those secrets are
unused on every scheduled run and just broaden the secret-exposure
surface.

Strip all 8 from the lane:
- env: block — 5 lines (CLIENT_ID + 4 AUTH_BROWSER_GITHUB_* helpers)
- Materialize provider storage state — 1 secret + its materialize block
- Materialize sensitive secrets — 2 secrets + their write_secret lines

Replace with explanatory comments pointing at BROWSER_CASES /
auth_registry.py so a future contributor doesn't re-add them by reflex
when github gets an OAuth flow.

Github coverage continues to live in SEEDED_CASES (auth-live-seeded
lane) which seeds the PAT directly — that lane's secrets are
unaffected.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(canary): align user-facing browser-cases list with auth_registry

Addresses @henrypark133's review on PR #2874: removing `github` from
BROWSER_CASES made `--mode browser --case github` invalid, but the
contract was still advertised in three places that operators read
when copying invocations:

- run_live_canary.py --help (`For browser mode: google, github, notion`)
- scripts/auth_live_canary/README.md (`github` listed under "Runs
  through Responses API and browser")
- scripts/live-canary/README.md (`CASES=google,github` example)
- scripts/live-canary/ACCOUNTS.md (full GitHub OAuth client + fixture
  + storage-state-secret sections still active, plus a Playwright
  storage-state recipe pointing at github.com/login)

Update each in lockstep:

- --help now says `For browser mode: google, notion. (github browser
  coverage is intentionally absent — the github WASM tool is PAT-only,
  not OAuth; see SEEDED_CASES instead.)`
- auth_live_canary/README — github entry now reads "Responses API
  only (PAT-only — not browser-OAuth)"; notion entry corrected to
  "Responses API and browser" (it was inaccurately listed as
  Responses API only).
- live-canary/README — example flips to `CASES=google,notion` with a
  one-line note pointing at auth_registry.py.
- live-canary/ACCOUNTS — drops the GitHub OAuth client + fixture
  sections, swaps the Playwright storage-state recipe target from
  github.com/login to accounts.google.com, drops
  AUTH_BROWSER_GITHUB_STORAGE_STATE_B64 from the CI-secrets list.

The argparse validator in run_live_canary.py already gives a clean
error if anyone passes `--mode browser --case github`:
"--case values ['github'] are not valid for --mode browser. Allowed:
['google', 'notion']", so the docs change is the user-facing fix.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(canary/telegram): split is_active into installed vs. active

Addresses Copilot AI's review on PR #2874: `is_telegram_active` only
checked that an extension named "telegram" appeared in
`/api/extensions`, returning True for an installed-but-inactive
extension (mid-setup, awaiting auth, activation_error). Two callers
(`telegram_round_trip._ensure_active`,
`routine_visibility_from_telegram._ensure_active_and_paired`) used
this as a precheck to skip `setup_telegram_channel()`, so a stale
inactive entry would short-circuit setup and the probe would then
fail mysteriously when the channel didn't respond.

Split into two helpers:

- `is_telegram_installed(...)` — original semantics (entry exists),
  used internally as a building block; not exported as a precheck.
- `wait_for_telegram_active(...)` — polls until the entry has
  `active=true` (the actual runtime-readiness signal — channel
  opened, hooks registered, credentials bound, per
  `.claude/rules/lifecycle.md`'s discovery-vs-activation rule).

Shared `_find_telegram` helper handles the three historical envelope
shapes the gateway has used (`extensions` / `items` / `installed`).

Update all 4 callers to use `wait_for_telegram_active`:
- telegram_channel_install.py
- telegram_round_trip.py (precheck + post-setup wait)
- routine_visibility_from_telegram.py (precheck + post-setup wait)
- manual_trigger_from_telegram.py (precheck + post-setup wait)

Verified: all 4 telegram probes still green back-to-back.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs(canary/periodic_reminder): align docstring with current behavior

Addresses Copilot AI's review on PR #2874: the module docstring still
described the Telegram delivery assertion as a "Phase 1B follow-up"
even though the scenario now sets verify_telegram=True and the
inline comment on the call site already explained the Phase 1B work
had landed. Future readers would assume Telegram verification was
missing from this probe.

Replace the docstring with a 5-step description of what the probe
actually does end-to-end:
1. Backdated cron routine inserted via libSQL
2. Routine engine cron-tick picks it up
3. Lightweight action runs against mock LLM → http sendMessage
4. IRONCLAW_TEST_HTTP_REMAP routes to telegram_mock
5. Asserts both terminal routine_runs status AND captured sendMessage

Also adds an explicit note that channel-install coverage (capability
patch + setup + pairing) lives in the sibling telegram_* scenarios —
this one covers the routine-driven sendMessage path and intentionally
hits api.telegram.org via the raw http tool rather than through the
installed channel.

Verified: probe still green.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(canary-report): rich failure blocks + cross-lane categorization + GH issues

Three additions to scripts/live-canary/notify_slack.py to make the
6h Slack report actionable instead of just informational:

1) **Per-lane rich failure block** — Haiku now extracts four
   structured fields when status==fail: test_name, error, root_cause,
   fix. The Slack section renders them in the issue-friendly shape
   the reviewer asked for:

       :x: auth-full (mock) — 11/13 passed, 1 failed in 213s
         Test: `test_wasm_tool_first_chat_auth_attempt_emits_auth_url`
         Error: SSE stream closed; auth_required event never arrived
         Root Cause: bridge gate not wired for installed-but-unauthed
                     extensions (#2868 fallout)
         Fix: route Extension::NeedsAuth through effect_adapter.rs

   For passing/skipped lanes the existing single-line `> reason` is
   preserved so the green-path Slack output is unchanged.

2) **Cross-lane "Summary by Category" block** — second Haiku pass
   over all failed-lane summaries that groups them by shared root
   cause (e.g. "WASM tool dispatch regression — Auth Full, Auth
   Smoke, Auth Live Seeded"). Only fires when there are 2+
   failures (single-failure runs are already obvious from the
   per-lane block). Rendered as a Slack mrkdwn bulleted list since
   Block Kit doesn't support real tables.

3) **Auto-opened GitHub issues** — opt-in via CANARY_CREATE_ISSUES=1
   env var (gated to scheduled runs only in live-canary.yml so
   workflow_dispatch debugging doesn't flood the tracker). For each
   failed lane:
   - Search for an OPEN issue with title `[canary] <lane>: <test>`.
   - If found: comment "another occurrence on <run_url>".
   - If not found: open a new issue with the rich body + labels
     `canary-failure` + `lane:<lane>`.

   Strategy chosen to avoid issue spam while still surfacing
   recurring failures. Uses GITHUB_TOKEN + the repo's existing
   `permissions: issues: write` block — no new secrets.

All three additions degrade silently — Haiku failure stamps
.notable but doesn't block the post; categorization failure produces
an "_(unavailable)_" placeholder; issue-creation errors are logged
to stderr only. The notifier still exits 0 in every failure path so
a flaky webhook can't fail the canary run.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* ci(canary-report): reuse AUTH_LIVE_GITHUB_TOKEN for issue creation

Swap the issue-creation token source from the built-in
secrets.GITHUB_TOKEN to the existing AUTH_LIVE_GITHUB_TOKEN PAT —
no new secrets to mint, and that PAT already covers
nearai/ironclaw operations.

Set as CANARY_ISSUES_TOKEN (the highest-priority env var in
notify_slack.py's --github-token precedence chain) so it wins over
GH_TOKEN / GITHUB_TOKEN if any of those are also present.

Verify the PAT has `issues: write` scope (Issues: read & write for
fine-grained PATs, repo scope for classic PATs). If it doesn't, the
notifier still degrades gracefully — the API call fails, the error
is logged to stderr, the canary run isn't blocked.

Co-Authored-By: Claude Opu…
theredspoon pushed a commit to theredspoon/ironclaw that referenced this pull request Jun 21, 2026
…earai#3234)

The v2-engine E2E group still listed test_v2_kernel_auth_preflight.py,
but that file was removed in nearai#2868 (engine-v2: callable-only available
actions) and replaced with test_v2_tool_activate_surface.py for the new
tool_activate / Activatable Integrations contract.

The Web E2E Full job is skipped on PR-level CI but runs in the merge
queue, so the bad path filter dequeued nearai#3197 and nearai#3203 with
"file or directory not found: test_v2_kernel_auth_preflight.py".
theredspoon pushed a commit to theredspoon/ironclaw that referenced this pull request Jun 21, 2026
…ange (nearai#3235)

* ci(e2e): replace deleted preflight test with tool_activate surface

The v2-engine E2E group still listed test_v2_kernel_auth_preflight.py,
but that file was removed in nearai#2868 (engine-v2: callable-only available
actions) and replaced with test_v2_tool_activate_surface.py for the new
tool_activate / Activatable Integrations contract.

The Web E2E Full job is skipped on PR-level CI but runs in the merge
queue, so the bad path filter dequeued nearai#3197 and nearai#3203 with
"file or directory not found: test_v2_kernel_auth_preflight.py".

* test(e2e): unblock Live Canary auth lanes after engine-v2 contract change

The Live Canary "Auth Smoke", "Auth Full", and "Auth Live Seeded" jobs
have failed every scheduled run since 2026-05-01 (when the canary cut
over to main). Three tests in test_v2_auth_oauth_matrix.py drive the
failures, all rooted in the engine-v2 callable-only contract from nearai#2868
that didn't exist when these tests were written.

## What was broken

`test_mcp_same_server_multi_user_via_browser`
After OAuth completes, sending "check mock mcp search" through each
user's browser opens an `approval` pending_gate on the first MCP tool
call (engine v2 default). The browser fixture has no auto-approve UI,
so the chat sat in `pending_gate` for the full 5-min Playwright
timeout — `expected_text_contains="Mock MCP search result"` could
never match because the assistant bubble never received any text.

`test_wasm_tool_oauth_refresh_on_demand`
Same shape: gmail call gates on `approval` before reaching the http
credential-injection layer that performs the OAuth refresh. Without
approving, refresh_count never went above 0, so the test failed with
"Timed out waiting for OAuth refresh request".

`test_wasm_tool_first_chat_auth_attempt_emits_auth_url`
Tested OLD engine-v2 behavior — that an LLM-emitted call to a not-yet-
authed extension would surface a `gate_required` Authentication event
with an auth URL. After nearai#2868, the engine returns "action 'gmail' is
not callable in this execution context" instead, and `tool_activate`
became the model-facing enablement path. The mock LLM is canned to
emit tool calls directly, so this scenario can't be reproduced from a
scripted LLM until the canned response is updated.

## Fixes

- `_wait_for_tool_call`: accept a `token` kwarg so multi-user tests can
  poll/approve through a per-user identity. Backwards-compatible.
- `test_mcp_same_server_multi_user_via_browser`: drive approval through
  the per-user API while waiting for the tool to land. Drop the broken
  `expected_text_contains` predicate and the tied "Mock MCP search
  result" text assertions; the bearer-token isolation assertion (what
  this test actually exists to prove) is retained and unaffected.
- `test_wasm_tool_oauth_refresh_on_demand`: insert a `_wait_for_tool_call`
  approval step between `_send_chat` and `_wait_for_refresh_request`
  so the http credential layer actually runs.
- `test_wasm_tool_first_chat_auth_attempt_emits_auth_url`: marked xfail
  with an inline reason pointing at nearai#2868 and the replacement coverage
  (`test_v2_tool_activate_surface.py`,
  `test_settings_first_gmail_auth_then_chat_runs`).
- Drop the now-unused `send_chat_and_wait_for_terminal_message` import.

## conftest fix

`ironclaw_server` now sets `SECRETS_MASTER_KEY` in the spawned env.
On macOS without it, `auto_generate_and_persist` blocks on a Keychain
authorization prompt that no one's home to click, so `wait_for_ready`
times out at 60s and the fixture kills the process with SIGKILL —
making any session-scoped browser test impossible to run locally.
On Linux, the keychain backend errors fast and the auto-generate
fallback writes to `.env`, so CI was unaffected. Setting the key
explicitly matches the pattern already used in
`auth_matrix_server`, `test_v2_engine_auth_cancel`,
`test_v2_tool_activate_surface`, etc.

## Verification

Local repro confirmed each failure mode (HTTP-only repro for the
non-browser tests, server-side log inspection for the multi-user
test). Reproduced the exact pending_gate=approval pattern, fixed it,
verified the assertion semantics still hold:

```
$ pytest tests/e2e/scenarios/test_v2_auth_oauth_matrix.py::test_wasm_tool_oauth_refresh_on_demand
PASSED in 6.74s

$ pytest tests/e2e/scenarios/test_v2_auth_oauth_matrix.py::test_wasm_tool_first_chat_auth_attempt_emits_auth_url
XFAIL in 93s

$ pytest tests/e2e/scenarios/test_v2_auth_oauth_matrix.py -v --timeout=120
12 passed, 2 skipped, 3 xfailed (browser tests errored locally;
they'll run cleanly in CI)
```

The browser-driven `test_mcp_same_server_multi_user_via_browser`
couldn't be exercised locally (chromium can't launch under this
shell sandbox), but the API + auto-approve flow it now relies on is
exercised by an HTTP-equivalent repro and matches the pattern used
in `test_settings_first_gmail_auth_then_chat_runs`.

* test(e2e): set LLM_API_KEY in auth_sse_server fixture

`test_auth_required_sse_without_duplicate_response` was failing in the
merge queue on every PR (most recently bouncing nearai#3197 and nearai#3203 from
the queue) because the `auth_sse_server` fixture never set
`LLM_API_KEY` in the spawned ironclaw env. After nearai#2572 added a missing-
API-key check to the openai_compatible config validator (Apr 22),
ironclaw rejected the env-supplied openai_compatible config, fell back
to the NearAI default, hit "missing session token", and failed the
turn before the github skill could even fire its 401.

The chat thus reached `state: Failed` with no tool calls and no
`onboarding_state/auth_required` event — which is exactly what the
test asserted on, hence the consistent failure.

Adding `LLM_API_KEY=mock-api-key` matches the value already used in
every other e2e fixture (auth_matrix, conftest's ironclaw_server,
v2_engine, etc.) and unblocks the assertion. Local run: PASSED in 8s.

* fix(gateway): suppress duplicate assistant bubble after streamed response

[skip-regression-check]

The SSE `response` handler unconditionally called addMessage('assistant',
data.content) even when stream_chunks had already populated and
finalized a bubble for the same response. This stayed invisible in the
common case but surfaced as a hard test failure under the path
test_switching_back_preserves_in_progress_turn:

1. Send "What is 2+2?" on thread A — stream chunks start filling an
   assistant bubble with "data-streaming".
2. Switch to thread B mid-stream — container clears (history reload).
3. Switch back to thread A — history rehydration shows the in-progress
   turn with no response yet, so 0 assistant bubbles in DOM.
4. Stream chunks continue to fire for A — appendToLastAssistant creates
   a new bubble and accumulates the response into it.
5. response event fires — flushes any remaining buffer, removes the
   data-streaming flag (good) — then addMessage('assistant', content)
   creates a SECOND identical bubble.

Result: locator(".message.assistant").filter(has_text="4") matches
two elements, Playwright strict mode rejects the wait_for, the test
fails. Outside the test, two identical bubbles render to the user.

Fix: only call addMessage in the response handler when there was no
in-flight streaming bubble. If one existed, the streamed content is
already correct (chunks accumulate `data.content` verbatim) and the
data-streaming flag has just been cleared. Non-streaming responses
(no chunks fired) still take the addMessage branch.

Regression coverage: tests/e2e/scenarios/test_message_persistence.py::
test_switching_back_preserves_in_progress_turn already reproduces this
exact scenario and was failing in the merge queue. With this fix it
passes; skip-regression-check used because the existing E2E test is
the regression test, and the gateway doesn't have a JS unit test
harness for SSE handler state.

* fix(gateway): dedupe history-rendered SSE responses

* test(e2e): set mock LLM API key in standalone fixtures

* fix(e2e): make v2 approval tests deterministic

* test(e2e): stabilize duplicate skill install assertion

* test(e2e): assert duplicate install stays ungated

* fix(skills): skip approval for disk-installed duplicates

* fix(v2): honor no-op skill installs without approval

* test(e2e): wait for pending send marker to clear

---------

Co-authored-by: Firat Sertgoz <f@nuff.tech>
theredspoon pushed a commit to theredspoon/ironclaw that referenced this pull request Jun 21, 2026
…arai#3157)

* fix(engine): inline gate await for Tier 0 + Tier 1 Approval gates

CodeAct scripts that hit a tool requiring approval surfaced as
`RuntimeError: execution paused by gate 'approval'` inside the script
instead of pausing for the user. The async tool-resolve path
converted `EngineError::GatePaused` into a Python exception; the
sync preflight path returned `need_approval` to the orchestrator,
which on resume re-ran the LLM step and re-executed any
non-idempotent earlier tool calls in the same script.

Replace both with a host-supplied `GateController` that pauses the
live execution in place. The Monty VM (Tier 1) and the Tier 0 batch
loop both stay alive across the user's approval; on `Approved` the
gated action re-executes (lease re-consumed, auto-approve installed
before delivery so subsequent gates short-circuit); on `Denied`
the script raises a typed `RuntimeError("user denied tool 'X': ...")`
that the script can catch.

Auth and External resume kinds keep the legacy thread re-entry path
- their resolution installs new state (credentials, callback
payloads) that only takes effect on the next run-through.

Boot sweep invalidates `Approval`-kind `PendingGate` rows from a
prior process so a stranded gate after restart fails fast instead
of taking the legacy path and re-executing earlier mutations.

Tests: 3 new regression tests in scripting.rs (approve / deny /
no-controller fallback) and 4 in gate_controller.rs covering the
resolution registry's one-shot, dropped-receiver, and unknown-
request semantics. Pre-existing failures `call_id_preserved_when_no_lease`
and `stop_thread_works` reproduce on staging without these changes.

Design: docs/plans/2026-05-01-codeact-inline-gate-await.md.

* test(engine): live regression for inline gate await with CodeAct

Two integration tests in engine_v2_gate_integration.rs that exercise
the inline gate-await flow end-to-end through `ThreadManager` →
`ExecutionLoop` → orchestrator → CodeAct → `EffectExecutor` →
`GateController`:

1. `codeact_inline_gate_await_resumes_user_reproducer` reproduces the
   exact reported bug shape: a CodeAct script issuing
   `await github_tool(action="search_issues_pull_requests", ...)` for
   "what are p1 bugs in nearai/ironclaw filed in last 7 days". The
   github_tool returns `EngineError::GatePaused` mid-execution; the
   test's `OneShotApprovingGateController` marks the effects mock
   approved and returns `Approved`; the engine retries inline; the
   tool succeeds; the script's `FINAL("Found 0 P1 bugs ...")` reaches
   the user. Asserts the controller saw exactly one pause request,
   github_tool was called twice, and the response is the script's
   FINAL — not the pre-fix `RuntimeError: execution paused by gate`.

2. `codeact_inline_gate_await_denial_does_not_retry` covers the deny
   path: controller returns `Denied { reason: "not now" }`. Asserts
   github_tool was called exactly once (no retry on denial), the
   typed `user denied tool 'github_tool': not now` message appears in
   the failure events, and the pre-fix `execution paused by gate`
   string does NOT appear anywhere.

Also restored a fmt-only line shape in `gate_controller.rs` from
`cargo fmt`.

* test(engine): use realistic mock issues in inline gate-await fixture

The pre-fix fixture returned `{"items": []}` which made the script's
FINAL emit "Found 0 P1 bugs in nearai/ironclaw" — misleading, since
the repo actually has open P1 issues (e.g. nearai#2818, nearai#2997). Update the
mock to return two such items and tighten the assertion to verify
the exact count flowed from tool result through CodeAct to FINAL().

* refactor(engine): require gate_controller, bound retry, drop V1 fallback

Removes the `Option<Arc<dyn GateController>>` foot-gun: the field's
`None` arm in the executors silently re-emitted the original
`"execution paused by gate 'approval'"` RuntimeError, which is exactly
the bug this PR exists to fix. With the field required and a named
`CancellingGateController` as the explicit drop-in for non-pausing
paths (post-resolution replay, mission protected writes, tests),
forgetting to wire a controller is a compile error.

Other follow-ups in the same change to keep them on one commit:

- `MAX_INLINE_GATE_RETRIES = 3` shared by `scripting::drive_inline_gate`
  (Tier 1 async output) and `structured::execute_with_inline_gate_retry`
  (Tier 0 mid-execution). A misbehaving tool that keeps gating after
  each approval surfaces a clean error instead of pinning a CPU.
- `denial_reason_for_resolution` helper centralizes the
  `GateResolution -> reason` mapping so denial messages can't drift
  between Tier 0 and Tier 1.
- `invalidate_stranded_approval_gates_evicts_only_approval_kind` unit
  test covers the boot sweep with a mixed Approval/Auth/External
  population.
- Existing `codeact_gate_without_controller_falls_back_to_runtime_error`
  test rewritten as `codeact_default_controller_cancels_approval_gates`
  to assert the inverted invariant: the legacy bug message must NEVER
  appear, even with the inert default controller.
- Design doc updated to reflect as-shipped shape (required field, the
  bounded-retry constant, denial helper, `max_duration` stays at 30 s).

* fix(engine): bound inline pause, propagate one-shot approval, race fix

Addresses review on PR nearai#3157 (serrrfirat + Copilot + gemini-code-assist).

Four blocking correctness fixes:

1. **Bounded BridgeGateController::pause.** The await on the resolution
   oneshot now races against `pending.expires_at`. Without this, a user
   ignoring the prompt past expiry would strand the engine: the DB row
   expires, the oneshot stays open, the VM keeps running. On expiry we
   discard the pending row, drop the registry entry, and return
   `Cancelled` so the VM unwinds cleanly.

2. **One-shot approval threaded through retry.** Add
   `call_approval_granted: bool` on `ThreadExecutionContext` (default
   false). Inline retry paths (`drive_inline_gate`,
   `execute_with_inline_gate_retry`,
   `execute_single_action_with_inline_retry`, scripting sync preflight)
   set it to true on the retry call so `EffectBridgeAdapter::execute_action`
   forwards it as `approval_already_granted=true` to the host's tool
   approval check. Mirrors the legacy `execute_resolved_pending_action`
   contract; without this, tools with `ApprovalRequirement::Always`
   gated again on every retry until the bound tripped, and `always=false`
   AskEachTime gates re-prompted immediately after approval.

3. **Per-execution context registration race fixed.**
   `set_execution_context` was called AFTER `handle_user_message().await`
   returned the thread_id — but the engine task is already running,
   so a fast tool gate could reach `pause()` before the entry existed
   and get `Cancelled`. Now the bridge calls
   `set_pre_execution_context` (per-user) BEFORE
   `handle_user_message`, then promotes to (user, thread)-keyed once
   thread_id is known. `pause()` falls back to the per-user entry on
   miss.

4. **Tier 0 parallel-batch coverage.** New
   `execute_single_action_with_inline_retry` wraps `execute_single_action`
   with the same bounded retry shape used by Tier 1. Both the
   single-runnable and multi-runnable branches of
   `handle_execute_actions_parallel` go through it, so simultaneous
   gates in a parallel batch no longer fall through to the legacy
   re-entry path.

Smaller fixes:

- PROJECTION lint annotations on the two `broadcast_for_user` sites
  added by this PR (gate_controller emit_gate_prompt; resolve_gate
  inline-await fast-path resolution event).

Test coverage:

- `GatingThenOkEffects` now records `context.call_approval_granted`
  per call; `codeact_gate_inline_await_approved_delivers_result`
  asserts the retry observes `true`. Locks in the one-shot approval
  propagation contract.

* test(engine): update gate integration tests for inline-await semantics

CI was failing on 5 tests in engine_v2_gate_integration.rs that asserted
the legacy `ThreadOutcome::GatePaused` unwind for `Approval` gates. With
inline-await + the required `CancellingGateController` (default), Approval
gates resolve inline and the thread completes (or fails) instead of
pausing — these tests were exercising the pre-PR flow that this PR
replaces.

- `gate_paused_transitions_thread_to_waiting` → renamed to
  `approval_gate_resolves_inline_via_controller`. Wires
  `AutoApprovingGateController`, asserts thread completes after inline
  approval, ApprovalRequested + ActionExecuted both recorded.
- `gate_paused_thread_resumes_to_completion` → renamed to
  `approval_denied_inline_completes_thread_with_failed_action`.
  Asserts the default `CancellingGateController` cancels the gate and
  the thread completes with a failed action (no stranded pending gate).
- `approval_chains_directly_into_auth_for_install_flow` rewritten:
  Approval handled inline by controller, Auth gate (still legacy path)
  bubbles up as ThreadOutcome::GatePaused, legacy auth-resume drives
  completion.
- `approval_resolution_executes_pending_call_directly` and
  `gate_resume_with_execution_obligation` reduced to no-op stubs with
  inline rationale documenting where the post-PR equivalent coverage
  lives. Removing entirely would erase the breadcrumb in git log.

Also threads the `ApprovalRequested` event through
`execute_single_action_with_inline_retry`: the wrapper now returns
`Vec<EventKind>` per call so the caller can append both the approval
prompt and the post-retry outcome to the thread event log. Without
this, observers saw only the final ActionExecuted/Failed event with
no record that a gate had fired.

Drops a few dead test helpers (`ApprovalTool`, `make_caps_with_approval_tool`,
unused imports) that the deleted assertions no longer reference.

Quality gate: fmt clean, clippy zero warnings, 513 engine lib + 443
bridge + 29 gate integration tests pass (2 pre-existing engine-lib
staging failures unrelated to this PR).

* test(engine): update skill_codeact integration tests for inline-await

Two more tests in engine_v2_skill_codeact.rs were asserting the
legacy `ThreadOutcome::GatePaused` → `resume_thread` flow for
`Approval` gates. With PR nearai#3157 the engine catches Approval gates
inline and the thread runs to completion in a single `join_thread`.

- `skill_prompt_context_survives_pause_and_resume`: wires
  `AutoApprovingHttpController`, asserts thread completes after inline
  approval, drops the now-impossible `resume_thread` step.
- `skill_prompt_context_survives_compaction_and_resume`: same; also
  drops the assertion on the persisted transcript at the *pause point*
  (no longer externally observable). The load-bearing post-compaction
  LLM-call assertions remain — they're the actual contract this test
  was protecting.

Adds an `AutoApprovingHttpController` test helper local to this file
that approves gates by marking the underlying `PausingHttpMockEffects`
approved before returning `Approved`.

* fix(engine): bridge cleanup on dropped sender + ActionFailed on preflight denial

Addresses Copilot review on PR nearai#3157.

- `BridgeGateController::pause`: when the resolution oneshot's sender
  is dropped (process shutdown / registry cleared), discard the
  `PendingGate` row before returning `Cancelled`. Without this, a
  stranded prompt remained visible in the UI and `pending_gates.insert`
  rejected duplicates for the same `(user, thread)` so a follow-up
  gate could not register. Same cleanup as the expiry branch already
  performed.

- `scripting.rs` sync-preflight denial path: emit `EventKind::ActionFailed`
  before resuming Monty with `RuntimeError`, so the thread event log
  is consistent with the other denial paths (`drive_inline_gate`,
  `structured.rs`). Auditing why a tool didn't run is now possible
  from the events alone.

- Delete the two empty `#[tokio::test]` stubs left in
  `engine_v2_gate_integration.rs`
  (`approval_resolution_executes_pending_call_directly_via_resolved_pending_action`,
  `gate_resume_with_execution_obligation`). The post-PR equivalent
  coverage is in `codeact_inline_gate_await_*` (this file) and the
  scripting unit tests; the rationale is preserved in `git log`
  (commit 87fe4fc) without an empty test slot misleading coverage
  signals.

* fix(engine): finish inline-await migration; remove legacy gate-paused for Tier 0 policy gates

Address PR nearai#3157 review:

- router.rs: clear pre_execution slot on handle_user_message error so a
  failed engine spawn doesn't leak a stale (user, conversation) entry
  that would mis-route the next gate prompt.

- router.rs:await_thread_outcome: on the 5-min request deadline, return
  BridgeOutcome::Pending instead of join_thread() — joining would block
  for up to the gate's 30-min expires_at when the parked task is in
  pause(). The PendingGate row stays live for resolution.

- gate_controller.rs: re-key pre_execution by (user_id, conversation_id)
  instead of user_id alone. Two concurrent conversations / browser
  tabs for the same user no longer clobber each other's slot. Plumbs
  conversation_id through ThreadExecutionContext and GatePauseRequest
  so pause() can match a gate to its originating conversation.

- gate_controller.rs: serialize concurrent inline gates per
  (user, thread) via a per-key tokio Mutex held across the
  PendingGateStore::insert + select-await window. A parallel batch
  where two tools both gate now queues the second behind the first
  rather than silently surfacing it as Cancelled on (user, thread)
  uniqueness collision.

- orchestrator.rs (Tier 1 / CodeAct): remove the legacy gate_paused JSON
  sentinel + thread re-entry path for Approval gates. Both
  __execute_action__ and __execute_actions_parallel__ now pause inline
  on PolicyDecision::RequireApproval (mirroring structured.rs) and
  route tool-raised gates through the existing
  execute_single_action_with_inline_retry wrapper. Authentication and
  External resume kinds keep the legacy re-entry path because their
  resolution installs new state that only takes effect on the next
  thread run-through.

Test fixtures across bridge/effect_adapter, action_projector, and the
gate/sandbox integration suites get the new conversation_id field
defaulted to None.

Refs: comments 3173757791, 3176248954, 3176248976, 3176249001, 3176249032

* fix(engine): defer post-Pending context cleanup; route parallel JoinSet through inline-retry

Address PR nearai#3157 review on commit d211bfc:

- router.rs: when await_thread_outcome returns BridgeOutcome::Pending and
  the engine task is still running (typically parked in
  BridgeGateController::pause), defer clear_execution_context to a
  spawned watcher task that polls is_running until the thread completes
  and then clears the (user, thread) context + gate-locks entry. Without
  this, the unconditional clear after a 5-min request timeout stranded
  the parked thread: the eventual gate resolution would call pause()
  for any subsequent gate with no registered context and surface as
  silent Cancelled. Watcher caps at 60 minutes (well past the 30-min
  PendingGate expiry) as a defensive safety bound.

- structured.rs: route the multi-runnable JoinSet branch in
  execute_action_calls through execute_with_inline_gate_retry, matching
  the single-runnable fast path. Without this, an Approval gate raised
  mid-execution in a parallel batch with >1 runnable tool call
  bubbled out as Err(GatePaused) and went through the legacy
  gate-paused / re-entry path, re-introducing the double-execution bug
  for already-completed sibling calls in the same batch. The signature
  of execute_action_calls switches leases from &LeaseManager to
  &Arc<LeaseManager> so the JoinSet tasks can clone an owned handle;
  all current callers are tests already constructing
  Arc::new(LeaseManager::new()), so this is a no-op call-site change.

Refs: comments 3176333657, 3176333685

* test(engine): fix call_id_preserved_when_no_lease MockEffects inventory

PR nearai#2868 added a callable-inventory check ahead of the lease lookup in
execute_action_calls. The test was constructing MockEffects with an
empty action list, so preflight short-circuited on "action is not
callable in this execution context" before reaching the lease branch
the test was actually trying to exercise. The assertion
error.contains(\"no lease\") then failed against the not-callable
message.

Populate the mock with test_action(\"web_search\") so the inventory
gate passes and the call reaches the lease check, matching the test's
documented intent.

* fix(engine): use gate-provided params in inline-await pause; drop Thread clone in parallel branches

- orchestrator.rs::execute_single_action_with_inline_retry now sources
  GatePauseRequest.parameters from the gate-paused payload (which
  reflects safety-layer transformations/redactions) rather than the
  original caller params, matching structured.rs::execute_with_inline_gate_retry.
- Both inline-retry helpers now take ThreadId + user_id instead of
  &Thread, so the parallel JoinSet branches no longer clone the full
  Thread (with message/event transcripts) per spawned task. The
  per-task ThreadExecutionContext clone is what the helpers actually
  need; in orchestrator.rs we build a base ctx once and override
  current_call_id per task instead of re-running thread_execution_context.

Addresses Copilot PR nearai#3157 review comments 3176594866, 3176594886, 3176594898.

* fix(engine): address serrrfirat review — audit event + stop-during-wait

Three latest serrrfirat review threads on PR nearai#3157:

* Medium: structured inline approval drops ApprovalRequested audit
  event (REAL — fixed). `execute_with_inline_gate_retry` previously
  swallowed `Err(GatePaused)` and returned only the post-retry
  outcome, so `classify_exec_result` never saw the gate and the
  `ApprovalRequested` event was lost. The orchestrator (Tier 1) path
  emits both events. Mirror that shape here:

  - `execute_with_inline_gate_retry` now returns
    `(Result<ActionResult, EngineError>, Vec<EventKind>)`. The Vec
    carries one `ApprovalRequested` per retry iteration that gated.
  - Slot type widens to `(ActionResult, EventKind, Vec<EventKind>)`
    so the merge phase flushes pre-terminal events before the
    terminal event in original-call order.
  - Two regression tests pin the contract (approved → Executed,
    denied → Failed); both assert ApprovalRequested precedes the
    terminal event.

* High: stop/cancel does not unblock thread parked in inline gate
  await (REAL — fixed). `BridgeGateController::pause()` selected
  only on the resolution oneshot + 30-min expiry; `stop_thread()`
  sent `ThreadSignal::Stop` but the parked engine task wasn't
  polling the signal channel.

  - Adds `GateController::cancel_thread(thread_id)` to the
    engine-side trait (default no-op).
  - `ThreadManager::stop_thread()` calls `cancel_thread()` BEFORE
    sending `ThreadSignal::Stop` so any parked `pause()` future
    wakes promptly with `GateResolution::Cancelled`.
  - `BridgeGateController` tracks in-flight pauses in
    `active_pauses: HashMap<ThreadId, HashSet<Uuid>>` and walks
    that set on cancel, delivering `Cancelled` via the existing
    `GateResolutions::try_deliver` channel and discarding pending
    rows via `PendingGateStore::discard_for_thread`.
  - `pause()` always untrack on exit (idempotent — `cancel_thread`
    may have already removed the entry).
  - Two new bridge-tier unit tests: parked pause wakes within 2s
    of cancel_thread; cancel_thread on a thread with no parked
    pause is a no-op.

* Medium: live inline gate waits are uncapped per user/process
  (PARTIAL — push back, follow-up). The reviewer's own comment
  notes this is a design-doc follow-up, not a blocker. The current
  bound is implicit: one pending gate per (user, thread) × the
  existing thread-creation budget × the 30-min expiry, which is
  enough to ship the inline-await substrate. Adding a per-user
  semaphore correctly requires designing the cap UX, fairness, and
  rejection error shape — out of scope for this PR. Add a TODO
  inside `pause()` pointing at the design doc and the
  follow-up issue, with the reasoning written down so the next
  contributor doesn't have to re-derive it.

Pre-existing on this branch (NOT introduced by these changes):
`runtime::manager::tests::stop_thread_works` flakes on the branch;
the original PR commit message acknowledges "stop_thread_works
reproduce on staging without these changes." Confirmed by stashing
this commit and running the test: still fails. Out of scope here.

Test totals after this commit:
  - executor::structured: 19 passing (incl. 2 new)
  - bridge::gate_controller: 6 passing (incl. 2 new)
  - cargo fmt + clippy --lib clean

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(engine): typed DenialOutcome + plug exhaustion-path lease leak

Three review-driven fixes to the inline gate-await path.

1. CancellingGateController and bridge expiry/shutdown surfaced as
   `RuntimeError("user denied tool 'X': cancelled")` — misleading
   because the user never saw a prompt. Replace `Option<String>`
   helper with a typed `DenialOutcome { DeniedByUser, Unavailable }`
   so cancelled/expired/no-handler gates render as "approval for
   tool 'X' unavailable: …" while real user denials keep the
   "user denied" framing. Updated all six call sites in
   scripting.rs / structured.rs / orchestrator.rs through the
   typed surface so wording can't drift.

2. Both `execute_with_inline_gate_retry` and the orchestrator's
   `execute_single_action_with_inline_retry` consumed a fresh
   lease use on the final approved iteration, then exited the
   loop without ever calling execute_action. The exhaustion-path
   error didn't trip the caller's refund check, so a misbehaving
   tool that gates after every approval would slowly drain
   `max_uses`. Refund the unused lease before returning.

3. Drop the dead `parameters: serde_json::Value` field on
   `PendingFuture::Tool` and the matching `_parameters` arg on
   `resolve_tool_future`. The gate's own parameter snapshot is the
   source of truth on retry; threading the original through made
   the signature read like there was a second source.

Quality gate: cargo fmt, cargo clippy --all --benches --tests
--examples --all-features (clean), cargo test -p ironclaw_engine
--lib (520 pass), cargo test --test engine_v2_gate_integration
(27 pass), cargo test --lib bridge:: (456 pass).

---------

Co-authored-by: Nikolay Pismenkov <nickpismenkov@gmail.com>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
theredspoon pushed a commit to theredspoon/ironclaw that referenced this pull request Jun 21, 2026
* fix(e2e): restore auth and approval coverage (#3430)

* test(e2e): avoid REPL auth retry race (#3437)

* feat: add pairing_approve tool for Slack binding via chat (#3396)

* feat: add pairing_approve tool for Slack binding via chat

Users can now paste their Slack pairing code in the IronClaw chat and
the LLM will call pairing_approve to bind their accounts. No need to
use the API directly.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: require approval before pairing + fix formatting

Address review comment: pairing_approve now requires UnlessAutoApproved
approval before executing, preventing accidental account binding.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* test: add regression test for pairing_approve tool

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: address review — Always approval, lock channel, add to protected list

1. Changed ApprovalRequirement to Always (not bypassable by auto-approve)
2. Locked channel to slack-relay constant (removed generic channel param)
3. Added pairing_approve to PROTECTED_TOOL_NAMES
4. Added tests: always-approval, protected-name, channel constant

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(web): isolate cross-tenant SSE/WS status events and thread access (#3390)

* fix(web): isolate cross-tenant SSE/WS status events and thread access

Plug a multi-tenant leak where unscoped `sse.broadcast(...)` calls
from `GatewayChannel::send_status`, sandbox `JobEvent` dispatch, the
WASM/Slack OAuth completion handlers, and any producer that lost
`metadata.user_id` along the way fan out to every connected
subscriber — exposing another tenant's tool calls, tool output,
onboarding state, and job lifecycle to anyone with an open SSE/WS.

Changes
- Extract `dispatch_status_event(sse, multi_tenant_mode, user_id, ev)`
  from `Channel::send_status`. In multi-tenant mode an unscoped event
  is dropped (with a WARN naming the producer to fix); single-tenant
  keeps the global broadcast since there is one subscriber population.
- `IncomingMessage::new` now defaults `metadata` to `{"user_id": ...}`,
  and `with_metadata` preserves the key so downstream `send_status`
  consumers always have an owner to scope by.
- WASM/Slack OAuth completion broadcasts route through
  `broadcast_for_user(&owner_id, ...)`. Sandbox `JobEvent` dispatch
  in `main.rs` respects `multi_tenant_mode` for the empty-`user_id`
  fallback.
- New pre-commit check #10 (`MULTITENANT`) flags unscoped
  `sse.broadcast(...)` lines without a `// multi-tenant-safe: <reason>`
  marker or a transport-only exemption. Marker regex accepts the marker
  anywhere in a `//` comment so compound annotations on a single line
  work.

Tests
- `src/channels/web/tests/status_event_isolation.rs` — 5 unit tests
  covering both modes and the per-variant drop invariant.
- `src/channels/web/platform/sse.rs` — 2 quadrant tests for the
  `subscribe_raw` filter (scoped/unscoped × matching/mismatched).
- `tests/thread_isolation_integration.rs` — 9 HTTP-level checks that
  Bob cannot reach Alice's chat history (paginated and not), threads
  list, engine v2 detail/steps/events, or Responses GET, plus an
  unauthenticated-rejection guard.
- 8 new self-test cases for the `MULTITENANT` script check, including
  the compound projection-exempt + multi-tenant-safe annotation case.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(web): pin cross-tenant boundaries on jobs, files, routines

Audit of the protected route surface for the same bug shape #3390 fixed
(handler that takes a user-controlled id and reads without an ownership
predicate) found that the implementations were correct but four
boundaries had no integration test. Lock them in before they regress.

- Sandbox job persisted-events history (`/api/jobs/{id}/events`):
  Bob → 404 on Alice's job; Alice → 200 on her own.
- Sandbox job workspace listing (`/api/jobs/{id}/files/list`):
  Bob → 404 on Alice's job.
- Sandbox job file read (`/api/jobs/{id}/files/read`): Bob → 404 on
  Alice's job; Alice → 200 on her own; Alice → 403/404 on
  `?path=../outside.txt` (path-traversal pin against the
  `canonicalize() + starts_with(base_canonical)` guard).
- Routine run history (`/api/routines/{id}/runs`): Bob → 404; Alice
  → 200 with at least one seeded run.

The OAuth-state and NEAR-nonce stores were also flagged in the audit
but neither is a real cross-tenant bug: both are pre-auth, single-use,
and the token IS the secret. Documenting here so a future audit
doesn't re-flag them.

Project-static (`/projects/{id}/...`) is left for a follow-up — it
relies on `ironclaw_base_dir()` which is a process-wide `LazyLock`,
making per-test override fragile in the integration runner.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(web): address PR #3390 review — forge-resistant metadata, OAuth toast routing

Addresses six review comments from gemini-code-assist, copilot, and
serrrfirat on PR #3390. False-positives and the perf nit on
`with_metadata` are explained in the reply thread, not changed in code.

- (HIGH, serrrfirat) `IncomingMessage::with_metadata` now ALWAYS sets
  `metadata.user_id` from `self.user_id`, dropping any caller-supplied
  value. A WASM channel emitting `{"user_id":"victim"}` via
  `apply_emitted_metadata` can no longer reroute downstream
  `ToolStarted` / `ToolResult` SSE events into another tenant's
  stream. New unit tests pin the forgery-resistance invariant.
- (MEDIUM, copilot + serrrfirat) Slack relay OAuth callback now
  broadcasts the completion toast to the resolved `oauth_user`
  (the IronClaw user who initiated the flow) rather than
  `state.owner_id`. In multi-tenant deployments those differ and the
  previous routing delivered the toast to the wrong browser tab.
  Extracted the lookup into `resolve_relay_oauth_user`; two unit
  tests cover the secret-present and secret-missing cases.
- (MEDIUM, gemini) `dispatch_status_event` treats empty-string
  `user_id` the same as `None` so producers that lost the field
  along the way fail-closed instead of falling through to a global
  broadcast in multi-tenant mode.
- (MEDIUM, gemini) `main.rs` sandbox JobEvent dispatch now reuses
  `dispatch_status_event` instead of duplicating the drop / WARN /
  broadcast policy. `dispatch_status_event` is bumped from
  `pub(crate)` to `pub` so the binary crate can call it.
- (LOW, copilot) Pre-commit `MULTITENANT` self-tests gain three
  cases (`state.sse.broadcast(`, `gw_state.sse.broadcast(`,
  annotated receiver-prefixed) to lock the existing boundary regex
  behaviour against future tightening.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(channels): preserve i64 metadata.user_id from Telegram in with_metadata

PR #3390's forge-resistance fix made `IncomingMessage::with_metadata`
*always* overwrite `metadata.user_id` with `self.user_id` as a String.
That broke the Telegram WASM channel: it persists Telegram's chat user
ID as `metadata.user_id: i64` and re-deserializes it into
`TelegramMessageMetadata { user_id: i64, ... }` in `on_respond` /
`on_status`. After the fix, `respond` blew up with
`invalid type: string "999", expected i64 at line 1 column 87`,
failing 3 Telegram integration tests in CI.

Narrow the carve-out: overwrite only when the existing `user_id` is a
String (or missing). Non-string values are channel-private and the
SSE routing layer reads via `as_str()` — non-strings already fail
closed in multi-tenant mode, so the forge threat (WASM emits
`{"user_id":"victim"}` as a string) is still mitigated, while
Telegram's i64 use case survives.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(web): redact WARN payload + harden dotdot traversal test (PR #3390)

Two follow-up fixes from the second pass of review on #3390:

1. `dispatch_status_event`'s WARN log used `?event`, which on
   `AppEvent::Response` / `Thinking` / `ToolResult` carries
   user-authored content into operator logs in a multi-tenant
   deployment. Replace with `event_kind = event.event_type()`
   (the wire-stable variant name) — enough to identify the
   misbehaving producer without leaking tenant data. Picked up via
   Copilot's review on `src/channels/web/mod.rs:666`.

2. `alice_job_file_read_rejects_dotdot_traversal` planted
   `outside.txt` under `outer.path()` (the `start_server_with_db`
   fixture's tempdir holding `test.db`) but probed
   `?path=../outside.txt` relative to `alice_proj` — a separate
   `tempfile::tempdir()` rooted at the OS temp directory. The two
   paths were unrelated, so the test could pass even if `..`
   traversal was permitted (probe just hit empty space). Build the
   directory tree by hand instead: `parent/alice_proj/` with the
   planted file at `parent/outside.txt`, so the probe deterministically
   resolves to the planted bytes. Add a body-content assertion that
   fails loudly if those bytes leak. Picked up via Copilot's review on
   `tests/cross_tenant_resource_isolation.rs:355`.

Plus a `cargo fmt` fix for `src/channels/channel.rs:1202` that was
breaking the Formatting CI check.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(web): address PR #3390 follow-ups — multi-tenant fallback WARN, exhaustive variant pin

- `resolve_relay_oauth_user`: take `multi_tenant_mode`; emit WARN when the
  `relay:{ext}:oauth_user` secret is missing in multi-tenant mode so the
  unrecoverable-initiator case surfaces in operator logs. Single-tenant
  fallback stays silent (owner == only user).
- `dispatch_status_event`: doc note clarifying the function is `pub` only
  for the sandbox JobEvent rx loop in `main.rs`; not part of a stable
  public API.
- `_compile_time_appevent_variant_check`: exhaustive-match helper paired
  with `unscoped_drop_holds_for_every_status_variant_in_multi_tenant`.
  Adding a new `AppEvent` variant now fails the test build, prompting an
  update to both the helper and the runtime leak-candidate list.
- `tests/thread_isolation_integration.rs`: honest scope note on the
  engine-v2 thread tests — they pin handler shape (404/empty for
  unknown id), not the cross-tenant ownership branch. Cross-tenant
  engine-v2 coverage requires an `ENGINE_STATE` test fixture; tracked
  as a follow-up in the comment block.

Tests: 8 unit (5 status_event_isolation + 3 resolve_relay_oauth_user)
and 9 thread_isolation_integration pass; clippy clean; pre-commit
safety scripts pass (regression suite 27 cases).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(web): unconditionally consume relay:{ext}:oauth_user secret in OAuth callback

The previous cleanup site lived inside the `if let Some(pairing_store)`
branch of the result block, which was unreachable on three failure
paths:

1. `pairing_store` is None (no identity pairing wired up)
2. The result block `?`-short-circuits before reaching the `if let`
   (e.g. `set_setting` fails, `activate_stored_relay` fails, an inner
   `relay_config()` / `list_connections()` errors)
3. The `if let` body itself errors before reaching the delete (e.g.
   `list_connections` returns no matching team)

Leaving the secret behind lets a subsequent OAuth callback for the
same extension read a stale initiating user and misroute the
completion toast — Copilot review on PR #3390 (comment id 3211833864).

Move the delete to right after `resolve_relay_oauth_user` returns,
where it always runs once the value has been captured, regardless of
downstream failure mode. Updated the inner comment to document that
the secret is already gone by the time the pairing branch reads
`oauth_user`.

Regression test: `test_relay_oauth_callback_consumes_oauth_user_secret_on_failure_path`
seeds the secret, fires the callback against a fixture with no real
relay backend (so the result block deterministically errors), and
asserts the secret is gone afterward. Pre-fix this would have left
the secret behind on the failure path.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(auth): tighten Telegram pairing UX and OAuth-failure recovery (#3317, #3319, #3320) (#3381)

* fix(auth): tighten Telegram pairing UX and OAuth-failure recovery (#3317, #3319, #3320)

Three Bug Bash P1 issues from the same user journey: setup → use → fail.
The unifying root cause was per-channel auth tested in isolation; cross-channel
flows (Telegram → Gmail OAuth → resume) had no coverage and three small leaks
combined into a stuck conversation.

#3317 — Telegram pairing reply now names every IronClaw surface explicitly
(web settings, agent chat, terminal). The agent submission parser learns
`approve <channel> <code>`, dispatched through a new bridge handler that
mirrors `POST /api/pairing/{channel}/approve`.

#3319 — OAuth callback failures now log a category + correlation ID so a
user-reported "I saw 400" maps to one log line. Adds the
`OauthCallbackFailure` enum and `oauth_failure_correlation_id` helper.

#3320 — Two cleanup gaps fixed: (a) `/clear` now drains
`pending_oauth_flows` for the user (otherwise stale flows linger 5min and
mask new auth attempts); (b) OAuth provider-error and exchange-failure paths
now auto-cancel the engine pending auth gate via `clear_engine_pending_auth`,
so the conversation isn't blocked waiting for a resume that will never arrive.

Tests: 5 new submission-parser tests, 2 new bridge-handler tests, and one
new OAuth callback test verifying the pending-flow drain on provider error.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(auth): cross-channel pairing claim coverage + canary lane (#3317)

Adds the structural coverage that was missing when #3317 shipped:

- E2E (`tests/e2e/scenarios/test_telegram_pairing_chat_claim.py`):
  three scenarios that drive the full Telegram pairing flow through
  the gateway. Asserts the bot reply names every IronClaw surface
  (web Settings, agent chat, terminal CLI), drives `approve telegram
  CODE` through `/api/chat/send` and verifies the paired user
  exchanges messages without re-prompting, and confirms invalid
  codes get a clear rejection instead of an LLM-improvised reply.

- Rust integration (`tests/telegram_pairing_chat_claim_integration.rs`):
  drives `Submission::PairingClaim` through a real `Agent` →
  `bridge::handle_pairing_claim` → `PairingStore::approve` chain
  using `TestRig` with engine v2 enabled. Covers the happy path
  (mints a code, claims it via chat, asserts `Pairing approved`)
  and the invalid-code rejection. The unit tests in `bridge/router`
  cover only the no-extension-manager and invalid-channel branches —
  this test exercises the wiring between submission parser, agent
  loop dispatch, and bridge handler that #3317 specifically broke.

- Canary (`scripts/live_canary/auth_registry.py`): adds the two
  user-visible scenarios to `AUTH_CHANNEL_TESTS` so the auth-channels
  lane (scheduled every 6h) catches the regression class in CI
  before any real user encounters it.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* docs: align pairing-claim and oauth-correlation comments with code (#3381)

Three Copilot review comments on PR #3381 flagged docstring/code drift in
already-merged PR #3317/#3319/#3320 changes. No behavior change — only
the doc strings move:

- `Submission::PairingClaim.code` and the inline `approve <channel> <code>`
  parser comment claimed the user's casing was preserved, but the parser
  builds the code from `lower` and the regression tests already lock in
  the lowercased shape (`code == "abc12345"`). Update both comments to
  describe the actual normalize-then-store contract.

- `oauth_failure_correlation_id` claimed the correlation appeared in the
  user-facing error subtitle, but the failure path renders
  `landing_html(label, false)` whose subtitle is fixed and never receives
  the correlation. Mark the helper as logs-only and note that plumbing
  the ID through the HTML is a follow-up.

[skip-regression-check] doc-only, behavior already covered by existing
pairing-claim parser tests in submission.rs.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore(telegram): bump channel registry to 0.2.11

The pairing-reply wording was updated in channels-src/telegram/, which
the version-check CI requires be matched by a registry version bump.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(telegram): fix import path in pairing chat claim e2e test

The scenarios/ folder is a Python package (has __init__.py), so a
flat `from test_telegram_e2e import …` fails with
ModuleNotFoundError during pytest collection. Switch to a relative
import that matches the package layout, and drop the unused
OWNER_USER_ID symbol while we're here.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(auth): address Copilot review on /clear OAuth drain + correlation doc

Two follow-ups on PR #3381's Copilot pass:

1. `Agent::process_clear` (engine v1 path) now drains in-flight OAuth
   flows for the clearing user, mirroring the engine-v2 cleanup added in
   `bridge::router::clear_engine_conversation`. Without this, `/clear`
   was a clean slate on v2 but v1 left ghost flows in
   `extension_manager.pending_oauth_flows()` until the 5-minute
   `OAUTH_FLOW_EXPIRY` ticked over — same regression class #3320 fixed
   on v2.

2. `oauth_failure_correlation_id`'s docstring previously said "redacted
   state fingerprint", but callers seed it with the raw `state` query
   value (or `flow.extension_name` for post-resolution failures).
   Updated the doc to describe the actual behaviour: an arbitrary seed
   that is hashed before any hex output, with a pointer to
   `redact_oauth_state_for_logs` for the log-safe fingerprint.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(e2e): make Telegram pairing chat-claim suite actually run

The scenario landed in PR #3317 was orphaned — never wired into any CI
lane and could not pass even when run by hand. Three structural
issues, all fixed here:

1. `install_telegram` now overlays the locally-built WASM (and matching
   capabilities file) on top of the registry-downloaded artifact when
   present. The pairing-reply wording lives inside the WASM binary, so
   without this overlay the test was asserting source-tree text against
   the previous release's bytes. The overlay is best-effort: when the
   local WASM is absent (CI groups that don't build the channel), the
   test that depends on it skips with a clear message and the canary
   lane in `scripts/live_canary/auth_registry.py` still covers the
   wording end-to-end against the deployed binary.

2. `Submission::PairingClaim` is handled out-of-band by the bridge
   layer; the response is delivered via `WebChannel::respond` →
   `AppEvent::Response` over SSE only — no `Turn` is persisted, so
   polling `/api/chat/history` could never see it. Refactored
   `test_chat_surface_approves_pairing_code` and
   `test_chat_surface_rejects_invalid_pairing_code` onto a
   `_send_and_collect_response` helper that opens the SSE stream first
   (so the broadcast doesn't fan out to zero subscribers) and matches
   on the `response` event for the test thread.

3. Wired the file into `e2e.yml`'s `extensions` group so the suite
   actually runs on every PR.

Verified locally: all three scenarios pass, plus the existing 22
Telegram e2e tests still green with the install-overlay change.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(bridge): bound and sanitize invalid-channel echo in pairing claim

Address Copilot review on PR #3381: `handle_pairing_claim`'s
invalid-channel branch was rendering the raw `channel` token (and the
underlying `IdentityError`, which itself echoes the offending input)
back to the user. Both routes are unbounded and could carry control
characters or markup since `channel` comes from chat input — a
hostile prompt could blow up the SSE / Telegram / TUI reply or smuggle
backticks/escape sequences through.

Cap the echo at 32 ASCII-alphanumeric (or `-`/`_`) characters and
replace the verbatim error with a fixed category description, so the
reply size and shape are bounded by what we render explicitly. Add a
regression test that drives a 200-char hostile blob (control chars +
backticks) through the handler and asserts the rendered reply stays
clean and short.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(bridge): include hyphens in invalid-channel error copy

Address Copilot review on PR #3381: the invalid-channel reply
listed "lowercase letters, digits, or underscores" as the valid
character set, but `ExtensionName::new` (and `web::features::pairing::
parse_channel`) intentionally accept hyphens too — they're folded to
underscores during canonicalization. A user typing `slack-relay` would
otherwise get an "invalid name" reply listing rules that contradict
the actual validator.

Updated the message to include hyphens with `telegram` and
`slack-relay` as concrete examples, and tightened the regression test
to assert against the user-controlled preview region between the
delimiter backticks rather than a global backtick count (which was
fragile to copy that includes example slugs in backticks).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(auth): address PR #3381 review on credential-scoped gate cleanup and Telegram surface promise

Three reviewer findings, one commit:

- OAuth provider-error and exchange-failure paths used
  `clear_engine_pending_auth(user, None)`, which discards every
  Authentication gate for the user. A failed Gmail callback could
  silently wipe an unrelated Slack/MCP gate waiting on a different
  thread. New `clear_engine_pending_auth_for_credential(user, credential)`
  helper in bridge::router scopes cleanup to the failed flow.
  Provider-error path tracks `removed_secret_name` alongside
  `removed_user_id` so the scoped variant is callable.
- Expired-flow branch in the OAuth callback handler had two bugs: it
  never cleared the engine pending auth gate (so the conversation sat
  blocked forever, same #3320 class the provider-error fix addresses),
  and the broader `clear_auth_mode` it called would re-discard via
  the unscoped helper anyway. Now calls the credential-scoped helper
  and the legacy-v1-only `clear_session_auth_mode_for_thread`.
- Telegram pairing reply advertised `approve telegram CODE` as
  usable "in any IronClaw chat (TUI / web / Telegram)", but an
  unpaired Telegram DM is intercepted by the allowlist gate before
  the agent parser sees the command — the user would just get
  another pairing reply. Reply now lists only the surfaces that
  actually work (web / TUI / CLI) and a comment explains why.

Regression coverage:

- `clear_engine_pending_auth_for_credential_only_clears_matching_credential`
  locks in helper scoping (Gmail/Slack two-gate scenario).
- `oauth_callback_expired_flow_clears_credential_scoped_engine_gate`
  drives the full callback through axum oneshot with engine state
  seeded; asserts the matching gate clears and the unrelated gate
  survives.
- E2E `test_telegram_dm_approve_command_is_intercepted_by_allowlist_gate`
  exercises the Telegram webhook path (not /api/chat/send) to lock in
  the channel-layer interception, wired into the auth canary lane.
- Existing E2E pairing-reply test gains an assertion that
  "TUI / web / Telegram" is *not* in the reply.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(oauth): mirror failure-cleanup contract on provider-error and reconcile stale comment

Copilot review on PR #3381 caught two real issues in the credential-scoped cleanup
landed in d45cf1bd8:

- Provider-error branch (`?error=access_denied`) returned the error page
  without broadcasting `OnboardingState::Failed` or clearing the legacy v1
  session `pending_auth`. The exchange-failure and expiry branches do both.
  Net effect: the auth card stayed spinning and the next user message was
  intercepted as a token. Now mirrors the other failure paths — keep the
  full `flow`, emit Failed SSE, clear v1 session, clear credential-scoped
  engine gate, then return the error page.
- Post-exchange comment said "failed callbacks should leave the gate
  visible for retry" — that was the pre-#3320 contract. Rewrote it to
  describe the new shape: each failure mode clears its own gate at the
  failure site; this section only handles legacy-v1 session cleanup that
  runs regardless of outcome. Also explains why we use
  `clear_session_auth_mode_for_thread` here instead of `clear_auth_mode`
  (the latter would re-clear the engine gate on the *success* path and
  break the `ExternalCallback` resume).

Regression: `test_oauth_callback_provider_error_broadcasts_onboarding_failed`
in the oauth tests module — drives an `?error=access_denied` callback with
a flow whose `sse_manager` is attached, asserts the receiver gets
`OnboardingState::Failed` with the provider's `error_description` as the
message body.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(common): describe paths and platform helpers in crate description (#3498)

Align the package `description` and the lib.rs crate-level doc with
the modules now exposed from `ironclaw_common` (paths, platform,
env_helpers, attachment), which #3387 lifted out of `src/`. The
previous wording predates that extraction and only mentioned "types
and utilities".

This is also the release-plumbing trigger for v0.28.1: release-plz
proposes a leaf bump on source-path changes, and once `ironclaw_common`
crosses to a new patch, the root `ironclaw` bump can be added on top
of the release-plz branch (same approach as commit 070cbede1 for
v0.28.0). See PR #3372 for the equivalent v0.28.0 trigger.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore: release

* chore(release): bump ironclaw to 0.28.1

cargo-semver-checks did not detect API-breaking changes in the
ironclaw_common 0.4.1 -> 0.4.2 leaf bump, so release-plz did not
cascade a bump into the root ironclaw package. Add the root version
bump and CHANGELOG entry manually so this release-plz PR produces
an ironclaw-v0.28.1 tag and triggers cargo-dist.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(llm): hide provider-specific auth, model fetch, and embeddings config behind facades (#3416)

* refactor(llm): hide provider-specific auth, model fetch, and embeddings config behind facades

External callers were reaching into provider-specific modules of
`ironclaw_llm` (`gemini_oauth::CredentialManager`,
`github_copilot_auth::*`, `OpenAiCodexSessionManager`,
`codex_auth::*`, `BedrockConfig`, etc.). Closes those leaks behind a
small set of verb-based public surfaces while keeping per-provider
behaviour inside the LLM crate.

Changes:

1. Extract `oauth_helpers.rs` into a new `ironclaw_oauth` crate. The
   loopback OAuth callback listener (port 9876, landing pages,
   `OAUTH_CALLBACK_HOST` rules) is shared by every IronClaw OAuth flow
   (NEAR AI session login, WASM tool auth, MCP) and never depended on
   `ironclaw_llm`. `src/auth/oauth.rs` now `pub use ironclaw_oauth::*`
   directly. `ironclaw_llm` no longer depends on `ironclaw_oauth` —
   the helper had zero internal callers.

2. Add `ironclaw_llm::auth` facade (`start_login`, `validate_token`,
   `default_headers`, `load_persisted_credentials`,
   `default_credentials_path`) with backend-agnostic types
   (`AuthPrompt`, `LoginRequest`, `AuthOutcome`, `PersistedCredentials`,
   `OpenAiCodexLoginOptions`, `AuthBackend`, `CredentialSource`).
   Privatize `gemini_oauth`, `github_copilot_auth`, `openai_codex_session`,
   `codex_auth` (`pub(crate) mod`). Migrate the wizard, the
   `ironclaw login --openai-codex` CLI subcommand, and the LLM config
   loader to the facade. Wizard introduces a single `WizardAuthPrompt`
   that handles device-code prompts + browser launch for all backends.

3. Add `ironclaw_llm::models::fetch_models_for(provider_id, &opts)`
   facade. Privatize `fetch_anthropic_models`, `fetch_openai_models`,
   `fetch_ollama_models`, `fetch_openai_compatible_models`,
   `is_openai_chat_model`, `openai_model_priority`, `sort_openai_models`.
   Wizard's per-backend match collapses to one call. Move classifier
   unit tests into `crates/ironclaw_llm/src/models.rs`; rewrite the
   two wizard fallback tests through the public API.

4. Decouple embeddings from `ironclaw_llm::BedrockConfig`. New
   `crate::workspace::BedrockEmbeddingSetup { region, profile }` carries
   only what `BedrockEmbeddings` actually needs. `EmbeddingsConfig::create_provider`
   and `BedrockEmbeddings::new` take the new type; callers translate from
   `LlmConfig.bedrock` at the boundary (`src/app.rs`, `src/cli/mod.rs`).

5. Add `ironclaw_llm::testing::nearai_test_config(model)` helper for
   tests that need a minimal `LlmConfig` shape (no retries, no caching,
   NEAR AI backend). Replaces two duplicated 30-line struct literals
   in the gateway settings hot-reload tests.

Boundary cleanup is behaviour-preserving: 4,932 main-binary unit tests,
729 ironclaw_llm unit tests, 4 ironclaw_oauth tests, 3 architecture
boundary tests all pass; `cargo clippy --all --benches --tests
--examples --all-features` is clean.

Three `pub` methods on `gemini_oauth::CredentialManager` /
`GeminiOauthProvider` (`get_valid_access_token`, `last_response_meta`,
`count_tokens`) and the `GeminiResponseMeta` struct are now reachable
only crate-internally and have no callers; marked `#[allow(dead_code)]`
with a comment rather than deleted to keep this PR purely a boundary
move (delete in a follow-up if no caller emerges).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* refactor(llm): promote dedicated backends into the registry; absorb config validation, defaults, and per-provider overrides into ironclaw_llm

Continues the LLM boundary cleanup from 0addf3ac2. After that commit
provider-specific auth, model fetch, and embeddings config lived behind
facades inside `ironclaw_llm`, but four backend-specific knowledge
sources still leaked out:

  1. Validation rules and default values for the dedicated-config
     backends (Bedrock cross-region prefixes, OpenAI Codex endpoints
     and client_id, Gemini OAuth credentials path defaults) lived
     inline in `src/config/llm.rs::resolve`.
  2. The dispatcher in `create_llm_provider` matched on backend strings
     ("nearai", "bedrock", ...) instead of a typed protocol value. The
     same booleans (`is_nearai`, `is_bedrock`, `is_gemini_oauth`,
     `is_openai_codex`) recurred across `src/config/llm.rs`,
     `src/app.rs`, `src/cli/models.rs`, and the wizard.
  3. The setup wizard had per-backend specialization in
     `step_inference_provider` and `run_provider_setup` (manual menu
     pushes for nearai/bedrock/codex/gemini_oauth, four dedicated
     `setup_*` entry points dispatched on string compares).
  4. `Settings` carried named `bedrock_region`, `bedrock_cross_region`,
     `bedrock_profile` columns even though no other dedicated backend
     had named columns and adding a new one would mean schema churn.

Layers A-D address each in turn:

* Layer A — `BedrockConfig::build`, `OpenAiCodexConfig::build`, and
  `GeminiOauthConfig::build` own validation + defaults inside the
  crate. `LlmConfigError` (`MissingRequired` / `InvalidValue`) carries
  the failures across the boundary, with a `From` impl into the
  binary's `ConfigError`. `src/config/llm.rs` calls the builders;
  named-string defaults are gone from the binary. The orphaned
  `tests/gemini_oauth_regression.rs` husk is deleted.

* Layer B — `ProviderProtocol` gains four new variants
  (`Bedrock`, `OpenAiCodex`, `GeminiOauth`, `NearAi`) plus a
  `has_dedicated_config()` predicate. The four dedicated-config
  backends (with all aliases) become first-class registry entries in
  `providers.json`, so `is_known()` / `model_env_var()` / the wizard /
  the gateway handler iterate the registry uniformly. The
  `is_nearai`/`is_bedrock`/`is_gemini_oauth`/`is_openai_codex` boolean
  spaghetti collapses to protocol comparisons. `OpenAiCodex` and
  `NearAi` carry explicit `#[serde(rename = "openai_codex" / "nearai",
  alias = ...)]` so the wire-stable adapter strings the gateway and
  frontend already use keep working. `LlmConfig::active_model_name()`
  is now consumed by `cli/doctor.rs` instead of an inlined partial
  dispatch.

* Layer C — `SetupHint` gains four credential-collection variants
  (`AwsCredentials`, `OAuthDeviceCode`, `FileBasedCredentials`,
  `SessionToken`). The wizard's `step_inference_provider` builds its
  menu from a single `registry.selectable()` iteration with generic
  env-detection (declared `api_key_env`, plus an Anthropic-specific
  OAuth fallback). `run_provider_setup` dispatches on the SetupHint
  variant; the remaining `def.id == "..."` checks live inside the
  `ApiKey` arm only because Anthropic and GitHub Copilot present a
  hybrid choice (API key OR OAuth) the simple `ApiKey` hint doesn't
  capture. The synthetic bedrock + nearai entries in
  `handlers/llm.rs::build_llm_providers` are deleted; a single
  registry-driven loop covers both. ADAPTER_LABELS in
  `static/js/surfaces/config.js` gains entries for the new protocols.

* Layer D — `LlmBuiltinOverride` gains a generic
  `extras: HashMap<String, String>` bag with `extra(key)` /
  `set_extra(key, value)` accessors. The bedrock resolver and wizard
  read/write through this bag; `Settings::migrate_legacy_provider_fields()`
  drains the named `bedrock_*` columns into `extras` on
  `Settings::load_from()` so existing `settings.json` files migrate
  losslessly. The named columns are kept (deprecated, marked with
  `#[serde(skip_serializing_if = "Option::is_none")]`) for one
  release; tracked for deletion in #3443.
  `strip_admin_only_llm_keys` and `llm_setting_requires_reload` now
  match dotted-path subkeys under `llm_builtin_overrides.*` so a
  write to e.g. `llm_builtin_overrides.bedrock.extras.region`
  triggers the right gating + chain reload.

Boundary cleanup is behaviour-preserving: 4,933 main-binary unit tests,
739 ironclaw_llm unit tests pass; `cargo clippy --all --benches --tests
--examples --all-features` is clean. New regression tests:
`crates/ironclaw_llm/src/config.rs` (6 builder tests),
`crates/ironclaw_llm/src/registry.rs::dedicated_config_backends_are_in_registry_and_selectable`,
and `src/setup/wizard.rs::legacy_bedrock_fields_migrate_into_extras_on_load`.

Three follow-ups tracked in #3443: delete the deprecated `bedrock_*`
named columns, move `BedrockEmbeddings` out of `src/workspace/` into
the LLM crate (last cargo-feature leak), and drive
`LlmConfig::active_model_name()` off `ProviderProtocol` instead of
backend strings.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test: add bug-bash regression-snapshot harness

Bug-bash fixtures pin specific open bugs to a deterministic snapshot.
When a bug is fixed, the snapshot diff is the reviewable proof; when
someone reintroduces the bug, the snapshot drifts and CI blocks the
merge.

This commit lands the harness plus the first recorded fixture for
issue #2541 (agent must call a tool, not answer from training data):

  tests/e2e_bug_bash_snapshots.rs
    `snapshot_summarization_uses_tools` replays the fixture, captures
    `ReplayOutcome`, and asserts the YAML snapshot. Gated on
    `feature = "libsql"`, same as other replay-snapshot tests.

  tests/fixtures/llm_traces/bug_bash/summarization_uses_tools.json
    Two-step recorded LLM trace (tool_call -> text) keyed off the
    user prompt via `request_hint.last_user_message_contains`.

  tests/fixtures/llm_traces/bug_bash/README.md
    Coverage map for #2540-#2546 (one recorded, six TODO) plus the
    `IRONCLAW_RECORD_TRACE` recording workflow.

  tests/snapshots/replay__bug_bash_summarization_uses_tools.snap
    Insta YAML snapshot pinning `tool_calls: [echo]`, 2 LLM calls,
    and the event-kind histogram. Drift = regression.

* fix(settings): preserve pre-existing extras during legacy bedrock migration

`migrate_legacy_provider_fields` claimed to be idempotent and to drain
named `bedrock_*` columns into `llm_builtin_overrides["bedrock"].extras`
once on load. The previous implementation drained correctly but used
`HashMap::insert` unconditionally, which means a settings file
carrying BOTH a legacy `bedrock_region` column AND an already-populated
`extras["region"]` (manual hand-edit, or a future writer emitting both
shapes during a transition) would silently downgrade to the legacy
value.

Guard each `set_extra` call with `entry.extra(key).is_none()` so the
new-shape value always wins. Clarify the docstring to state this
explicitly.

Add three regression tests in `settings::tests`:

- `legacy_bedrock_migration_round_trips_through_save` — legacy JSON ->
  load_from -> serialize -> reload, asserts the deprecated columns are
  not re-emitted and extras survive the round trip.
- `legacy_bedrock_migration_preserves_existing_extras` — file with both
  shapes; asserts the pre-existing extras value is kept and absent
  extras are still backfilled from legacy fields.
- `legacy_bedrock_migration_is_idempotent_in_memory` — calling the
  migration twice on the same Settings is a no-op (compares serialized
  shape, since LlmBuiltinOverride does not derive PartialEq).

* fix(pr-3416): address PR review — migration on DB/TOML, admin-key gate, codex login, credential_kind/has_credentials

Addresses comments from gemini-code-assist, Copilot, and serrrfirat on PR #3416.

## Bugs

**Legacy bedrock fields not migrated on DB/TOML loads** (serrrfirat, High).
`Settings::load_from` (JSON) ran `migrate_legacy_provider_fields`, but
`from_db_map` and `load_toml` did not. Existing operators with
`bedrock_*` settings persisted in the DB or `config.toml` would silently
lose their AWS region/profile/cross-region after upgrade because the
resolver now reads only from `llm_builtin_overrides["bedrock"].extras`.
Both loaders now call the migration; added round-trip tests for each.

**Admin-only key write gate had narrower matching than read gate**
(Copilot, High). `strip_admin_only_llm_keys` matches both exact keys
and dotted subpaths under admin-only roots; `is_admin_only_setting_key`
in the web settings handler used `.contains(&key)` only. A non-admin
could write `llm_builtin_overrides.bedrock.extras.region` directly,
bypassing the gate. Promoted `is_admin_only_llm_key` to `pub(crate)`,
made the web write-side gate call it, added regression tests covering
dotted subpaths.

**`ironclaw login --openai-codex` dropped TOML/DB config** (Copilot,
High). The pre-refactor code resolved `Config::from_env` and used
`config.llm.openai_codex` so endpoint / client-id / session-path
overrides committed via TOML or DB stuck. The post-refactor code only
read env vars via `OpenAiCodexLoginOptions::from_env`. Added
`OpenAiCodexLoginOptions::from_resolved_config(&OpenAiCodexConfig)`;
the login command now prefers the resolved config when present and
falls back to env-only when `Config::from_env` itself fails (fresh
machine, no DB).

**Dedicated-auth backends marked configured without credentials**
(serrrfirat, Medium). `nearai` / `gemini_oauth` / `openai_codex` ship
`api_key_required: false` because they don't authenticate via a bearer
API key. The frontend `isProviderConfigured` treated that as "no
credentials needed" and rendered the Use button on a fresh install,
where clicking could trigger an interactive device-code OAuth from
inside a settings request.

Added `credential_kind` (wire-stable snake_case discriminator matching
`SetupHint::kind()`, e.g. `session_token`, `o_auth_device_code`,
`file_based_credentials`, `aws_credentials`) and `has_credentials`
(backend-authoritative; checks AWS env vars for Bedrock, codex session
file existence, file-based credential path expansion + existence) to
the web LLM providers payload. Frontend `isProviderConfigured` /
`providerMissingReason` now gate non-api-key kinds on `has_credentials`.

## Nits

**`fetch_models_for` doc overclaimed "Always returns something"**
(Copilot). The generic openai-compatible branch returns `vec![]` when
`base_url` is empty. Updated the docstring to call this out so callers
know to handle the empty case.

**`AuthError::Other` used for "validation not applicable"** (Gemini
bot). Added a dedicated `AuthError::TokenValidationNotSupported { backend }`
variant; `validate_token` now returns it for Gemini / OpenAiCodex
instead of stringly-formatted `Other`.

**Bug-bash regression-harness URLs pointed at `near/ironclaw`**
(Copilot, x2). The canonical tracker is `nearai/ironclaw`. Rewrote
all seven URLs in `tests/fixtures/llm_traces/bug_bash/README.md` and
the one in `tests/e2e_bug_bash_snapshots.rs`.

## Declined

The Gemini bot's MalformedConfig suggestion at
`crates/ironclaw_llm/src/models.rs:46` was not adopted: the call site is
the openai-compatible model-listing path, not a security-sensitive
request. The fetcher early-returns `vec![]` on empty `base_url` — no
URL parsing happens — and the docstring tightening above covers the
observable surprise. Promoting it to a typed error would change the
public-facing `fetch_models_for` signature for no behavioural gain.

## Tests

- `cargo fmt --check` clean
- `cargo clippy --all --benches --tests --examples --all-features` zero warnings
- `cargo test --lib` 4,941 / 4,941 pass
- `cargo test --features libsql --test e2e_bug_bash_snapshots` 1 / 1 pass
- New regression tests:
  - `settings::tests::legacy_bedrock_fields_migrate_into_extras_on_db_load`
  - `settings::tests::legacy_bedrock_fields_migrate_into_extras_on_toml_load`
  - `channels::web::features::settings::tests::test_admin_only_setting_keys_cover_dotted_subpaths`
  - `channels::web::handlers::llm::tests::test_llm_providers_expose_credential_kind_and_has_credentials`
  - `channels::web::handlers::llm::tests::test_nearai_has_credentials_true_when_session_token_loaded`

* fix(pr-3416): tighten Bedrock/Codex has_credentials probes; collapse set_extra into one .into()

- `backend_has_credentials` for AWS now requires `AWS_PROFILE` OR
  (`AWS_ACCESS_KEY_ID` AND `AWS_SECRET_ACCESS_KEY`). The lone
  `AWS_ACCESS_KEY_ID` / `AWS_SESSION_TOKEN` arms previously flipped
  has_credentials true even though the AWS SDK can't sign without the
  secret key, so the UI was rendering Bedrock as configured on hosts
  that would fail at first call.
- `backend_has_credentials` for OpenAI Codex now honours
  `OPENAI_CODEX_SESSION_PATH` via `read_env` before falling back to
  the default session path under `~/.ironclaw/`. Users with a custom
  session location were seeing "not configured" despite a valid login.
- New regression tests `test_bedrock_partial_aws_env_reports_not_configured`
  and `test_openai_codex_honours_session_path_env` drive the
  `build_llm_providers` call site (not just the helper) so both gaps
  stay closed.
- Tidied `LlmBuiltinOverride::set_extra` to convert the key once and
  reuse it across the remove/insert branches; behaviour identical.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore(providers): default nearai model to "auto"

Switch the nearai registry entry's `default_model` from
`claude-sonnet-4-5` to `auto`, NEAR AI's server-side routing alias.
New installs without `NEARAI_MODEL` set now get auto-routed instead
of being pinned to a specific Anthropic model.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Make Skills E2E lifecycle deterministic (#3309)

* test(e2e): make skills lifecycle deterministic

* test(e2e): address skills review comments (#3309)

* test(e2e): unxfail two auth-matrix tests now that contracts match (#3589)

Both xfails in tests/e2e/scenarios/test_v2_auth_oauth_matrix.py are
stale and pass against current code:

- test_wasm_tool_first_chat_auth_attempt_emits_auth_url
  Marked xfail in #3235 because the engine-v2 callable-only contract
  (#2868) stopped emitting an auth gate on direct LLM-driven tool
  calls. PR #3157 (auth-preflight + inline-await) restored the
  behavior the test asserts: when the LLM emits a direct call to a
  not-yet-authed extension, the bridge raises an Authentication gate
  with auth_url populated (src/bridge/effect_adapter.rs:1356-1392).
  Marker removed; test passes.

- test_settings_first_custom_mcp_auth_then_chat_runs
  The xfail reason claimed post-auth tool-output propagation was
  broken. Real cause: engine-v2 gates the first MCP tool call on
  `approval` and the browser fixture has no auto-approve UI, so the
  chat sat in pending_gate forever. Same shape as the bugs fixed in
  #3235 for test_wasm_tool_oauth_refresh_on_demand and
  test_mcp_same_server_multi_user_via_browser. Inserted
  _wait_for_tool_call between _send_chat and _wait_for_response_contains
  to drive approval through the API; test passes.

Verified locally: both tests pass back-to-back in 27s on a fresh
auth_matrix_server.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(extensions): restore chat-driven tool_install + fix double-invoke + auto-approve footgun (#3533) (#3559)

* fix(extensions): restore chat-driven tool_install + fix double-invoke + auto-approve footgun (#3533)

"Connect my telegram" was giving the user two options and not actually
installing anything because three layered issues had accumulated since
engine v2:

1. **`tool_install` was hidden from the agent** (#2868). The unified
   `tool_activate` it was meant to be subsumed by was later removed in
   #3166, but the hidden-from-callable-surface gate stayed. Restored
   by dropping `hidden_from_model_callable_surface` from
   `bridge::action_projector`. User consent is mediated by the tool's
   own `ApprovalRequirement::UnlessAutoApproved` and the seeded
   `AskEachTime` permission.

2. **Two competing Telegram registry entries** (`telegram` channel and
   `telegram_mtproto` tool) both surfaced in the agent prompt's
   `Activatable Integrations` section. The LLM correctly enumerated
   them as "Option 1" and "Option 2" instead of installing the
   canonical bot channel. Added a `hidden: bool` field to
   `ExtensionManifest` / `RegistryEntry`, set `telegram_mtproto` to
   `hidden: true`, and filter hidden entries out of the
   "available-but-not-installed" appendix in `ExtensionManager::list`.
   Hidden entries remain installable by explicit name.

3. **Updated the agent prompt** so `Activatable Integrations` instructs
   the model to call `tool_install(name="<name>")` directly rather than
   describing manual UI steps.

Fixes the double-`tool_install` invocation that surfaced once the agent
could install from chat:

- **`InlineGate` discarded cached output.** The bridge raised an
  Authentication gate after `tool_install` succeeded, and the
  inline-await retry re-executed the action (re-downloading the WASM
  bundle) instead of returning the already-computed output. Added
  `resume_output: Option<serde_json::Value>` to `InlineGate`; on
  approval, return the cached output if present. Mirror fix in the
  orchestrator's `execute_single_action_with_inline_retry` (reading
  `result_json["resume_output"]`) and the structured-batch retry path.
- **`effect_adapter::auth_gate_from_extension_result`** now passes
  `Some(output_value.clone())` as the gate's `resume_output` so the
  retry has cached state to short-circuit on.
- **OAuth callback double-fired.** `oauth_callback_handler` now skips
  the `ExternalCallback` re-entry when the inline-await path already
  woke a parked waiter — eliminates the "thread already running" race.
- **`resolve_inline_gates_for_credential`** now also discards matching
  Authentication rows from `pending_gates` so the row doesn't linger
  in `HistoryResponse.pending_gate` after inline resolution.

Fixes the auto-approve footgun:

- **`ToolPermissionSnapshot::resolve_permission`** now collapses DB
  values that match the seeded default to `explicit = None`. Before
  this, the boot-time `seed_tool_permissions` write of `tool_install ->
  AskEachTime` was indistinguishable from a user-explicit override,
  causing `effect_adapter::enforce_tool_permission`'s `is_explicit_ask`
  check to refuse `AGENT_AUTO_APPROVE_TOOLS=true`. Real overrides
  (`AlwaysAllow`, `Disabled`) still surface as `Some(...)`.

Tests
- Unit: 4980/4980 pass (host) + 525/525 pass (engine).
- Unit: new tests in `bridge::tool_permissions::tests` lock seeded-vs-
  explicit collapse; new test in `bridge::action_projector::tests`
  asserts `tool_install` is callable; new manifest hidden-flag tests
  in `registry::manifest::tests` and `extensions::manager::tests`.
- E2E: removed `@pytest.mark.xfail` on
  `test_chat_first_gmail_installs_prompts_and_retries` (now passes
  end-to-end via the chat-driven install path). Added
  `test_chat_install_approval_then_auth_card` driving the
  explicit-approval variant with a single Approve click (no Always
  workaround needed) — wired into the `auth-full` canary lane.
- Mock LLM: extended the gmail-install-then-retry pattern to recognize
  both the legacy "Extension not installed:" and the post-#3533 "is
  not callable in this execution context" error strings, and to retry
  `gmail(action="list_messages")` after a successful `tool_install`.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(permissions): address #3559 review (permission bypass, lease accounting, hidden search filter)

Five fixes from the #3559 review (4× Copilot doc nits + 3× serrrfirat
security/correctness findings):

1. **Permission bypass (High).** Pre-#3559's `resolve_permission`
   collapsed any DB row whose value matched the seeded default to
   `explicit = None`, so a user who deliberately set `tool_install =
   AskEachTime` had their explicit choice silently dropped and
   `AGENT_AUTO_APPROVE_TOOLS=true` bypassed the gate. Provenance is now
   handled at write time: `seed_tool_permissions` is gone and a
   one-shot, sentinel-gated migration (`cleanup_ghost_seeded_tool_permissions`)
   deletes existing ghost-seeded rows at startup. With no ghost rows,
   the resolver treats every DB row as user-explicit and honors it.

2. **Lease/event accounting on `resume_output` replay (Medium).**
   Inline-gate handlers in `structured.rs`, `scripting.rs`
   (`resolve_tool_future` + `drive_inline_gate` retry loop), and
   `orchestrator.rs` refunded the lease use the action just consumed,
   then returned the cached `resume_output` on approval without
   re-consuming — netting successful side-effecting actions to zero
   lease uses. Skip the refund when the gate carries cached output.

3. **Hidden registry filter on `tool_search` (Medium).**
   `RegistryCatalog::search` did not filter `hidden: true` entries,
   so `telegram_mtproto` could resurface through the search path and
   reintroduce the "two Telegram options" outcome that #3533 fixes
   for the default-list path. Added the filter and a regression test.

4-7. Copilot doc nits: outdated `_set_tool_permission` docstring;
   misleading "bridge-side auto-install implemented" comment in
   `mock_llm.py`; `tool_install` described as "non-agent surface" in
   `src/bridge/CLAUDE.md` while a paragraph below says the model
   calls it directly; dangling `issue #3533 / PR —` placeholders
   in both CLAUDE.md docs.

Regression tests:
- `bridge::tool_permissions::user_explicit_value_matching_seeded_default_stays_explicit` —
  the original Copilot/serrrfirat bug case.
- `app::cleanup_ghost_seeded_tool_permissions_removes_seed_matching_rows` —
  idempotent migration + sentinel.
- `extensions::registry::test_search_skips_hidden_entries` — hidden
  entries excluded from search but still installable by exact name.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(#3559): caller-level regression coverage for review findings 1 & 2

Two follow-up regression tests for the #3559 security review, plus a
real bug surfaced by the first one.

1. `executor::structured::resume_output_replay_consumes_exactly_one_lease_use`
   exercises the post-execution Authentication gate inline-retry path
   with `max_uses=1` and asserts:
   - Cached output is returned as a successful `ActionResult`.
   - Exactly one `ActionExecuted` event is emitted.
   - The lease budget is exhausted after one execution (refund-skip
     keeps the consumption from being undone).

   Writing this test surfaced a real bug: the structured cached-output
   branch pushed `ActionExecuted` into `emitted_events`, and the
   caller's `classify_exec_result` emitted ANOTHER terminal
   `ActionExecuted` for the same Ok result — double-emit for one
   action. Tier 1 (`scripting::drive_inline_gate`) and Tier 1 alt
   (`orchestrator::execute_action_with_inline_gate`) emit themselves
   because their callers don't run an Ok-branch classifier; structured
   was the outlier. Dropped the redundant push; the classifier emits
   the single canonical event.

2. `bridge::effect_adapter::explicit_ask_each_time_for_seeded_default_tool_still_gates`
   drives `execute_action` end-to-end (the side-effecting caller) with
   a tool whose `name()` matches a seeded-`AskEachTime` baseline
   (`tool_install`) and an explicit `AskEachTime` user override. The
   resolver collapse-to-implicit bug would have shown up here — not
   just in the helper-level test that already exists in
   `bridge::tool_permissions::tests`. Per `.claude/rules/testing.md`
   "Test Through the Caller, Not Just the Helper".

   Added `SeededAskEachTimeTestTool` as a `tool_install`-named test
   fixture with `requires_approval: UnlessAutoApproved` to mirror the
   real tool's contract.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* chore: release

* feat(engine): IRONCLAW_DISABLE_CODEACT flag to disable v2 CodeAct (#3665)

* add flag to disable codeact on engine v2

* fmt

* fix(engine): keep compact actions reachable when CodeAct is disabled

With IRONCLAW_DISABLE_CODEACT=true the structured-tools prompt told
the model to use the provider's tool_calls interface for every action,
but the bridge filtered the provider tool list down to
emits_full_schema_tool(). Most tools default to CompactToolInfo
(mission_create, gmail_send, notion_search, ...), so they appeared in
the prompt as "available" while being absent from the provider tool
list — i.e. unreachable. Addresses serrrfirat's review on PR #3665.

Fix coordinates both halves of the surface:

- src/bridge/llm_adapter.rs: in disabled-CodeAct mode, drop the
  emits_full_schema_tool() filter and emit every action into the
  provider tool list with its full schema.
- crates/ironclaw_engine/src/executor/prompt.rs: in disabled-CodeAct
  mode, skip the "## Enabled Tools" section. The compact-form listing
  with the tool_info(detail="schema") instruction is meaningless when
  the provider already sends full schemas, and would just duplicate
  the surface. "## Activatable Integrations" stays — the model still
  needs to know what tool_install can target.

Test seam: build_codeact_system_prompt_inner now takes disable_codeact
as an explicit parameter, called once at the public entry points. This
lets prompt tests exercise both branches without process-global env
mutation.

Tests:
- executor::prompt::tests::disabled_codeact_omits_enabled_tools_section_and_keeps_activatable
- bridge::llm_adapter::tests::complete_emits_compact_actions_when_codeact_disabled
- existing complete_with_tools_only_emits_full_schema_provider_tools
  now serialized via lock_env() so env mutation in the new test
  can't leak across parallel runs.

cargo test -p ironclaw_engine --lib: 527 passed
cargo test --lib bridge::: 469 passed
cargo clippy -p ironclaw_engine --all-targets -- -D warnings: clean
cargo clippy --lib --tests -- -D warnings: clean
cargo fmt --check: clean

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Emil Bogomolov <emil.bogomolov@near.ai>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* Fix markdown_to_mrkdwn to avoid converting emphasis inside generated <… (#3532)

* agent: Fix markdown_to_mrkdwn to avoid converting emphasis inside g…

* agent: Fix rustfmt/clippy CI failure by removing extra blank line b…

* agent: slack: fix markdown_to_mrkdwn replacement order to satisfy p…

* slack: protect generated links and sanitize sentinels in markdown_to_mrkdwn

Two issues raised on PR #3532 review:

1. Emphasis inside generated `<url|text>` was still rewritten because the
   global `**`/`~~` → `*`/`~` substitution ran after link materialization.
   Push the generated link span into the same protected arena used for
   Slack-native `<...>` constructs so subsequent global replacements can't
   reach inside it. Matches the PR's stated goal.

2. Untrusted input containing the private-use sentinel chars
   (U+E000 / U+E001) could forge a protected-span reference and pull in
   another span's content. Strip those chars from input up front.

Adds regression tests for both. Bumps registry/channels/slack.json
0.3.2 → 0.3.3 to satisfy the channel-source version-bump check.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* slack: escape link labels, drop pipe/gt URLs, expand nested sentinels

Addresses two follow-up review concerns on PR #3532:

Copilot review: `<url|text>` was built by string concatenation, so
`|` or `>` inside the URL would corrupt the entity, and `<` / `>`
inside the label would open/close a Slack span and break the link.
The label now escapes `<` → `&lt;` and `>` → `&gt;` (Slack's documented
literal-character form); a URL containing `<`, `>`, or `|` falls back
to leaving the original markdown form intact (those chars are not
valid URL characters per RFC 3986 anyway).

Latent nested-sentinel bug introduced by the previous fix: a markdown
link whose label contained a Slack-native `<...>` span (e.g.
`[<@U1> hi](url)`) ended up with the inner sentinel buried inside the
arena entry for the outer link span. The final restore pass advances
past the outer sentinel without rescanning what it just emitted, so
the raw U+E000/U+E001 characters would leak into the output. URL and
label are now pre-expanded before the link span is pushed.

While here, factor the duplicated restore loop into
`expand_protected_spans`, reused by both the pre-link expansion and
the final restore, and lift the sentinel constants to file scope.

Adds three regression tests covering label-bracket escaping, the URL
pipe/gt fallback, and the nested-span case. Bumps
registry/channels/slack.json 0.3.3 → 0.3.4.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* feat(gateway): add logs download button (#3588)

* feat(web): support externally-provided tools in Responses API (#3122)

* feat(web): support externally-provided tools in Responses API

Lets callers of `/v1/responses` (and `/api/v1/responses`) declare their
own `function`-typed tools and feed back results via
`function_call_output` items, matching the OpenAI Responses wire shape.

Since IronClaw's engine has no per-request tool surface, integration
happens at the prompt level: the catalog is rendered as
`<external-tools>` in the user message and the agent signals a call by
ending its response with a fenced ```` ```tool_call ```` block. When
that fence is recognised, the reply is split into a leading `Message`
plus a `function_call` `ResponseOutputItem`.

Validation rejects unsupported tool types (`web_search`, `file_search`,
`code_interpreter`) and tools missing `name` with 400, with two new
integration tests covering both paths.

* refactor(responses-api): switch external tools to engine v2 native path

Replace the prompt-level fence protocol from PR #3122 with engine v2
native tool calls: caller-supplied `tools[]` are surfaced as real
LLM-callable actions, the engine pauses with `ResumeKind::External`
when one is invoked, and the bridge router projects the pause to a
new `AppEvent::ExternalToolCall` carrying the OpenAI-shaped
`function_call` wire fields.

The integration is small because v2 already has the right primitives:

- `ResumeKind::External { callback_id }` and
  `GateResolution::ExternalCallback { payload }` already existed for
  OAuth-style callbacks.
- `agent_loop.rs:1480` already routes Responses API messages to
  `handle_with_engine` when `ENGINE_V2=true`, so no v2 migration of
  the endpoint itself is needed.
- `EffectBridgeAdapter::execute_action` is the single chokepoint
  where caller tools can be detected before they reach the dispatch
  pipeline.

Changes:

- New `src/bridge/external_tools.rs` (`ExternalToolCatalog`) — per-thread
  registry of caller-supplied `ActionDef`s, plus the `ext_tool:`
  callback-id helpers used to disambiguate external-tool pauses from
  OAuth/pairing pauses (which also use `ResumeKind::External`).
- `EffectBridgeAdapter` consults the catalog: any name in it is
  short-circuited to a `GatePaused { resume_kind: External {
  callback_id: ext_tool:<call_id> } }` before any registry dispatch,
  and `available_action_inventory` merges the catalog into the
  LLM-visible action surface (internal beats external on collision).
- `Submission::ExternalCallback` gains an optional `payload` field;
  `bridge::handle_external_callback` plumbs it into
  `GateResolution::ExternalCallback { payload }`. Fallback predicate
  `gate_resume_is_external` lets non-auth External pauses (i.e.
  caller-tool resumes) resolve through the same handler.
- New `AppEvent::ExternalToolCall` projected by `notify_pending_gate`
  when a paused gate carries an `ext_tool:` callback id; OAuth/
  pairing flows keep flowing through the existing `GateRequired`
  channel.
- `responses_api.rs` is gutted of the prompt rendering and fence
  parsing (`render_external_tools_preamble`, `extract_trailing_tool_call`,
  `parse_external_tool_call`, `ParsedToolCall`, `external_tool_names`
  accumulator field, and the `TOOL_CALL_FENCE` constants). The handler
  now: rejects `tools[]` when `ENGINE_V2=false`, registers caller
  tools in the catalog under the resolved thread id, detects resume
  requests (`previous_response_id` + `function_call_output` items in
  `input`) and submits them as `Submission::ExternalCallback` with
  the outputs as the resolution payload, and surfaces
  `AppEvent::ExternalToolCall` as a `function_call` `ResponseOutputItem`
  in both streaming (`output_item.added`+`done`) and non-streaming.
- All existing OAuth/pairing `ExternalCallback` constructors updated
  to pass `payload: None` (no behaviour change).
- Fence-protocol unit tests removed; replaced with coverage for the
  new `responses_tools_to_action_defs` converter and the accumulator's
  `ExternalToolCall` arm.

Existing 9 integration tests in `tests/responses_api_path_prefix.rs`
still pass.

Note for reviewers:
- The accumulator-side text response no longer tries to split the
  reply on a fenced `tool_call` block. The wire shape that callers
  receive for caller-tool invocations is purely event-driven now.
- Internal vs external collision is handled silently by the dedup in
  `available_action_inventory` (internal wins). A request-time
  rejection for shadowing names is a follow-up — the current behavior
  is safe (the LLM only sees the internal version) but could surprise
  a caller who expects their tool to run.

* test(responses-api): cover ENGINE_V2-off and resume-without-pending-gate

Two new integration tests for behaviours added by the engine-native
external-tool refactor:

- `external_tools_rejected_when_engine_v2_disabled`: a request with
  caller-supplied `tools[]` while `ENGINE_V2` is off must 400 with a
  message naming the flag, not silently fall through.
- `resume_without_pending_gate_returns_400`: a request with
  `function_call_output` items and a `previous_response_id` that
  doesn't correspond to a live external-tool gate must 400, not start
  a fresh turn against the (unrelated) thread.

Both tests drive the full router (`start_test_server` + bearer auth)
per `.claude/rules/testing.md` "Test Through the Caller".

* test(responses-api): integration tests + drop unsafe env mutation

Three groups of changes:

1. **Drop unsafe env-var mutation in tests.** `responses_api.rs` no
   longer reads `ENGINE_V2` directly: it keys off the presence of the
   live `ExternalToolCatalog` (initialized by `init_engine`) as the
   "engine v2 is up" signal. The path-prefix test that exercises the
   no-engine branch no longer needs `unsafe { std::env::remove_var }`
   — the absence of `init_engine` in `TestGatewayBuilder` is what
   makes the catalog absent, which is what makes the request reject.

2.…
theredspoon pushed a commit to theredspoon/ironclaw that referenced this pull request Jun 21, 2026
…arai#3589)

Both xfails in tests/e2e/scenarios/test_v2_auth_oauth_matrix.py are
stale and pass against current code:

- test_wasm_tool_first_chat_auth_attempt_emits_auth_url
  Marked xfail in nearai#3235 because the engine-v2 callable-only contract
  (nearai#2868) stopped emitting an auth gate on direct LLM-driven tool
  calls. PR nearai#3157 (auth-preflight + inline-await) restored the
  behavior the test asserts: when the LLM emits a direct call to a
  not-yet-authed extension, the bridge raises an Authentication gate
  with auth_url populated (src/bridge/effect_adapter.rs:1356-1392).
  Marker removed; test passes.

- test_settings_first_custom_mcp_auth_then_chat_runs
  The xfail reason claimed post-auth tool-output propagation was
  broken. Real cause: engine-v2 gates the first MCP tool call on
  `approval` and the browser fixture has no auto-approve UI, so the
  chat sat in pending_gate forever. Same shape as the bugs fixed in
  nearai#3235 for test_wasm_tool_oauth_refresh_on_demand and
  test_mcp_same_server_multi_user_via_browser. Inserted
  _wait_for_tool_call between _send_chat and _wait_for_response_contains
  to drive approval through the API; test passes.

Verified locally: both tests pass back-to-back in 27s on a fresh
auth_matrix_server.

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
theredspoon pushed a commit to theredspoon/ironclaw that referenced this pull request Jun 21, 2026
… + auto-approve footgun (nearai#3533) (nearai#3559)

* fix(extensions): restore chat-driven tool_install + fix double-invoke + auto-approve footgun (nearai#3533)

"Connect my telegram" was giving the user two options and not actually
installing anything because three layered issues had accumulated since
engine v2:

1. **`tool_install` was hidden from the agent** (nearai#2868). The unified
   `tool_activate` it was meant to be subsumed by was later removed in
   nearai#3166, but the hidden-from-callable-surface gate stayed. Restored
   by dropping `hidden_from_model_callable_surface` from
   `bridge::action_projector`. User consent is mediated by the tool's
   own `ApprovalRequirement::UnlessAutoApproved` and the seeded
   `AskEachTime` permission.

2. **Two competing Telegram registry entries** (`telegram` channel and
   `telegram_mtproto` tool) both surfaced in the agent prompt's
   `Activatable Integrations` section. The LLM correctly enumerated
   them as "Option 1" and "Option 2" instead of installing the
   canonical bot channel. Added a `hidden: bool` field to
   `ExtensionManifest` / `RegistryEntry`, set `telegram_mtproto` to
   `hidden: true`, and filter hidden entries out of the
   "available-but-not-installed" appendix in `ExtensionManager::list`.
   Hidden entries remain installable by explicit name.

3. **Updated the agent prompt** so `Activatable Integrations` instructs
   the model to call `tool_install(name="<name>")` directly rather than
   describing manual UI steps.

Fixes the double-`tool_install` invocation that surfaced once the agent
could install from chat:

- **`InlineGate` discarded cached output.** The bridge raised an
  Authentication gate after `tool_install` succeeded, and the
  inline-await retry re-executed the action (re-downloading the WASM
  bundle) instead of returning the already-computed output. Added
  `resume_output: Option<serde_json::Value>` to `InlineGate`; on
  approval, return the cached output if present. Mirror fix in the
  orchestrator's `execute_single_action_with_inline_retry` (reading
  `result_json["resume_output"]`) and the structured-batch retry path.
- **`effect_adapter::auth_gate_from_extension_result`** now passes
  `Some(output_value.clone())` as the gate's `resume_output` so the
  retry has cached state to short-circuit on.
- **OAuth callback double-fired.** `oauth_callback_handler` now skips
  the `ExternalCallback` re-entry when the inline-await path already
  woke a parked waiter — eliminates the "thread already running" race.
- **`resolve_inline_gates_for_credential`** now also discards matching
  Authentication rows from `pending_gates` so the row doesn't linger
  in `HistoryResponse.pending_gate` after inline resolution.

Fixes the auto-approve footgun:

- **`ToolPermissionSnapshot::resolve_permission`** now collapses DB
  values that match the seeded default to `explicit = None`. Before
  this, the boot-time `seed_tool_permissions` write of `tool_install ->
  AskEachTime` was indistinguishable from a user-explicit override,
  causing `effect_adapter::enforce_tool_permission`'s `is_explicit_ask`
  check to refuse `AGENT_AUTO_APPROVE_TOOLS=true`. Real overrides
  (`AlwaysAllow`, `Disabled`) still surface as `Some(...)`.

Tests
- Unit: 4980/4980 pass (host) + 525/525 pass (engine).
- Unit: new tests in `bridge::tool_permissions::tests` lock seeded-vs-
  explicit collapse; new test in `bridge::action_projector::tests`
  asserts `tool_install` is callable; new manifest hidden-flag tests
  in `registry::manifest::tests` and `extensions::manager::tests`.
- E2E: removed `@pytest.mark.xfail` on
  `test_chat_first_gmail_installs_prompts_and_retries` (now passes
  end-to-end via the chat-driven install path). Added
  `test_chat_install_approval_then_auth_card` driving the
  explicit-approval variant with a single Approve click (no Always
  workaround needed) — wired into the `auth-full` canary lane.
- Mock LLM: extended the gmail-install-then-retry pattern to recognize
  both the legacy "Extension not installed:" and the post-nearai#3533 "is
  not callable in this execution context" error strings, and to retry
  `gmail(action="list_messages")` after a successful `tool_install`.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(permissions): address nearai#3559 review (permission bypass, lease accounting, hidden search filter)

Five fixes from the nearai#3559 review (4× Copilot doc nits + 3× serrrfirat
security/correctness findings):

1. **Permission bypass (High).** Pre-nearai#3559's `resolve_permission`
   collapsed any DB row whose value matched the seeded default to
   `explicit = None`, so a user who deliberately set `tool_install =
   AskEachTime` had their explicit choice silently dropped and
   `AGENT_AUTO_APPROVE_TOOLS=true` bypassed the gate. Provenance is now
   handled at write time: `seed_tool_permissions` is gone and a
   one-shot, sentinel-gated migration (`cleanup_ghost_seeded_tool_permissions`)
   deletes existing ghost-seeded rows at startup. With no ghost rows,
   the resolver treats every DB row as user-explicit and honors it.

2. **Lease/event accounting on `resume_output` replay (Medium).**
   Inline-gate handlers in `structured.rs`, `scripting.rs`
   (`resolve_tool_future` + `drive_inline_gate` retry loop), and
   `orchestrator.rs` refunded the lease use the action just consumed,
   then returned the cached `resume_output` on approval without
   re-consuming — netting successful side-effecting actions to zero
   lease uses. Skip the refund when the gate carries cached output.

3. **Hidden registry filter on `tool_search` (Medium).**
   `RegistryCatalog::search` did not filter `hidden: true` entries,
   so `telegram_mtproto` could resurface through the search path and
   reintroduce the "two Telegram options" outcome that nearai#3533 fixes
   for the default-list path. Added the filter and a regression test.

4-7. Copilot doc nits: outdated `_set_tool_permission` docstring;
   misleading "bridge-side auto-install implemented" comment in
   `mock_llm.py`; `tool_install` described as "non-agent surface" in
   `src/bridge/CLAUDE.md` while a paragraph below says the model
   calls it directly; dangling `issue nearai#3533 / PR —` placeholders
   in both CLAUDE.md docs.

Regression tests:
- `bridge::tool_permissions::user_explicit_value_matching_seeded_default_stays_explicit` —
  the original Copilot/serrrfirat bug case.
- `app::cleanup_ghost_seeded_tool_permissions_removes_seed_matching_rows` —
  idempotent migration + sentinel.
- `extensions::registry::test_search_skips_hidden_entries` — hidden
  entries excluded from search but still installable by exact name.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(nearai#3559): caller-level regression coverage for review findings 1 & 2

Two follow-up regression tests for the nearai#3559 security review, plus a
real bug surfaced by the first one.

1. `executor::structured::resume_output_replay_consumes_exactly_one_lease_use`
   exercises the post-execution Authentication gate inline-retry path
   with `max_uses=1` and asserts:
   - Cached output is returned as a successful `ActionResult`.
   - Exactly one `ActionExecuted` event is emitted.
   - The lease budget is exhausted after one execution (refund-skip
     keeps the consumption from being undone).

   Writing this test surfaced a real bug: the structured cached-output
   branch pushed `ActionExecuted` into `emitted_events`, and the
   caller's `classify_exec_result` emitted ANOTHER terminal
   `ActionExecuted` for the same Ok result — double-emit for one
   action. Tier 1 (`scripting::drive_inline_gate`) and Tier 1 alt
   (`orchestrator::execute_action_with_inline_gate`) emit themselves
   because their callers don't run an Ok-branch classifier; structured
   was the outlier. Dropped the redundant push; the classifier emits
   the single canonical event.

2. `bridge::effect_adapter::explicit_ask_each_time_for_seeded_default_tool_still_gates`
   drives `execute_action` end-to-end (the side-effecting caller) with
   a tool whose `name()` matches a seeded-`AskEachTime` baseline
   (`tool_install`) and an explicit `AskEachTime` user override. The
   resolver collapse-to-implicit bug would have shown up here — not
   just in the helper-level test that already exists in
   `bridge::tool_permissions::tests`. Per `.claude/rules/testing.md`
   "Test Through the Caller, Not Just the Helper".

   Added `SeededAskEachTimeTestTool` as a `tool_install`-named test
   fixture with `requires_approval: UnlessAutoApproved` to mirror the
   real tool's contract.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: medium Business logic, config, or moderate-risk modules scope: agent Agent core (agent loop, router, scheduler) scope: channel/web Web gateway channel scope: docs Documentation scope: extensions Extension management scope: llm LLM integration scope: tool/builtin Built-in tools scope: tool/wasm WASM tool sandbox scope: tool Tool infrastructure size: XL 500+ changed lines

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants