engine-v2: make available_actions callable-only for blocked providers - #2868
Conversation
There was a problem hiding this comment.
Pull request overview
Aligns engine-v2 available_actions() with the “callable-only” surface policy by filtering out extension-backed provider actions when the provider is installed-but-blocked (needs auth/setup/inactive), and adds coverage to prevent regressions.
Changes:
- Removed the special-case bypass that kept
NeedsAuthprovider tools inavailable_actionswithinActionProjector. - Added
ActionProjectortests covering omission forNeedsAuth,NeedsSetup,Inactive, and routed-only channels. - Added
EffectBridgeAdapter-level tests using an installed-provider fixture to assert blocked installed provider actions are omitted fromavailable_actions.
Reviewed changes
Copilot reviewed 2 out of 2 changed files in this pull request and generated 1 comment.
| File | Description |
|---|---|
src/bridge/effect_adapter.rs |
Adds adapter-level fixture/tests for omitting installed-but-blocked provider actions from available_actions. |
src/bridge/action_projector.rs |
Removes NeedsAuth bypass and expands projector-level tests for blocked/routed providers. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
There was a problem hiding this comment.
Code Review
This pull request updates the ActionProjector to omit tools that are inactive or require authentication or setup from the available_actions list. It also introduces comprehensive test helpers and integration tests in action_projector.rs and effect_adapter.rs to verify these filtering rules. A review comment identifies a resource leak in a test helper where std::mem::forget prevents temporary directory cleanup, suggesting a refactor to return the TempDir object instead.
* fix(engine): align prompt metadata refresh with resume state * fix(engine): finish prompt refresh compaction coverage (#2869) * fix(engine): preserve prompt refresh on resume (#2869) * Add engine v2 action discovery metadata (#2876) * Add engine v2 action discovery metadata * fix(engine): address action discovery review (#2876) * fix(engine): address follow-up review comments (#2876) * fix(engine): satisfy clippy in orchestrator lookup * fix(engine): propagate action snapshots in executor paths (#2876) * fix(bridge): restrict tool_info to callable actions (#2876) * [codex] Finish engine v2 deferred action inventory cleanup (#2889) * Add deferred action inventory groundwork * fix(engine): address deferred action inventory follow-up * fix(engine): address deferred inventory review feedback * test: fix fmt and clippy failures
There was a problem hiding this comment.
Pull request overview
Copilot reviewed 39 out of 39 changed files in this pull request and generated 3 comments.
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
# Conflicts: # src/bridge/action_projector.rs
henrypark133
left a comment
There was a problem hiding this comment.
Review: callable-only surface and prompt/inventory alignment look good
I did a full pass on the stacked engine-v2 changes here and did not find any verified blocker-level issues.
What looks good:
- The callable-surface contract is much cleaner now: blocked provider actions move out of
available_actions, while background capability state stays model-visible through the canonical capabilities section. - The snapshot plumbing is materially better.
tool_info, structured execution, scripting, and orchestrator paths now read the same callable snapshot instead of drifting on alias handling or live registry state. - The prompt refresh work is directionally right. Resume/compaction now rebuild the engine-owned system prompt while preserving appended step-zero context like prior knowledge and active skills.
- The new caller-level regression coverage is the right shape for this area, especially around resume/compaction behavior and
tool_infonot mutating the next-step callable set.
Low-priority notes:
refresh_system_prompt()still fetchesavailable_action_inventory()even though the current prompt builder ignores that inventory. That looks like avoidable start/resume work unless a deferred inventory section is about to land.project_tool_action()currently emits discovery metadata for every callable tool even when it only repeats the callable name and has no summary/schema override. Tightening that would trim some per-turn payload size.
Verification:
cargo test --test engine_v2_gate_integration tool_info_does_not_gate_callable_tool_into_next_llm_callable_set -- --nocapturecargo test --test engine_v2_skill_codeact skill_prompt_context_survives_pause_and_resume -- --nocapturecargo test --test engine_v2_skill_codeact skill_prompt_context_survives_compaction_and_resume -- --nocapture
…ronclaw into v2-engine-callable-only-cleanup
Same pipe-deadlock fix as scripts/live_canary/common.py f59981d, applied to tests/e2e/scenarios/test_v2_auth_oauth_matrix.py's _start_auth_matrix_server. The auth-matrix fixture spawns ironclaw with stdout=PIPE + stderr=PIPE and never drains them, so under sustained log volume the kernel pipe buffer fills, ironclaw blocks on its next stdout write, and any test that relies on subsequent gateway responses (auth gate emission, SSE events, chat replies) hangs until pytest-timeout fires. This fix doesn't make the auth-full lane's failing test pass — the real bug is engine-v2 silently dropping `auth_required` SSE events for unauthenticated extensions (introduced by #2868). But it makes the failure mode debuggable: gateway log is captured to /tmp/ironclaw-auth-matrix-gateway.log (overridable via IRONCLAW_AUTH_MATRIX_LOG env), and RUST_LOG passes through from the test runner so we can crank up verbosity without rebuilding. Without this change, the failing test's log was empty after the extension-install line; with this change you see the engine-v2 trace summary that surfaces the actual NotCallable-without-auth-gate bug. That diagnostic visibility is the value here. - _drain_stream_to_file: asyncio drainer mirroring common.py's sync threading version - _start_auth_matrix_server: drain stdout/stderr to log_path - _shutdown_auth_matrix_server: cancel drain_tasks for clean exit - env: RUST_LOG forwarding so debug runs work Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…+ GH issues
Three additions to scripts/live-canary/notify_slack.py to make the
6h Slack report actionable instead of just informational:
1) **Per-lane rich failure block** — Haiku now extracts four
structured fields when status==fail: test_name, error, root_cause,
fix. The Slack section renders them in the issue-friendly shape
the reviewer asked for:
❌ auth-full (mock) — 11/13 passed, 1 failed in 213s
Test: `test_wasm_tool_first_chat_auth_attempt_emits_auth_url`
Error: SSE stream closed; auth_required event never arrived
Root Cause: bridge gate not wired for installed-but-unauthed
extensions (#2868 fallout)
Fix: route Extension::NeedsAuth through effect_adapter.rs
For passing/skipped lanes the existing single-line `> reason` is
preserved so the green-path Slack output is unchanged.
2) **Cross-lane "Summary by Category" block** — second Haiku pass
over all failed-lane summaries that groups them by shared root
cause (e.g. "WASM tool dispatch regression — Auth Full, Auth
Smoke, Auth Live Seeded"). Only fires when there are 2+
failures (single-failure runs are already obvious from the
per-lane block). Rendered as a Slack mrkdwn bulleted list since
Block Kit doesn't support real tables.
3) **Auto-opened GitHub issues** — opt-in via CANARY_CREATE_ISSUES=1
env var (gated to scheduled runs only in live-canary.yml so
workflow_dispatch debugging doesn't flood the tracker). For each
failed lane:
- Search for an OPEN issue with title `[canary] <lane>: <test>`.
- If found: comment "another occurrence on <run_url>".
- If not found: open a new issue with the rich body + labels
`canary-failure` + `lane:<lane>`.
Strategy chosen to avoid issue spam while still surfacing
recurring failures. Uses GITHUB_TOKEN + the repo's existing
`permissions: issues: write` block — no new secrets.
All three additions degrade silently — Haiku failure stamps
.notable but doesn't block the post; categorization failure produces
an "_(unavailable)_" placeholder; issue-creation errors are logged
to stderr only. The notifier still exits 0 in every failure path so
a flaky webhook can't fail the canary run.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(oauth): remove pending flow on provider-error callback
The /oauth/callback handler's ?error= branch (RFC 6749 §4.1.2.1
provider-side failures — user cancels consent, scope denied, etc.)
returned the error page immediately without removing the flow from
ext_mgr.pending_oauth_flows(). The ghost entry then lingered until
the 5-minute expiry sweep, and any subsequent auth dance for the
same (extension, user) pair had to dedupe against it.
Mirror the happy-path cleanup: decode the state param, remove the
keyed flow, then return the error page.
Surfaced during live-canary auth-full repro: after
test_wasm_tool_oauth_provider_error_leaves_extension_unauthed ran,
the stale flow sat in the shared auth_matrix_server fixture.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(e2e): widen auth OAuth matrix timeouts for CI load
Four tests in live-canary auth-full were failing in CI with
`Page.wait_for_function: Timeout 60000ms exceeded`,
`ClientConnectionError('Connection closed')`, and
`Timed out waiting for OAuth refresh request` — all inside 60/20s
deadlines that are tuned for a dev laptop and don't leave margin
for ubuntu-latest's 2-vCPU runner under full suite load.
Raise the per-call deadlines so the inner budgets fit comfortably
inside pyproject.toml's 120s per-test cap:
_wait_for_refresh_request default: 20.0s -> 60.0s
_wait_for_auth_event call site: 60 -> 90
_wait_for_auth_prompt call site: 60 -> 90
send_chat_and_wait_for_terminal_message call sites: 60000 -> 90000
_wait_for_mock_google_tokens call site: 60.0 -> 90.0
_wait_for_response_contains (gmail) call site: 60.0 -> 90.0
Strictly widening; no passing test is slowed, no semantics change.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(canary): Haiku-powered Slack report job
Replace the team's raw Slack subscription (firehose of workflow
notifications) with one curated per-run summary:
Canary: 9 passed, 1 failed of 10 lanes
:x: auth-full (mock) — 12/13 passed, 1 failed in 350s
> test_wasm_tool_first_chat_auth_attempt_emits_auth_url timed
> out waiting for auth_required SSE event on the fresh thread
tools: shell, http_request, gmail (~6 calls)
...
commit `abc1234` • <github run link>
New `canary-report` job (needs: every lane, if: always) downloads
all lane artifacts, parses junit + summary + log tail per lane, and
asks claude-haiku-4-5 to return a compact JSON per lane
({status, reason, tool_calls_total, tools_used, notable}). That's
aggregated into a single Slack block message and posted via
incoming webhook.
Safety shape:
- Script exits 0 even on Haiku/Slack failure so the notifier never
masks the underlying canary signal.
- Missing ANTHROPIC_API_KEY falls back to raw junit-only phrasing.
- Slack POST failure falls back to plain-text "X/Y lanes failed"
with the GH run URL so the channel still hears something.
- No new Python deps — pure stdlib (urllib.request, xml.etree).
- 20 KB log-tail cap per lane to keep Haiku token usage bounded.
Secrets:
- ANTHROPIC_API_KEY (already present, used by provider-matrix)
- SLACK_WEBHOOK_URL (new — create an incoming webhook in Slack
and add as repo secret; notifier prints to stdout otherwise)
Testing:
- Trigger manually via Actions -> "Live Canary" -> "Run workflow"
with any single lane; canary-report runs after regardless of
which lanes executed.
- Run locally with --dry-run to preview the Slack payload.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(canary): post_json error handling + robust Haiku JSON extraction
Address gemini-code-assist review on scripts/live-canary/notify_slack.py:
1. `post_json` unreachable error branch: `urllib.request.urlopen`
raises `urllib.error.HTTPError` for 4xx/5xx before reaching the
`if resp.status >= 300` check, so the error body was never
surfaced. Wrap in try/except and read the body from the
HTTPError instance — that's where Anthropic's "invalid API key"
/ "rate limited" detail lives.
2. Haiku JSON extraction was fragile: `startswith("```")` assumed
the response had no prose preamble and only handled one fence
shape. Replace with `re.search(r"\{.*\}", text, re.DOTALL)` so
we pick the outermost JSON object regardless of any wrapper
markdown or leading/trailing text. Greedy + DOTALL is correct
for the single top-level object our schema requires.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(e2e): raise pytest timeout + bump multi-user chat wait to 180s
The CI run on feat/canary-report surfaced that 90s was still not
enough for test_mcp_same_server_multi_user_via_browser on
ubuntu-latest — it timed out at the inner Playwright
wait_for_function deadline with "Timeout 90000ms exceeded" after
118s of total test time.
The test opens two browser contexts + two SSE streams and drives a
full chat turn per user in sequence. Under 2-vCPU contention the
compound pipeline genuinely takes over 90s.
- tests/e2e/pyproject.toml: timeout 120 -> 240 (pytest-level cap)
- test_v2_auth_oauth_matrix.py: send_chat_and_wait_for_terminal_message
call sites 90000 -> 180000 (two owner/member turns, each budgeted
for one runner-slow turn)
180s < 240s, so the inner deadline fires first with the useful
Playwright traceback instead of the generic pytest SIGTERM.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(e2e): fix pytest-timeout CLI override + widen Mode-C deadlines
The previous commit (c3c9bbab) raised tests/e2e/pyproject.toml's
timeout from 120 to 240, but the auth canary runs the suite via
scripts/auth_canary/run_canary.py which hardcodes
`--timeout=120` on the pytest command line. The CLI flag wins
over pyproject's ini_options, so the 240 bump was invisible to
the auth lanes. That's why auth-smoke on the canary `all` run
still failed with "Timeout (>120.0s) from pytest-timeout" even
after our 180s inner widening — the outer CLI cap was firing at
120s first.
Fix the override and widen the two remaining Mode-C deadlines
that blew in the same run:
scripts/auth_canary/run_canary.py: --timeout=120 -> 240
_wait_for_refresh_request default: 60.0 -> 120.0
(test_wasm_tool_oauth_refresh_on_demand and
test_mcp_oauth_refresh_on_demand both use the default)
test_settings_first_gmail_auth_then_chat_runs call sites:
_wait_for_mock_google_tokens 90.0 -> 120.0
_wait_for_response_contains 90.0 -> 120.0
All remain comfortably under the new 240s pytest-level cap so a
real hang still fails fast with a useful traceback.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(e2e): opt-in text-match predicate for multi-user browser test
Ship the structural fix that was overdue. Repeated budget bumps on
send_chat_and_wait_for_terminal_message weren't holding under
ubuntu-latest "all"-mode parallelism — 120s, 180s both exceeded on
test_mcp_same_server_multi_user_via_browser. The underlying race is
in the JS predicate: it waits for the assistant bubble AND the
data-streaming attribute cleared AND the chat input re-enabled.
Under 2-vCPU contention an SSE reconnect can drop the final
attribute-clearing delta, and the compound predicate never flips
even though the response text arrived long ago.
Add an opt-in `expected_text_contains` parameter. When supplied,
the predicate succeeds the moment the expected substring appears in
the new assistant message — regardless of data-streaming or input
state. Callers that already assert on specific response text (the
existing MCP / gmail tests) can now short-circuit the race without
compromising correctness: the test's own content assertions remain
the gate.
Default behavior unchanged for the ~30 existing call sites across
test_chat.py, test_sse_reconnect.py, test_tool_approval.py,
test_portfolio.py, test_message_persistence.py, test_agent_loop_recovery.py,
test_pending_user_messages.py, test_widget_customization.py.
Applied to the two multi-user call sites with
expected_text_contains="Mock MCP search result" — that's exactly
what the test's next two assertions verify.
Local run of the flaky test alone: 40s, green.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* ci(canary): move auth-smoke to self-hosted runner
Multi-user browser test (test_mcp_same_server_multi_user_via_browser)
consistently exceeds the Playwright budget on GH ubuntu-latest under
the 2-vCPU parallelism pressure of an "all" canary run — a single
compound chat turn burns >180s, with each budget bump we apply it
ratchets the flake, not the fix.
Pilot move onto the [self-hosted, ironclaw-live] runner that
private-oauth already uses. Same runner label means no new
infrastructure required; if the self-hosted box has Python 3.12 and
Playwright browsers installed (or can provision them via the existing
setup-python + scripts/live-canary/run.sh's `PLAYWRIGHT_INSTALL=with-deps`
flow), this is a zero-code-change canary fix.
If the pilot works, auth-full is the next candidate. If the runner
queues become a bottleneck, we'd scale to multiple workers under
the same label rather than revert to ubuntu-latest.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* ci(canary): revert auth-smoke to ubuntu-latest + widen budgets to 300s/360s
Railway self-hosted runner ('railway-private-oauth' on a small Docker
container) turned out to be no faster than GH ubuntu-latest for the
multi-user browser flow — both take ~194–196s for
test_mcp_same_server_multi_user_via_browser. The runner container is
evidently provisioned at a similar vCPU allocation, so the move
bought nothing.
Revert to ubuntu-latest (parallel canary shape preserved; avoids
serialising auth lanes behind private-oauth on the single
self-hosted worker) and widen deadlines for the last CI-load hop:
test_v2_auth_oauth_matrix.py multi-user call sites:
Playwright wait_for_function 180000 -> 300000 ms
scripts/auth_canary/run_canary.py:
--timeout=240 -> 360 (outer pytest cap)
tests/e2e/pyproject.toml:
timeout = 240 -> 360
300s inner fits inside the new 360s outer with 60s margin. Local
run of the same test alone completes in ~40s, so we have plenty
of headroom against real hangs still surfacing fast with a
useful traceback.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* disable report
* scripts(auth-canary): add Google storage-state bootstrap helper
The auth-browser-consent lane drives Google's real OAuth consent UI in
Playwright, but Google's risk engine routinely interrupts the flow with
a "Verify it's you" challenge that handle_google_popup cannot solve, so
the test stalls on the password screen.
Bypass: log in once interactively in Playwright Chromium, save cookies
+ localStorage to a storage_state.json, point AUTH_BROWSER_GOOGLE_-
STORAGE_STATE_PATH at it. Subsequent canary runs spawn contexts with
that state preloaded, so the popup arrives at consent with no login or
challenge in the way.
- scripts/auth_live_canary/bootstrap_google_storage_state.py: new
one-shot interactive helper that writes
~/.ironclaw/auth-canary/google_storage_state.json by default
- scripts/auth_live_canary/README.md: document the bypass under
"Browser-consent Google challenge bypass"
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* canary(auth-browser-consent): fix Google account-picker + chat drift
The auth-browser-consent google case was failing on two distinct
issues, the first masking the second:
1) Account picker. When AUTH_BROWSER_GOOGLE_STORAGE_STATE_PATH is set
(the recommended path — username/password automation gets blocked
by Google's risk engine), Google's OAuth popup lands on a "Choose
an account" picker before the consent screen. handle_google_popup
only knew how to fill email + password and click Continue/Allow,
so the popup sat on the picker until complete_provider_auth's
120s callback wait timed out. Added a picker-detection step that
tries selectors in order — username text, [data-identifier], and
a generic "any visible @-bearing text not equal to 'Use another
account'" XPath — and clicks the first hit, with debug logging
so future regressions surface in the run output.
2) Tool-name and response-text drift. After the OAuth fix unblocked
the rest of the probe, browser_chat still failed because:
- case.expected_tool_name was "gmail", but the gateway records
the tool call under its WASM module name "gmail_tool"
- case.expected_text was "Gmail" (case-sensitive), but real LLM
responses to "check gmail unread" against an empty inbox vary
("Your inbox is clear...", "Inbox is empty", etc.) and rarely
emit literal "Gmail"
Updated BROWSER_CASES["google"] to expected_tool_name="gmail_tool"
and expected_text="inbox", and made the browser_chat assertion's
text comparison case-insensitive so the canary doesn't depend on
exact wording.
After both fixes the auth-browser-consent google lane runs green:
✓ browser_oauth (popup -> /oauth/callback)
✓ browser_chat (assistant references inbox)
✓ responses_api (real Gmail tool call)
Not addressed here: BROWSER_CASES["github"] likely has the same
expected_tool_name drift ("github" vs probably "github_tool"); needs
verification with real GitHub OAuth creds before changing.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* canary(auth-browser-consent): robust account-picker fallback + browser channel
Two follow-ups discovered during local debugging of the auth-browser-
consent google lane:
1) Account-picker fallback was matching hidden <style> blocks. The XPath
`//*[contains(text(), '@') ...]` matched any element whose text
contains `@`, which includes <style> tags carrying CSS at-rules
(@font-face, @media). Replaced the XPath with role-based locators
(get_by_role link/button) filtered by an email regex — only
interactive elements match, no false positives from style blocks.
Verified locally that the fallback now clicks the right account row
even when AUTH_BROWSER_GOOGLE_USERNAME is unset.
2) Bootstrap script: Google's anti-automation blocks Playwright's
default Chromium (Chrome for Testing) at sign-in with "This browser
or app may not be secure". Added a --browser flag with a default of
firefox (Marionette is less aggressively fingerprinted than CDP),
plus chrome (system Google Chrome) and chromium (override) options.
For accounts where Google blocks even those — typically brand-new
Gmails or accounts with high risk scores — the fallback path is to
launch Chrome manually with --remote-debugging-port and connect via
playwright.chromium.connect_over_cdp; documented in the README.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* canary(auth-live-canary): include observed extension state in timeout error
When `wait_for_extension_state` times out the bare error
"Timed out waiting for extension state: gmail" is unhelpful for
diagnosing CI failures, since CI artifacts don't capture IronClaw's
gateway logs — there's no way to tell whether the extension never
appeared, appeared but never authenticated, or authenticated but
never activated.
Track the last-observed extension on each poll and surface
authenticated/active in the timeout message. After this change a
failed run says e.g.
"Timed out waiting for extension state: gmail (expected
authenticated=True, active=True; last observed: authenticated=False,
active=False)", which immediately separates token-exchange failures
from activation-state-machine bugs.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* canary(auth-live-canary): widen chat-wait deadlines 120s -> 300s
The auth-browser-consent google probe completed OAuth + extension
activation successfully on CI but timed out at the next step
(send_chat_and_wait_for_terminal_message), with the agent stuck on
"Thinking (step 1)" for the full 120s budget. Local runs on the
same code path complete the chat in ~36s, but ubuntu-latest 2-vCPU
runners under cold-start load (gateway restart, mock LLM bootstrap,
WASM tool first-invocation) need substantially more headroom.
300s matches the precedent set by `d8765714 ci(canary): revert
auth-smoke to ubuntu-latest + widen budgets to 300s/360s` for the
auth-smoke lane on the same runner class.
Both call sites widened — the seeded Responses-API probe at line 221
and the browser_oauth probe at line 800.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* canary(common): drain gateway/mock_llm stdout pipes (was deadlocking CI)
scripts/live_canary/common.py spawns the IronClaw gateway and the
mock LLM with stdout=PIPE + stderr=STDOUT, reads one line of mock_llm
output to discover its bound port, then never reads from either pipe
again. On Linux the kernel pipe buffer caps at 64 KiB; once a
sustained chat request fills it with `RUST_LOG=info` output, the
child blocks on its next stdout write and the request handler
freezes mid-response.
That's why every auth-browser-consent CI run got stuck on
"Thinking (step 1)..." for the full chat-wait budget while the same
test passes locally — macOS pipe buffers are larger and the test
completes before the buffer fills.
Fix: spawn a daemon thread per subprocess that drains the pipe to a
log file under the run's output_dir. Two wins:
- Pipes never fill, child never blocks.
- gateway.log and mock_llm.log become CI artifacts, so the next
failure that doesn't have a clear runner-side error message is
immediately debuggable from IronClaw's own logs.
Verified locally that the lane still passes after the change and
both log files are produced. Locally each is < 10 KiB; CI runs may
be larger but well under any artifact size limit.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* canary: pin LLM backend via settings API + add LLM_API_KEY (root cause of CI freeze)
The auth-browser-consent google lane has been freezing on CI at
"Thinking (step 1)..." for the full chat-wait budget. Gateway logs
captured by the previous commit's pipe drainer reveal the smoking
gun:
ERROR Configured LLM backend is not usable.
backend=openai_compatible reason=missing API key
WARN LLM_BACKEND env var is set but DB setting takes priority.
db_value=nearai env_value=openai_compatible
WARN Active LLM backend fell back to NearAI default
attempted=openai_compatible active=nearai
Two compounding issues:
1. The openai_compatible provider refuses to instantiate without an
API key, even though the mock LLM ignores the value. Fix: set
`LLM_API_KEY=mock-api-key` in `build_gateway_env`, matching what
`tests/e2e/conftest.py` already does for the e2e suite.
2. IronClaw's DB-stored LLM settings take priority over env vars,
and the freshly-seeded canary DB defaults `llm_backend` to
`nearai`. So even with a clean env, the agent fell back to NearAI
and entered an interactive auth flow that hangs indefinitely in
CI (the "Thinking" never ends). This is the exact trap
`tests/e2e/CLAUDE.md` documents: "do not rely on env-vs-DB
precedence … pin the provider explicitly through /api/settings/...".
Fix: pin `llm_backend`, `openai_compatible_base_url`, and
`selected_model` via PUT /api/settings/<key> immediately after the
gateway becomes healthy.
Also revert the BROWSER_CASES["google"] case I touched earlier:
when NearAI was driving it emitted the WASM canonical tool name
(`gmail_tool`), but the mock LLM (now correctly driving) emits the
tool name it knows from its mapping (`gmail`). Restoring the original
`expected_tool_name="gmail"` / `expected_text="gmail"` matches what
the mock LLM actually produces.
Verified locally: all three browser_oauth / browser_chat /
responses_api probes now pass.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* canary(auth-live-canary): revert chat-wait deadline 300s -> 120s
The 300s widening at 98abeebe was a band-aid attempt to work around
the actual root cause (subprocess pipe deadlock + DB-overrides-env
LLM backend), which were both fixed at f59981d3 and 8733d3c0
respectively. With those fixes the chat completes in ~35s on CI, so
the 300s budget is overkill — revert to the original 120s, which
gives ~3.5x headroom over the observed steady-state and matches the
deadline shape used elsewhere in the e2e suite.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* ci(canary): rename github oauth secrets to dodge GITHUB_ prefix block
GitHub Actions reserves the GITHUB_ prefix for auto-generated repo
secrets (GITHUB_TOKEN, etc.) and rejects user-created secrets that
start with it: "Secret names must not start with GITHUB_". The
existing references to GITHUB_OAUTH_CLIENT_ID and GITHUB_OAUTH_-
CLIENT_SECRET in this workflow couldn't be backed by actual secrets
for that reason — the OAuth-client config was effectively unset for
the github browser-consent case, which is why it was silently
filtered out by configured_browser_cases().
Decouple the secret name from the env var name: store the secrets
under the AUTH_BROWSER_GITHUB_CLIENT_ID / AUTH_BROWSER_GITHUB_CLIENT_-
SECRET names (matching the AUTH_BROWSER_GITHUB_* convention used by
the other github canary fixture vars), and re-export them here under
the GITHUB_OAUTH_CLIENT_ID / _SECRET env names that
auth_registry.py and the WASM github tool expect.
No code changes needed in auth_registry.py / scripts/auth_live_-
canary/ — they continue to read GITHUB_OAUTH_CLIENT_ID/_SECRET from
the environment as before.
Operator action: create the OAuth app on GitHub (Settings →
Developer settings → OAuth Apps → New OAuth App) and store the
resulting credentials at:
AUTH_BROWSER_GITHUB_CLIENT_ID
AUTH_BROWSER_GITHUB_CLIENT_SECRET
(not GITHUB_OAUTH_CLIENT_ID / _SECRET, which GitHub will reject).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* canary(auth-browser-consent): drop github case (tool is PAT-only, not OAuth)
CI run 25022303491 surfaced that `Activate /api/extensions/github/-
activate` returns `{success: false, awaiting_token: true,
message: "Create a Personal Access Token..."}` with no `auth_url`,
which the browser-consent probe needs in order to drive the OAuth
popup.
Confirmed via `registry/tools/github.json`:
"auth_summary": {
"method": "manual", <- PAT paste, not OAuth
"secrets": ["github_token"],
"setup_url": "https://github.com/settings/tokens"
}
The github WASM tool's source capabilities JSON does carry an `oauth`
block, but the released v0.2.3 artifact (referenced from the registry)
ships with the manual-auth path. Until a release flips
`auth_summary.method` to "oauth" — and the github extension actually
returns an `auth_url` from /activate — there's nothing for the
browser-consent probe to do.
- Drop the `github` entry from BROWSER_CASES with a comment pointing
at the criterion for re-adding it.
- Drop the github-specific filter in `configured_browser_cases` since
the case is gone (no risk of an env-aware code path that quietly
skips github when secrets are present-but-mismatched).
GitHub coverage is unchanged in SEEDED_CASES, which seeds the PAT
directly via `AUTH_LIVE_GITHUB_TOKEN` and exercises real
`/v1/responses` + browser tool calls — that lane already works.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* canary(auth-browser-consent): tick notion's trust-URL checkbox before Continue
CI run 25023708895 surfaced the notion case timing out at "Timed out
waiting for notion OAuth callback page". The popup screenshot shows
Notion MCP's consent screen with:
- Workspace correctly auto-selected (storage state worked)
- A yellow warning: "I recognize and trust this URL"
- An unchecked checkbox next to that text
- A grayed-out (disabled) Continue button
The button is gated behind the checkbox. handle_notion_popup
clicked the disabled Continue and silently no-op'd, so the
complete_provider_auth loop waited the full 120s for /oauth/callback
that never arrived.
Add a checkbox-detection step before the Continue click:
popup.get_by_text(re.compile("I recognize and trust this URL", I))
.first.click(timeout=3000)
Includes debug print statements (matching the auth-canary pattern
established for google's account picker) so future Notion UI
changes are immediately visible in test-output.log.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(e2e): drain ironclaw subprocess pipes in auth-matrix fixture
Same pipe-deadlock fix as scripts/live_canary/common.py f59981d3,
applied to tests/e2e/scenarios/test_v2_auth_oauth_matrix.py's
_start_auth_matrix_server. The auth-matrix fixture spawns ironclaw
with stdout=PIPE + stderr=PIPE and never drains them, so under
sustained log volume the kernel pipe buffer fills, ironclaw blocks
on its next stdout write, and any test that relies on subsequent
gateway responses (auth gate emission, SSE events, chat replies)
hangs until pytest-timeout fires.
This fix doesn't make the auth-full lane's failing test pass — the
real bug is engine-v2 silently dropping `auth_required` SSE events
for unauthenticated extensions (introduced by #2868). But it makes
the failure mode debuggable: gateway log is captured to
/tmp/ironclaw-auth-matrix-gateway.log (overridable via
IRONCLAW_AUTH_MATRIX_LOG env), and RUST_LOG passes through from the
test runner so we can crank up verbosity without rebuilding.
Without this change, the failing test's log was empty after the
extension-install line; with this change you see the engine-v2
trace summary that surfaces the actual NotCallable-without-auth-gate
bug. That diagnostic visibility is the value here.
- _drain_stream_to_file: asyncio drainer mirroring common.py's sync
threading version
- _start_auth_matrix_server: drain stdout/stderr to log_path
- _shutdown_auth_matrix_server: cancel drain_tasks for clean exit
- env: RUST_LOG forwarding so debug runs work
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* canary(workflow): add Telegram Bot API mock
Foundation piece for the new workflow-canary lane that exercises
multi-tool / multi-channel user workflows from issue #1044 (Telegram +
routines + Sheets/Calendar/Gmail end-to-end). Models the same
single-port aiohttp-based mock pattern used by tests/e2e/mock_llm.py.
Endpoints:
- /bot{token}/{getMe,getUpdates,sendMessage,sendChatAction,
setWebhook,deleteWebhook,getFile} — the subset IronClaw's WASM
telegram tool + channels-src/telegram actually call. Tokens are
accepted without validation; the canary doesn't need to test
Telegram's auth — just IronClaw's flow against a Bot API shape.
- /__mock/inject_message — push a simulated incoming user message
onto the next getUpdates response, so scenarios can drive a
Telegram → IronClaw round-trip without a real Telegram account.
- /__mock/sent_messages — drain the queue of every sendMessage /
sendChatAction IronClaw emitted, for end-to-end assertions.
- /__mock/reset — clear all state between probes.
IronClaw routes its API calls through this mock via
IRONCLAW_TEST_HTTP_REMAP=api.telegram.org=<mock_url>, the same
mechanism the auth-live-canary uses for Gmail/Calendar/Sheets mocks.
Smoke-tested: getMe → success, inject_message → getUpdates returns
the injected message, sendMessage → bot response shape + recorded
in sent_messages.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* canary(workflow): land workflow-canary lane with periodic-reminder scenario
Phase 1A of the workflow-canary system from issue #1044. Adds a new
canary lane that exercises the routine engine + cron-fire path, the
foundation that the remaining four scripts (Telegram → Sheets,
Calendar prep, HN monitor, CRM tracker) will layer on.
Components:
- scripts/workflow_canary/routines.py — direct libSQL helpers for
inserting a lightweight cron routine with a backdated next_fire_at
and polling routine_runs for terminal status (ok / attention /
failed). Backdating beats wall-clock cron in tests by 30+ s per
probe and is the same shape auth-live-seeded uses for
expire_secret_in_db.
- scripts/workflow_canary/run_workflow_canary.py — entrypoint that
starts the Telegram mock, calls common.start_gateway_stack with
workflow-tuned env (ROUTINES_ENABLED=true, ROUTINES_CRON_INTERVAL=2,
IRONCLAW_TEST_HTTP_REMAP=api.telegram.org=<mock>), and runs
scenario modules. CLI mirrors run_live_canary.py.
- scripts/workflow_canary/scenarios/periodic_reminder.py — Script 4
Phase 1A: insert lightweight routine → wait for engine to fire →
assert run row reaches a terminal status. Verified locally: 1
probe, 1 fire, status=attention.
Plumbing:
- .github/workflows/live-canary.yml — new workflow-canary job + lane
added to the workflow_dispatch choice list and the canary-report
aggregator's needs:.
- scripts/live-canary/run.sh — workflow-canary case dispatches to
run_workflow_canary.py.
Phase 1B follow-ups in subsequent commits:
- Telegram channel install + bot-token seeding (needs admin auth or
direct encrypted-secrets DB write)
- Verify Telegram sendMessage was emitted to the mock during the
routine fire (covered by mock telegram's /__mock/sent_messages)
- Scripts 1, 3, 5 (Sheets / HN / Gmail-CRM)
- Script 2 (Calendar prep with web search)
Local verification:
$ tests/e2e/.venv/bin/python scripts/workflow_canary/run_workflow_canary.py \
--skip-build --skip-python-bootstrap
[workflow-canary] mock telegram listening at http://127.0.0.1:51139
[periodic_reminder] inserted routine ..., next_fire_at backdated 60s
[periodic_reminder] routine fired: status=attention
[workflow-canary] all 1 probe(s) passed.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* canary(workflow): land all 5 issue #1044 scenarios + scenario README
Layer Scripts 1, 2, 3, 5 onto the foundation shipped in 16278ea9, so
the workflow-canary lane covers all five user-workflow scripts from
issue #1044. Each scenario delegates to a shared
`run_routine_probe()` helper that captures the Phase 1A shape: insert
a Lightweight cron routine with a script-specific prompt → backdate
next_fire_at → poll routine_runs for terminal status.
Scenarios added:
- bug_logger.py (Script 1 — Telegram bugs → Google Sheet)
- calendar_prep.py (Script 2 — Calendar prep → Telegram, Reporter: Nick)
- hn_monitor.py (Script 3 — Hacker News → Telegram, Reporter: Emil)
- crm_tracker.py (Script 5 — Gmail → Sheets CRM, Reporter: Cameron)
Plus periodic_reminder.py (Script 4, Reporter: Henry) refactored to
also use run_routine_probe.
scenarios/_common.py centralizes the routine plumbing — each scenario
file is now ~30 lines of routine-name + prompt + Phase 1B follow-up
notes. The Phase 1B follow-up plan (Telegram channel install, mock
Sheets writes, mock Calendar reads, mock HN scrape, LLM email
classification, dedup verification) is documented inline in each
scenario's docstring AND in the new scripts/workflow_canary/README.md.
Local verification: all 5 probes green in ~2 s each.
$ tests/e2e/.venv/bin/python scripts/workflow_canary/run_workflow_canary.py \
--skip-build --skip-python-bootstrap
[workflow-canary] === Script 1 — Telegram → Google Sheet Bug Logger ===
[workflow-canary] === Script 2 — Calendar Prep Assistant ===
[workflow-canary] === Script 3 — Hacker News Keyword Monitor ===
[workflow-canary] === Script 4 — Periodic Reminder via Telegram ===
[workflow-canary] === Script 5 — Email → CRM Inbound Tracker ===
[workflow-canary] all 5 probe(s) passed.
What this catches:
- Routine engine cron-tick path (spawn_cron_ticker → check_cron_triggers)
- RoutineAction::Lightweight execution
- DB serialization of action_config / trigger_config
- Mock-LLM round-trip latency under cron scheduling
- routines.next_fire_at → routine_runs status state machine
What it doesn't catch yet (per-scenario Phase 1B work, documented in
README + scenario docstrings):
- Telegram channel install + sendMessage assertion
- Mock Sheets / Calendar / Gmail / HN write+read semantics
- LLM-driven structured classification (CRM)
- Cross-fire dedup verification
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* canary(workflow): scaffold Phase 1B telegram-side-effect verification
Lays the groundwork for verifying mock-Telegram side effects from
each scenario's routine fire — but gates the verification off until
a separate engine bug is fixed.
What's added:
- tests/e2e/mock_llm.py: new TOOL_CALL_PATTERNS entry that matches
``[CANARY-WORKFLOW-<key>]`` in any prompt and emits a deterministic
http tool call to api.telegram.org/.../sendMessage with a
per-scenario ack text.
- scripts/workflow_canary/scenarios/_common.py: each scenario now
composes its prompt as
``<prompt_intro>\n\n[CANARY-WORKFLOW-<key>]`` so the matcher fires.
When ``verify_telegram=True``, the helper polls
/__mock/sent_messages for up to 5 s and asserts the expected ack
was captured. Default is ``verify_telegram=False`` (Phase 1A
parity) — see below.
- scripts/workflow_canary/telegram_mock.py: aiohttp request-logger
middleware so the canary's stdout shows every inbound request,
giving operators a one-line answer to "did the gateway's HTTP
remap actually reach the mock?".
- scripts/workflow_canary/scenarios/{bug_logger,calendar_prep,
hn_monitor,periodic_reminder,crm_tracker}.py: scenarios pass
``mock_telegram_url=mock_telegram_url`` and ``prompt_intro=...``
ready for verify_telegram to flip on.
What's gated off and why:
The mock-Telegram verification path requires
``IRONCLAW_TEST_HTTP_REMAP=api.telegram.org=<mock>`` to route
the http tool's sendMessage call into the mock. The remap is
correctly registered at gateway startup
(src/app.rs::http_interceptor + src/http_intercept.rs), but the
ToolContext built inside the routine engine's Lightweight action
loop does NOT inherit the global ``http_interceptor`` slot. Result:
the http tool reaches into the real network for api.telegram.org
(returning a 401 since the bot token is fake) and the mock never
sees the request — confirmed via the new request-logger middleware
showing zero non-internal hits.
That's a real engine bug in routine-driven tool dispatch — the
http_interceptor needs to propagate through the routine action's
ToolContext just like it does for chat-driven tool dispatch. Out of
scope for this canary PR; tracked as a follow-up. Once fixed, flip
the default in ``run_routine_probe`` and every scenario's
verify_telegram check activates with no further changes.
Local verification: all 5 probes still green at the Phase 1A level.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* canary(workflow): re-exec under venv after bootstrap (fix CI 'No module named httpx')
CI run 25028445222 failed on the workflow-canary lane with:
[workflow-canary] mock telegram listening at http://...
[workflow-canary] error: No module named 'httpx'
Root cause: run_workflow_canary.py was missing the bootstrap-then-
reexec pattern that scripts/auth_live_canary/run_live_canary.py
uses (line 1229+). bootstrap_python() creates the venv and installs
tests/e2e/'s pyproject deps (which include httpx + aiohttp), but
the parent process keeps executing under whatever interpreter
invoked it — typically the system Python on CI runners, which
doesn't have httpx. The scenario module's `import httpx` at top
level then fails immediately.
Fix: copy the auth-live-canary reexec pattern. main() now:
1. If not --skip-python-bootstrap AND WORKFLOW_CANARY_REEXEC is
unset: bootstrap the venv, install playwright, build cargo,
then subprocess-spawn ourselves under the venv python with
--skip-python-bootstrap and WORKFLOW_CANARY_REEXEC=1 so this
branch isn't re-entered.
2. The reexecuted process sees skip_python_bootstrap=True and runs
the actual canary against the venv interpreter that has all
deps available.
Local sanity check: still passes (--skip-build --skip-python-bootstrap
short-circuits the bootstrap, both branches behave identically when
the venv already exists).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(routine-engine): propagate http_interceptor into Lightweight tool dispatch
The chat path's tool dispatch correctly receives the global
HTTP interceptor (e.g., the `IRONCLAW_TEST_HTTP_REMAP` debug-only
host remapper installed in `src/app.rs::http_interceptor`), but the
routine engine's Lightweight action path constructed its
`JobContext` from scratch with `..Default::default()`, leaving
`http_interceptor: None`. Tools called from a routine therefore
reached the real network even when the rest of the system was
configured to route through mocks.
Plumb the interceptor through:
- `RoutineEngine` gains an `http_interceptor` field
- `RoutineEngine::new` takes it as the 11th argument
- `EngineContext` carries it across the spawn boundary
- `JobContext` construction at the Lightweight action site copies
it from the engine context
Threading complete: AgentDeps → RoutineEngine → EngineContext →
JobContext → http tool. Same shape the chat path already uses.
Test rigs updated: `tests/support/test_rig.rs` and
`tests/e2e_routine_heartbeat.rs` (10 call sites total) pass `None`
for the new arg, matching their existing minimal stack model.
Build clean against `--no-default-features --features libsql`.
Why this matters: with the interceptor lost, every workflow-canary
probe's http tool dispatch reached real api.telegram.org and 401'd
on the fake token — leaving the mock Telegram bot empty and the
canary's send-side assertions unverifiable. With the fix, the
interceptor honors the IRONCLAW_TEST_HTTP_REMAP and the workflow
canary's Phase 1B verification activates immediately.
Activates in this commit:
- scripts/workflow_canary/scenarios/_common.py default flips to
`verify_telegram=True`
- All 5 scenarios (bug_logger, calendar_prep, hn_monitor,
periodic_reminder, crm_tracker) now assert that the mock
Telegram bot received the per-scenario ack message
`[canary-workflow:<key>] ack`
Local verification:
$ tests/e2e/.venv/bin/python scripts/workflow_canary/run_workflow_canary.py \
--skip-build --skip-python-bootstrap
[workflow-canary] === Script 1 — Telegram → Google Sheet Bug Logger ===
... (all 5 scenarios) ...
[workflow-canary] all 5 probe(s) passed.
$ grep "POST /bot" artifacts/workflow-canary/telegram_mock.log | wc -l
5 # one per scenario, distinct ack text per probe
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* canary(workflow): add manual_trigger + lifecycle + dedup_cooldown probes
Three new scenarios covering issue #1044 assertions that the existing
5 cron-fire probes don't reach. Each scenario tests a distinct
back-end mechanism that real users hit:
- **manual_trigger** (Scripts 3 PHASE 2.1 + 3 PHASE 4.2 + 4 PHASE 4.2)
Inserts a routine WITHOUT backdating next_fire_at, so the only path
to a fire is the manual-trigger API. POSTs
/api/routines/<id>/trigger, asserts response carries a run_id, polls
routine_runs for terminal status, then verifies mock Telegram
captured the per-scenario ack. Catches regressions in
RoutineEngine::fire_manual end-to-end.
- **lifecycle** (Scripts 1 PHASE 5 + 4 PHASE 5) — three sub-probes:
1. disabled-blocks-fires: insert with enabled=False + backdate;
assert no routine_runs row appears within 8 s window.
2. enable-resumes-fires: toggle enabled=true via API, backdate,
assert fire reaches terminal status.
3. delete-removes-routine: confirm /api/routines lists it, DELETE,
confirm it's gone.
Catches regressions in toggle handler, delete handler, and the
engine's enabled-flag respect during cron tick selection.
- **dedup_cooldown** (Scripts 1 PHASE 4.4 + 3 PHASE 3.2 + 5 PHASE 5.5)
Insert with cooldown_secs=30; first fire lands within ~5 s; immediate
re-backdate; assert ONLY ONE run row exists after 8 s. Catches
regressions in cooldown enforcement during check_cron_triggers.
This is the closest engine-level correlate to the user-script
"no duplicate rows / alerts / messages" assertions, which are
application-level dedup that lives outside the canary's
deterministic-mock surface.
Plumbing:
- routines.py: trigger_routine_via_api / toggle_routine_via_api /
delete_routine_via_api / list_routines_via_api helpers (all auth-
bearer, JSON in/out, raise_for_status).
- routines.py: insert_lightweight_cron_routine grew `cooldown_secs`
+ `enabled` parameters; defaults preserve existing behavior.
- run_workflow_canary.py: registered the three new scenario keys.
Local verification — all 10 probes (5 original + 5 new sub-probes
across 3 new scenarios) green:
✅ bug_logger / calendar_prep / hn_monitor / periodic_reminder /
crm_tracker (existing — Telegram ack capture)
✅ manual_trigger (548ms)
✅ lifecycle_disable (8004ms — full no-fire window)
✅ lifecycle_toggle (1543ms)
✅ lifecycle_delete (56ms)
✅ dedup_cooldown (10017ms — first fire + 8s no-fire window)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* canary(workflow): add NL-driven routine_create + routine_update probes
Two scenarios that close issue #1044's chat-driven assertions
(Script 1 PHASE 3.1, Script 2 PHASE 3.1, Script 3 PHASE 2.1,
Script 4 PHASE 2.1 + 5.1, Script 5 PHASE 4.1):
- **nl_routine_create**: opens a thread via /api/chat/thread/new,
posts an NL message tagged [CANARY-WORKFLOW-NL-CREATE], waits for
the agent to dispatch routine_create, then verifies the routines
row landed in libSQL AND is visible via GET /api/routines.
- **nl_schedule_update**: pre-seeds a target routine
(canary-nl-update-target), posts an NL message tagged
[CANARY-WORKFLOW-NL-UPDATE], waits for the agent to dispatch
routine_update with a new schedule, then verifies trigger_config
changed in libSQL. Asserts on schedule-changed (not exact match)
because the engine normalizes 5-field cron → 7-field internal
form ("0 */5 * * *" → "0 0 */5 * * * *").
Plumbing:
- Two new TOOL_CALL_PATTERNS entries in tests/e2e/mock_llm.py
matched in priority order (specific NL-CREATE / NL-UPDATE
sentinels checked BEFORE the generic [CANARY-WORKFLOW-<key>]
http-tool fallback, since the canary's own routines emit the
generic pattern from inside their action prompts).
- Helper additions in scripts/workflow_canary/routines.py:
_open_thread / _send_chat / _read_routine / _wait_for_*.
Local verification — all 12 probes green:
✅ bug_logger / calendar_prep / hn_monitor / periodic_reminder /
crm_tracker (5 cron-fire + telegram-ack)
✅ manual_trigger (POST /api/routines/<id>/trigger)
✅ lifecycle_disable / lifecycle_toggle / lifecycle_delete
✅ dedup_cooldown (cooldown_secs suppresses second fire)
✅ nl_routine_create (chat → routine_create tool)
✅ nl_schedule_update (chat → routine_update tool)
What's still deferred to follow-up PRs (per-provider mocks, each
~1-3 days of work — see scripts/workflow_canary/README.md):
- Mock Google Sheets (Scripts 1 + 5 dedicated assertions)
- Mock Google Calendar (Script 2)
- Mock Hacker News (Script 3)
- LLM-driven email classification with seeded inbox (Script 5)
- Telegram channel install + bot-token validation flow (Scripts 1-5
PHASE 1)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(workflow-canary): Phase 1 — mock Sheets + bug_logger Sheet-write probe
Adds scripts/workflow_canary/sheets_mock.py: single-port aiohttp Google
Sheets v4 mock supporting POST /v4/spreadsheets, values:append, values
get, plus /__mock/ test hooks for seeding, draining, and resetting.
The append handler enforces values=list-of-lists (returns the canonical
"expected a sequence" 400) so the canary catches the issue #1044 FAIL
CRITERIA shape.
Wires the mock into run_workflow_canary.py:
- generic _spawn_mock helper for telegram_mock + sheets_mock
- IRONCLAW_TEST_HTTP_REMAP carries comma-separated entries for
api.telegram.org and sheets.googleapis.com
- mock_sheets_url passed through to every scenario's run() kwargs
Rewrites scenarios/bug_logger.py to drop the run_routine_probe Telegram
fallback in favor of a Sheet-write end-to-end assertion: pre-seed the
spreadsheet, fire the routine with [CANARY-WORKFLOW-SHEET-APPEND], wait
for the appended row, validate shape (timestamp / message / source).
Mock LLM: new TOOL_CALL_PATTERNS entry that matches the SHEET-APPEND
sentinel and emits an http POST values:append with a hardcoded canary
row.
All 12 probes still pass locally.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(workflow-canary): Phase 2-4 — Calendar / HN / Gmail / web_search mocks + e2e probes
Phase 2 (Calendar): scripts/workflow_canary/calendar_mock.py — Google
Calendar v3 events surface (list / insert / get / delete) with seed
hooks. calendar_prep_e2e seeds one canary event, fires the routine,
asserts events.list was hit and Telegram received the prep briefing
referencing the seeded event title.
Phase 3 (Hacker News): scripts/workflow_canary/hn_mock.py — /newest
HTML fixture with seeded "Show HN" posts (canary-distinct
``<!-- canary-hn-feed -->`` marker). hn_monitor_e2e re-seeds posts,
asserts /newest GET landed and Telegram summary references both
seeded posts.
Phase 4 (CRM tracker): scripts/workflow_canary/gmail_mock.py +
web_search_mock.py — Gmail v1 messages.list/.get + Brave Search v3.
crm_tracker_e2e seeds 1 lead + 1 newsletter + 1 receipt; asserts
exactly ONE row appended to the CRM sheet (only the lead) with all
6 expected columns + Telegram ack referencing 1 lead.
Mock LLM TOOL_CALL_PATTERNS gain three parallel-call entries
([CANARY-WORKFLOW-CAL-LIST] → http GET events.list + http POST
sendMessage; [CANARY-WORKFLOW-HN-FETCH] → GET /newest + sendMessage;
[CANARY-WORKFLOW-CRM-CLASSIFY] → Gmail GET + Sheets append + Telegram
ack). Parallel emit is required because the engine's lightweight
loop dedups same-tool re-dispatch (see match_tool_call:1178).
run_workflow_canary.py now spawns six mock subprocesses; remap covers
api.telegram.org, sheets.googleapis.com, www.googleapis.com,
news.ycombinator.com, gmail.googleapis.com, api.search.brave.com.
All 12 existing probes pass + 3 phase 2-4 probes upgrade from
side-effect-only to full content-correctness assertions.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(workflow-canary): Phase 5 — Telegram channel install + round-trip
scripts/workflow_canary/telegram_setup.py: install + capability patch
+ setup helpers (mirrors tests/e2e/scenarios/test_telegram_e2e.py
patch_capabilities + activate flow). Adds pair_telegram_user that
sends an "hello" webhook, extracts the pairing code from
mock_telegram, and approves it via /api/pairing/telegram/approve.
scripts/live_canary/common.py: GatewayStack now exposes http_url
(HTTP-channel webhook port) + channels_dir (WASM_CHANNELS_DIR)
so workflow-canary scenarios can drive the Telegram channel install
+ patch + webhook flow.
run_workflow_canary.py: passes IRONCLAW_TEST_TELEGRAM_API_BASE_URL
so the hardcoded validate_telegram_bot_token getMe call (in
src/extensions/manager.rs) routes to mock_telegram. The bot-token
validate path bypasses the standard IRONCLAW_TEST_HTTP_REMAP flow,
hence the additional env override.
New scenarios:
- telegram_channel_install: install + patch caps + setup + assert
channel reaches Active state. Catches "HTTP 404 on valid token"
regression (Script 4 PHASE 1.1).
- telegram_round_trip: post inbound webhook → assert mock_telegram
receives an outbound sendMessage with the actual chat_id (NOT
'default'). Catches the chat_id 'default' regression.
- routine_visibility_from_telegram: pair user, ask for routines,
assert agent replies on the paired chat_id. Covers Scripts 1-4
PHASE "routine visibility from Telegram" assertions.
- manual_trigger_from_telegram: pair user, hit /api/routines/<id>/
trigger, assert routine fires through lightweight loop and ack
reaches the paired chat_id. Covers Script 4 PHASE 4.2.
All 16 probes pass locally.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(workflow-canary): Phase 6 — first_immediate_run + log_assertions
scripts/workflow_canary/scenarios/first_immediate_run.py: insert a
routine with a "0 * * * *" schedule + fire_immediately=True; assert
the first run reaches terminal status within 10s. Catches "first
check is delayed to next hour" regression (Script 3 PHASE 2.1).
scripts/workflow_canary/scenarios/log_assertions.py: scan
gateway.log at the end of the lane for known fail-criterion regex
patterns: chat_id 'default', parsed naive timestamp without timezone,
retry after None, expected a sequence. Catches log regressions across
all 5 issue #1044 scripts simultaneously.
Auth-recovery (token revocation → auth_required SSE) is deferred to
the auth-live-canary lane; it requires a working OAuth setup to
revoke, which is outside this lane's mock-only scope.
All 18 probes pass locally.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(workflow-canary): Phase 7 — cron timing + idempotent toggle + README
scripts/workflow_canary/scenarios/cron_timing_accuracy.py: insert a
routine, set next_fire_at to "now + 5s" explicitly, assert the engine
fires within ±10s of the set boundary. Catches "cron skipped a cycle"
+ "fires never trigger" regressions (Scripts 3 PHASE 3.1, 4 PHASE 3.4).
scripts/workflow_canary/scenarios/idempotent_disable_enable.py:
double-toggle disable then double-toggle enable, assert both halves
are no-ops; finally backdate, fire once, then disable + backdate again
and assert no NEW runs land in the next 6s. Catches "disable doesn't
take effect" + "enable triggers a phantom run" regressions
(Script 1 PHASE 5.1 / 5.2).
scripts/workflow_canary/README.md: rewritten to reflect 20-probe
coverage matrix across phases 1–7 with mock surface + scenarios
inventory.
Final canary state: 20 probes across 7 phases, all green locally.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(workflow-canary): close gaps — wire web_search + add auth_recovery
[CANARY-WORKFLOW-CAL-LIST] now emits a parallel triplet (calendar
events.list + web_search company lookup + telegram sendMessage).
calendar_prep asserts mock_web_search captured the lookup with the
expected company-name query parameter, completing the Script 2
"company background + recent news" assertion from issue #1044.
scripts/workflow_canary/scenarios/auth_recovery.py: drives a chat
that triggers an unauthenticated gmail tool call, asserts the agent
surfaces a graceful response — chat send returns 202 (not 5xx),
thread settles, history contains no Error 400 / Internal Server
Error / panicked / Traceback fragments. Catches the regression
shape from Script 2 PHASE 5 fail criteria without requiring a real
OAuth handshake (full token-revocation coverage stays in
auth-live-canary).
21 probes total, all green locally.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* ci(canary): run every 6h + re-enable Slack report
Schedule: cron flips from "0 2 * * *" (once daily at 02:00 UTC) to
"0 */6 * * *" (4× daily at 00/06/12/18 UTC). All twelve job-level
`if:` guards updated in lockstep so each lane still gates on the
schedule string.
Slack report: drop the `if: false` hardcode on the canary-report
job's notify step and replace with a schedule + workflow_dispatch
gate. The notifier (scripts/live-canary/notify_slack.py) already
exits 0 on Haiku/Slack failures so a flaky webhook can't mask lane
status. PR-triggered runs (currently none, but possible via
workflow_run) skip the post to keep noise out of the channel.
Both ANTHROPIC_API_KEY and SLACK_WEBHOOK_URL repo secrets are
already populated.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(canary-report): parse workflow-canary results.json shape
The notifier reads `auth-canary-junit.xml` for JUnit-emitting lanes
(auth-smoke, auth-full, auth-channels, auth-live-seeded,
auth-browser-consent). The workflow-canary lane writes its own
`results.json` instead — one entry per probe with `success: bool`,
`latency_ms`, `details`. The notifier had no parser for that shape, so
the workflow-canary slot in Slack rendered as a useless
`:grey_question: 0/0 passed, 0 failed` line.
Add `parse_results_json` mirroring the JUnit parser's contract:
`passed = sum(success)`, `failed = sum(!success)`, each failed probe
becomes a `(provider/mode, error-or-summary)` entry on
`junit_failures` so the Slack reason field renders the same way as an
auth-canary failure. Latencies sum to `duration_s`. Both parsers run
on every lane dir; first one whose file exists wins (auth-canary lanes
emit XML only, workflow-canary lane emits JSON only — no overlap).
Validated by re-running the notifier locally against the downloaded
artifact from CI run 25033224036:
before: ":grey_question: workflow-canary (mock) — 0/0 passed"
after: ":white_check_mark: workflow-canary (mock) — 21/21 passed,
0 failed in 69s"
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(canary-report): log notifier progress for diagnosability
Until now `notify_slack.py` was silent on the success path, which made
it impossible to verify from CI logs alone whether Haiku enrichment
actually ran. Add four stderr lines covering each phase:
[notify_slack] discovered N lane dir(s): lane1/provider1, ...
[notify_slack] lane/provider: tests=N passed=N failed=N skipped=N status=...
[notify_slack] haiku enriched X/N lane(s)
[notify_slack] posted Slack message for N lane(s)
Lines stay terse and structured so they're greppable from `gh run
view --log`. Haiku-failure tracking inspects `r.notable` — `run_haiku`
stamps it with `haiku call failed:` / `haiku returned no JSON object`
/ `haiku JSON parse failed` on the three failure paths.
Confirmed from local dry-run against the artifact downloaded from
the previous CI run (which had the results.json parser): tests=21,
passed=21, failed=0, status=pass.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(canary/workflow-canary): forward SCENARIO into --scenario
Addresses @henrypark133's review on PR #2874: the workflow-canary lane
of `scripts/live-canary/run.sh` ignored `${SCENARIO}` and always ran
the full 21-probe suite. The matching workflow_dispatch job didn't
export `inputs.scenario` either, so manual dispatch with a scenario
filter went nowhere. Targeted local reruns / debugging hit the same
gap.
run.sh: translate `${SCENARIO}` (comma-list supported) into one or
more `--scenario <name>` flags on `run_workflow_canary.py`. Empty
SCENARIO falls through to the full suite. Guards the array splat for
bash 3.2 / macOS where `${arr[@]}` on an empty array under `set -u`
explodes.
live-canary.yml: add `SCENARIO: ${{ inputs.scenario }}` to the
Workflow Canary job's env so workflow_dispatch reaches run.sh.
Verified:
tests/e2e/.venv/bin/python \
scripts/workflow_canary/run_workflow_canary.py \
--skip-build --skip-python-bootstrap \
--scenario telegram_round_trip
→ "all 1 probe(s) passed"
(full suite without the flag still runs all 21 probes)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(canary/workflow-canary): align nl_schedule_update on 'every 6 hours'
Addresses Copilot AI's review on PR #2874: the docstring claimed
"every 5 minutes" while EXPECTED_NEW_SCHEDULE / mock LLM emitted
"0 */5 * * *" (every 5 hours), and the chat prompt the canary sent
said "every 5 hours". Three different cadences across one probe.
Pick "every 6 hours" consistently:
- Docstring narrative: "every 6 hours"
- Constant: EXPECTED_NEW_SCHEDULE = "0 */6 * * *"
- Chat prompt: "fire every 6 hours"
- mock_llm.py routine_update args: schedule = "0 */6 * * *"
Verified locally: nl_schedule_update probe still green.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(canary/auth-browser-consent): drop stale GitHub secret exposure
Addresses @henrypark133's review on PR #2874: the auth-browser-consent
job kept exporting 8 GitHub-related secrets (GITHUB_OAUTH_CLIENT_ID,
GITHUB_OAUTH_CLIENT_SECRET, AUTH_BROWSER_GITHUB_OWNER / _REPO /
_ISSUE_NUMBER / _USERNAME / _PASSWORD / _STORAGE_STATE_B64) even
though the lane no longer drives a GitHub OAuth flow. BROWSER_CASES
in `scripts/live_canary/auth_registry.py` was reduced to {google,
notion} when github was reclassified as PAT-only — those secrets are
unused on every scheduled run and just broaden the secret-exposure
surface.
Strip all 8 from the lane:
- env: block — 5 lines (CLIENT_ID + 4 AUTH_BROWSER_GITHUB_* helpers)
- Materialize provider storage state — 1 secret + its materialize block
- Materialize sensitive secrets — 2 secrets + their write_secret lines
Replace with explanatory comments pointing at BROWSER_CASES /
auth_registry.py so a future contributor doesn't re-add them by reflex
when github gets an OAuth flow.
Github coverage continues to live in SEEDED_CASES (auth-live-seeded
lane) which seeds the PAT directly — that lane's secrets are
unaffected.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* docs(canary): align user-facing browser-cases list with auth_registry
Addresses @henrypark133's review on PR #2874: removing `github` from
BROWSER_CASES made `--mode browser --case github` invalid, but the
contract was still advertised in three places that operators read
when copying invocations:
- run_live_canary.py --help (`For browser mode: google, github, notion`)
- scripts/auth_live_canary/README.md (`github` listed under "Runs
through Responses API and browser")
- scripts/live-canary/README.md (`CASES=google,github` example)
- scripts/live-canary/ACCOUNTS.md (full GitHub OAuth client + fixture
+ storage-state-secret sections still active, plus a Playwright
storage-state recipe pointing at github.com/login)
Update each in lockstep:
- --help now says `For browser mode: google, notion. (github browser
coverage is intentionally absent — the github WASM tool is PAT-only,
not OAuth; see SEEDED_CASES instead.)`
- auth_live_canary/README — github entry now reads "Responses API
only (PAT-only — not browser-OAuth)"; notion entry corrected to
"Responses API and browser" (it was inaccurately listed as
Responses API only).
- live-canary/README — example flips to `CASES=google,notion` with a
one-line note pointing at auth_registry.py.
- live-canary/ACCOUNTS — drops the GitHub OAuth client + fixture
sections, swaps the Playwright storage-state recipe target from
github.com/login to accounts.google.com, drops
AUTH_BROWSER_GITHUB_STORAGE_STATE_B64 from the CI-secrets list.
The argparse validator in run_live_canary.py already gives a clean
error if anyone passes `--mode browser --case github`:
"--case values ['github'] are not valid for --mode browser. Allowed:
['google', 'notion']", so the docs change is the user-facing fix.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(canary/telegram): split is_active into installed vs. active
Addresses Copilot AI's review on PR #2874: `is_telegram_active` only
checked that an extension named "telegram" appeared in
`/api/extensions`, returning True for an installed-but-inactive
extension (mid-setup, awaiting auth, activation_error). Two callers
(`telegram_round_trip._ensure_active`,
`routine_visibility_from_telegram._ensure_active_and_paired`) used
this as a precheck to skip `setup_telegram_channel()`, so a stale
inactive entry would short-circuit setup and the probe would then
fail mysteriously when the channel didn't respond.
Split into two helpers:
- `is_telegram_installed(...)` — original semantics (entry exists),
used internally as a building block; not exported as a precheck.
- `wait_for_telegram_active(...)` — polls until the entry has
`active=true` (the actual runtime-readiness signal — channel
opened, hooks registered, credentials bound, per
`.claude/rules/lifecycle.md`'s discovery-vs-activation rule).
Shared `_find_telegram` helper handles the three historical envelope
shapes the gateway has used (`extensions` / `items` / `installed`).
Update all 4 callers to use `wait_for_telegram_active`:
- telegram_channel_install.py
- telegram_round_trip.py (precheck + post-setup wait)
- routine_visibility_from_telegram.py (precheck + post-setup wait)
- manual_trigger_from_telegram.py (precheck + post-setup wait)
Verified: all 4 telegram probes still green back-to-back.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* docs(canary/periodic_reminder): align docstring with current behavior
Addresses Copilot AI's review on PR #2874: the module docstring still
described the Telegram delivery assertion as a "Phase 1B follow-up"
even though the scenario now sets verify_telegram=True and the
inline comment on the call site already explained the Phase 1B work
had landed. Future readers would assume Telegram verification was
missing from this probe.
Replace the docstring with a 5-step description of what the probe
actually does end-to-end:
1. Backdated cron routine inserted via libSQL
2. Routine engine cron-tick picks it up
3. Lightweight action runs against mock LLM → http sendMessage
4. IRONCLAW_TEST_HTTP_REMAP routes to telegram_mock
5. Asserts both terminal routine_runs status AND captured sendMessage
Also adds an explicit note that channel-install coverage (capability
patch + setup + pairing) lives in the sibling telegram_* scenarios —
this one covers the routine-driven sendMessage path and intentionally
hits api.telegram.org via the raw http tool rather than through the
installed channel.
Verified: probe still green.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(canary-report): rich failure blocks + cross-lane categorization + GH issues
Three additions to scripts/live-canary/notify_slack.py to make the
6h Slack report actionable instead of just informational:
1) **Per-lane rich failure block** — Haiku now extracts four
structured fields when status==fail: test_name, error, root_cause,
fix. The Slack section renders them in the issue-friendly shape
the reviewer asked for:
:x: auth-full (mock) — 11/13 passed, 1 failed in 213s
Test: `test_wasm_tool_first_chat_auth_attempt_emits_auth_url`
Error: SSE stream closed; auth_required event never arrived
Root Cause: bridge gate not wired for installed-but-unauthed
extensions (#2868 fallout)
Fix: route Extension::NeedsAuth through effect_adapter.rs
For passing/skipped lanes the existing single-line `> reason` is
preserved so the green-path Slack output is unchanged.
2) **Cross-lane "Summary by Category" block** — second Haiku pass
over all failed-lane summaries that groups them by shared root
cause (e.g. "WASM tool dispatch regression — Auth Full, Auth
Smoke, Auth Live Seeded"). Only fires when there are 2+
failures (single-failure runs are already obvious from the
per-lane block). Rendered as a Slack mrkdwn bulleted list since
Block Kit doesn't support real tables.
3) **Auto-opened GitHub issues** — opt-in via CANARY_CREATE_ISSUES=1
env var (gated to scheduled runs only in live-canary.yml so
workflow_dispatch debugging doesn't flood the tracker). For each
failed lane:
- Search for an OPEN issue with title `[canary] <lane>: <test>`.
- If found: comment "another occurrence on <run_url>".
- If not found: open a new issue with the rich body + labels
`canary-failure` + `lane:<lane>`.
Strategy chosen to avoid issue spam while still surfacing
recurring failures. Uses GITHUB_TOKEN + the repo's existing
`permissions: issues: write` block — no new secrets.
All three additions degrade silently — Haiku failure stamps
.notable but doesn't block the post; categorization failure produces
an "_(unavailable)_" placeholder; issue-creation errors are logged
to stderr only. The notifier still exits 0 in every failure path so
a flaky webhook can't fail the canary run.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* ci(canary-report): reuse AUTH_LIVE_GITHUB_TOKEN for issue creation
Swap the issue-creation token source from the built-in
secrets.GITHUB_TOKEN to the existing AUTH_LIVE_GITHUB_TOKEN PAT —
no new secrets to mint, and that PAT already covers
nearai/ironclaw operations.
Set as CANARY_ISSUES_TOKEN (the highest-priority env var in
notify_slack.py's --github-token precedence chain) so it wins over
GH_TOKEN / GITHUB_TOKEN if any of those are also present.
Verify the PAT has `issues: write` scope (Issues: read & write for
fine-grained PATs, repo scope for classic PATs). If it doesn't, the
notifier still degrades gracefully — the API call fails, the error
is logged to stderr, the canary run isn't blocked.
Co-Authored-By: Claude Opu…
PR #2868 added a callable-inventory check ahead of the lease lookup in execute_action_calls. The test was constructing MockEffects with an empty action list, so preflight short-circuited on "action is not callable in this execution context" before reaching the lease branch the test was actually trying to exercise. The assertion error.contains(\"no lease\") then failed against the not-callable message. Populate the mock with test_action(\"web_search\") so the inventory gate passes and the call reaches the lease check, matching the test's documented intent.
…3234) The v2-engine E2E group still listed test_v2_kernel_auth_preflight.py, but that file was removed in #2868 (engine-v2: callable-only available actions) and replaced with test_v2_tool_activate_surface.py for the new tool_activate / Activatable Integrations contract. The Web E2E Full job is skipped on PR-level CI but runs in the merge queue, so the bad path filter dequeued #3197 and #3203 with "file or directory not found: test_v2_kernel_auth_preflight.py".
…ange (#3235) * ci(e2e): replace deleted preflight test with tool_activate surface The v2-engine E2E group still listed test_v2_kernel_auth_preflight.py, but that file was removed in #2868 (engine-v2: callable-only available actions) and replaced with test_v2_tool_activate_surface.py for the new tool_activate / Activatable Integrations contract. The Web E2E Full job is skipped on PR-level CI but runs in the merge queue, so the bad path filter dequeued #3197 and #3203 with "file or directory not found: test_v2_kernel_auth_preflight.py". * test(e2e): unblock Live Canary auth lanes after engine-v2 contract change The Live Canary "Auth Smoke", "Auth Full", and "Auth Live Seeded" jobs have failed every scheduled run since 2026-05-01 (when the canary cut over to main). Three tests in test_v2_auth_oauth_matrix.py drive the failures, all rooted in the engine-v2 callable-only contract from #2868 that didn't exist when these tests were written. ## What was broken `test_mcp_same_server_multi_user_via_browser` After OAuth completes, sending "check mock mcp search" through each user's browser opens an `approval` pending_gate on the first MCP tool call (engine v2 default). The browser fixture has no auto-approve UI, so the chat sat in `pending_gate` for the full 5-min Playwright timeout — `expected_text_contains="Mock MCP search result"` could never match because the assistant bubble never received any text. `test_wasm_tool_oauth_refresh_on_demand` Same shape: gmail call gates on `approval` before reaching the http credential-injection layer that performs the OAuth refresh. Without approving, refresh_count never went above 0, so the test failed with "Timed out waiting for OAuth refresh request". `test_wasm_tool_first_chat_auth_attempt_emits_auth_url` Tested OLD engine-v2 behavior — that an LLM-emitted call to a not-yet- authed extension would surface a `gate_required` Authentication event with an auth URL. After #2868, the engine returns "action 'gmail' is not callable in this execution context" instead, and `tool_activate` became the model-facing enablement path. The mock LLM is canned to emit tool calls directly, so this scenario can't be reproduced from a scripted LLM until the canned response is updated. ## Fixes - `_wait_for_tool_call`: accept a `token` kwarg so multi-user tests can poll/approve through a per-user identity. Backwards-compatible. - `test_mcp_same_server_multi_user_via_browser`: drive approval through the per-user API while waiting for the tool to land. Drop the broken `expected_text_contains` predicate and the tied "Mock MCP search result" text assertions; the bearer-token isolation assertion (what this test actually exists to prove) is retained and unaffected. - `test_wasm_tool_oauth_refresh_on_demand`: insert a `_wait_for_tool_call` approval step between `_send_chat` and `_wait_for_refresh_request` so the http credential layer actually runs. - `test_wasm_tool_first_chat_auth_attempt_emits_auth_url`: marked xfail with an inline reason pointing at #2868 and the replacement coverage (`test_v2_tool_activate_surface.py`, `test_settings_first_gmail_auth_then_chat_runs`). - Drop the now-unused `send_chat_and_wait_for_terminal_message` import. ## conftest fix `ironclaw_server` now sets `SECRETS_MASTER_KEY` in the spawned env. On macOS without it, `auto_generate_and_persist` blocks on a Keychain authorization prompt that no one's home to click, so `wait_for_ready` times out at 60s and the fixture kills the process with SIGKILL — making any session-scoped browser test impossible to run locally. On Linux, the keychain backend errors fast and the auto-generate fallback writes to `.env`, so CI was unaffected. Setting the key explicitly matches the pattern already used in `auth_matrix_server`, `test_v2_engine_auth_cancel`, `test_v2_tool_activate_surface`, etc. ## Verification Local repro confirmed each failure mode (HTTP-only repro for the non-browser tests, server-side log inspection for the multi-user test). Reproduced the exact pending_gate=approval pattern, fixed it, verified the assertion semantics still hold: ``` $ pytest tests/e2e/scenarios/test_v2_auth_oauth_matrix.py::test_wasm_tool_oauth_refresh_on_demand PASSED in 6.74s $ pytest tests/e2e/scenarios/test_v2_auth_oauth_matrix.py::test_wasm_tool_first_chat_auth_attempt_emits_auth_url XFAIL in 93s $ pytest tests/e2e/scenarios/test_v2_auth_oauth_matrix.py -v --timeout=120 12 passed, 2 skipped, 3 xfailed (browser tests errored locally; they'll run cleanly in CI) ``` The browser-driven `test_mcp_same_server_multi_user_via_browser` couldn't be exercised locally (chromium can't launch under this shell sandbox), but the API + auto-approve flow it now relies on is exercised by an HTTP-equivalent repro and matches the pattern used in `test_settings_first_gmail_auth_then_chat_runs`. * test(e2e): set LLM_API_KEY in auth_sse_server fixture `test_auth_required_sse_without_duplicate_response` was failing in the merge queue on every PR (most recently bouncing #3197 and #3203 from the queue) because the `auth_sse_server` fixture never set `LLM_API_KEY` in the spawned ironclaw env. After #2572 added a missing- API-key check to the openai_compatible config validator (Apr 22), ironclaw rejected the env-supplied openai_compatible config, fell back to the NearAI default, hit "missing session token", and failed the turn before the github skill could even fire its 401. The chat thus reached `state: Failed` with no tool calls and no `onboarding_state/auth_required` event — which is exactly what the test asserted on, hence the consistent failure. Adding `LLM_API_KEY=mock-api-key` matches the value already used in every other e2e fixture (auth_matrix, conftest's ironclaw_server, v2_engine, etc.) and unblocks the assertion. Local run: PASSED in 8s. * fix(gateway): suppress duplicate assistant bubble after streamed response [skip-regression-check] The SSE `response` handler unconditionally called addMessage('assistant', data.content) even when stream_chunks had already populated and finalized a bubble for the same response. This stayed invisible in the common case but surfaced as a hard test failure under the path test_switching_back_preserves_in_progress_turn: 1. Send "What is 2+2?" on thread A — stream chunks start filling an assistant bubble with "data-streaming". 2. Switch to thread B mid-stream — container clears (history reload). 3. Switch back to thread A — history rehydration shows the in-progress turn with no response yet, so 0 assistant bubbles in DOM. 4. Stream chunks continue to fire for A — appendToLastAssistant creates a new bubble and accumulates the response into it. 5. response event fires — flushes any remaining buffer, removes the data-streaming flag (good) — then addMessage('assistant', content) creates a SECOND identical bubble. Result: locator(".message.assistant").filter(has_text="4") matches two elements, Playwright strict mode rejects the wait_for, the test fails. Outside the test, two identical bubbles render to the user. Fix: only call addMessage in the response handler when there was no in-flight streaming bubble. If one existed, the streamed content is already correct (chunks accumulate `data.content` verbatim) and the data-streaming flag has just been cleared. Non-streaming responses (no chunks fired) still take the addMessage branch. Regression coverage: tests/e2e/scenarios/test_message_persistence.py:: test_switching_back_preserves_in_progress_turn already reproduces this exact scenario and was failing in the merge queue. With this fix it passes; skip-regression-check used because the existing E2E test is the regression test, and the gateway doesn't have a JS unit test harness for SSE handler state. * fix(gateway): dedupe history-rendered SSE responses * test(e2e): set mock LLM API key in standalone fixtures * fix(e2e): make v2 approval tests deterministic * test(e2e): stabilize duplicate skill install assertion * test(e2e): assert duplicate install stays ungated * fix(skills): skip approval for disk-installed duplicates * fix(v2): honor no-op skill installs without approval * test(e2e): wait for pending send marker to clear --------- Co-authored-by: Firat Sertgoz <f@nuff.tech>
…arai#3157) * fix(engine): inline gate await for Tier 0 + Tier 1 Approval gates CodeAct scripts that hit a tool requiring approval surfaced as `RuntimeError: execution paused by gate 'approval'` inside the script instead of pausing for the user. The async tool-resolve path converted `EngineError::GatePaused` into a Python exception; the sync preflight path returned `need_approval` to the orchestrator, which on resume re-ran the LLM step and re-executed any non-idempotent earlier tool calls in the same script. Replace both with a host-supplied `GateController` that pauses the live execution in place. The Monty VM (Tier 1) and the Tier 0 batch loop both stay alive across the user's approval; on `Approved` the gated action re-executes (lease re-consumed, auto-approve installed before delivery so subsequent gates short-circuit); on `Denied` the script raises a typed `RuntimeError("user denied tool 'X': ...")` that the script can catch. Auth and External resume kinds keep the legacy thread re-entry path - their resolution installs new state (credentials, callback payloads) that only takes effect on the next run-through. Boot sweep invalidates `Approval`-kind `PendingGate` rows from a prior process so a stranded gate after restart fails fast instead of taking the legacy path and re-executing earlier mutations. Tests: 3 new regression tests in scripting.rs (approve / deny / no-controller fallback) and 4 in gate_controller.rs covering the resolution registry's one-shot, dropped-receiver, and unknown- request semantics. Pre-existing failures `call_id_preserved_when_no_lease` and `stop_thread_works` reproduce on staging without these changes. Design: docs/plans/2026-05-01-codeact-inline-gate-await.md. * test(engine): live regression for inline gate await with CodeAct Two integration tests in engine_v2_gate_integration.rs that exercise the inline gate-await flow end-to-end through `ThreadManager` → `ExecutionLoop` → orchestrator → CodeAct → `EffectExecutor` → `GateController`: 1. `codeact_inline_gate_await_resumes_user_reproducer` reproduces the exact reported bug shape: a CodeAct script issuing `await github_tool(action="search_issues_pull_requests", ...)` for "what are p1 bugs in nearai/ironclaw filed in last 7 days". The github_tool returns `EngineError::GatePaused` mid-execution; the test's `OneShotApprovingGateController` marks the effects mock approved and returns `Approved`; the engine retries inline; the tool succeeds; the script's `FINAL("Found 0 P1 bugs ...")` reaches the user. Asserts the controller saw exactly one pause request, github_tool was called twice, and the response is the script's FINAL — not the pre-fix `RuntimeError: execution paused by gate`. 2. `codeact_inline_gate_await_denial_does_not_retry` covers the deny path: controller returns `Denied { reason: "not now" }`. Asserts github_tool was called exactly once (no retry on denial), the typed `user denied tool 'github_tool': not now` message appears in the failure events, and the pre-fix `execution paused by gate` string does NOT appear anywhere. Also restored a fmt-only line shape in `gate_controller.rs` from `cargo fmt`. * test(engine): use realistic mock issues in inline gate-await fixture The pre-fix fixture returned `{"items": []}` which made the script's FINAL emit "Found 0 P1 bugs in nearai/ironclaw" — misleading, since the repo actually has open P1 issues (e.g. nearai#2818, nearai#2997). Update the mock to return two such items and tighten the assertion to verify the exact count flowed from tool result through CodeAct to FINAL(). * refactor(engine): require gate_controller, bound retry, drop V1 fallback Removes the `Option<Arc<dyn GateController>>` foot-gun: the field's `None` arm in the executors silently re-emitted the original `"execution paused by gate 'approval'"` RuntimeError, which is exactly the bug this PR exists to fix. With the field required and a named `CancellingGateController` as the explicit drop-in for non-pausing paths (post-resolution replay, mission protected writes, tests), forgetting to wire a controller is a compile error. Other follow-ups in the same change to keep them on one commit: - `MAX_INLINE_GATE_RETRIES = 3` shared by `scripting::drive_inline_gate` (Tier 1 async output) and `structured::execute_with_inline_gate_retry` (Tier 0 mid-execution). A misbehaving tool that keeps gating after each approval surfaces a clean error instead of pinning a CPU. - `denial_reason_for_resolution` helper centralizes the `GateResolution -> reason` mapping so denial messages can't drift between Tier 0 and Tier 1. - `invalidate_stranded_approval_gates_evicts_only_approval_kind` unit test covers the boot sweep with a mixed Approval/Auth/External population. - Existing `codeact_gate_without_controller_falls_back_to_runtime_error` test rewritten as `codeact_default_controller_cancels_approval_gates` to assert the inverted invariant: the legacy bug message must NEVER appear, even with the inert default controller. - Design doc updated to reflect as-shipped shape (required field, the bounded-retry constant, denial helper, `max_duration` stays at 30 s). * fix(engine): bound inline pause, propagate one-shot approval, race fix Addresses review on PR nearai#3157 (serrrfirat + Copilot + gemini-code-assist). Four blocking correctness fixes: 1. **Bounded BridgeGateController::pause.** The await on the resolution oneshot now races against `pending.expires_at`. Without this, a user ignoring the prompt past expiry would strand the engine: the DB row expires, the oneshot stays open, the VM keeps running. On expiry we discard the pending row, drop the registry entry, and return `Cancelled` so the VM unwinds cleanly. 2. **One-shot approval threaded through retry.** Add `call_approval_granted: bool` on `ThreadExecutionContext` (default false). Inline retry paths (`drive_inline_gate`, `execute_with_inline_gate_retry`, `execute_single_action_with_inline_retry`, scripting sync preflight) set it to true on the retry call so `EffectBridgeAdapter::execute_action` forwards it as `approval_already_granted=true` to the host's tool approval check. Mirrors the legacy `execute_resolved_pending_action` contract; without this, tools with `ApprovalRequirement::Always` gated again on every retry until the bound tripped, and `always=false` AskEachTime gates re-prompted immediately after approval. 3. **Per-execution context registration race fixed.** `set_execution_context` was called AFTER `handle_user_message().await` returned the thread_id — but the engine task is already running, so a fast tool gate could reach `pause()` before the entry existed and get `Cancelled`. Now the bridge calls `set_pre_execution_context` (per-user) BEFORE `handle_user_message`, then promotes to (user, thread)-keyed once thread_id is known. `pause()` falls back to the per-user entry on miss. 4. **Tier 0 parallel-batch coverage.** New `execute_single_action_with_inline_retry` wraps `execute_single_action` with the same bounded retry shape used by Tier 1. Both the single-runnable and multi-runnable branches of `handle_execute_actions_parallel` go through it, so simultaneous gates in a parallel batch no longer fall through to the legacy re-entry path. Smaller fixes: - PROJECTION lint annotations on the two `broadcast_for_user` sites added by this PR (gate_controller emit_gate_prompt; resolve_gate inline-await fast-path resolution event). Test coverage: - `GatingThenOkEffects` now records `context.call_approval_granted` per call; `codeact_gate_inline_await_approved_delivers_result` asserts the retry observes `true`. Locks in the one-shot approval propagation contract. * test(engine): update gate integration tests for inline-await semantics CI was failing on 5 tests in engine_v2_gate_integration.rs that asserted the legacy `ThreadOutcome::GatePaused` unwind for `Approval` gates. With inline-await + the required `CancellingGateController` (default), Approval gates resolve inline and the thread completes (or fails) instead of pausing — these tests were exercising the pre-PR flow that this PR replaces. - `gate_paused_transitions_thread_to_waiting` → renamed to `approval_gate_resolves_inline_via_controller`. Wires `AutoApprovingGateController`, asserts thread completes after inline approval, ApprovalRequested + ActionExecuted both recorded. - `gate_paused_thread_resumes_to_completion` → renamed to `approval_denied_inline_completes_thread_with_failed_action`. Asserts the default `CancellingGateController` cancels the gate and the thread completes with a failed action (no stranded pending gate). - `approval_chains_directly_into_auth_for_install_flow` rewritten: Approval handled inline by controller, Auth gate (still legacy path) bubbles up as ThreadOutcome::GatePaused, legacy auth-resume drives completion. - `approval_resolution_executes_pending_call_directly` and `gate_resume_with_execution_obligation` reduced to no-op stubs with inline rationale documenting where the post-PR equivalent coverage lives. Removing entirely would erase the breadcrumb in git log. Also threads the `ApprovalRequested` event through `execute_single_action_with_inline_retry`: the wrapper now returns `Vec<EventKind>` per call so the caller can append both the approval prompt and the post-retry outcome to the thread event log. Without this, observers saw only the final ActionExecuted/Failed event with no record that a gate had fired. Drops a few dead test helpers (`ApprovalTool`, `make_caps_with_approval_tool`, unused imports) that the deleted assertions no longer reference. Quality gate: fmt clean, clippy zero warnings, 513 engine lib + 443 bridge + 29 gate integration tests pass (2 pre-existing engine-lib staging failures unrelated to this PR). * test(engine): update skill_codeact integration tests for inline-await Two more tests in engine_v2_skill_codeact.rs were asserting the legacy `ThreadOutcome::GatePaused` → `resume_thread` flow for `Approval` gates. With PR nearai#3157 the engine catches Approval gates inline and the thread runs to completion in a single `join_thread`. - `skill_prompt_context_survives_pause_and_resume`: wires `AutoApprovingHttpController`, asserts thread completes after inline approval, drops the now-impossible `resume_thread` step. - `skill_prompt_context_survives_compaction_and_resume`: same; also drops the assertion on the persisted transcript at the *pause point* (no longer externally observable). The load-bearing post-compaction LLM-call assertions remain — they're the actual contract this test was protecting. Adds an `AutoApprovingHttpController` test helper local to this file that approves gates by marking the underlying `PausingHttpMockEffects` approved before returning `Approved`. * fix(engine): bridge cleanup on dropped sender + ActionFailed on preflight denial Addresses Copilot review on PR nearai#3157. - `BridgeGateController::pause`: when the resolution oneshot's sender is dropped (process shutdown / registry cleared), discard the `PendingGate` row before returning `Cancelled`. Without this, a stranded prompt remained visible in the UI and `pending_gates.insert` rejected duplicates for the same `(user, thread)` so a follow-up gate could not register. Same cleanup as the expiry branch already performed. - `scripting.rs` sync-preflight denial path: emit `EventKind::ActionFailed` before resuming Monty with `RuntimeError`, so the thread event log is consistent with the other denial paths (`drive_inline_gate`, `structured.rs`). Auditing why a tool didn't run is now possible from the events alone. - Delete the two empty `#[tokio::test]` stubs left in `engine_v2_gate_integration.rs` (`approval_resolution_executes_pending_call_directly_via_resolved_pending_action`, `gate_resume_with_execution_obligation`). The post-PR equivalent coverage is in `codeact_inline_gate_await_*` (this file) and the scripting unit tests; the rationale is preserved in `git log` (commit 87fe4fc) without an empty test slot misleading coverage signals. * fix(engine): finish inline-await migration; remove legacy gate-paused for Tier 0 policy gates Address PR nearai#3157 review: - router.rs: clear pre_execution slot on handle_user_message error so a failed engine spawn doesn't leak a stale (user, conversation) entry that would mis-route the next gate prompt. - router.rs:await_thread_outcome: on the 5-min request deadline, return BridgeOutcome::Pending instead of join_thread() — joining would block for up to the gate's 30-min expires_at when the parked task is in pause(). The PendingGate row stays live for resolution. - gate_controller.rs: re-key pre_execution by (user_id, conversation_id) instead of user_id alone. Two concurrent conversations / browser tabs for the same user no longer clobber each other's slot. Plumbs conversation_id through ThreadExecutionContext and GatePauseRequest so pause() can match a gate to its originating conversation. - gate_controller.rs: serialize concurrent inline gates per (user, thread) via a per-key tokio Mutex held across the PendingGateStore::insert + select-await window. A parallel batch where two tools both gate now queues the second behind the first rather than silently surfacing it as Cancelled on (user, thread) uniqueness collision. - orchestrator.rs (Tier 1 / CodeAct): remove the legacy gate_paused JSON sentinel + thread re-entry path for Approval gates. Both __execute_action__ and __execute_actions_parallel__ now pause inline on PolicyDecision::RequireApproval (mirroring structured.rs) and route tool-raised gates through the existing execute_single_action_with_inline_retry wrapper. Authentication and External resume kinds keep the legacy re-entry path because their resolution installs new state that only takes effect on the next thread run-through. Test fixtures across bridge/effect_adapter, action_projector, and the gate/sandbox integration suites get the new conversation_id field defaulted to None. Refs: comments 3173757791, 3176248954, 3176248976, 3176249001, 3176249032 * fix(engine): defer post-Pending context cleanup; route parallel JoinSet through inline-retry Address PR nearai#3157 review on commit d211bfc: - router.rs: when await_thread_outcome returns BridgeOutcome::Pending and the engine task is still running (typically parked in BridgeGateController::pause), defer clear_execution_context to a spawned watcher task that polls is_running until the thread completes and then clears the (user, thread) context + gate-locks entry. Without this, the unconditional clear after a 5-min request timeout stranded the parked thread: the eventual gate resolution would call pause() for any subsequent gate with no registered context and surface as silent Cancelled. Watcher caps at 60 minutes (well past the 30-min PendingGate expiry) as a defensive safety bound. - structured.rs: route the multi-runnable JoinSet branch in execute_action_calls through execute_with_inline_gate_retry, matching the single-runnable fast path. Without this, an Approval gate raised mid-execution in a parallel batch with >1 runnable tool call bubbled out as Err(GatePaused) and went through the legacy gate-paused / re-entry path, re-introducing the double-execution bug for already-completed sibling calls in the same batch. The signature of execute_action_calls switches leases from &LeaseManager to &Arc<LeaseManager> so the JoinSet tasks can clone an owned handle; all current callers are tests already constructing Arc::new(LeaseManager::new()), so this is a no-op call-site change. Refs: comments 3176333657, 3176333685 * test(engine): fix call_id_preserved_when_no_lease MockEffects inventory PR nearai#2868 added a callable-inventory check ahead of the lease lookup in execute_action_calls. The test was constructing MockEffects with an empty action list, so preflight short-circuited on "action is not callable in this execution context" before reaching the lease branch the test was actually trying to exercise. The assertion error.contains(\"no lease\") then failed against the not-callable message. Populate the mock with test_action(\"web_search\") so the inventory gate passes and the call reaches the lease check, matching the test's documented intent. * fix(engine): use gate-provided params in inline-await pause; drop Thread clone in parallel branches - orchestrator.rs::execute_single_action_with_inline_retry now sources GatePauseRequest.parameters from the gate-paused payload (which reflects safety-layer transformations/redactions) rather than the original caller params, matching structured.rs::execute_with_inline_gate_retry. - Both inline-retry helpers now take ThreadId + user_id instead of &Thread, so the parallel JoinSet branches no longer clone the full Thread (with message/event transcripts) per spawned task. The per-task ThreadExecutionContext clone is what the helpers actually need; in orchestrator.rs we build a base ctx once and override current_call_id per task instead of re-running thread_execution_context. Addresses Copilot PR nearai#3157 review comments 3176594866, 3176594886, 3176594898. * fix(engine): address serrrfirat review — audit event + stop-during-wait Three latest serrrfirat review threads on PR nearai#3157: * Medium: structured inline approval drops ApprovalRequested audit event (REAL — fixed). `execute_with_inline_gate_retry` previously swallowed `Err(GatePaused)` and returned only the post-retry outcome, so `classify_exec_result` never saw the gate and the `ApprovalRequested` event was lost. The orchestrator (Tier 1) path emits both events. Mirror that shape here: - `execute_with_inline_gate_retry` now returns `(Result<ActionResult, EngineError>, Vec<EventKind>)`. The Vec carries one `ApprovalRequested` per retry iteration that gated. - Slot type widens to `(ActionResult, EventKind, Vec<EventKind>)` so the merge phase flushes pre-terminal events before the terminal event in original-call order. - Two regression tests pin the contract (approved → Executed, denied → Failed); both assert ApprovalRequested precedes the terminal event. * High: stop/cancel does not unblock thread parked in inline gate await (REAL — fixed). `BridgeGateController::pause()` selected only on the resolution oneshot + 30-min expiry; `stop_thread()` sent `ThreadSignal::Stop` but the parked engine task wasn't polling the signal channel. - Adds `GateController::cancel_thread(thread_id)` to the engine-side trait (default no-op). - `ThreadManager::stop_thread()` calls `cancel_thread()` BEFORE sending `ThreadSignal::Stop` so any parked `pause()` future wakes promptly with `GateResolution::Cancelled`. - `BridgeGateController` tracks in-flight pauses in `active_pauses: HashMap<ThreadId, HashSet<Uuid>>` and walks that set on cancel, delivering `Cancelled` via the existing `GateResolutions::try_deliver` channel and discarding pending rows via `PendingGateStore::discard_for_thread`. - `pause()` always untrack on exit (idempotent — `cancel_thread` may have already removed the entry). - Two new bridge-tier unit tests: parked pause wakes within 2s of cancel_thread; cancel_thread on a thread with no parked pause is a no-op. * Medium: live inline gate waits are uncapped per user/process (PARTIAL — push back, follow-up). The reviewer's own comment notes this is a design-doc follow-up, not a blocker. The current bound is implicit: one pending gate per (user, thread) × the existing thread-creation budget × the 30-min expiry, which is enough to ship the inline-await substrate. Adding a per-user semaphore correctly requires designing the cap UX, fairness, and rejection error shape — out of scope for this PR. Add a TODO inside `pause()` pointing at the design doc and the follow-up issue, with the reasoning written down so the next contributor doesn't have to re-derive it. Pre-existing on this branch (NOT introduced by these changes): `runtime::manager::tests::stop_thread_works` flakes on the branch; the original PR commit message acknowledges "stop_thread_works reproduce on staging without these changes." Confirmed by stashing this commit and running the test: still fails. Out of scope here. Test totals after this commit: - executor::structured: 19 passing (incl. 2 new) - bridge::gate_controller: 6 passing (incl. 2 new) - cargo fmt + clippy --lib clean Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(engine): typed DenialOutcome + plug exhaustion-path lease leak Three review-driven fixes to the inline gate-await path. 1. CancellingGateController and bridge expiry/shutdown surfaced as `RuntimeError("user denied tool 'X': cancelled")` — misleading because the user never saw a prompt. Replace `Option<String>` helper with a typed `DenialOutcome { DeniedByUser, Unavailable }` so cancelled/expired/no-handler gates render as "approval for tool 'X' unavailable: …" while real user denials keep the "user denied" framing. Updated all six call sites in scripting.rs / structured.rs / orchestrator.rs through the typed surface so wording can't drift. 2. Both `execute_with_inline_gate_retry` and the orchestrator's `execute_single_action_with_inline_retry` consumed a fresh lease use on the final approved iteration, then exited the loop without ever calling execute_action. The exhaustion-path error didn't trip the caller's refund check, so a misbehaving tool that gates after every approval would slowly drain `max_uses`. Refund the unused lease before returning. 3. Drop the dead `parameters: serde_json::Value` field on `PendingFuture::Tool` and the matching `_parameters` arg on `resolve_tool_future`. The gate's own parameter snapshot is the source of truth on retry; threading the original through made the signature read like there was a second source. Quality gate: cargo fmt, cargo clippy --all --benches --tests --examples --all-features (clean), cargo test -p ironclaw_engine --lib (520 pass), cargo test --test engine_v2_gate_integration (27 pass), cargo test --lib bridge:: (456 pass). --------- Co-authored-by: Nikolay Pismenkov <nickpismenkov@gmail.com> Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…arai#3589) Both xfails in tests/e2e/scenarios/test_v2_auth_oauth_matrix.py are stale and pass against current code: - test_wasm_tool_first_chat_auth_attempt_emits_auth_url Marked xfail in nearai#3235 because the engine-v2 callable-only contract (nearai#2868) stopped emitting an auth gate on direct LLM-driven tool calls. PR nearai#3157 (auth-preflight + inline-await) restored the behavior the test asserts: when the LLM emits a direct call to a not-yet-authed extension, the bridge raises an Authentication gate with auth_url populated (src/bridge/effect_adapter.rs:1356-1392). Marker removed; test passes. - test_settings_first_custom_mcp_auth_then_chat_runs The xfail reason claimed post-auth tool-output propagation was broken. Real cause: engine-v2 gates the first MCP tool call on `approval` and the browser fixture has no auto-approve UI, so the chat sat in pending_gate forever. Same shape as the bugs fixed in nearai#3235 for test_wasm_tool_oauth_refresh_on_demand and test_mcp_same_server_multi_user_via_browser. Inserted _wait_for_tool_call between _send_chat and _wait_for_response_contains to drive approval through the API; test passes. Verified locally: both tests pass back-to-back in 27s on a fresh auth_matrix_server. Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
… + auto-approve footgun (#3533) (#3559) * fix(extensions): restore chat-driven tool_install + fix double-invoke + auto-approve footgun (#3533) "Connect my telegram" was giving the user two options and not actually installing anything because three layered issues had accumulated since engine v2: 1. **`tool_install` was hidden from the agent** (#2868). The unified `tool_activate` it was meant to be subsumed by was later removed in #3166, but the hidden-from-callable-surface gate stayed. Restored by dropping `hidden_from_model_callable_surface` from `bridge::action_projector`. User consent is mediated by the tool's own `ApprovalRequirement::UnlessAutoApproved` and the seeded `AskEachTime` permission. 2. **Two competing Telegram registry entries** (`telegram` channel and `telegram_mtproto` tool) both surfaced in the agent prompt's `Activatable Integrations` section. The LLM correctly enumerated them as "Option 1" and "Option 2" instead of installing the canonical bot channel. Added a `hidden: bool` field to `ExtensionManifest` / `RegistryEntry`, set `telegram_mtproto` to `hidden: true`, and filter hidden entries out of the "available-but-not-installed" appendix in `ExtensionManager::list`. Hidden entries remain installable by explicit name. 3. **Updated the agent prompt** so `Activatable Integrations` instructs the model to call `tool_install(name="<name>")` directly rather than describing manual UI steps. Fixes the double-`tool_install` invocation that surfaced once the agent could install from chat: - **`InlineGate` discarded cached output.** The bridge raised an Authentication gate after `tool_install` succeeded, and the inline-await retry re-executed the action (re-downloading the WASM bundle) instead of returning the already-computed output. Added `resume_output: Option<serde_json::Value>` to `InlineGate`; on approval, return the cached output if present. Mirror fix in the orchestrator's `execute_single_action_with_inline_retry` (reading `result_json["resume_output"]`) and the structured-batch retry path. - **`effect_adapter::auth_gate_from_extension_result`** now passes `Some(output_value.clone())` as the gate's `resume_output` so the retry has cached state to short-circuit on. - **OAuth callback double-fired.** `oauth_callback_handler` now skips the `ExternalCallback` re-entry when the inline-await path already woke a parked waiter — eliminates the "thread already running" race. - **`resolve_inline_gates_for_credential`** now also discards matching Authentication rows from `pending_gates` so the row doesn't linger in `HistoryResponse.pending_gate` after inline resolution. Fixes the auto-approve footgun: - **`ToolPermissionSnapshot::resolve_permission`** now collapses DB values that match the seeded default to `explicit = None`. Before this, the boot-time `seed_tool_permissions` write of `tool_install -> AskEachTime` was indistinguishable from a user-explicit override, causing `effect_adapter::enforce_tool_permission`'s `is_explicit_ask` check to refuse `AGENT_AUTO_APPROVE_TOOLS=true`. Real overrides (`AlwaysAllow`, `Disabled`) still surface as `Some(...)`. Tests - Unit: 4980/4980 pass (host) + 525/525 pass (engine). - Unit: new tests in `bridge::tool_permissions::tests` lock seeded-vs- explicit collapse; new test in `bridge::action_projector::tests` asserts `tool_install` is callable; new manifest hidden-flag tests in `registry::manifest::tests` and `extensions::manager::tests`. - E2E: removed `@pytest.mark.xfail` on `test_chat_first_gmail_installs_prompts_and_retries` (now passes end-to-end via the chat-driven install path). Added `test_chat_install_approval_then_auth_card` driving the explicit-approval variant with a single Approve click (no Always workaround needed) — wired into the `auth-full` canary lane. - Mock LLM: extended the gmail-install-then-retry pattern to recognize both the legacy "Extension not installed:" and the post-#3533 "is not callable in this execution context" error strings, and to retry `gmail(action="list_messages")` after a successful `tool_install`. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(permissions): address #3559 review (permission bypass, lease accounting, hidden search filter) Five fixes from the #3559 review (4× Copilot doc nits + 3× serrrfirat security/correctness findings): 1. **Permission bypass (High).** Pre-#3559's `resolve_permission` collapsed any DB row whose value matched the seeded default to `explicit = None`, so a user who deliberately set `tool_install = AskEachTime` had their explicit choice silently dropped and `AGENT_AUTO_APPROVE_TOOLS=true` bypassed the gate. Provenance is now handled at write time: `seed_tool_permissions` is gone and a one-shot, sentinel-gated migration (`cleanup_ghost_seeded_tool_permissions`) deletes existing ghost-seeded rows at startup. With no ghost rows, the resolver treats every DB row as user-explicit and honors it. 2. **Lease/event accounting on `resume_output` replay (Medium).** Inline-gate handlers in `structured.rs`, `scripting.rs` (`resolve_tool_future` + `drive_inline_gate` retry loop), and `orchestrator.rs` refunded the lease use the action just consumed, then returned the cached `resume_output` on approval without re-consuming — netting successful side-effecting actions to zero lease uses. Skip the refund when the gate carries cached output. 3. **Hidden registry filter on `tool_search` (Medium).** `RegistryCatalog::search` did not filter `hidden: true` entries, so `telegram_mtproto` could resurface through the search path and reintroduce the "two Telegram options" outcome that #3533 fixes for the default-list path. Added the filter and a regression test. 4-7. Copilot doc nits: outdated `_set_tool_permission` docstring; misleading "bridge-side auto-install implemented" comment in `mock_llm.py`; `tool_install` described as "non-agent surface" in `src/bridge/CLAUDE.md` while a paragraph below says the model calls it directly; dangling `issue #3533 / PR —` placeholders in both CLAUDE.md docs. Regression tests: - `bridge::tool_permissions::user_explicit_value_matching_seeded_default_stays_explicit` — the original Copilot/serrrfirat bug case. - `app::cleanup_ghost_seeded_tool_permissions_removes_seed_matching_rows` — idempotent migration + sentinel. - `extensions::registry::test_search_skips_hidden_entries` — hidden entries excluded from search but still installable by exact name. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(#3559): caller-level regression coverage for review findings 1 & 2 Two follow-up regression tests for the #3559 security review, plus a real bug surfaced by the first one. 1. `executor::structured::resume_output_replay_consumes_exactly_one_lease_use` exercises the post-execution Authentication gate inline-retry path with `max_uses=1` and asserts: - Cached output is returned as a successful `ActionResult`. - Exactly one `ActionExecuted` event is emitted. - The lease budget is exhausted after one execution (refund-skip keeps the consumption from being undone). Writing this test surfaced a real bug: the structured cached-output branch pushed `ActionExecuted` into `emitted_events`, and the caller's `classify_exec_result` emitted ANOTHER terminal `ActionExecuted` for the same Ok result — double-emit for one action. Tier 1 (`scripting::drive_inline_gate`) and Tier 1 alt (`orchestrator::execute_action_with_inline_gate`) emit themselves because their callers don't run an Ok-branch classifier; structured was the outlier. Dropped the redundant push; the classifier emits the single canonical event. 2. `bridge::effect_adapter::explicit_ask_each_time_for_seeded_default_tool_still_gates` drives `execute_action` end-to-end (the side-effecting caller) with a tool whose `name()` matches a seeded-`AskEachTime` baseline (`tool_install`) and an explicit `AskEachTime` user override. The resolver collapse-to-implicit bug would have shown up here — not just in the helper-level test that already exists in `bridge::tool_permissions::tests`. Per `.claude/rules/testing.md` "Test Through the Caller, Not Just the Helper". Added `SeededAskEachTimeTestTool` as a `tool_install`-named test fixture with `requires_approval: UnlessAutoApproved` to mirror the real tool's contract. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(e2e): restore auth and approval coverage (#3430)
* test(e2e): avoid REPL auth retry race (#3437)
* feat: add pairing_approve tool for Slack binding via chat (#3396)
* feat: add pairing_approve tool for Slack binding via chat
Users can now paste their Slack pairing code in the IronClaw chat and
the LLM will call pairing_approve to bind their accounts. No need to
use the API directly.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: require approval before pairing + fix formatting
Address review comment: pairing_approve now requires UnlessAutoApproved
approval before executing, preventing accidental account binding.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* test: add regression test for pairing_approve tool
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: address review — Always approval, lock channel, add to protected list
1. Changed ApprovalRequirement to Always (not bypassable by auto-approve)
2. Locked channel to slack-relay constant (removed generic channel param)
3. Added pairing_approve to PROTECTED_TOOL_NAMES
4. Added tests: always-approval, protected-name, channel constant
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(web): isolate cross-tenant SSE/WS status events and thread access (#3390)
* fix(web): isolate cross-tenant SSE/WS status events and thread access
Plug a multi-tenant leak where unscoped `sse.broadcast(...)` calls
from `GatewayChannel::send_status`, sandbox `JobEvent` dispatch, the
WASM/Slack OAuth completion handlers, and any producer that lost
`metadata.user_id` along the way fan out to every connected
subscriber — exposing another tenant's tool calls, tool output,
onboarding state, and job lifecycle to anyone with an open SSE/WS.
Changes
- Extract `dispatch_status_event(sse, multi_tenant_mode, user_id, ev)`
from `Channel::send_status`. In multi-tenant mode an unscoped event
is dropped (with a WARN naming the producer to fix); single-tenant
keeps the global broadcast since there is one subscriber population.
- `IncomingMessage::new` now defaults `metadata` to `{"user_id": ...}`,
and `with_metadata` preserves the key so downstream `send_status`
consumers always have an owner to scope by.
- WASM/Slack OAuth completion broadcasts route through
`broadcast_for_user(&owner_id, ...)`. Sandbox `JobEvent` dispatch
in `main.rs` respects `multi_tenant_mode` for the empty-`user_id`
fallback.
- New pre-commit check #10 (`MULTITENANT`) flags unscoped
`sse.broadcast(...)` lines without a `// multi-tenant-safe: <reason>`
marker or a transport-only exemption. Marker regex accepts the marker
anywhere in a `//` comment so compound annotations on a single line
work.
Tests
- `src/channels/web/tests/status_event_isolation.rs` — 5 unit tests
covering both modes and the per-variant drop invariant.
- `src/channels/web/platform/sse.rs` — 2 quadrant tests for the
`subscribe_raw` filter (scoped/unscoped × matching/mismatched).
- `tests/thread_isolation_integration.rs` — 9 HTTP-level checks that
Bob cannot reach Alice's chat history (paginated and not), threads
list, engine v2 detail/steps/events, or Responses GET, plus an
unauthenticated-rejection guard.
- 8 new self-test cases for the `MULTITENANT` script check, including
the compound projection-exempt + multi-tenant-safe annotation case.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(web): pin cross-tenant boundaries on jobs, files, routines
Audit of the protected route surface for the same bug shape #3390 fixed
(handler that takes a user-controlled id and reads without an ownership
predicate) found that the implementations were correct but four
boundaries had no integration test. Lock them in before they regress.
- Sandbox job persisted-events history (`/api/jobs/{id}/events`):
Bob → 404 on Alice's job; Alice → 200 on her own.
- Sandbox job workspace listing (`/api/jobs/{id}/files/list`):
Bob → 404 on Alice's job.
- Sandbox job file read (`/api/jobs/{id}/files/read`): Bob → 404 on
Alice's job; Alice → 200 on her own; Alice → 403/404 on
`?path=../outside.txt` (path-traversal pin against the
`canonicalize() + starts_with(base_canonical)` guard).
- Routine run history (`/api/routines/{id}/runs`): Bob → 404; Alice
→ 200 with at least one seeded run.
The OAuth-state and NEAR-nonce stores were also flagged in the audit
but neither is a real cross-tenant bug: both are pre-auth, single-use,
and the token IS the secret. Documenting here so a future audit
doesn't re-flag them.
Project-static (`/projects/{id}/...`) is left for a follow-up — it
relies on `ironclaw_base_dir()` which is a process-wide `LazyLock`,
making per-test override fragile in the integration runner.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(web): address PR #3390 review — forge-resistant metadata, OAuth toast routing
Addresses six review comments from gemini-code-assist, copilot, and
serrrfirat on PR #3390. False-positives and the perf nit on
`with_metadata` are explained in the reply thread, not changed in code.
- (HIGH, serrrfirat) `IncomingMessage::with_metadata` now ALWAYS sets
`metadata.user_id` from `self.user_id`, dropping any caller-supplied
value. A WASM channel emitting `{"user_id":"victim"}` via
`apply_emitted_metadata` can no longer reroute downstream
`ToolStarted` / `ToolResult` SSE events into another tenant's
stream. New unit tests pin the forgery-resistance invariant.
- (MEDIUM, copilot + serrrfirat) Slack relay OAuth callback now
broadcasts the completion toast to the resolved `oauth_user`
(the IronClaw user who initiated the flow) rather than
`state.owner_id`. In multi-tenant deployments those differ and the
previous routing delivered the toast to the wrong browser tab.
Extracted the lookup into `resolve_relay_oauth_user`; two unit
tests cover the secret-present and secret-missing cases.
- (MEDIUM, gemini) `dispatch_status_event` treats empty-string
`user_id` the same as `None` so producers that lost the field
along the way fail-closed instead of falling through to a global
broadcast in multi-tenant mode.
- (MEDIUM, gemini) `main.rs` sandbox JobEvent dispatch now reuses
`dispatch_status_event` instead of duplicating the drop / WARN /
broadcast policy. `dispatch_status_event` is bumped from
`pub(crate)` to `pub` so the binary crate can call it.
- (LOW, copilot) Pre-commit `MULTITENANT` self-tests gain three
cases (`state.sse.broadcast(`, `gw_state.sse.broadcast(`,
annotated receiver-prefixed) to lock the existing boundary regex
behaviour against future tightening.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(channels): preserve i64 metadata.user_id from Telegram in with_metadata
PR #3390's forge-resistance fix made `IncomingMessage::with_metadata`
*always* overwrite `metadata.user_id` with `self.user_id` as a String.
That broke the Telegram WASM channel: it persists Telegram's chat user
ID as `metadata.user_id: i64` and re-deserializes it into
`TelegramMessageMetadata { user_id: i64, ... }` in `on_respond` /
`on_status`. After the fix, `respond` blew up with
`invalid type: string "999", expected i64 at line 1 column 87`,
failing 3 Telegram integration tests in CI.
Narrow the carve-out: overwrite only when the existing `user_id` is a
String (or missing). Non-string values are channel-private and the
SSE routing layer reads via `as_str()` — non-strings already fail
closed in multi-tenant mode, so the forge threat (WASM emits
`{"user_id":"victim"}` as a string) is still mitigated, while
Telegram's i64 use case survives.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(web): redact WARN payload + harden dotdot traversal test (PR #3390)
Two follow-up fixes from the second pass of review on #3390:
1. `dispatch_status_event`'s WARN log used `?event`, which on
`AppEvent::Response` / `Thinking` / `ToolResult` carries
user-authored content into operator logs in a multi-tenant
deployment. Replace with `event_kind = event.event_type()`
(the wire-stable variant name) — enough to identify the
misbehaving producer without leaking tenant data. Picked up via
Copilot's review on `src/channels/web/mod.rs:666`.
2. `alice_job_file_read_rejects_dotdot_traversal` planted
`outside.txt` under `outer.path()` (the `start_server_with_db`
fixture's tempdir holding `test.db`) but probed
`?path=../outside.txt` relative to `alice_proj` — a separate
`tempfile::tempdir()` rooted at the OS temp directory. The two
paths were unrelated, so the test could pass even if `..`
traversal was permitted (probe just hit empty space). Build the
directory tree by hand instead: `parent/alice_proj/` with the
planted file at `parent/outside.txt`, so the probe deterministically
resolves to the planted bytes. Add a body-content assertion that
fails loudly if those bytes leak. Picked up via Copilot's review on
`tests/cross_tenant_resource_isolation.rs:355`.
Plus a `cargo fmt` fix for `src/channels/channel.rs:1202` that was
breaking the Formatting CI check.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(web): address PR #3390 follow-ups — multi-tenant fallback WARN, exhaustive variant pin
- `resolve_relay_oauth_user`: take `multi_tenant_mode`; emit WARN when the
`relay:{ext}:oauth_user` secret is missing in multi-tenant mode so the
unrecoverable-initiator case surfaces in operator logs. Single-tenant
fallback stays silent (owner == only user).
- `dispatch_status_event`: doc note clarifying the function is `pub` only
for the sandbox JobEvent rx loop in `main.rs`; not part of a stable
public API.
- `_compile_time_appevent_variant_check`: exhaustive-match helper paired
with `unscoped_drop_holds_for_every_status_variant_in_multi_tenant`.
Adding a new `AppEvent` variant now fails the test build, prompting an
update to both the helper and the runtime leak-candidate list.
- `tests/thread_isolation_integration.rs`: honest scope note on the
engine-v2 thread tests — they pin handler shape (404/empty for
unknown id), not the cross-tenant ownership branch. Cross-tenant
engine-v2 coverage requires an `ENGINE_STATE` test fixture; tracked
as a follow-up in the comment block.
Tests: 8 unit (5 status_event_isolation + 3 resolve_relay_oauth_user)
and 9 thread_isolation_integration pass; clippy clean; pre-commit
safety scripts pass (regression suite 27 cases).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(web): unconditionally consume relay:{ext}:oauth_user secret in OAuth callback
The previous cleanup site lived inside the `if let Some(pairing_store)`
branch of the result block, which was unreachable on three failure
paths:
1. `pairing_store` is None (no identity pairing wired up)
2. The result block `?`-short-circuits before reaching the `if let`
(e.g. `set_setting` fails, `activate_stored_relay` fails, an inner
`relay_config()` / `list_connections()` errors)
3. The `if let` body itself errors before reaching the delete (e.g.
`list_connections` returns no matching team)
Leaving the secret behind lets a subsequent OAuth callback for the
same extension read a stale initiating user and misroute the
completion toast — Copilot review on PR #3390 (comment id 3211833864).
Move the delete to right after `resolve_relay_oauth_user` returns,
where it always runs once the value has been captured, regardless of
downstream failure mode. Updated the inner comment to document that
the secret is already gone by the time the pairing branch reads
`oauth_user`.
Regression test: `test_relay_oauth_callback_consumes_oauth_user_secret_on_failure_path`
seeds the secret, fires the callback against a fixture with no real
relay backend (so the result block deterministically errors), and
asserts the secret is gone afterward. Pre-fix this would have left
the secret behind on the failure path.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(auth): tighten Telegram pairing UX and OAuth-failure recovery (#3317, #3319, #3320) (#3381)
* fix(auth): tighten Telegram pairing UX and OAuth-failure recovery (#3317, #3319, #3320)
Three Bug Bash P1 issues from the same user journey: setup → use → fail.
The unifying root cause was per-channel auth tested in isolation; cross-channel
flows (Telegram → Gmail OAuth → resume) had no coverage and three small leaks
combined into a stuck conversation.
#3317 — Telegram pairing reply now names every IronClaw surface explicitly
(web settings, agent chat, terminal). The agent submission parser learns
`approve <channel> <code>`, dispatched through a new bridge handler that
mirrors `POST /api/pairing/{channel}/approve`.
#3319 — OAuth callback failures now log a category + correlation ID so a
user-reported "I saw 400" maps to one log line. Adds the
`OauthCallbackFailure` enum and `oauth_failure_correlation_id` helper.
#3320 — Two cleanup gaps fixed: (a) `/clear` now drains
`pending_oauth_flows` for the user (otherwise stale flows linger 5min and
mask new auth attempts); (b) OAuth provider-error and exchange-failure paths
now auto-cancel the engine pending auth gate via `clear_engine_pending_auth`,
so the conversation isn't blocked waiting for a resume that will never arrive.
Tests: 5 new submission-parser tests, 2 new bridge-handler tests, and one
new OAuth callback test verifying the pending-flow drain on provider error.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(auth): cross-channel pairing claim coverage + canary lane (#3317)
Adds the structural coverage that was missing when #3317 shipped:
- E2E (`tests/e2e/scenarios/test_telegram_pairing_chat_claim.py`):
three scenarios that drive the full Telegram pairing flow through
the gateway. Asserts the bot reply names every IronClaw surface
(web Settings, agent chat, terminal CLI), drives `approve telegram
CODE` through `/api/chat/send` and verifies the paired user
exchanges messages without re-prompting, and confirms invalid
codes get a clear rejection instead of an LLM-improvised reply.
- Rust integration (`tests/telegram_pairing_chat_claim_integration.rs`):
drives `Submission::PairingClaim` through a real `Agent` →
`bridge::handle_pairing_claim` → `PairingStore::approve` chain
using `TestRig` with engine v2 enabled. Covers the happy path
(mints a code, claims it via chat, asserts `Pairing approved`)
and the invalid-code rejection. The unit tests in `bridge/router`
cover only the no-extension-manager and invalid-channel branches —
this test exercises the wiring between submission parser, agent
loop dispatch, and bridge handler that #3317 specifically broke.
- Canary (`scripts/live_canary/auth_registry.py`): adds the two
user-visible scenarios to `AUTH_CHANNEL_TESTS` so the auth-channels
lane (scheduled every 6h) catches the regression class in CI
before any real user encounters it.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* docs: align pairing-claim and oauth-correlation comments with code (#3381)
Three Copilot review comments on PR #3381 flagged docstring/code drift in
already-merged PR #3317/#3319/#3320 changes. No behavior change — only
the doc strings move:
- `Submission::PairingClaim.code` and the inline `approve <channel> <code>`
parser comment claimed the user's casing was preserved, but the parser
builds the code from `lower` and the regression tests already lock in
the lowercased shape (`code == "abc12345"`). Update both comments to
describe the actual normalize-then-store contract.
- `oauth_failure_correlation_id` claimed the correlation appeared in the
user-facing error subtitle, but the failure path renders
`landing_html(label, false)` whose subtitle is fixed and never receives
the correlation. Mark the helper as logs-only and note that plumbing
the ID through the HTML is a follow-up.
[skip-regression-check] doc-only, behavior already covered by existing
pairing-claim parser tests in submission.rs.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* chore(telegram): bump channel registry to 0.2.11
The pairing-reply wording was updated in channels-src/telegram/, which
the version-check CI requires be matched by a registry version bump.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(telegram): fix import path in pairing chat claim e2e test
The scenarios/ folder is a Python package (has __init__.py), so a
flat `from test_telegram_e2e import …` fails with
ModuleNotFoundError during pytest collection. Switch to a relative
import that matches the package layout, and drop the unused
OWNER_USER_ID symbol while we're here.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(auth): address Copilot review on /clear OAuth drain + correlation doc
Two follow-ups on PR #3381's Copilot pass:
1. `Agent::process_clear` (engine v1 path) now drains in-flight OAuth
flows for the clearing user, mirroring the engine-v2 cleanup added in
`bridge::router::clear_engine_conversation`. Without this, `/clear`
was a clean slate on v2 but v1 left ghost flows in
`extension_manager.pending_oauth_flows()` until the 5-minute
`OAUTH_FLOW_EXPIRY` ticked over — same regression class #3320 fixed
on v2.
2. `oauth_failure_correlation_id`'s docstring previously said "redacted
state fingerprint", but callers seed it with the raw `state` query
value (or `flow.extension_name` for post-resolution failures).
Updated the doc to describe the actual behaviour: an arbitrary seed
that is hashed before any hex output, with a pointer to
`redact_oauth_state_for_logs` for the log-safe fingerprint.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(e2e): make Telegram pairing chat-claim suite actually run
The scenario landed in PR #3317 was orphaned — never wired into any CI
lane and could not pass even when run by hand. Three structural
issues, all fixed here:
1. `install_telegram` now overlays the locally-built WASM (and matching
capabilities file) on top of the registry-downloaded artifact when
present. The pairing-reply wording lives inside the WASM binary, so
without this overlay the test was asserting source-tree text against
the previous release's bytes. The overlay is best-effort: when the
local WASM is absent (CI groups that don't build the channel), the
test that depends on it skips with a clear message and the canary
lane in `scripts/live_canary/auth_registry.py` still covers the
wording end-to-end against the deployed binary.
2. `Submission::PairingClaim` is handled out-of-band by the bridge
layer; the response is delivered via `WebChannel::respond` →
`AppEvent::Response` over SSE only — no `Turn` is persisted, so
polling `/api/chat/history` could never see it. Refactored
`test_chat_surface_approves_pairing_code` and
`test_chat_surface_rejects_invalid_pairing_code` onto a
`_send_and_collect_response` helper that opens the SSE stream first
(so the broadcast doesn't fan out to zero subscribers) and matches
on the `response` event for the test thread.
3. Wired the file into `e2e.yml`'s `extensions` group so the suite
actually runs on every PR.
Verified locally: all three scenarios pass, plus the existing 22
Telegram e2e tests still green with the install-overlay change.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(bridge): bound and sanitize invalid-channel echo in pairing claim
Address Copilot review on PR #3381: `handle_pairing_claim`'s
invalid-channel branch was rendering the raw `channel` token (and the
underlying `IdentityError`, which itself echoes the offending input)
back to the user. Both routes are unbounded and could carry control
characters or markup since `channel` comes from chat input — a
hostile prompt could blow up the SSE / Telegram / TUI reply or smuggle
backticks/escape sequences through.
Cap the echo at 32 ASCII-alphanumeric (or `-`/`_`) characters and
replace the verbatim error with a fixed category description, so the
reply size and shape are bounded by what we render explicitly. Add a
regression test that drives a 200-char hostile blob (control chars +
backticks) through the handler and asserts the rendered reply stays
clean and short.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(bridge): include hyphens in invalid-channel error copy
Address Copilot review on PR #3381: the invalid-channel reply
listed "lowercase letters, digits, or underscores" as the valid
character set, but `ExtensionName::new` (and `web::features::pairing::
parse_channel`) intentionally accept hyphens too — they're folded to
underscores during canonicalization. A user typing `slack-relay` would
otherwise get an "invalid name" reply listing rules that contradict
the actual validator.
Updated the message to include hyphens with `telegram` and
`slack-relay` as concrete examples, and tightened the regression test
to assert against the user-controlled preview region between the
delimiter backticks rather than a global backtick count (which was
fragile to copy that includes example slugs in backticks).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(auth): address PR #3381 review on credential-scoped gate cleanup and Telegram surface promise
Three reviewer findings, one commit:
- OAuth provider-error and exchange-failure paths used
`clear_engine_pending_auth(user, None)`, which discards every
Authentication gate for the user. A failed Gmail callback could
silently wipe an unrelated Slack/MCP gate waiting on a different
thread. New `clear_engine_pending_auth_for_credential(user, credential)`
helper in bridge::router scopes cleanup to the failed flow.
Provider-error path tracks `removed_secret_name` alongside
`removed_user_id` so the scoped variant is callable.
- Expired-flow branch in the OAuth callback handler had two bugs: it
never cleared the engine pending auth gate (so the conversation sat
blocked forever, same #3320 class the provider-error fix addresses),
and the broader `clear_auth_mode` it called would re-discard via
the unscoped helper anyway. Now calls the credential-scoped helper
and the legacy-v1-only `clear_session_auth_mode_for_thread`.
- Telegram pairing reply advertised `approve telegram CODE` as
usable "in any IronClaw chat (TUI / web / Telegram)", but an
unpaired Telegram DM is intercepted by the allowlist gate before
the agent parser sees the command — the user would just get
another pairing reply. Reply now lists only the surfaces that
actually work (web / TUI / CLI) and a comment explains why.
Regression coverage:
- `clear_engine_pending_auth_for_credential_only_clears_matching_credential`
locks in helper scoping (Gmail/Slack two-gate scenario).
- `oauth_callback_expired_flow_clears_credential_scoped_engine_gate`
drives the full callback through axum oneshot with engine state
seeded; asserts the matching gate clears and the unrelated gate
survives.
- E2E `test_telegram_dm_approve_command_is_intercepted_by_allowlist_gate`
exercises the Telegram webhook path (not /api/chat/send) to lock in
the channel-layer interception, wired into the auth canary lane.
- Existing E2E pairing-reply test gains an assertion that
"TUI / web / Telegram" is *not* in the reply.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(oauth): mirror failure-cleanup contract on provider-error and reconcile stale comment
Copilot review on PR #3381 caught two real issues in the credential-scoped cleanup
landed in d45cf1bd8:
- Provider-error branch (`?error=access_denied`) returned the error page
without broadcasting `OnboardingState::Failed` or clearing the legacy v1
session `pending_auth`. The exchange-failure and expiry branches do both.
Net effect: the auth card stayed spinning and the next user message was
intercepted as a token. Now mirrors the other failure paths — keep the
full `flow`, emit Failed SSE, clear v1 session, clear credential-scoped
engine gate, then return the error page.
- Post-exchange comment said "failed callbacks should leave the gate
visible for retry" — that was the pre-#3320 contract. Rewrote it to
describe the new shape: each failure mode clears its own gate at the
failure site; this section only handles legacy-v1 session cleanup that
runs regardless of outcome. Also explains why we use
`clear_session_auth_mode_for_thread` here instead of `clear_auth_mode`
(the latter would re-clear the engine gate on the *success* path and
break the `ExternalCallback` resume).
Regression: `test_oauth_callback_provider_error_broadcasts_onboarding_failed`
in the oauth tests module — drives an `?error=access_denied` callback with
a flow whose `sse_manager` is attached, asserts the receiver gets
`OnboardingState::Failed` with the provider's `error_description` as the
message body.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(common): describe paths and platform helpers in crate description (#3498)
Align the package `description` and the lib.rs crate-level doc with
the modules now exposed from `ironclaw_common` (paths, platform,
env_helpers, attachment), which #3387 lifted out of `src/`. The
previous wording predates that extraction and only mentioned "types
and utilities".
This is also the release-plumbing trigger for v0.28.1: release-plz
proposes a leaf bump on source-path changes, and once `ironclaw_common`
crosses to a new patch, the root `ironclaw` bump can be added on top
of the release-plz branch (same approach as commit 9e69f22d2 for
v0.28.0). See PR #3372 for the equivalent v0.28.0 trigger.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* chore: release
* chore(release): bump ironclaw to 0.28.1
cargo-semver-checks did not detect API-breaking changes in the
ironclaw_common 0.4.1 -> 0.4.2 leaf bump, so release-plz did not
cascade a bump into the root ironclaw package. Add the root version
bump and CHANGELOG entry manually so this release-plz PR produces
an ironclaw-v0.28.1 tag and triggers cargo-dist.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor(llm): hide provider-specific auth, model fetch, and embeddings config behind facades (#3416)
* refactor(llm): hide provider-specific auth, model fetch, and embeddings config behind facades
External callers were reaching into provider-specific modules of
`ironclaw_llm` (`gemini_oauth::CredentialManager`,
`github_copilot_auth::*`, `OpenAiCodexSessionManager`,
`codex_auth::*`, `BedrockConfig`, etc.). Closes those leaks behind a
small set of verb-based public surfaces while keeping per-provider
behaviour inside the LLM crate.
Changes:
1. Extract `oauth_helpers.rs` into a new `ironclaw_oauth` crate. The
loopback OAuth callback listener (port 9876, landing pages,
`OAUTH_CALLBACK_HOST` rules) is shared by every IronClaw OAuth flow
(NEAR AI session login, WASM tool auth, MCP) and never depended on
`ironclaw_llm`. `src/auth/oauth.rs` now `pub use ironclaw_oauth::*`
directly. `ironclaw_llm` no longer depends on `ironclaw_oauth` —
the helper had zero internal callers.
2. Add `ironclaw_llm::auth` facade (`start_login`, `validate_token`,
`default_headers`, `load_persisted_credentials`,
`default_credentials_path`) with backend-agnostic types
(`AuthPrompt`, `LoginRequest`, `AuthOutcome`, `PersistedCredentials`,
`OpenAiCodexLoginOptions`, `AuthBackend`, `CredentialSource`).
Privatize `gemini_oauth`, `github_copilot_auth`, `openai_codex_session`,
`codex_auth` (`pub(crate) mod`). Migrate the wizard, the
`ironclaw login --openai-codex` CLI subcommand, and the LLM config
loader to the facade. Wizard introduces a single `WizardAuthPrompt`
that handles device-code prompts + browser launch for all backends.
3. Add `ironclaw_llm::models::fetch_models_for(provider_id, &opts)`
facade. Privatize `fetch_anthropic_models`, `fetch_openai_models`,
`fetch_ollama_models`, `fetch_openai_compatible_models`,
`is_openai_chat_model`, `openai_model_priority`, `sort_openai_models`.
Wizard's per-backend match collapses to one call. Move classifier
unit tests into `crates/ironclaw_llm/src/models.rs`; rewrite the
two wizard fallback tests through the public API.
4. Decouple embeddings from `ironclaw_llm::BedrockConfig`. New
`crate::workspace::BedrockEmbeddingSetup { region, profile }` carries
only what `BedrockEmbeddings` actually needs. `EmbeddingsConfig::create_provider`
and `BedrockEmbeddings::new` take the new type; callers translate from
`LlmConfig.bedrock` at the boundary (`src/app.rs`, `src/cli/mod.rs`).
5. Add `ironclaw_llm::testing::nearai_test_config(model)` helper for
tests that need a minimal `LlmConfig` shape (no retries, no caching,
NEAR AI backend). Replaces two duplicated 30-line struct literals
in the gateway settings hot-reload tests.
Boundary cleanup is behaviour-preserving: 4,932 main-binary unit tests,
729 ironclaw_llm unit tests, 4 ironclaw_oauth tests, 3 architecture
boundary tests all pass; `cargo clippy --all --benches --tests
--examples --all-features` is clean.
Three `pub` methods on `gemini_oauth::CredentialManager` /
`GeminiOauthProvider` (`get_valid_access_token`, `last_response_meta`,
`count_tokens`) and the `GeminiResponseMeta` struct are now reachable
only crate-internally and have no callers; marked `#[allow(dead_code)]`
with a comment rather than deleted to keep this PR purely a boundary
move (delete in a follow-up if no caller emerges).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor(llm): promote dedicated backends into the registry; absorb config validation, defaults, and per-provider overrides into ironclaw_llm
Continues the LLM boundary cleanup from 0addf3ac2. After that commit
provider-specific auth, model fetch, and embeddings config lived behind
facades inside `ironclaw_llm`, but four backend-specific knowledge
sources still leaked out:
1. Validation rules and default values for the dedicated-config
backends (Bedrock cross-region prefixes, OpenAI Codex endpoints
and client_id, Gemini OAuth credentials path defaults) lived
inline in `src/config/llm.rs::resolve`.
2. The dispatcher in `create_llm_provider` matched on backend strings
("nearai", "bedrock", ...) instead of a typed protocol value. The
same booleans (`is_nearai`, `is_bedrock`, `is_gemini_oauth`,
`is_openai_codex`) recurred across `src/config/llm.rs`,
`src/app.rs`, `src/cli/models.rs`, and the wizard.
3. The setup wizard had per-backend specialization in
`step_inference_provider` and `run_provider_setup` (manual menu
pushes for nearai/bedrock/codex/gemini_oauth, four dedicated
`setup_*` entry points dispatched on string compares).
4. `Settings` carried named `bedrock_region`, `bedrock_cross_region`,
`bedrock_profile` columns even though no other dedicated backend
had named columns and adding a new one would mean schema churn.
Layers A-D address each in turn:
* Layer A — `BedrockConfig::build`, `OpenAiCodexConfig::build`, and
`GeminiOauthConfig::build` own validation + defaults inside the
crate. `LlmConfigError` (`MissingRequired` / `InvalidValue`) carries
the failures across the boundary, with a `From` impl into the
binary's `ConfigError`. `src/config/llm.rs` calls the builders;
named-string defaults are gone from the binary. The orphaned
`tests/gemini_oauth_regression.rs` husk is deleted.
* Layer B — `ProviderProtocol` gains four new variants
(`Bedrock`, `OpenAiCodex`, `GeminiOauth`, `NearAi`) plus a
`has_dedicated_config()` predicate. The four dedicated-config
backends (with all aliases) become first-class registry entries in
`providers.json`, so `is_known()` / `model_env_var()` / the wizard /
the gateway handler iterate the registry uniformly. The
`is_nearai`/`is_bedrock`/`is_gemini_oauth`/`is_openai_codex` boolean
spaghetti collapses to protocol comparisons. `OpenAiCodex` and
`NearAi` carry explicit `#[serde(rename = "openai_codex" / "nearai",
alias = ...)]` so the wire-stable adapter strings the gateway and
frontend already use keep working. `LlmConfig::active_model_name()`
is now consumed by `cli/doctor.rs` instead of an inlined partial
dispatch.
* Layer C — `SetupHint` gains four credential-collection variants
(`AwsCredentials`, `OAuthDeviceCode`, `FileBasedCredentials`,
`SessionToken`). The wizard's `step_inference_provider` builds its
menu from a single `registry.selectable()` iteration with generic
env-detection (declared `api_key_env`, plus an Anthropic-specific
OAuth fallback). `run_provider_setup` dispatches on the SetupHint
variant; the remaining `def.id == "..."` checks live inside the
`ApiKey` arm only because Anthropic and GitHub Copilot present a
hybrid choice (API key OR OAuth) the simple `ApiKey` hint doesn't
capture. The synthetic bedrock + nearai entries in
`handlers/llm.rs::build_llm_providers` are deleted; a single
registry-driven loop covers both. ADAPTER_LABELS in
`static/js/surfaces/config.js` gains entries for the new protocols.
* Layer D — `LlmBuiltinOverride` gains a generic
`extras: HashMap<String, String>` bag with `extra(key)` /
`set_extra(key, value)` accessors. The bedrock resolver and wizard
read/write through this bag; `Settings::migrate_legacy_provider_fields()`
drains the named `bedrock_*` columns into `extras` on
`Settings::load_from()` so existing `settings.json` files migrate
losslessly. The named columns are kept (deprecated, marked with
`#[serde(skip_serializing_if = "Option::is_none")]`) for one
release; tracked for deletion in #3443.
`strip_admin_only_llm_keys` and `llm_setting_requires_reload` now
match dotted-path subkeys under `llm_builtin_overrides.*` so a
write to e.g. `llm_builtin_overrides.bedrock.extras.region`
triggers the right gating + chain reload.
Boundary cleanup is behaviour-preserving: 4,933 main-binary unit tests,
739 ironclaw_llm unit tests pass; `cargo clippy --all --benches --tests
--examples --all-features` is clean. New regression tests:
`crates/ironclaw_llm/src/config.rs` (6 builder tests),
`crates/ironclaw_llm/src/registry.rs::dedicated_config_backends_are_in_registry_and_selectable`,
and `src/setup/wizard.rs::legacy_bedrock_fields_migrate_into_extras_on_load`.
Three follow-ups tracked in #3443: delete the deprecated `bedrock_*`
named columns, move `BedrockEmbeddings` out of `src/workspace/` into
the LLM crate (last cargo-feature leak), and drive
`LlmConfig::active_model_name()` off `ProviderProtocol` instead of
backend strings.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test: add bug-bash regression-snapshot harness
Bug-bash fixtures pin specific open bugs to a deterministic snapshot.
When a bug is fixed, the snapshot diff is the reviewable proof; when
someone reintroduces the bug, the snapshot drifts and CI blocks the
merge.
This commit lands the harness plus the first recorded fixture for
issue #2541 (agent must call a tool, not answer from training data):
tests/e2e_bug_bash_snapshots.rs
`snapshot_summarization_uses_tools` replays the fixture, captures
`ReplayOutcome`, and asserts the YAML snapshot. Gated on
`feature = "libsql"`, same as other replay-snapshot tests.
tests/fixtures/llm_traces/bug_bash/summarization_uses_tools.json
Two-step recorded LLM trace (tool_call -> text) keyed off the
user prompt via `request_hint.last_user_message_contains`.
tests/fixtures/llm_traces/bug_bash/README.md
Coverage map for #2540-#2546 (one recorded, six TODO) plus the
`IRONCLAW_RECORD_TRACE` recording workflow.
tests/snapshots/replay__bug_bash_summarization_uses_tools.snap
Insta YAML snapshot pinning `tool_calls: [echo]`, 2 LLM calls,
and the event-kind histogram. Drift = regression.
* fix(settings): preserve pre-existing extras during legacy bedrock migration
`migrate_legacy_provider_fields` claimed to be idempotent and to drain
named `bedrock_*` columns into `llm_builtin_overrides["bedrock"].extras`
once on load. The previous implementation drained correctly but used
`HashMap::insert` unconditionally, which means a settings file
carrying BOTH a legacy `bedrock_region` column AND an already-populated
`extras["region"]` (manual hand-edit, or a future writer emitting both
shapes during a transition) would silently downgrade to the legacy
value.
Guard each `set_extra` call with `entry.extra(key).is_none()` so the
new-shape value always wins. Clarify the docstring to state this
explicitly.
Add three regression tests in `settings::tests`:
- `legacy_bedrock_migration_round_trips_through_save` — legacy JSON ->
load_from -> serialize -> reload, asserts the deprecated columns are
not re-emitted and extras survive the round trip.
- `legacy_bedrock_migration_preserves_existing_extras` — file with both
shapes; asserts the pre-existing extras value is kept and absent
extras are still backfilled from legacy fields.
- `legacy_bedrock_migration_is_idempotent_in_memory` — calling the
migration twice on the same Settings is a no-op (compares serialized
shape, since LlmBuiltinOverride does not derive PartialEq).
* fix(pr-3416): address PR review — migration on DB/TOML, admin-key gate, codex login, credential_kind/has_credentials
Addresses comments from gemini-code-assist, Copilot, and serrrfirat on PR #3416.
## Bugs
**Legacy bedrock fields not migrated on DB/TOML loads** (serrrfirat, High).
`Settings::load_from` (JSON) ran `migrate_legacy_provider_fields`, but
`from_db_map` and `load_toml` did not. Existing operators with
`bedrock_*` settings persisted in the DB or `config.toml` would silently
lose their AWS region/profile/cross-region after upgrade because the
resolver now reads only from `llm_builtin_overrides["bedrock"].extras`.
Both loaders now call the migration; added round-trip tests for each.
**Admin-only key write gate had narrower matching than read gate**
(Copilot, High). `strip_admin_only_llm_keys` matches both exact keys
and dotted subpaths under admin-only roots; `is_admin_only_setting_key`
in the web settings handler used `.contains(&key)` only. A non-admin
could write `llm_builtin_overrides.bedrock.extras.region` directly,
bypassing the gate. Promoted `is_admin_only_llm_key` to `pub(crate)`,
made the web write-side gate call it, added regression tests covering
dotted subpaths.
**`ironclaw login --openai-codex` dropped TOML/DB config** (Copilot,
High). The pre-refactor code resolved `Config::from_env` and used
`config.llm.openai_codex` so endpoint / client-id / session-path
overrides committed via TOML or DB stuck. The post-refactor code only
read env vars via `OpenAiCodexLoginOptions::from_env`. Added
`OpenAiCodexLoginOptions::from_resolved_config(&OpenAiCodexConfig)`;
the login command now prefers the resolved config when present and
falls back to env-only when `Config::from_env` itself fails (fresh
machine, no DB).
**Dedicated-auth backends marked configured without credentials**
(serrrfirat, Medium). `nearai` / `gemini_oauth` / `openai_codex` ship
`api_key_required: false` because they don't authenticate via a bearer
API key. The frontend `isProviderConfigured` treated that as "no
credentials needed" and rendered the Use button on a fresh install,
where clicking could trigger an interactive device-code OAuth from
inside a settings request.
Added `credential_kind` (wire-stable snake_case discriminator matching
`SetupHint::kind()`, e.g. `session_token`, `o_auth_device_code`,
`file_based_credentials`, `aws_credentials`) and `has_credentials`
(backend-authoritative; checks AWS env vars for Bedrock, codex session
file existence, file-based credential path expansion + existence) to
the web LLM providers payload. Frontend `isProviderConfigured` /
`providerMissingReason` now gate non-api-key kinds on `has_credentials`.
## Nits
**`fetch_models_for` doc overclaimed "Always returns something"**
(Copilot). The generic openai-compatible branch returns `vec![]` when
`base_url` is empty. Updated the docstring to call this out so callers
know to handle the empty case.
**`AuthError::Other` used for "validation not applicable"** (Gemini
bot). Added a dedicated `AuthError::TokenValidationNotSupported { backend }`
variant; `validate_token` now returns it for Gemini / OpenAiCodex
instead of stringly-formatted `Other`.
**Bug-bash regression-harness URLs pointed at `near/ironclaw`**
(Copilot, x2). The canonical tracker is `nearai/ironclaw`. Rewrote
all seven URLs in `tests/fixtures/llm_traces/bug_bash/README.md` and
the one in `tests/e2e_bug_bash_snapshots.rs`.
## Declined
The Gemini bot's MalformedConfig suggestion at
`crates/ironclaw_llm/src/models.rs:46` was not adopted: the call site is
the openai-compatible model-listing path, not a security-sensitive
request. The fetcher early-returns `vec![]` on empty `base_url` — no
URL parsing happens — and the docstring tightening above covers the
observable surprise. Promoting it to a typed error would change the
public-facing `fetch_models_for` signature for no behavioural gain.
## Tests
- `cargo fmt --check` clean
- `cargo clippy --all --benches --tests --examples --all-features` zero warnings
- `cargo test --lib` 4,941 / 4,941 pass
- `cargo test --features libsql --test e2e_bug_bash_snapshots` 1 / 1 pass
- New regression tests:
- `settings::tests::legacy_bedrock_fields_migrate_into_extras_on_db_load`
- `settings::tests::legacy_bedrock_fields_migrate_into_extras_on_toml_load`
- `channels::web::features::settings::tests::test_admin_only_setting_keys_cover_dotted_subpaths`
- `channels::web::handlers::llm::tests::test_llm_providers_expose_credential_kind_and_has_credentials`
- `channels::web::handlers::llm::tests::test_nearai_has_credentials_true_when_session_token_loaded`
* fix(pr-3416): tighten Bedrock/Codex has_credentials probes; collapse set_extra into one .into()
- `backend_has_credentials` for AWS now requires `AWS_PROFILE` OR
(`AWS_ACCESS_KEY_ID` AND `AWS_SECRET_ACCESS_KEY`). The lone
`AWS_ACCESS_KEY_ID` / `AWS_SESSION_TOKEN` arms previously flipped
has_credentials true even though the AWS SDK can't sign without the
secret key, so the UI was rendering Bedrock as configured on hosts
that would fail at first call.
- `backend_has_credentials` for OpenAI Codex now honours
`OPENAI_CODEX_SESSION_PATH` via `read_env` before falling back to
the default session path under `~/.ironclaw/`. Users with a custom
session location were seeing "not configured" despite a valid login.
- New regression tests `test_bedrock_partial_aws_env_reports_not_configured`
and `test_openai_codex_honours_session_path_env` drive the
`build_llm_providers` call site (not just the helper) so both gaps
stay closed.
- Tidied `LlmBuiltinOverride::set_extra` to convert the key once and
reuse it across the remove/insert branches; behaviour identical.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* chore(providers): default nearai model to "auto"
Switch the nearai registry entry's `default_model` from
`claude-sonnet-4-5` to `auto`, NEAR AI's server-side routing alias.
New installs without `NEARAI_MODEL` set now get auto-routed instead
of being pinned to a specific Anthropic model.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* Make Skills E2E lifecycle deterministic (#3309)
* test(e2e): make skills lifecycle deterministic
* test(e2e): address skills review comments (#3309)
* test(e2e): unxfail two auth-matrix tests now that contracts match (#3589)
Both xfails in tests/e2e/scenarios/test_v2_auth_oauth_matrix.py are
stale and pass against current code:
- test_wasm_tool_first_chat_auth_attempt_emits_auth_url
Marked xfail in #3235 because the engine-v2 callable-only contract
(#2868) stopped emitting an auth gate on direct LLM-driven tool
calls. PR #3157 (auth-preflight + inline-await) restored the
behavior the test asserts: when the LLM emits a direct call to a
not-yet-authed extension, the bridge raises an Authentication gate
with auth_url populated (src/bridge/effect_adapter.rs:1356-1392).
Marker removed; test passes.
- test_settings_first_custom_mcp_auth_then_chat_runs
The xfail reason claimed post-auth tool-output propagation was
broken. Real cause: engine-v2 gates the first MCP tool call on
`approval` and the browser fixture has no auto-approve UI, so the
chat sat in pending_gate forever. Same shape as the bugs fixed in
#3235 for test_wasm_tool_oauth_refresh_on_demand and
test_mcp_same_server_multi_user_via_browser. Inserted
_wait_for_tool_call between _send_chat and _wait_for_response_contains
to drive approval through the API; test passes.
Verified locally: both tests pass back-to-back in 27s on a fresh
auth_matrix_server.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(extensions): restore chat-driven tool_install + fix double-invoke + auto-approve footgun (#3533) (#3559)
* fix(extensions): restore chat-driven tool_install + fix double-invoke + auto-approve footgun (#3533)
"Connect my telegram" was giving the user two options and not actually
installing anything because three layered issues had accumulated since
engine v2:
1. **`tool_install` was hidden from the agent** (#2868). The unified
`tool_activate` it was meant to be subsumed by was later removed in
#3166, but the hidden-from-callable-surface gate stayed. Restored
by dropping `hidden_from_model_callable_surface` from
`bridge::action_projector`. User consent is mediated by the tool's
own `ApprovalRequirement::UnlessAutoApproved` and the seeded
`AskEachTime` permission.
2. **Two competing Telegram registry entries** (`telegram` channel and
`telegram_mtproto` tool) both surfaced in the agent prompt's
`Activatable Integrations` section. The LLM correctly enumerated
them as "Option 1" and "Option 2" instead of installing the
canonical bot channel. Added a `hidden: bool` field to
`ExtensionManifest` / `RegistryEntry`, set `telegram_mtproto` to
`hidden: true`, and filter hidden entries out of the
"available-but-not-installed" appendix in `ExtensionManager::list`.
Hidden entries remain installable by explicit name.
3. **Updated the agent prompt** so `Activatable Integrations` instructs
the model to call `tool_install(name="<name>")` directly rather than
describing manual UI steps.
Fixes the double-`tool_install` invocation that surfaced once the agent
could install from chat:
- **`InlineGate` discarded cached output.** The bridge raised an
Authentication gate after `tool_install` succeeded, and the
inline-await retry re-executed the action (re-downloading the WASM
bundle) instead of returning the already-computed output. Added
`resume_output: Option<serde_json::Value>` to `InlineGate`; on
approval, return the cached output if present. Mirror fix in the
orchestrator's `execute_single_action_with_inline_retry` (reading
`result_json["resume_output"]`) and the structured-batch retry path.
- **`effect_adapter::auth_gate_from_extension_result`** now passes
`Some(output_value.clone())` as the gate's `resume_output` so the
retry has cached state to short-circuit on.
- **OAuth callback double-fired.** `oauth_callback_handler` now skips
the `ExternalCallback` re-entry when the inline-await path already
woke a parked waiter — eliminates the "thread already running" race.
- **`resolve_inline_gates_for_credential`** now also discards matching
Authentication rows from `pending_gates` so the row doesn't linger
in `HistoryResponse.pending_gate` after inline resolution.
Fixes the auto-approve footgun:
- **`ToolPermissionSnapshot::resolve_permission`** now collapses DB
values that match the seeded default to `explicit = None`. Before
this, the boot-time `seed_tool_permissions` write of `tool_install ->
AskEachTime` was indistinguishable from a user-explicit override,
causing `effect_adapter::enforce_tool_permission`'s `is_explicit_ask`
check to refuse `AGENT_AUTO_APPROVE_TOOLS=true`. Real overrides
(`AlwaysAllow`, `Disabled`) still surface as `Some(...)`.
Tests
- Unit: 4980/4980 pass (host) + 525/525 pass (engine).
- Unit: new tests in `bridge::tool_permissions::tests` lock seeded-vs-
explicit collapse; new test in `bridge::action_projector::tests`
asserts `tool_install` is callable; new manifest hidden-flag tests
in `registry::manifest::tests` and `extensions::manager::tests`.
- E2E: removed `@pytest.mark.xfail` on
`test_chat_first_gmail_installs_prompts_and_retries` (now passes
end-to-end via the chat-driven install path). Added
`test_chat_install_approval_then_auth_card` driving the
explicit-approval variant with a single Approve click (no Always
workaround needed) — wired into the `auth-full` canary lane.
- Mock LLM: extended the gmail-install-then-retry pattern to recognize
both the legacy "Extension not installed:" and the post-#3533 "is
not callable in this execution context" error strings, and to retry
`gmail(action="list_messages")` after a successful `tool_install`.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(permissions): address #3559 review (permission bypass, lease accounting, hidden search filter)
Five fixes from the #3559 review (4× Copilot doc nits + 3× serrrfirat
security/correctness findings):
1. **Permission bypass (High).** Pre-#3559's `resolve_permission`
collapsed any DB row whose value matched the seeded default to
`explicit = None`, so a user who deliberately set `tool_install =
AskEachTime` had their explicit choice silently dropped and
`AGENT_AUTO_APPROVE_TOOLS=true` bypassed the gate. Provenance is now
handled at write time: `seed_tool_permissions` is gone and a
one-shot, sentinel-gated migration (`cleanup_ghost_seeded_tool_permissions`)
deletes existing ghost-seeded rows at startup. With no ghost rows,
the resolver treats every DB row as user-explicit and honors it.
2. **Lease/event accounting on `resume_output` replay (Medium).**
Inline-gate handlers in `structured.rs`, `scripting.rs`
(`resolve_tool_future` + `drive_inline_gate` retry loop), and
`orchestrator.rs` refunded the lease use the action just consumed,
then returned the cached `resume_output` on approval without
re-consuming — netting successful side-effecting actions to zero
lease uses. Skip the refund when the gate carries cached output.
3. **Hidden registry filter on `tool_search` (Medium).**
`RegistryCatalog::search` did not filter `hidden: true` entries,
so `telegram_mtproto` could resurface through the search path and
reintroduce the "two Telegram options" outcome that #3533 fixes
for the default-list path. Added the filter and a regression test.
4-7. Copilot doc nits: outdated `_set_tool_permission` docstring;
misleading "bridge-side auto-install implemented" comment in
`mock_llm.py`; `tool_install` described as "non-agent surface" in
`src/bridge/CLAUDE.md` while a paragraph below says the model
calls it directly; dangling `issue #3533 / PR —` placeholders
in both CLAUDE.md docs.
Regression tests:
- `bridge::tool_permissions::user_explicit_value_matching_seeded_default_stays_explicit` —
the original Copilot/serrrfirat bug case.
- `app::cleanup_ghost_seeded_tool_permissions_removes_seed_matching_rows` —
idempotent migration + sentinel.
- `extensions::registry::test_search_skips_hidden_entries` — hidden
entries excluded from search but still installable by exact name.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(#3559): caller-level regression coverage for review findings 1 & 2
Two follow-up regression tests for the #3559 security review, plus a
real bug surfaced by the first one.
1. `executor::structured::resume_output_replay_consumes_exactly_one_lease_use`
exercises the post-execution Authentication gate inline-retry path
with `max_uses=1` and asserts:
- Cached output is returned as a successful `ActionResult`.
- Exactly one `ActionExecuted` event is emitted.
- The lease budget is exhausted after one execution (refund-skip
keeps the consumption from being undone).
Writing this test surfaced a real bug: the structured cached-output
branch pushed `ActionExecuted` into `emitted_events`, and the
caller's `classify_exec_result` emitted ANOTHER terminal
`ActionExecuted` for the same Ok result — double-emit for one
action. Tier 1 (`scripting::drive_inline_gate`) and Tier 1 alt
(`orchestrator::execute_action_with_inline_gate`) emit themselves
because their callers don't run an Ok-branch classifier; structured
was the outlier. Dropped the redundant push; the classifier emits
the single canonical event.
2. `bridge::effect_adapter::explicit_ask_each_time_for_seeded_default_tool_still_gates`
drives `execute_action` end-to-end (the side-effecting caller) with
a tool whose `name()` matches a seeded-`AskEachTime` baseline
(`tool_install`) and an explicit `AskEachTime` user override. The
resolver collapse-to-implicit bug would have shown up here — not
just in the helper-level test that already exists in
`bridge::tool_permissions::tests`. Per `.claude/rules/testing.md`
"Test Through the Caller, Not Just the Helper".
Added `SeededAskEachTimeTestTool` as a `tool_install`-named test
fixture with `requires_approval: UnlessAutoApproved` to mirror the
real tool's contract.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* chore: release
* feat(engine): IRONCLAW_DISABLE_CODEACT flag to disable v2 CodeAct (#3665)
* add flag to disable codeact on engine v2
* fmt
* fix(engine): keep compact actions reachable when CodeAct is disabled
With IRONCLAW_DISABLE_CODEACT=true the structured-tools prompt told
the model to use the provider's tool_calls interface for every action,
but the bridge filtered the provider tool list down to
emits_full_schema_tool(). Most tools default to CompactToolInfo
(mission_create, gmail_send, notion_search, ...), so they appeared in
the prompt as "available" while being absent from the provider tool
list — i.e. unreachable. Addresses serrrfirat's review on PR #3665.
Fix coordinates both halves of the surface:
- src/bridge/llm_adapter.rs: in disabled-CodeAct mode, drop the
emits_full_schema_tool() filter and emit every action into the
provider tool list with its full schema.
- crates/ironclaw_engine/src/executor/prompt.rs: in disabled-CodeAct
mode, skip the "## Enabled Tools" section. The compact-form listing
with the tool_info(detail="schema") instruction is meaningless when
the provider already sends full schemas, and would just duplicate
the surface. "## Activatable Integrations" stays — the model still
needs to know what tool_install can target.
Test seam: build_codeact_system_prompt_inner now takes disable_codeact
as an explicit parameter, called once at the public entry points. This
lets prompt tests exercise both branches without process-global env
mutation.
Tests:
- executor::prompt::tests::disabled_codeact_omits_enabled_tools_section_and_keeps_activatable
- bridge::llm_adapter::tests::complete_emits_compact_actions_when_codeact_disabled
- existing complete_with_tools_only_emits_full_schema_provider_tools
now serialized via lock_env() so env mutation in the new test
can't leak across parallel runs.
cargo test -p ironclaw_engine --lib: 527 passed
cargo test --lib bridge::: 469 passed
cargo clippy -p ironclaw_engine --all-targets -- -D warnings: clean
cargo clippy --lib --tests -- -D warnings: clean
cargo fmt --check: clean
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Emil Bogomolov <emil.bogomolov@near.ai>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* Fix markdown_to_mrkdwn to avoid converting emphasis inside generated <… (#3532)
* agent: Fix markdown_to_mrkdwn to avoid converting emphasis inside g…
* agent: Fix rustfmt/clippy CI failure by removing extra blank line b…
* agent: slack: fix markdown_to_mrkdwn replacement order to satisfy p…
* slack: protect generated links and sanitize sentinels in markdown_to_mrkdwn
Two issues raised on PR #3532 review:
1. Emphasis inside generated `<url|text>` was still rewritten because the
global `**`/`~~` → `*`/`~` substitution ran after link materialization.
Push the generated link span into the same protected arena used for
Slack-native `<...>` constructs so subsequent global replacements can't
reach inside it. Matches the PR's stated goal.
2. Untrusted input containing the private-use sentinel chars
(U+E000 / U+E001) could forge a protected-span reference and pull in
another span's content. Strip those chars from input up front.
Adds regression tests for both. Bumps registry/channels/slack.json
0.3.2 → 0.3.3 to satisfy the channel-source version-bump check.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* slack: escape link labels, drop pipe/gt URLs, expand nested sentinels
Addresses two follow-up review concerns on PR #3532:
Copilot review: `<url|text>` was built by string concatenation, so
`|` or `>` inside the URL would corrupt the entity, and `<` / `>`
inside the label would open/close a Slack span and break the link.
The label now escapes `<` → `<` and `>` → `>` (Slack's documented
literal-character form); a URL containing `<`, `>`, or `|` falls back
to leaving the original markdown form intact (those chars are not
valid URL characters per RFC 3986 anyway).
Latent nested-sentinel bug introduced by the previous fix: a markdown
link whose label contained a Slack-native `<...>` span (e.g.
`[<@U1> hi](url)`) ended up with the inner sentinel buried inside the
arena entry for the outer link span. The final restore pass advances
past the outer sentinel without rescanning what it just emitted, so
the raw U+E000/U+E001 characters would leak into the output. URL and
label are now pre-expanded before the link span is pushed.
While here, factor the duplicated restore loop into
`expand_protected_spans`, reused by both the pre-link expansion and
the final restore, and lift the sentinel constants to file scope.
Adds three regression tests covering label-bracket escaping, the URL
pipe/gt fallback, and the nested-span case. Bumps
registry/channels/slack.json 0.3.3 → 0.3.4.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(gateway): add logs download button (#3588)
* feat(web): support externally-provided tools in Responses API (#3122)
* feat(web): support externally-provided tools in Responses API
Lets callers of `/v1/responses` (and `/api/v1/responses`) declare their
own `function`-typed tools and feed back results via
`function_call_output` items, matching the OpenAI Responses wire shape.
Since IronClaw's engine has no per-request tool surface, integration
happens at the prompt level: the catalog is rendered as
`<external-tools>` in the user message and the agent signals a call by
ending its response with a fenced ```` ```tool_call ```` block. When
that fence is recognised, the reply is split into a leading `Message`
plus a `function_call` `ResponseOutputItem`.
Validation rejects unsupported tool types (`web_search`, `file_search`,
`code_interpreter`) and tools missing `name` with 400, with two new
integration tests covering both paths.
* refactor(responses-api): switch external tools to engine v2 native path
Replace the prompt-level fence protocol from PR #3122 with engine v2
native tool calls: caller-supplied `tools[]` are surfaced as real
LLM-callable actions, the engine pauses with `ResumeKind::External`
when one is invoked, and the bridge router projects the pause to a
new `AppEvent::ExternalToolCall` carrying the OpenAI-shaped
`function_call` wire fields.
The integration is small because v2 already has the right primitives:
- `ResumeKind::External { callback_id }` and
`GateResolution::ExternalCallback { payload }` already existed for
OAuth-style callbacks.
- `agent_loop.rs:1480` already routes Responses API messages to
`handle_with_engine` when `ENGINE_V2=true`, so no v2 migration of
the endpoint itself is needed.
- `EffectBridgeAdapter::execute_action` is the single chokepoint
where caller tools can be detected before they reach the dispatch
pipeline.
Changes:
- New `src/bridge/external_tools.rs` (`ExternalToolCatalog`) — per-thread
registry of caller-supplied `ActionDef`s, plus the `ext_tool:`
callback-id helpers used to disambiguate external-tool pauses from
OAuth/pairing pauses (which also use `ResumeKind::External`).
- `EffectBridgeAdapter` consults the catalog: any name in it is
short-circuited to a `GatePaused { resume_kind: External {
callback_id: ext_tool:<call_id> } }` before any registry dispatch,
and `available_action_inventory` merges the catalog into the
LLM-visible action surface (internal beats external on collision).
- `Submission::ExternalCallback` gains an optional `payload` field;
`bridge::handle_external_callback` plumbs it into
`GateResolution::ExternalCallback { payload }`. Fallback predicate
`gate_resume_is_external` lets non-auth External pauses (i.e.
caller-tool resumes) resolve through the same handler.
- New `AppEvent::ExternalToolCall` projected by `notify_pending_gate`
when a paused gate carries an `ext_tool:` callback id; OAuth/
pairing flows keep flowing through the existing `GateRequired`
channel.
- `responses_api.rs` is gutted of the prompt rendering and fence
parsing (`render_external_tools_preamble`, `extract_trailing_tool_call`,
`parse_external_tool_call`, `ParsedToolCall`, `external_tool_names`
accumulator field, and the `TOOL_CALL_FENCE` constants). The handler
now: rejects `tools[]` when `ENGINE_V2=false`, registers caller
tools in the catalog under the resolved thread id, detects resume
requests (`previous_response_id` + `function_call_output` items in
`input`) and submits them as `Submission::ExternalCallback` with
the outputs as the resolution payload, and surfaces
`AppEvent::ExternalToolCall` as a `function_call` `ResponseOutputItem`
in both streaming (`output_item.added`+`done`) and non-streaming.
- All existing OAuth/pairing `ExternalCallback` constructors updated
to pass `payload: None` (no behaviour change).
- Fence-protocol unit tests removed; replaced with coverage for the
new `responses_tools_to_action_defs` converter and the accumulator's
`ExternalToolCall` arm.
Existing 9 integration tests in `tests/responses_api_path_prefix.rs`
still pass.
Note for reviewers:
- The accumulator-side text response no longer tries to split the
reply on a fenced `tool_call` block. The wire shape that callers
receive for caller-tool invocations is purely event-driven now.
- Internal vs external collision is handled silently by the dedup in
`available_action_inventory` (internal wins). A request-time
rejection for shadowing names is a follow-up — the current behavior
is safe (the LLM only sees the internal version) but could surprise
a caller who expects their tool to run.
* test(responses-api): cover ENGINE_V2-off and resume-without-pending-gate
Two new integration tests for behaviours added by the engine-native
external-tool refactor:
- `external_tools_rejected_when_engine_v2_disabled`: a request with
caller-supplied `tools[]` while `ENGINE_V2` is off must 400 with a
message naming the flag, not silently fall through.
- `resume_without_pending_gate_returns_400`: a request with
`function_call_output` items and a `previous_response_id` that
doesn't correspond to a live external-tool gate must 400, not start
a fresh turn against the (unrelated) thread.
Both tests drive the full router (`start_test_server` + bearer auth)
per `.claude/rules/testing.md` "Test Through the Caller".
* test(responses-api): integration tests + drop unsafe env mutation
Three groups of changes:
1. **Drop unsafe env-var mutation in tests.** `responses_api.rs` no
longer reads `ENGINE_V2` directly: it keys off the presence of the
live `ExternalToolCatalog` (initialized by `init_engine`) as the
"engine v2 is up" signal. The path-prefix test that exercises the
no-engine branch no longer needs `unsafe { std::env::remove_var }`
— the absence of `init_engine` in `TestGatewayBuilder` is what
makes the catalog absent, which is what makes the request reject.
2.…
…nearai#2868) * engine-v2: make available_actions callable-only for blocked providers * fix(engine): address review fixture tempdir leak (nearai#2868) * engine-v2: refresh canonical prompt metadata on resume (nearai#2869) * fix(engine): align prompt metadata refresh with resume state * fix(engine): finish prompt refresh compaction coverage (nearai#2869) * fix(engine): preserve prompt refresh on resume (nearai#2869) * Add engine v2 action discovery metadata (nearai#2876) * Add engine v2 action discovery metadata * fix(engine): address action discovery review (nearai#2876) * fix(engine): address follow-up review comments (nearai#2876) * fix(engine): satisfy clippy in orchestrator lookup * fix(engine): propagate action snapshots in executor paths (nearai#2876) * fix(bridge): restrict tool_info to callable actions (nearai#2876) * [codex] Finish engine v2 deferred action inventory cleanup (nearai#2889) * Add deferred action inventory groundwork * fix(engine): address deferred action inventory follow-up * fix(engine): address deferred inventory review feedback * test: fix fmt and clippy failures * engine-v2: trim unused callable discovery payload * tests: restore env vars in review-fix cases * engine-v2: populate callable snapshots consistently * Unify v2 integration enablement on tool_activate * engine-v2: tighten tool_info inventory and approvals * llm: normalize tool_info hint syntax * engine-v2: tighten tool_activate install approval lookup * tests: align gmail settings-first flow with approval contract * engine-v2: fix remaining tool surface review issues * engine-v2: restore auto-approve defaults * fix(engine): align v2 tool permissions with defaults * fix(engine): close v2 callable snapshot gaps * fix(bridge): label latent-only providers accurately
* fix(oauth): remove pending flow on provider-error callback
The /oauth/callback handler's ?error= branch (RFC 6749 §4.1.2.1
provider-side failures — user cancels consent, scope denied, etc.)
returned the error page immediately without removing the flow from
ext_mgr.pending_oauth_flows(). The ghost entry then lingered until
the 5-minute expiry sweep, and any subsequent auth dance for the
same (extension, user) pair had to dedupe against it.
Mirror the happy-path cleanup: decode the state param, remove the
keyed flow, then return the error page.
Surfaced during live-canary auth-full repro: after
test_wasm_tool_oauth_provider_error_leaves_extension_unauthed ran,
the stale flow sat in the shared auth_matrix_server fixture.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(e2e): widen auth OAuth matrix timeouts for CI load
Four tests in live-canary auth-full were failing in CI with
`Page.wait_for_function: Timeout 60000ms exceeded`,
`ClientConnectionError('Connection closed')`, and
`Timed out waiting for OAuth refresh request` — all inside 60/20s
deadlines that are tuned for a dev laptop and don't leave margin
for ubuntu-latest's 2-vCPU runner under full suite load.
Raise the per-call deadlines so the inner budgets fit comfortably
inside pyproject.toml's 120s per-test cap:
_wait_for_refresh_request default: 20.0s -> 60.0s
_wait_for_auth_event call site: 60 -> 90
_wait_for_auth_prompt call site: 60 -> 90
send_chat_and_wait_for_terminal_message call sites: 60000 -> 90000
_wait_for_mock_google_tokens call site: 60.0 -> 90.0
_wait_for_response_contains (gmail) call site: 60.0 -> 90.0
Strictly widening; no passing test is slowed, no semantics change.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(canary): Haiku-powered Slack report job
Replace the team's raw Slack subscription (firehose of workflow
notifications) with one curated per-run summary:
Canary: 9 passed, 1 failed of 10 lanes
:x: auth-full (mock) — 12/13 passed, 1 failed in 350s
> test_wasm_tool_first_chat_auth_attempt_emits_auth_url timed
> out waiting for auth_required SSE event on the fresh thread
tools: shell, http_request, gmail (~6 calls)
...
commit `abc1234` • <github run link>
New `canary-report` job (needs: every lane, if: always) downloads
all lane artifacts, parses junit + summary + log tail per lane, and
asks claude-haiku-4-5 to return a compact JSON per lane
({status, reason, tool_calls_total, tools_used, notable}). That's
aggregated into a single Slack block message and posted via
incoming webhook.
Safety shape:
- Script exits 0 even on Haiku/Slack failure so the notifier never
masks the underlying canary signal.
- Missing ANTHROPIC_API_KEY falls back to raw junit-only phrasing.
- Slack POST failure falls back to plain-text "X/Y lanes failed"
with the GH run URL so the channel still hears something.
- No new Python deps — pure stdlib (urllib.request, xml.etree).
- 20 KB log-tail cap per lane to keep Haiku token usage bounded.
Secrets:
- ANTHROPIC_API_KEY (already present, used by provider-matrix)
- SLACK_WEBHOOK_URL (new — create an incoming webhook in Slack
and add as repo secret; notifier prints to stdout otherwise)
Testing:
- Trigger manually via Actions -> "Live Canary" -> "Run workflow"
with any single lane; canary-report runs after regardless of
which lanes executed.
- Run locally with --dry-run to preview the Slack payload.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(canary): post_json error handling + robust Haiku JSON extraction
Address gemini-code-assist review on scripts/live-canary/notify_slack.py:
1. `post_json` unreachable error branch: `urllib.request.urlopen`
raises `urllib.error.HTTPError` for 4xx/5xx before reaching the
`if resp.status >= 300` check, so the error body was never
surfaced. Wrap in try/except and read the body from the
HTTPError instance — that's where Anthropic's "invalid API key"
/ "rate limited" detail lives.
2. Haiku JSON extraction was fragile: `startswith("```")` assumed
the response had no prose preamble and only handled one fence
shape. Replace with `re.search(r"\{.*\}", text, re.DOTALL)` so
we pick the outermost JSON object regardless of any wrapper
markdown or leading/trailing text. Greedy + DOTALL is correct
for the single top-level object our schema requires.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(e2e): raise pytest timeout + bump multi-user chat wait to 180s
The CI run on feat/canary-report surfaced that 90s was still not
enough for test_mcp_same_server_multi_user_via_browser on
ubuntu-latest — it timed out at the inner Playwright
wait_for_function deadline with "Timeout 90000ms exceeded" after
118s of total test time.
The test opens two browser contexts + two SSE streams and drives a
full chat turn per user in sequence. Under 2-vCPU contention the
compound pipeline genuinely takes over 90s.
- tests/e2e/pyproject.toml: timeout 120 -> 240 (pytest-level cap)
- test_v2_auth_oauth_matrix.py: send_chat_and_wait_for_terminal_message
call sites 90000 -> 180000 (two owner/member turns, each budgeted
for one runner-slow turn)
180s < 240s, so the inner deadline fires first with the useful
Playwright traceback instead of the generic pytest SIGTERM.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(e2e): fix pytest-timeout CLI override + widen Mode-C deadlines
The previous commit (c3c9bbab) raised tests/e2e/pyproject.toml's
timeout from 120 to 240, but the auth canary runs the suite via
scripts/auth_canary/run_canary.py which hardcodes
`--timeout=120` on the pytest command line. The CLI flag wins
over pyproject's ini_options, so the 240 bump was invisible to
the auth lanes. That's why auth-smoke on the canary `all` run
still failed with "Timeout (>120.0s) from pytest-timeout" even
after our 180s inner widening — the outer CLI cap was firing at
120s first.
Fix the override and widen the two remaining Mode-C deadlines
that blew in the same run:
scripts/auth_canary/run_canary.py: --timeout=120 -> 240
_wait_for_refresh_request default: 60.0 -> 120.0
(test_wasm_tool_oauth_refresh_on_demand and
test_mcp_oauth_refresh_on_demand both use the default)
test_settings_first_gmail_auth_then_chat_runs call sites:
_wait_for_mock_google_tokens 90.0 -> 120.0
_wait_for_response_contains 90.0 -> 120.0
All remain comfortably under the new 240s pytest-level cap so a
real hang still fails fast with a useful traceback.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(e2e): opt-in text-match predicate for multi-user browser test
Ship the structural fix that was overdue. Repeated budget bumps on
send_chat_and_wait_for_terminal_message weren't holding under
ubuntu-latest "all"-mode parallelism — 120s, 180s both exceeded on
test_mcp_same_server_multi_user_via_browser. The underlying race is
in the JS predicate: it waits for the assistant bubble AND the
data-streaming attribute cleared AND the chat input re-enabled.
Under 2-vCPU contention an SSE reconnect can drop the final
attribute-clearing delta, and the compound predicate never flips
even though the response text arrived long ago.
Add an opt-in `expected_text_contains` parameter. When supplied,
the predicate succeeds the moment the expected substring appears in
the new assistant message — regardless of data-streaming or input
state. Callers that already assert on specific response text (the
existing MCP / gmail tests) can now short-circuit the race without
compromising correctness: the test's own content assertions remain
the gate.
Default behavior unchanged for the ~30 existing call sites across
test_chat.py, test_sse_reconnect.py, test_tool_approval.py,
test_portfolio.py, test_message_persistence.py, test_agent_loop_recovery.py,
test_pending_user_messages.py, test_widget_customization.py.
Applied to the two multi-user call sites with
expected_text_contains="Mock MCP search result" — that's exactly
what the test's next two assertions verify.
Local run of the flaky test alone: 40s, green.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* ci(canary): move auth-smoke to self-hosted runner
Multi-user browser test (test_mcp_same_server_multi_user_via_browser)
consistently exceeds the Playwright budget on GH ubuntu-latest under
the 2-vCPU parallelism pressure of an "all" canary run — a single
compound chat turn burns >180s, with each budget bump we apply it
ratchets the flake, not the fix.
Pilot move onto the [self-hosted, ironclaw-live] runner that
private-oauth already uses. Same runner label means no new
infrastructure required; if the self-hosted box has Python 3.12 and
Playwright browsers installed (or can provision them via the existing
setup-python + scripts/live-canary/run.sh's `PLAYWRIGHT_INSTALL=with-deps`
flow), this is a zero-code-change canary fix.
If the pilot works, auth-full is the next candidate. If the runner
queues become a bottleneck, we'd scale to multiple workers under
the same label rather than revert to ubuntu-latest.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* ci(canary): revert auth-smoke to ubuntu-latest + widen budgets to 300s/360s
Railway self-hosted runner ('railway-private-oauth' on a small Docker
container) turned out to be no faster than GH ubuntu-latest for the
multi-user browser flow — both take ~194–196s for
test_mcp_same_server_multi_user_via_browser. The runner container is
evidently provisioned at a similar vCPU allocation, so the move
bought nothing.
Revert to ubuntu-latest (parallel canary shape preserved; avoids
serialising auth lanes behind private-oauth on the single
self-hosted worker) and widen deadlines for the last CI-load hop:
test_v2_auth_oauth_matrix.py multi-user call sites:
Playwright wait_for_function 180000 -> 300000 ms
scripts/auth_canary/run_canary.py:
--timeout=240 -> 360 (outer pytest cap)
tests/e2e/pyproject.toml:
timeout = 240 -> 360
300s inner fits inside the new 360s outer with 60s margin. Local
run of the same test alone completes in ~40s, so we have plenty
of headroom against real hangs still surfacing fast with a
useful traceback.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* disable report
* scripts(auth-canary): add Google storage-state bootstrap helper
The auth-browser-consent lane drives Google's real OAuth consent UI in
Playwright, but Google's risk engine routinely interrupts the flow with
a "Verify it's you" challenge that handle_google_popup cannot solve, so
the test stalls on the password screen.
Bypass: log in once interactively in Playwright Chromium, save cookies
+ localStorage to a storage_state.json, point AUTH_BROWSER_GOOGLE_-
STORAGE_STATE_PATH at it. Subsequent canary runs spawn contexts with
that state preloaded, so the popup arrives at consent with no login or
challenge in the way.
- scripts/auth_live_canary/bootstrap_google_storage_state.py: new
one-shot interactive helper that writes
~/.ironclaw/auth-canary/google_storage_state.json by default
- scripts/auth_live_canary/README.md: document the bypass under
"Browser-consent Google challenge bypass"
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* canary(auth-browser-consent): fix Google account-picker + chat drift
The auth-browser-consent google case was failing on two distinct
issues, the first masking the second:
1) Account picker. When AUTH_BROWSER_GOOGLE_STORAGE_STATE_PATH is set
(the recommended path — username/password automation gets blocked
by Google's risk engine), Google's OAuth popup lands on a "Choose
an account" picker before the consent screen. handle_google_popup
only knew how to fill email + password and click Continue/Allow,
so the popup sat on the picker until complete_provider_auth's
120s callback wait timed out. Added a picker-detection step that
tries selectors in order — username text, [data-identifier], and
a generic "any visible @-bearing text not equal to 'Use another
account'" XPath — and clicks the first hit, with debug logging
so future regressions surface in the run output.
2) Tool-name and response-text drift. After the OAuth fix unblocked
the rest of the probe, browser_chat still failed because:
- case.expected_tool_name was "gmail", but the gateway records
the tool call under its WASM module name "gmail_tool"
- case.expected_text was "Gmail" (case-sensitive), but real LLM
responses to "check gmail unread" against an empty inbox vary
("Your inbox is clear...", "Inbox is empty", etc.) and rarely
emit literal "Gmail"
Updated BROWSER_CASES["google"] to expected_tool_name="gmail_tool"
and expected_text="inbox", and made the browser_chat assertion's
text comparison case-insensitive so the canary doesn't depend on
exact wording.
After both fixes the auth-browser-consent google lane runs green:
✓ browser_oauth (popup -> /oauth/callback)
✓ browser_chat (assistant references inbox)
✓ responses_api (real Gmail tool call)
Not addressed here: BROWSER_CASES["github"] likely has the same
expected_tool_name drift ("github" vs probably "github_tool"); needs
verification with real GitHub OAuth creds before changing.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* canary(auth-browser-consent): robust account-picker fallback + browser channel
Two follow-ups discovered during local debugging of the auth-browser-
consent google lane:
1) Account-picker fallback was matching hidden <style> blocks. The XPath
`//*[contains(text(), '@') ...]` matched any element whose text
contains `@`, which includes <style> tags carrying CSS at-rules
(@font-face, @media). Replaced the XPath with role-based locators
(get_by_role link/button) filtered by an email regex — only
interactive elements match, no false positives from style blocks.
Verified locally that the fallback now clicks the right account row
even when AUTH_BROWSER_GOOGLE_USERNAME is unset.
2) Bootstrap script: Google's anti-automation blocks Playwright's
default Chromium (Chrome for Testing) at sign-in with "This browser
or app may not be secure". Added a --browser flag with a default of
firefox (Marionette is less aggressively fingerprinted than CDP),
plus chrome (system Google Chrome) and chromium (override) options.
For accounts where Google blocks even those — typically brand-new
Gmails or accounts with high risk scores — the fallback path is to
launch Chrome manually with --remote-debugging-port and connect via
playwright.chromium.connect_over_cdp; documented in the README.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* canary(auth-live-canary): include observed extension state in timeout error
When `wait_for_extension_state` times out the bare error
"Timed out waiting for extension state: gmail" is unhelpful for
diagnosing CI failures, since CI artifacts don't capture IronClaw's
gateway logs — there's no way to tell whether the extension never
appeared, appeared but never authenticated, or authenticated but
never activated.
Track the last-observed extension on each poll and surface
authenticated/active in the timeout message. After this change a
failed run says e.g.
"Timed out waiting for extension state: gmail (expected
authenticated=True, active=True; last observed: authenticated=False,
active=False)", which immediately separates token-exchange failures
from activation-state-machine bugs.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* canary(auth-live-canary): widen chat-wait deadlines 120s -> 300s
The auth-browser-consent google probe completed OAuth + extension
activation successfully on CI but timed out at the next step
(send_chat_and_wait_for_terminal_message), with the agent stuck on
"Thinking (step 1)" for the full 120s budget. Local runs on the
same code path complete the chat in ~36s, but ubuntu-latest 2-vCPU
runners under cold-start load (gateway restart, mock LLM bootstrap,
WASM tool first-invocation) need substantially more headroom.
300s matches the precedent set by `d8765714 ci(canary): revert
auth-smoke to ubuntu-latest + widen budgets to 300s/360s` for the
auth-smoke lane on the same runner class.
Both call sites widened — the seeded Responses-API probe at line 221
and the browser_oauth probe at line 800.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* canary(common): drain gateway/mock_llm stdout pipes (was deadlocking CI)
scripts/live_canary/common.py spawns the IronClaw gateway and the
mock LLM with stdout=PIPE + stderr=STDOUT, reads one line of mock_llm
output to discover its bound port, then never reads from either pipe
again. On Linux the kernel pipe buffer caps at 64 KiB; once a
sustained chat request fills it with `RUST_LOG=info` output, the
child blocks on its next stdout write and the request handler
freezes mid-response.
That's why every auth-browser-consent CI run got stuck on
"Thinking (step 1)..." for the full chat-wait budget while the same
test passes locally — macOS pipe buffers are larger and the test
completes before the buffer fills.
Fix: spawn a daemon thread per subprocess that drains the pipe to a
log file under the run's output_dir. Two wins:
- Pipes never fill, child never blocks.
- gateway.log and mock_llm.log become CI artifacts, so the next
failure that doesn't have a clear runner-side error message is
immediately debuggable from IronClaw's own logs.
Verified locally that the lane still passes after the change and
both log files are produced. Locally each is < 10 KiB; CI runs may
be larger but well under any artifact size limit.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* canary: pin LLM backend via settings API + add LLM_API_KEY (root cause of CI freeze)
The auth-browser-consent google lane has been freezing on CI at
"Thinking (step 1)..." for the full chat-wait budget. Gateway logs
captured by the previous commit's pipe drainer reveal the smoking
gun:
ERROR Configured LLM backend is not usable.
backend=openai_compatible reason=missing API key
WARN LLM_BACKEND env var is set but DB setting takes priority.
db_value=nearai env_value=openai_compatible
WARN Active LLM backend fell back to NearAI default
attempted=openai_compatible active=nearai
Two compounding issues:
1. The openai_compatible provider refuses to instantiate without an
API key, even though the mock LLM ignores the value. Fix: set
`LLM_API_KEY=mock-api-key` in `build_gateway_env`, matching what
`tests/e2e/conftest.py` already does for the e2e suite.
2. IronClaw's DB-stored LLM settings take priority over env vars,
and the freshly-seeded canary DB defaults `llm_backend` to
`nearai`. So even with a clean env, the agent fell back to NearAI
and entered an interactive auth flow that hangs indefinitely in
CI (the "Thinking" never ends). This is the exact trap
`tests/e2e/CLAUDE.md` documents: "do not rely on env-vs-DB
precedence … pin the provider explicitly through /api/settings/...".
Fix: pin `llm_backend`, `openai_compatible_base_url`, and
`selected_model` via PUT /api/settings/<key> immediately after the
gateway becomes healthy.
Also revert the BROWSER_CASES["google"] case I touched earlier:
when NearAI was driving it emitted the WASM canonical tool name
(`gmail_tool`), but the mock LLM (now correctly driving) emits the
tool name it knows from its mapping (`gmail`). Restoring the original
`expected_tool_name="gmail"` / `expected_text="gmail"` matches what
the mock LLM actually produces.
Verified locally: all three browser_oauth / browser_chat /
responses_api probes now pass.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* canary(auth-live-canary): revert chat-wait deadline 300s -> 120s
The 300s widening at 98abeebe was a band-aid attempt to work around
the actual root cause (subprocess pipe deadlock + DB-overrides-env
LLM backend), which were both fixed at f59981d3 and 8733d3c0
respectively. With those fixes the chat completes in ~35s on CI, so
the 300s budget is overkill — revert to the original 120s, which
gives ~3.5x headroom over the observed steady-state and matches the
deadline shape used elsewhere in the e2e suite.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* ci(canary): rename github oauth secrets to dodge GITHUB_ prefix block
GitHub Actions reserves the GITHUB_ prefix for auto-generated repo
secrets (GITHUB_TOKEN, etc.) and rejects user-created secrets that
start with it: "Secret names must not start with GITHUB_". The
existing references to GITHUB_OAUTH_CLIENT_ID and GITHUB_OAUTH_-
CLIENT_SECRET in this workflow couldn't be backed by actual secrets
for that reason — the OAuth-client config was effectively unset for
the github browser-consent case, which is why it was silently
filtered out by configured_browser_cases().
Decouple the secret name from the env var name: store the secrets
under the AUTH_BROWSER_GITHUB_CLIENT_ID / AUTH_BROWSER_GITHUB_CLIENT_-
SECRET names (matching the AUTH_BROWSER_GITHUB_* convention used by
the other github canary fixture vars), and re-export them here under
the GITHUB_OAUTH_CLIENT_ID / _SECRET env names that
auth_registry.py and the WASM github tool expect.
No code changes needed in auth_registry.py / scripts/auth_live_-
canary/ — they continue to read GITHUB_OAUTH_CLIENT_ID/_SECRET from
the environment as before.
Operator action: create the OAuth app on GitHub (Settings →
Developer settings → OAuth Apps → New OAuth App) and store the
resulting credentials at:
AUTH_BROWSER_GITHUB_CLIENT_ID
AUTH_BROWSER_GITHUB_CLIENT_SECRET
(not GITHUB_OAUTH_CLIENT_ID / _SECRET, which GitHub will reject).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* canary(auth-browser-consent): drop github case (tool is PAT-only, not OAuth)
CI run 25022303491 surfaced that `Activate /api/extensions/github/-
activate` returns `{success: false, awaiting_token: true,
message: "Create a Personal Access Token..."}` with no `auth_url`,
which the browser-consent probe needs in order to drive the OAuth
popup.
Confirmed via `registry/tools/github.json`:
"auth_summary": {
"method": "manual", <- PAT paste, not OAuth
"secrets": ["github_token"],
"setup_url": "https://github.com/settings/tokens"
}
The github WASM tool's source capabilities JSON does carry an `oauth`
block, but the released v0.2.3 artifact (referenced from the registry)
ships with the manual-auth path. Until a release flips
`auth_summary.method` to "oauth" — and the github extension actually
returns an `auth_url` from /activate — there's nothing for the
browser-consent probe to do.
- Drop the `github` entry from BROWSER_CASES with a comment pointing
at the criterion for re-adding it.
- Drop the github-specific filter in `configured_browser_cases` since
the case is gone (no risk of an env-aware code path that quietly
skips github when secrets are present-but-mismatched).
GitHub coverage is unchanged in SEEDED_CASES, which seeds the PAT
directly via `AUTH_LIVE_GITHUB_TOKEN` and exercises real
`/v1/responses` + browser tool calls — that lane already works.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* canary(auth-browser-consent): tick notion's trust-URL checkbox before Continue
CI run 25023708895 surfaced the notion case timing out at "Timed out
waiting for notion OAuth callback page". The popup screenshot shows
Notion MCP's consent screen with:
- Workspace correctly auto-selected (storage state worked)
- A yellow warning: "I recognize and trust this URL"
- An unchecked checkbox next to that text
- A grayed-out (disabled) Continue button
The button is gated behind the checkbox. handle_notion_popup
clicked the disabled Continue and silently no-op'd, so the
complete_provider_auth loop waited the full 120s for /oauth/callback
that never arrived.
Add a checkbox-detection step before the Continue click:
popup.get_by_text(re.compile("I recognize and trust this URL", I))
.first.click(timeout=3000)
Includes debug print statements (matching the auth-canary pattern
established for google's account picker) so future Notion UI
changes are immediately visible in test-output.log.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(e2e): drain ironclaw subprocess pipes in auth-matrix fixture
Same pipe-deadlock fix as scripts/live_canary/common.py f59981d3,
applied to tests/e2e/scenarios/test_v2_auth_oauth_matrix.py's
_start_auth_matrix_server. The auth-matrix fixture spawns ironclaw
with stdout=PIPE + stderr=PIPE and never drains them, so under
sustained log volume the kernel pipe buffer fills, ironclaw blocks
on its next stdout write, and any test that relies on subsequent
gateway responses (auth gate emission, SSE events, chat replies)
hangs until pytest-timeout fires.
This fix doesn't make the auth-full lane's failing test pass — the
real bug is engine-v2 silently dropping `auth_required` SSE events
for unauthenticated extensions (introduced by #2868). But it makes
the failure mode debuggable: gateway log is captured to
/tmp/ironclaw-auth-matrix-gateway.log (overridable via
IRONCLAW_AUTH_MATRIX_LOG env), and RUST_LOG passes through from the
test runner so we can crank up verbosity without rebuilding.
Without this change, the failing test's log was empty after the
extension-install line; with this change you see the engine-v2
trace summary that surfaces the actual NotCallable-without-auth-gate
bug. That diagnostic visibility is the value here.
- _drain_stream_to_file: asyncio drainer mirroring common.py's sync
threading version
- _start_auth_matrix_server: drain stdout/stderr to log_path
- _shutdown_auth_matrix_server: cancel drain_tasks for clean exit
- env: RUST_LOG forwarding so debug runs work
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* canary(workflow): add Telegram Bot API mock
Foundation piece for the new workflow-canary lane that exercises
multi-tool / multi-channel user workflows from issue #1044 (Telegram +
routines + Sheets/Calendar/Gmail end-to-end). Models the same
single-port aiohttp-based mock pattern used by tests/e2e/mock_llm.py.
Endpoints:
- /bot{token}/{getMe,getUpdates,sendMessage,sendChatAction,
setWebhook,deleteWebhook,getFile} — the subset IronClaw's WASM
telegram tool + channels-src/telegram actually call. Tokens are
accepted without validation; the canary doesn't need to test
Telegram's auth — just IronClaw's flow against a Bot API shape.
- /__mock/inject_message — push a simulated incoming user message
onto the next getUpdates response, so scenarios can drive a
Telegram → IronClaw round-trip without a real Telegram account.
- /__mock/sent_messages — drain the queue of every sendMessage /
sendChatAction IronClaw emitted, for end-to-end assertions.
- /__mock/reset — clear all state between probes.
IronClaw routes its API calls through this mock via
IRONCLAW_TEST_HTTP_REMAP=api.telegram.org=<mock_url>, the same
mechanism the auth-live-canary uses for Gmail/Calendar/Sheets mocks.
Smoke-tested: getMe → success, inject_message → getUpdates returns
the injected message, sendMessage → bot response shape + recorded
in sent_messages.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* canary(workflow): land workflow-canary lane with periodic-reminder scenario
Phase 1A of the workflow-canary system from issue #1044. Adds a new
canary lane that exercises the routine engine + cron-fire path, the
foundation that the remaining four scripts (Telegram → Sheets,
Calendar prep, HN monitor, CRM tracker) will layer on.
Components:
- scripts/workflow_canary/routines.py — direct libSQL helpers for
inserting a lightweight cron routine with a backdated next_fire_at
and polling routine_runs for terminal status (ok / attention /
failed). Backdating beats wall-clock cron in tests by 30+ s per
probe and is the same shape auth-live-seeded uses for
expire_secret_in_db.
- scripts/workflow_canary/run_workflow_canary.py — entrypoint that
starts the Telegram mock, calls common.start_gateway_stack with
workflow-tuned env (ROUTINES_ENABLED=true, ROUTINES_CRON_INTERVAL=2,
IRONCLAW_TEST_HTTP_REMAP=api.telegram.org=<mock>), and runs
scenario modules. CLI mirrors run_live_canary.py.
- scripts/workflow_canary/scenarios/periodic_reminder.py — Script 4
Phase 1A: insert lightweight routine → wait for engine to fire →
assert run row reaches a terminal status. Verified locally: 1
probe, 1 fire, status=attention.
Plumbing:
- .github/workflows/live-canary.yml — new workflow-canary job + lane
added to the workflow_dispatch choice list and the canary-report
aggregator's needs:.
- scripts/live-canary/run.sh — workflow-canary case dispatches to
run_workflow_canary.py.
Phase 1B follow-ups in subsequent commits:
- Telegram channel install + bot-token seeding (needs admin auth or
direct encrypted-secrets DB write)
- Verify Telegram sendMessage was emitted to the mock during the
routine fire (covered by mock telegram's /__mock/sent_messages)
- Scripts 1, 3, 5 (Sheets / HN / Gmail-CRM)
- Script 2 (Calendar prep with web search)
Local verification:
$ tests/e2e/.venv/bin/python scripts/workflow_canary/run_workflow_canary.py \
--skip-build --skip-python-bootstrap
[workflow-canary] mock telegram listening at http://127.0.0.1:51139
[periodic_reminder] inserted routine ..., next_fire_at backdated 60s
[periodic_reminder] routine fired: status=attention
[workflow-canary] all 1 probe(s) passed.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* canary(workflow): land all 5 issue #1044 scenarios + scenario README
Layer Scripts 1, 2, 3, 5 onto the foundation shipped in 16278ea9, so
the workflow-canary lane covers all five user-workflow scripts from
issue #1044. Each scenario delegates to a shared
`run_routine_probe()` helper that captures the Phase 1A shape: insert
a Lightweight cron routine with a script-specific prompt → backdate
next_fire_at → poll routine_runs for terminal status.
Scenarios added:
- bug_logger.py (Script 1 — Telegram bugs → Google Sheet)
- calendar_prep.py (Script 2 — Calendar prep → Telegram, Reporter: Nick)
- hn_monitor.py (Script 3 — Hacker News → Telegram, Reporter: Emil)
- crm_tracker.py (Script 5 — Gmail → Sheets CRM, Reporter: Cameron)
Plus periodic_reminder.py (Script 4, Reporter: Henry) refactored to
also use run_routine_probe.
scenarios/_common.py centralizes the routine plumbing — each scenario
file is now ~30 lines of routine-name + prompt + Phase 1B follow-up
notes. The Phase 1B follow-up plan (Telegram channel install, mock
Sheets writes, mock Calendar reads, mock HN scrape, LLM email
classification, dedup verification) is documented inline in each
scenario's docstring AND in the new scripts/workflow_canary/README.md.
Local verification: all 5 probes green in ~2 s each.
$ tests/e2e/.venv/bin/python scripts/workflow_canary/run_workflow_canary.py \
--skip-build --skip-python-bootstrap
[workflow-canary] === Script 1 — Telegram → Google Sheet Bug Logger ===
[workflow-canary] === Script 2 — Calendar Prep Assistant ===
[workflow-canary] === Script 3 — Hacker News Keyword Monitor ===
[workflow-canary] === Script 4 — Periodic Reminder via Telegram ===
[workflow-canary] === Script 5 — Email → CRM Inbound Tracker ===
[workflow-canary] all 5 probe(s) passed.
What this catches:
- Routine engine cron-tick path (spawn_cron_ticker → check_cron_triggers)
- RoutineAction::Lightweight execution
- DB serialization of action_config / trigger_config
- Mock-LLM round-trip latency under cron scheduling
- routines.next_fire_at → routine_runs status state machine
What it doesn't catch yet (per-scenario Phase 1B work, documented in
README + scenario docstrings):
- Telegram channel install + sendMessage assertion
- Mock Sheets / Calendar / Gmail / HN write+read semantics
- LLM-driven structured classification (CRM)
- Cross-fire dedup verification
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* canary(workflow): scaffold Phase 1B telegram-side-effect verification
Lays the groundwork for verifying mock-Telegram side effects from
each scenario's routine fire — but gates the verification off until
a separate engine bug is fixed.
What's added:
- tests/e2e/mock_llm.py: new TOOL_CALL_PATTERNS entry that matches
``[CANARY-WORKFLOW-<key>]`` in any prompt and emits a deterministic
http tool call to api.telegram.org/.../sendMessage with a
per-scenario ack text.
- scripts/workflow_canary/scenarios/_common.py: each scenario now
composes its prompt as
``<prompt_intro>\n\n[CANARY-WORKFLOW-<key>]`` so the matcher fires.
When ``verify_telegram=True``, the helper polls
/__mock/sent_messages for up to 5 s and asserts the expected ack
was captured. Default is ``verify_telegram=False`` (Phase 1A
parity) — see below.
- scripts/workflow_canary/telegram_mock.py: aiohttp request-logger
middleware so the canary's stdout shows every inbound request,
giving operators a one-line answer to "did the gateway's HTTP
remap actually reach the mock?".
- scripts/workflow_canary/scenarios/{bug_logger,calendar_prep,
hn_monitor,periodic_reminder,crm_tracker}.py: scenarios pass
``mock_telegram_url=mock_telegram_url`` and ``prompt_intro=...``
ready for verify_telegram to flip on.
What's gated off and why:
The mock-Telegram verification path requires
``IRONCLAW_TEST_HTTP_REMAP=api.telegram.org=<mock>`` to route
the http tool's sendMessage call into the mock. The remap is
correctly registered at gateway startup
(src/app.rs::http_interceptor + src/http_intercept.rs), but the
ToolContext built inside the routine engine's Lightweight action
loop does NOT inherit the global ``http_interceptor`` slot. Result:
the http tool reaches into the real network for api.telegram.org
(returning a 401 since the bot token is fake) and the mock never
sees the request — confirmed via the new request-logger middleware
showing zero non-internal hits.
That's a real engine bug in routine-driven tool dispatch — the
http_interceptor needs to propagate through the routine action's
ToolContext just like it does for chat-driven tool dispatch. Out of
scope for this canary PR; tracked as a follow-up. Once fixed, flip
the default in ``run_routine_probe`` and every scenario's
verify_telegram check activates with no further changes.
Local verification: all 5 probes still green at the Phase 1A level.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* canary(workflow): re-exec under venv after bootstrap (fix CI 'No module named httpx')
CI run 25028445222 failed on the workflow-canary lane with:
[workflow-canary] mock telegram listening at http://...
[workflow-canary] error: No module named 'httpx'
Root cause: run_workflow_canary.py was missing the bootstrap-then-
reexec pattern that scripts/auth_live_canary/run_live_canary.py
uses (line 1229+). bootstrap_python() creates the venv and installs
tests/e2e/'s pyproject deps (which include httpx + aiohttp), but
the parent process keeps executing under whatever interpreter
invoked it — typically the system Python on CI runners, which
doesn't have httpx. The scenario module's `import httpx` at top
level then fails immediately.
Fix: copy the auth-live-canary reexec pattern. main() now:
1. If not --skip-python-bootstrap AND WORKFLOW_CANARY_REEXEC is
unset: bootstrap the venv, install playwright, build cargo,
then subprocess-spawn ourselves under the venv python with
--skip-python-bootstrap and WORKFLOW_CANARY_REEXEC=1 so this
branch isn't re-entered.
2. The reexecuted process sees skip_python_bootstrap=True and runs
the actual canary against the venv interpreter that has all
deps available.
Local sanity check: still passes (--skip-build --skip-python-bootstrap
short-circuits the bootstrap, both branches behave identically when
the venv already exists).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(routine-engine): propagate http_interceptor into Lightweight tool dispatch
The chat path's tool dispatch correctly receives the global
HTTP interceptor (e.g., the `IRONCLAW_TEST_HTTP_REMAP` debug-only
host remapper installed in `src/app.rs::http_interceptor`), but the
routine engine's Lightweight action path constructed its
`JobContext` from scratch with `..Default::default()`, leaving
`http_interceptor: None`. Tools called from a routine therefore
reached the real network even when the rest of the system was
configured to route through mocks.
Plumb the interceptor through:
- `RoutineEngine` gains an `http_interceptor` field
- `RoutineEngine::new` takes it as the 11th argument
- `EngineContext` carries it across the spawn boundary
- `JobContext` construction at the Lightweight action site copies
it from the engine context
Threading complete: AgentDeps → RoutineEngine → EngineContext →
JobContext → http tool. Same shape the chat path already uses.
Test rigs updated: `tests/support/test_rig.rs` and
`tests/e2e_routine_heartbeat.rs` (10 call sites total) pass `None`
for the new arg, matching their existing minimal stack model.
Build clean against `--no-default-features --features libsql`.
Why this matters: with the interceptor lost, every workflow-canary
probe's http tool dispatch reached real api.telegram.org and 401'd
on the fake token — leaving the mock Telegram bot empty and the
canary's send-side assertions unverifiable. With the fix, the
interceptor honors the IRONCLAW_TEST_HTTP_REMAP and the workflow
canary's Phase 1B verification activates immediately.
Activates in this commit:
- scripts/workflow_canary/scenarios/_common.py default flips to
`verify_telegram=True`
- All 5 scenarios (bug_logger, calendar_prep, hn_monitor,
periodic_reminder, crm_tracker) now assert that the mock
Telegram bot received the per-scenario ack message
`[canary-workflow:<key>] ack`
Local verification:
$ tests/e2e/.venv/bin/python scripts/workflow_canary/run_workflow_canary.py \
--skip-build --skip-python-bootstrap
[workflow-canary] === Script 1 — Telegram → Google Sheet Bug Logger ===
... (all 5 scenarios) ...
[workflow-canary] all 5 probe(s) passed.
$ grep "POST /bot" artifacts/workflow-canary/telegram_mock.log | wc -l
5 # one per scenario, distinct ack text per probe
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* canary(workflow): add manual_trigger + lifecycle + dedup_cooldown probes
Three new scenarios covering issue #1044 assertions that the existing
5 cron-fire probes don't reach. Each scenario tests a distinct
back-end mechanism that real users hit:
- **manual_trigger** (Scripts 3 PHASE 2.1 + 3 PHASE 4.2 + 4 PHASE 4.2)
Inserts a routine WITHOUT backdating next_fire_at, so the only path
to a fire is the manual-trigger API. POSTs
/api/routines/<id>/trigger, asserts response carries a run_id, polls
routine_runs for terminal status, then verifies mock Telegram
captured the per-scenario ack. Catches regressions in
RoutineEngine::fire_manual end-to-end.
- **lifecycle** (Scripts 1 PHASE 5 + 4 PHASE 5) — three sub-probes:
1. disabled-blocks-fires: insert with enabled=False + backdate;
assert no routine_runs row appears within 8 s window.
2. enable-resumes-fires: toggle enabled=true via API, backdate,
assert fire reaches terminal status.
3. delete-removes-routine: confirm /api/routines lists it, DELETE,
confirm it's gone.
Catches regressions in toggle handler, delete handler, and the
engine's enabled-flag respect during cron tick selection.
- **dedup_cooldown** (Scripts 1 PHASE 4.4 + 3 PHASE 3.2 + 5 PHASE 5.5)
Insert with cooldown_secs=30; first fire lands within ~5 s; immediate
re-backdate; assert ONLY ONE run row exists after 8 s. Catches
regressions in cooldown enforcement during check_cron_triggers.
This is the closest engine-level correlate to the user-script
"no duplicate rows / alerts / messages" assertions, which are
application-level dedup that lives outside the canary's
deterministic-mock surface.
Plumbing:
- routines.py: trigger_routine_via_api / toggle_routine_via_api /
delete_routine_via_api / list_routines_via_api helpers (all auth-
bearer, JSON in/out, raise_for_status).
- routines.py: insert_lightweight_cron_routine grew `cooldown_secs`
+ `enabled` parameters; defaults preserve existing behavior.
- run_workflow_canary.py: registered the three new scenario keys.
Local verification — all 10 probes (5 original + 5 new sub-probes
across 3 new scenarios) green:
✅ bug_logger / calendar_prep / hn_monitor / periodic_reminder /
crm_tracker (existing — Telegram ack capture)
✅ manual_trigger (548ms)
✅ lifecycle_disable (8004ms — full no-fire window)
✅ lifecycle_toggle (1543ms)
✅ lifecycle_delete (56ms)
✅ dedup_cooldown (10017ms — first fire + 8s no-fire window)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* canary(workflow): add NL-driven routine_create + routine_update probes
Two scenarios that close issue #1044's chat-driven assertions
(Script 1 PHASE 3.1, Script 2 PHASE 3.1, Script 3 PHASE 2.1,
Script 4 PHASE 2.1 + 5.1, Script 5 PHASE 4.1):
- **nl_routine_create**: opens a thread via /api/chat/thread/new,
posts an NL message tagged [CANARY-WORKFLOW-NL-CREATE], waits for
the agent to dispatch routine_create, then verifies the routines
row landed in libSQL AND is visible via GET /api/routines.
- **nl_schedule_update**: pre-seeds a target routine
(canary-nl-update-target), posts an NL message tagged
[CANARY-WORKFLOW-NL-UPDATE], waits for the agent to dispatch
routine_update with a new schedule, then verifies trigger_config
changed in libSQL. Asserts on schedule-changed (not exact match)
because the engine normalizes 5-field cron → 7-field internal
form ("0 */5 * * *" → "0 0 */5 * * * *").
Plumbing:
- Two new TOOL_CALL_PATTERNS entries in tests/e2e/mock_llm.py
matched in priority order (specific NL-CREATE / NL-UPDATE
sentinels checked BEFORE the generic [CANARY-WORKFLOW-<key>]
http-tool fallback, since the canary's own routines emit the
generic pattern from inside their action prompts).
- Helper additions in scripts/workflow_canary/routines.py:
_open_thread / _send_chat / _read_routine / _wait_for_*.
Local verification — all 12 probes green:
✅ bug_logger / calendar_prep / hn_monitor / periodic_reminder /
crm_tracker (5 cron-fire + telegram-ack)
✅ manual_trigger (POST /api/routines/<id>/trigger)
✅ lifecycle_disable / lifecycle_toggle / lifecycle_delete
✅ dedup_cooldown (cooldown_secs suppresses second fire)
✅ nl_routine_create (chat → routine_create tool)
✅ nl_schedule_update (chat → routine_update tool)
What's still deferred to follow-up PRs (per-provider mocks, each
~1-3 days of work — see scripts/workflow_canary/README.md):
- Mock Google Sheets (Scripts 1 + 5 dedicated assertions)
- Mock Google Calendar (Script 2)
- Mock Hacker News (Script 3)
- LLM-driven email classification with seeded inbox (Script 5)
- Telegram channel install + bot-token validation flow (Scripts 1-5
PHASE 1)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(workflow-canary): Phase 1 — mock Sheets + bug_logger Sheet-write probe
Adds scripts/workflow_canary/sheets_mock.py: single-port aiohttp Google
Sheets v4 mock supporting POST /v4/spreadsheets, values:append, values
get, plus /__mock/ test hooks for seeding, draining, and resetting.
The append handler enforces values=list-of-lists (returns the canonical
"expected a sequence" 400) so the canary catches the issue #1044 FAIL
CRITERIA shape.
Wires the mock into run_workflow_canary.py:
- generic _spawn_mock helper for telegram_mock + sheets_mock
- IRONCLAW_TEST_HTTP_REMAP carries comma-separated entries for
api.telegram.org and sheets.googleapis.com
- mock_sheets_url passed through to every scenario's run() kwargs
Rewrites scenarios/bug_logger.py to drop the run_routine_probe Telegram
fallback in favor of a Sheet-write end-to-end assertion: pre-seed the
spreadsheet, fire the routine with [CANARY-WORKFLOW-SHEET-APPEND], wait
for the appended row, validate shape (timestamp / message / source).
Mock LLM: new TOOL_CALL_PATTERNS entry that matches the SHEET-APPEND
sentinel and emits an http POST values:append with a hardcoded canary
row.
All 12 probes still pass locally.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(workflow-canary): Phase 2-4 — Calendar / HN / Gmail / web_search mocks + e2e probes
Phase 2 (Calendar): scripts/workflow_canary/calendar_mock.py — Google
Calendar v3 events surface (list / insert / get / delete) with seed
hooks. calendar_prep_e2e seeds one canary event, fires the routine,
asserts events.list was hit and Telegram received the prep briefing
referencing the seeded event title.
Phase 3 (Hacker News): scripts/workflow_canary/hn_mock.py — /newest
HTML fixture with seeded "Show HN" posts (canary-distinct
``<!-- canary-hn-feed -->`` marker). hn_monitor_e2e re-seeds posts,
asserts /newest GET landed and Telegram summary references both
seeded posts.
Phase 4 (CRM tracker): scripts/workflow_canary/gmail_mock.py +
web_search_mock.py — Gmail v1 messages.list/.get + Brave Search v3.
crm_tracker_e2e seeds 1 lead + 1 newsletter + 1 receipt; asserts
exactly ONE row appended to the CRM sheet (only the lead) with all
6 expected columns + Telegram ack referencing 1 lead.
Mock LLM TOOL_CALL_PATTERNS gain three parallel-call entries
([CANARY-WORKFLOW-CAL-LIST] → http GET events.list + http POST
sendMessage; [CANARY-WORKFLOW-HN-FETCH] → GET /newest + sendMessage;
[CANARY-WORKFLOW-CRM-CLASSIFY] → Gmail GET + Sheets append + Telegram
ack). Parallel emit is required because the engine's lightweight
loop dedups same-tool re-dispatch (see match_tool_call:1178).
run_workflow_canary.py now spawns six mock subprocesses; remap covers
api.telegram.org, sheets.googleapis.com, www.googleapis.com,
news.ycombinator.com, gmail.googleapis.com, api.search.brave.com.
All 12 existing probes pass + 3 phase 2-4 probes upgrade from
side-effect-only to full content-correctness assertions.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(workflow-canary): Phase 5 — Telegram channel install + round-trip
scripts/workflow_canary/telegram_setup.py: install + capability patch
+ setup helpers (mirrors tests/e2e/scenarios/test_telegram_e2e.py
patch_capabilities + activate flow). Adds pair_telegram_user that
sends an "hello" webhook, extracts the pairing code from
mock_telegram, and approves it via /api/pairing/telegram/approve.
scripts/live_canary/common.py: GatewayStack now exposes http_url
(HTTP-channel webhook port) + channels_dir (WASM_CHANNELS_DIR)
so workflow-canary scenarios can drive the Telegram channel install
+ patch + webhook flow.
run_workflow_canary.py: passes IRONCLAW_TEST_TELEGRAM_API_BASE_URL
so the hardcoded validate_telegram_bot_token getMe call (in
src/extensions/manager.rs) routes to mock_telegram. The bot-token
validate path bypasses the standard IRONCLAW_TEST_HTTP_REMAP flow,
hence the additional env override.
New scenarios:
- telegram_channel_install: install + patch caps + setup + assert
channel reaches Active state. Catches "HTTP 404 on valid token"
regression (Script 4 PHASE 1.1).
- telegram_round_trip: post inbound webhook → assert mock_telegram
receives an outbound sendMessage with the actual chat_id (NOT
'default'). Catches the chat_id 'default' regression.
- routine_visibility_from_telegram: pair user, ask for routines,
assert agent replies on the paired chat_id. Covers Scripts 1-4
PHASE "routine visibility from Telegram" assertions.
- manual_trigger_from_telegram: pair user, hit /api/routines/<id>/
trigger, assert routine fires through lightweight loop and ack
reaches the paired chat_id. Covers Script 4 PHASE 4.2.
All 16 probes pass locally.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(workflow-canary): Phase 6 — first_immediate_run + log_assertions
scripts/workflow_canary/scenarios/first_immediate_run.py: insert a
routine with a "0 * * * *" schedule + fire_immediately=True; assert
the first run reaches terminal status within 10s. Catches "first
check is delayed to next hour" regression (Script 3 PHASE 2.1).
scripts/workflow_canary/scenarios/log_assertions.py: scan
gateway.log at the end of the lane for known fail-criterion regex
patterns: chat_id 'default', parsed naive timestamp without timezone,
retry after None, expected a sequence. Catches log regressions across
all 5 issue #1044 scripts simultaneously.
Auth-recovery (token revocation → auth_required SSE) is deferred to
the auth-live-canary lane; it requires a working OAuth setup to
revoke, which is outside this lane's mock-only scope.
All 18 probes pass locally.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(workflow-canary): Phase 7 — cron timing + idempotent toggle + README
scripts/workflow_canary/scenarios/cron_timing_accuracy.py: insert a
routine, set next_fire_at to "now + 5s" explicitly, assert the engine
fires within ±10s of the set boundary. Catches "cron skipped a cycle"
+ "fires never trigger" regressions (Scripts 3 PHASE 3.1, 4 PHASE 3.4).
scripts/workflow_canary/scenarios/idempotent_disable_enable.py:
double-toggle disable then double-toggle enable, assert both halves
are no-ops; finally backdate, fire once, then disable + backdate again
and assert no NEW runs land in the next 6s. Catches "disable doesn't
take effect" + "enable triggers a phantom run" regressions
(Script 1 PHASE 5.1 / 5.2).
scripts/workflow_canary/README.md: rewritten to reflect 20-probe
coverage matrix across phases 1–7 with mock surface + scenarios
inventory.
Final canary state: 20 probes across 7 phases, all green locally.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(workflow-canary): close gaps — wire web_search + add auth_recovery
[CANARY-WORKFLOW-CAL-LIST] now emits a parallel triplet (calendar
events.list + web_search company lookup + telegram sendMessage).
calendar_prep asserts mock_web_search captured the lookup with the
expected company-name query parameter, completing the Script 2
"company background + recent news" assertion from issue #1044.
scripts/workflow_canary/scenarios/auth_recovery.py: drives a chat
that triggers an unauthenticated gmail tool call, asserts the agent
surfaces a graceful response — chat send returns 202 (not 5xx),
thread settles, history contains no Error 400 / Internal Server
Error / panicked / Traceback fragments. Catches the regression
shape from Script 2 PHASE 5 fail criteria without requiring a real
OAuth handshake (full token-revocation coverage stays in
auth-live-canary).
21 probes total, all green locally.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* ci(canary): run every 6h + re-enable Slack report
Schedule: cron flips from "0 2 * * *" (once daily at 02:00 UTC) to
"0 */6 * * *" (4× daily at 00/06/12/18 UTC). All twelve job-level
`if:` guards updated in lockstep so each lane still gates on the
schedule string.
Slack report: drop the `if: false` hardcode on the canary-report
job's notify step and replace with a schedule + workflow_dispatch
gate. The notifier (scripts/live-canary/notify_slack.py) already
exits 0 on Haiku/Slack failures so a flaky webhook can't mask lane
status. PR-triggered runs (currently none, but possible via
workflow_run) skip the post to keep noise out of the channel.
Both ANTHROPIC_API_KEY and SLACK_WEBHOOK_URL repo secrets are
already populated.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(canary-report): parse workflow-canary results.json shape
The notifier reads `auth-canary-junit.xml` for JUnit-emitting lanes
(auth-smoke, auth-full, auth-channels, auth-live-seeded,
auth-browser-consent). The workflow-canary lane writes its own
`results.json` instead — one entry per probe with `success: bool`,
`latency_ms`, `details`. The notifier had no parser for that shape, so
the workflow-canary slot in Slack rendered as a useless
`:grey_question: 0/0 passed, 0 failed` line.
Add `parse_results_json` mirroring the JUnit parser's contract:
`passed = sum(success)`, `failed = sum(!success)`, each failed probe
becomes a `(provider/mode, error-or-summary)` entry on
`junit_failures` so the Slack reason field renders the same way as an
auth-canary failure. Latencies sum to `duration_s`. Both parsers run
on every lane dir; first one whose file exists wins (auth-canary lanes
emit XML only, workflow-canary lane emits JSON only — no overlap).
Validated by re-running the notifier locally against the downloaded
artifact from CI run 25033224036:
before: ":grey_question: workflow-canary (mock) — 0/0 passed"
after: ":white_check_mark: workflow-canary (mock) — 21/21 passed,
0 failed in 69s"
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(canary-report): log notifier progress for diagnosability
Until now `notify_slack.py` was silent on the success path, which made
it impossible to verify from CI logs alone whether Haiku enrichment
actually ran. Add four stderr lines covering each phase:
[notify_slack] discovered N lane dir(s): lane1/provider1, ...
[notify_slack] lane/provider: tests=N passed=N failed=N skipped=N status=...
[notify_slack] haiku enriched X/N lane(s)
[notify_slack] posted Slack message for N lane(s)
Lines stay terse and structured so they're greppable from `gh run
view --log`. Haiku-failure tracking inspects `r.notable` — `run_haiku`
stamps it with `haiku call failed:` / `haiku returned no JSON object`
/ `haiku JSON parse failed` on the three failure paths.
Confirmed from local dry-run against the artifact downloaded from
the previous CI run (which had the results.json parser): tests=21,
passed=21, failed=0, status=pass.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(canary/workflow-canary): forward SCENARIO into --scenario
Addresses @henrypark133's review on PR #2874: the workflow-canary lane
of `scripts/live-canary/run.sh` ignored `${SCENARIO}` and always ran
the full 21-probe suite. The matching workflow_dispatch job didn't
export `inputs.scenario` either, so manual dispatch with a scenario
filter went nowhere. Targeted local reruns / debugging hit the same
gap.
run.sh: translate `${SCENARIO}` (comma-list supported) into one or
more `--scenario <name>` flags on `run_workflow_canary.py`. Empty
SCENARIO falls through to the full suite. Guards the array splat for
bash 3.2 / macOS where `${arr[@]}` on an empty array under `set -u`
explodes.
live-canary.yml: add `SCENARIO: ${{ inputs.scenario }}` to the
Workflow Canary job's env so workflow_dispatch reaches run.sh.
Verified:
tests/e2e/.venv/bin/python \
scripts/workflow_canary/run_workflow_canary.py \
--skip-build --skip-python-bootstrap \
--scenario telegram_round_trip
→ "all 1 probe(s) passed"
(full suite without the flag still runs all 21 probes)
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(canary/workflow-canary): align nl_schedule_update on 'every 6 hours'
Addresses Copilot AI's review on PR #2874: the docstring claimed
"every 5 minutes" while EXPECTED_NEW_SCHEDULE / mock LLM emitted
"0 */5 * * *" (every 5 hours), and the chat prompt the canary sent
said "every 5 hours". Three different cadences across one probe.
Pick "every 6 hours" consistently:
- Docstring narrative: "every 6 hours"
- Constant: EXPECTED_NEW_SCHEDULE = "0 */6 * * *"
- Chat prompt: "fire every 6 hours"
- mock_llm.py routine_update args: schedule = "0 */6 * * *"
Verified locally: nl_schedule_update probe still green.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(canary/auth-browser-consent): drop stale GitHub secret exposure
Addresses @henrypark133's review on PR #2874: the auth-browser-consent
job kept exporting 8 GitHub-related secrets (GITHUB_OAUTH_CLIENT_ID,
GITHUB_OAUTH_CLIENT_SECRET, AUTH_BROWSER_GITHUB_OWNER / _REPO /
_ISSUE_NUMBER / _USERNAME / _PASSWORD / _STORAGE_STATE_B64) even
though the lane no longer drives a GitHub OAuth flow. BROWSER_CASES
in `scripts/live_canary/auth_registry.py` was reduced to {google,
notion} when github was reclassified as PAT-only — those secrets are
unused on every scheduled run and just broaden the secret-exposure
surface.
Strip all 8 from the lane:
- env: block — 5 lines (CLIENT_ID + 4 AUTH_BROWSER_GITHUB_* helpers)
- Materialize provider storage state — 1 secret + its materialize block
- Materialize sensitive secrets — 2 secrets + their write_secret lines
Replace with explanatory comments pointing at BROWSER_CASES /
auth_registry.py so a future contributor doesn't re-add them by reflex
when github gets an OAuth flow.
Github coverage continues to live in SEEDED_CASES (auth-live-seeded
lane) which seeds the PAT directly — that lane's secrets are
unaffected.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* docs(canary): align user-facing browser-cases list with auth_registry
Addresses @henrypark133's review on PR #2874: removing `github` from
BROWSER_CASES made `--mode browser --case github` invalid, but the
contract was still advertised in three places that operators read
when copying invocations:
- run_live_canary.py --help (`For browser mode: google, github, notion`)
- scripts/auth_live_canary/README.md (`github` listed under "Runs
through Responses API and browser")
- scripts/live-canary/README.md (`CASES=google,github` example)
- scripts/live-canary/ACCOUNTS.md (full GitHub OAuth client + fixture
+ storage-state-secret sections still active, plus a Playwright
storage-state recipe pointing at github.com/login)
Update each in lockstep:
- --help now says `For browser mode: google, notion. (github browser
coverage is intentionally absent — the github WASM tool is PAT-only,
not OAuth; see SEEDED_CASES instead.)`
- auth_live_canary/README — github entry now reads "Responses API
only (PAT-only — not browser-OAuth)"; notion entry corrected to
"Responses API and browser" (it was inaccurately listed as
Responses API only).
- live-canary/README — example flips to `CASES=google,notion` with a
one-line note pointing at auth_registry.py.
- live-canary/ACCOUNTS — drops the GitHub OAuth client + fixture
sections, swaps the Playwright storage-state recipe target from
github.com/login to accounts.google.com, drops
AUTH_BROWSER_GITHUB_STORAGE_STATE_B64 from the CI-secrets list.
The argparse validator in run_live_canary.py already gives a clean
error if anyone passes `--mode browser --case github`:
"--case values ['github'] are not valid for --mode browser. Allowed:
['google', 'notion']", so the docs change is the user-facing fix.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(canary/telegram): split is_active into installed vs. active
Addresses Copilot AI's review on PR #2874: `is_telegram_active` only
checked that an extension named "telegram" appeared in
`/api/extensions`, returning True for an installed-but-inactive
extension (mid-setup, awaiting auth, activation_error). Two callers
(`telegram_round_trip._ensure_active`,
`routine_visibility_from_telegram._ensure_active_and_paired`) used
this as a precheck to skip `setup_telegram_channel()`, so a stale
inactive entry would short-circuit setup and the probe would then
fail mysteriously when the channel didn't respond.
Split into two helpers:
- `is_telegram_installed(...)` — original semantics (entry exists),
used internally as a building block; not exported as a precheck.
- `wait_for_telegram_active(...)` — polls until the entry has
`active=true` (the actual runtime-readiness signal — channel
opened, hooks registered, credentials bound, per
`.claude/rules/lifecycle.md`'s discovery-vs-activation rule).
Shared `_find_telegram` helper handles the three historical envelope
shapes the gateway has used (`extensions` / `items` / `installed`).
Update all 4 callers to use `wait_for_telegram_active`:
- telegram_channel_install.py
- telegram_round_trip.py (precheck + post-setup wait)
- routine_visibility_from_telegram.py (precheck + post-setup wait)
- manual_trigger_from_telegram.py (precheck + post-setup wait)
Verified: all 4 telegram probes still green back-to-back.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* docs(canary/periodic_reminder): align docstring with current behavior
Addresses Copilot AI's review on PR #2874: the module docstring still
described the Telegram delivery assertion as a "Phase 1B follow-up"
even though the scenario now sets verify_telegram=True and the
inline comment on the call site already explained the Phase 1B work
had landed. Future readers would assume Telegram verification was
missing from this probe.
Replace the docstring with a 5-step description of what the probe
actually does end-to-end:
1. Backdated cron routine inserted via libSQL
2. Routine engine cron-tick picks it up
3. Lightweight action runs against mock LLM → http sendMessage
4. IRONCLAW_TEST_HTTP_REMAP routes to telegram_mock
5. Asserts both terminal routine_runs status AND captured sendMessage
Also adds an explicit note that channel-install coverage (capability
patch + setup + pairing) lives in the sibling telegram_* scenarios —
this one covers the routine-driven sendMessage path and intentionally
hits api.telegram.org via the raw http tool rather than through the
installed channel.
Verified: probe still green.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(canary-report): rich failure blocks + cross-lane categorization + GH issues
Three additions to scripts/live-canary/notify_slack.py to make the
6h Slack report actionable instead of just informational:
1) **Per-lane rich failure block** — Haiku now extracts four
structured fields when status==fail: test_name, error, root_cause,
fix. The Slack section renders them in the issue-friendly shape
the reviewer asked for:
:x: auth-full (mock) — 11/13 passed, 1 failed in 213s
Test: `test_wasm_tool_first_chat_auth_attempt_emits_auth_url`
Error: SSE stream closed; auth_required event never arrived
Root Cause: bridge gate not wired for installed-but-unauthed
extensions (#2868 fallout)
Fix: route Extension::NeedsAuth through effect_adapter.rs
For passing/skipped lanes the existing single-line `> reason` is
preserved so the green-path Slack output is unchanged.
2) **Cross-lane "Summary by Category" block** — second Haiku pass
over all failed-lane summaries that groups them by shared root
cause (e.g. "WASM tool dispatch regression — Auth Full, Auth
Smoke, Auth Live Seeded"). Only fires when there are 2+
failures (single-failure runs are already obvious from the
per-lane block). Rendered as a Slack mrkdwn bulleted list since
Block Kit doesn't support real tables.
3) **Auto-opened GitHub issues** — opt-in via CANARY_CREATE_ISSUES=1
env var (gated to scheduled runs only in live-canary.yml so
workflow_dispatch debugging doesn't flood the tracker). For each
failed lane:
- Search for an OPEN issue with title `[canary] <lane>: <test>`.
- If found: comment "another occurrence on <run_url>".
- If not found: open a new issue with the rich body + labels
`canary-failure` + `lane:<lane>`.
Strategy chosen to avoid issue spam while still surfacing
recurring failures. Uses GITHUB_TOKEN + the repo's existing
`permissions: issues: write` block — no new secrets.
All three additions degrade silently — Haiku failure stamps
.notable but doesn't block the post; categorization failure produces
an "_(unavailable)_" placeholder; issue-creation errors are logged
to stderr only. The notifier still exits 0 in every failure path so
a flaky webhook can't fail the canary run.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* ci(canary-report): reuse AUTH_LIVE_GITHUB_TOKEN for issue creation
Swap the issue-creation token source from the built-in
secrets.GITHUB_TOKEN to the existing AUTH_LIVE_GITHUB_TOKEN PAT —
no new secrets to mint, and that PAT already covers
nearai/ironclaw operations.
Set as CANARY_ISSUES_TOKEN (the highest-priority env var in
notify_slack.py's --github-token precedence chain) so it wins over
GH_TOKEN / GITHUB_TOKEN if any of those are also present.
Verify the PAT has `issues: write` scope (Issues: read & write for
fine-grained PATs, repo scope for classic PATs). If it doesn't, the
notifier still degrades gracefully — the API call fails, the error
is logged to stderr, the canary run isn't blocked.
Co-Authored-By: Claude Opu…
…earai#3234) The v2-engine E2E group still listed test_v2_kernel_auth_preflight.py, but that file was removed in nearai#2868 (engine-v2: callable-only available actions) and replaced with test_v2_tool_activate_surface.py for the new tool_activate / Activatable Integrations contract. The Web E2E Full job is skipped on PR-level CI but runs in the merge queue, so the bad path filter dequeued nearai#3197 and nearai#3203 with "file or directory not found: test_v2_kernel_auth_preflight.py".
…ange (nearai#3235) * ci(e2e): replace deleted preflight test with tool_activate surface The v2-engine E2E group still listed test_v2_kernel_auth_preflight.py, but that file was removed in nearai#2868 (engine-v2: callable-only available actions) and replaced with test_v2_tool_activate_surface.py for the new tool_activate / Activatable Integrations contract. The Web E2E Full job is skipped on PR-level CI but runs in the merge queue, so the bad path filter dequeued nearai#3197 and nearai#3203 with "file or directory not found: test_v2_kernel_auth_preflight.py". * test(e2e): unblock Live Canary auth lanes after engine-v2 contract change The Live Canary "Auth Smoke", "Auth Full", and "Auth Live Seeded" jobs have failed every scheduled run since 2026-05-01 (when the canary cut over to main). Three tests in test_v2_auth_oauth_matrix.py drive the failures, all rooted in the engine-v2 callable-only contract from nearai#2868 that didn't exist when these tests were written. ## What was broken `test_mcp_same_server_multi_user_via_browser` After OAuth completes, sending "check mock mcp search" through each user's browser opens an `approval` pending_gate on the first MCP tool call (engine v2 default). The browser fixture has no auto-approve UI, so the chat sat in `pending_gate` for the full 5-min Playwright timeout — `expected_text_contains="Mock MCP search result"` could never match because the assistant bubble never received any text. `test_wasm_tool_oauth_refresh_on_demand` Same shape: gmail call gates on `approval` before reaching the http credential-injection layer that performs the OAuth refresh. Without approving, refresh_count never went above 0, so the test failed with "Timed out waiting for OAuth refresh request". `test_wasm_tool_first_chat_auth_attempt_emits_auth_url` Tested OLD engine-v2 behavior — that an LLM-emitted call to a not-yet- authed extension would surface a `gate_required` Authentication event with an auth URL. After nearai#2868, the engine returns "action 'gmail' is not callable in this execution context" instead, and `tool_activate` became the model-facing enablement path. The mock LLM is canned to emit tool calls directly, so this scenario can't be reproduced from a scripted LLM until the canned response is updated. ## Fixes - `_wait_for_tool_call`: accept a `token` kwarg so multi-user tests can poll/approve through a per-user identity. Backwards-compatible. - `test_mcp_same_server_multi_user_via_browser`: drive approval through the per-user API while waiting for the tool to land. Drop the broken `expected_text_contains` predicate and the tied "Mock MCP search result" text assertions; the bearer-token isolation assertion (what this test actually exists to prove) is retained and unaffected. - `test_wasm_tool_oauth_refresh_on_demand`: insert a `_wait_for_tool_call` approval step between `_send_chat` and `_wait_for_refresh_request` so the http credential layer actually runs. - `test_wasm_tool_first_chat_auth_attempt_emits_auth_url`: marked xfail with an inline reason pointing at nearai#2868 and the replacement coverage (`test_v2_tool_activate_surface.py`, `test_settings_first_gmail_auth_then_chat_runs`). - Drop the now-unused `send_chat_and_wait_for_terminal_message` import. ## conftest fix `ironclaw_server` now sets `SECRETS_MASTER_KEY` in the spawned env. On macOS without it, `auto_generate_and_persist` blocks on a Keychain authorization prompt that no one's home to click, so `wait_for_ready` times out at 60s and the fixture kills the process with SIGKILL — making any session-scoped browser test impossible to run locally. On Linux, the keychain backend errors fast and the auto-generate fallback writes to `.env`, so CI was unaffected. Setting the key explicitly matches the pattern already used in `auth_matrix_server`, `test_v2_engine_auth_cancel`, `test_v2_tool_activate_surface`, etc. ## Verification Local repro confirmed each failure mode (HTTP-only repro for the non-browser tests, server-side log inspection for the multi-user test). Reproduced the exact pending_gate=approval pattern, fixed it, verified the assertion semantics still hold: ``` $ pytest tests/e2e/scenarios/test_v2_auth_oauth_matrix.py::test_wasm_tool_oauth_refresh_on_demand PASSED in 6.74s $ pytest tests/e2e/scenarios/test_v2_auth_oauth_matrix.py::test_wasm_tool_first_chat_auth_attempt_emits_auth_url XFAIL in 93s $ pytest tests/e2e/scenarios/test_v2_auth_oauth_matrix.py -v --timeout=120 12 passed, 2 skipped, 3 xfailed (browser tests errored locally; they'll run cleanly in CI) ``` The browser-driven `test_mcp_same_server_multi_user_via_browser` couldn't be exercised locally (chromium can't launch under this shell sandbox), but the API + auto-approve flow it now relies on is exercised by an HTTP-equivalent repro and matches the pattern used in `test_settings_first_gmail_auth_then_chat_runs`. * test(e2e): set LLM_API_KEY in auth_sse_server fixture `test_auth_required_sse_without_duplicate_response` was failing in the merge queue on every PR (most recently bouncing nearai#3197 and nearai#3203 from the queue) because the `auth_sse_server` fixture never set `LLM_API_KEY` in the spawned ironclaw env. After nearai#2572 added a missing- API-key check to the openai_compatible config validator (Apr 22), ironclaw rejected the env-supplied openai_compatible config, fell back to the NearAI default, hit "missing session token", and failed the turn before the github skill could even fire its 401. The chat thus reached `state: Failed` with no tool calls and no `onboarding_state/auth_required` event — which is exactly what the test asserted on, hence the consistent failure. Adding `LLM_API_KEY=mock-api-key` matches the value already used in every other e2e fixture (auth_matrix, conftest's ironclaw_server, v2_engine, etc.) and unblocks the assertion. Local run: PASSED in 8s. * fix(gateway): suppress duplicate assistant bubble after streamed response [skip-regression-check] The SSE `response` handler unconditionally called addMessage('assistant', data.content) even when stream_chunks had already populated and finalized a bubble for the same response. This stayed invisible in the common case but surfaced as a hard test failure under the path test_switching_back_preserves_in_progress_turn: 1. Send "What is 2+2?" on thread A — stream chunks start filling an assistant bubble with "data-streaming". 2. Switch to thread B mid-stream — container clears (history reload). 3. Switch back to thread A — history rehydration shows the in-progress turn with no response yet, so 0 assistant bubbles in DOM. 4. Stream chunks continue to fire for A — appendToLastAssistant creates a new bubble and accumulates the response into it. 5. response event fires — flushes any remaining buffer, removes the data-streaming flag (good) — then addMessage('assistant', content) creates a SECOND identical bubble. Result: locator(".message.assistant").filter(has_text="4") matches two elements, Playwright strict mode rejects the wait_for, the test fails. Outside the test, two identical bubbles render to the user. Fix: only call addMessage in the response handler when there was no in-flight streaming bubble. If one existed, the streamed content is already correct (chunks accumulate `data.content` verbatim) and the data-streaming flag has just been cleared. Non-streaming responses (no chunks fired) still take the addMessage branch. Regression coverage: tests/e2e/scenarios/test_message_persistence.py:: test_switching_back_preserves_in_progress_turn already reproduces this exact scenario and was failing in the merge queue. With this fix it passes; skip-regression-check used because the existing E2E test is the regression test, and the gateway doesn't have a JS unit test harness for SSE handler state. * fix(gateway): dedupe history-rendered SSE responses * test(e2e): set mock LLM API key in standalone fixtures * fix(e2e): make v2 approval tests deterministic * test(e2e): stabilize duplicate skill install assertion * test(e2e): assert duplicate install stays ungated * fix(skills): skip approval for disk-installed duplicates * fix(v2): honor no-op skill installs without approval * test(e2e): wait for pending send marker to clear --------- Co-authored-by: Firat Sertgoz <f@nuff.tech>
…arai#3157) * fix(engine): inline gate await for Tier 0 + Tier 1 Approval gates CodeAct scripts that hit a tool requiring approval surfaced as `RuntimeError: execution paused by gate 'approval'` inside the script instead of pausing for the user. The async tool-resolve path converted `EngineError::GatePaused` into a Python exception; the sync preflight path returned `need_approval` to the orchestrator, which on resume re-ran the LLM step and re-executed any non-idempotent earlier tool calls in the same script. Replace both with a host-supplied `GateController` that pauses the live execution in place. The Monty VM (Tier 1) and the Tier 0 batch loop both stay alive across the user's approval; on `Approved` the gated action re-executes (lease re-consumed, auto-approve installed before delivery so subsequent gates short-circuit); on `Denied` the script raises a typed `RuntimeError("user denied tool 'X': ...")` that the script can catch. Auth and External resume kinds keep the legacy thread re-entry path - their resolution installs new state (credentials, callback payloads) that only takes effect on the next run-through. Boot sweep invalidates `Approval`-kind `PendingGate` rows from a prior process so a stranded gate after restart fails fast instead of taking the legacy path and re-executing earlier mutations. Tests: 3 new regression tests in scripting.rs (approve / deny / no-controller fallback) and 4 in gate_controller.rs covering the resolution registry's one-shot, dropped-receiver, and unknown- request semantics. Pre-existing failures `call_id_preserved_when_no_lease` and `stop_thread_works` reproduce on staging without these changes. Design: docs/plans/2026-05-01-codeact-inline-gate-await.md. * test(engine): live regression for inline gate await with CodeAct Two integration tests in engine_v2_gate_integration.rs that exercise the inline gate-await flow end-to-end through `ThreadManager` → `ExecutionLoop` → orchestrator → CodeAct → `EffectExecutor` → `GateController`: 1. `codeact_inline_gate_await_resumes_user_reproducer` reproduces the exact reported bug shape: a CodeAct script issuing `await github_tool(action="search_issues_pull_requests", ...)` for "what are p1 bugs in nearai/ironclaw filed in last 7 days". The github_tool returns `EngineError::GatePaused` mid-execution; the test's `OneShotApprovingGateController` marks the effects mock approved and returns `Approved`; the engine retries inline; the tool succeeds; the script's `FINAL("Found 0 P1 bugs ...")` reaches the user. Asserts the controller saw exactly one pause request, github_tool was called twice, and the response is the script's FINAL — not the pre-fix `RuntimeError: execution paused by gate`. 2. `codeact_inline_gate_await_denial_does_not_retry` covers the deny path: controller returns `Denied { reason: "not now" }`. Asserts github_tool was called exactly once (no retry on denial), the typed `user denied tool 'github_tool': not now` message appears in the failure events, and the pre-fix `execution paused by gate` string does NOT appear anywhere. Also restored a fmt-only line shape in `gate_controller.rs` from `cargo fmt`. * test(engine): use realistic mock issues in inline gate-await fixture The pre-fix fixture returned `{"items": []}` which made the script's FINAL emit "Found 0 P1 bugs in nearai/ironclaw" — misleading, since the repo actually has open P1 issues (e.g. nearai#2818, nearai#2997). Update the mock to return two such items and tighten the assertion to verify the exact count flowed from tool result through CodeAct to FINAL(). * refactor(engine): require gate_controller, bound retry, drop V1 fallback Removes the `Option<Arc<dyn GateController>>` foot-gun: the field's `None` arm in the executors silently re-emitted the original `"execution paused by gate 'approval'"` RuntimeError, which is exactly the bug this PR exists to fix. With the field required and a named `CancellingGateController` as the explicit drop-in for non-pausing paths (post-resolution replay, mission protected writes, tests), forgetting to wire a controller is a compile error. Other follow-ups in the same change to keep them on one commit: - `MAX_INLINE_GATE_RETRIES = 3` shared by `scripting::drive_inline_gate` (Tier 1 async output) and `structured::execute_with_inline_gate_retry` (Tier 0 mid-execution). A misbehaving tool that keeps gating after each approval surfaces a clean error instead of pinning a CPU. - `denial_reason_for_resolution` helper centralizes the `GateResolution -> reason` mapping so denial messages can't drift between Tier 0 and Tier 1. - `invalidate_stranded_approval_gates_evicts_only_approval_kind` unit test covers the boot sweep with a mixed Approval/Auth/External population. - Existing `codeact_gate_without_controller_falls_back_to_runtime_error` test rewritten as `codeact_default_controller_cancels_approval_gates` to assert the inverted invariant: the legacy bug message must NEVER appear, even with the inert default controller. - Design doc updated to reflect as-shipped shape (required field, the bounded-retry constant, denial helper, `max_duration` stays at 30 s). * fix(engine): bound inline pause, propagate one-shot approval, race fix Addresses review on PR nearai#3157 (serrrfirat + Copilot + gemini-code-assist). Four blocking correctness fixes: 1. **Bounded BridgeGateController::pause.** The await on the resolution oneshot now races against `pending.expires_at`. Without this, a user ignoring the prompt past expiry would strand the engine: the DB row expires, the oneshot stays open, the VM keeps running. On expiry we discard the pending row, drop the registry entry, and return `Cancelled` so the VM unwinds cleanly. 2. **One-shot approval threaded through retry.** Add `call_approval_granted: bool` on `ThreadExecutionContext` (default false). Inline retry paths (`drive_inline_gate`, `execute_with_inline_gate_retry`, `execute_single_action_with_inline_retry`, scripting sync preflight) set it to true on the retry call so `EffectBridgeAdapter::execute_action` forwards it as `approval_already_granted=true` to the host's tool approval check. Mirrors the legacy `execute_resolved_pending_action` contract; without this, tools with `ApprovalRequirement::Always` gated again on every retry until the bound tripped, and `always=false` AskEachTime gates re-prompted immediately after approval. 3. **Per-execution context registration race fixed.** `set_execution_context` was called AFTER `handle_user_message().await` returned the thread_id — but the engine task is already running, so a fast tool gate could reach `pause()` before the entry existed and get `Cancelled`. Now the bridge calls `set_pre_execution_context` (per-user) BEFORE `handle_user_message`, then promotes to (user, thread)-keyed once thread_id is known. `pause()` falls back to the per-user entry on miss. 4. **Tier 0 parallel-batch coverage.** New `execute_single_action_with_inline_retry` wraps `execute_single_action` with the same bounded retry shape used by Tier 1. Both the single-runnable and multi-runnable branches of `handle_execute_actions_parallel` go through it, so simultaneous gates in a parallel batch no longer fall through to the legacy re-entry path. Smaller fixes: - PROJECTION lint annotations on the two `broadcast_for_user` sites added by this PR (gate_controller emit_gate_prompt; resolve_gate inline-await fast-path resolution event). Test coverage: - `GatingThenOkEffects` now records `context.call_approval_granted` per call; `codeact_gate_inline_await_approved_delivers_result` asserts the retry observes `true`. Locks in the one-shot approval propagation contract. * test(engine): update gate integration tests for inline-await semantics CI was failing on 5 tests in engine_v2_gate_integration.rs that asserted the legacy `ThreadOutcome::GatePaused` unwind for `Approval` gates. With inline-await + the required `CancellingGateController` (default), Approval gates resolve inline and the thread completes (or fails) instead of pausing — these tests were exercising the pre-PR flow that this PR replaces. - `gate_paused_transitions_thread_to_waiting` → renamed to `approval_gate_resolves_inline_via_controller`. Wires `AutoApprovingGateController`, asserts thread completes after inline approval, ApprovalRequested + ActionExecuted both recorded. - `gate_paused_thread_resumes_to_completion` → renamed to `approval_denied_inline_completes_thread_with_failed_action`. Asserts the default `CancellingGateController` cancels the gate and the thread completes with a failed action (no stranded pending gate). - `approval_chains_directly_into_auth_for_install_flow` rewritten: Approval handled inline by controller, Auth gate (still legacy path) bubbles up as ThreadOutcome::GatePaused, legacy auth-resume drives completion. - `approval_resolution_executes_pending_call_directly` and `gate_resume_with_execution_obligation` reduced to no-op stubs with inline rationale documenting where the post-PR equivalent coverage lives. Removing entirely would erase the breadcrumb in git log. Also threads the `ApprovalRequested` event through `execute_single_action_with_inline_retry`: the wrapper now returns `Vec<EventKind>` per call so the caller can append both the approval prompt and the post-retry outcome to the thread event log. Without this, observers saw only the final ActionExecuted/Failed event with no record that a gate had fired. Drops a few dead test helpers (`ApprovalTool`, `make_caps_with_approval_tool`, unused imports) that the deleted assertions no longer reference. Quality gate: fmt clean, clippy zero warnings, 513 engine lib + 443 bridge + 29 gate integration tests pass (2 pre-existing engine-lib staging failures unrelated to this PR). * test(engine): update skill_codeact integration tests for inline-await Two more tests in engine_v2_skill_codeact.rs were asserting the legacy `ThreadOutcome::GatePaused` → `resume_thread` flow for `Approval` gates. With PR nearai#3157 the engine catches Approval gates inline and the thread runs to completion in a single `join_thread`. - `skill_prompt_context_survives_pause_and_resume`: wires `AutoApprovingHttpController`, asserts thread completes after inline approval, drops the now-impossible `resume_thread` step. - `skill_prompt_context_survives_compaction_and_resume`: same; also drops the assertion on the persisted transcript at the *pause point* (no longer externally observable). The load-bearing post-compaction LLM-call assertions remain — they're the actual contract this test was protecting. Adds an `AutoApprovingHttpController` test helper local to this file that approves gates by marking the underlying `PausingHttpMockEffects` approved before returning `Approved`. * fix(engine): bridge cleanup on dropped sender + ActionFailed on preflight denial Addresses Copilot review on PR nearai#3157. - `BridgeGateController::pause`: when the resolution oneshot's sender is dropped (process shutdown / registry cleared), discard the `PendingGate` row before returning `Cancelled`. Without this, a stranded prompt remained visible in the UI and `pending_gates.insert` rejected duplicates for the same `(user, thread)` so a follow-up gate could not register. Same cleanup as the expiry branch already performed. - `scripting.rs` sync-preflight denial path: emit `EventKind::ActionFailed` before resuming Monty with `RuntimeError`, so the thread event log is consistent with the other denial paths (`drive_inline_gate`, `structured.rs`). Auditing why a tool didn't run is now possible from the events alone. - Delete the two empty `#[tokio::test]` stubs left in `engine_v2_gate_integration.rs` (`approval_resolution_executes_pending_call_directly_via_resolved_pending_action`, `gate_resume_with_execution_obligation`). The post-PR equivalent coverage is in `codeact_inline_gate_await_*` (this file) and the scripting unit tests; the rationale is preserved in `git log` (commit 87fe4fc) without an empty test slot misleading coverage signals. * fix(engine): finish inline-await migration; remove legacy gate-paused for Tier 0 policy gates Address PR nearai#3157 review: - router.rs: clear pre_execution slot on handle_user_message error so a failed engine spawn doesn't leak a stale (user, conversation) entry that would mis-route the next gate prompt. - router.rs:await_thread_outcome: on the 5-min request deadline, return BridgeOutcome::Pending instead of join_thread() — joining would block for up to the gate's 30-min expires_at when the parked task is in pause(). The PendingGate row stays live for resolution. - gate_controller.rs: re-key pre_execution by (user_id, conversation_id) instead of user_id alone. Two concurrent conversations / browser tabs for the same user no longer clobber each other's slot. Plumbs conversation_id through ThreadExecutionContext and GatePauseRequest so pause() can match a gate to its originating conversation. - gate_controller.rs: serialize concurrent inline gates per (user, thread) via a per-key tokio Mutex held across the PendingGateStore::insert + select-await window. A parallel batch where two tools both gate now queues the second behind the first rather than silently surfacing it as Cancelled on (user, thread) uniqueness collision. - orchestrator.rs (Tier 1 / CodeAct): remove the legacy gate_paused JSON sentinel + thread re-entry path for Approval gates. Both __execute_action__ and __execute_actions_parallel__ now pause inline on PolicyDecision::RequireApproval (mirroring structured.rs) and route tool-raised gates through the existing execute_single_action_with_inline_retry wrapper. Authentication and External resume kinds keep the legacy re-entry path because their resolution installs new state that only takes effect on the next thread run-through. Test fixtures across bridge/effect_adapter, action_projector, and the gate/sandbox integration suites get the new conversation_id field defaulted to None. Refs: comments 3173757791, 3176248954, 3176248976, 3176249001, 3176249032 * fix(engine): defer post-Pending context cleanup; route parallel JoinSet through inline-retry Address PR nearai#3157 review on commit d211bfc: - router.rs: when await_thread_outcome returns BridgeOutcome::Pending and the engine task is still running (typically parked in BridgeGateController::pause), defer clear_execution_context to a spawned watcher task that polls is_running until the thread completes and then clears the (user, thread) context + gate-locks entry. Without this, the unconditional clear after a 5-min request timeout stranded the parked thread: the eventual gate resolution would call pause() for any subsequent gate with no registered context and surface as silent Cancelled. Watcher caps at 60 minutes (well past the 30-min PendingGate expiry) as a defensive safety bound. - structured.rs: route the multi-runnable JoinSet branch in execute_action_calls through execute_with_inline_gate_retry, matching the single-runnable fast path. Without this, an Approval gate raised mid-execution in a parallel batch with >1 runnable tool call bubbled out as Err(GatePaused) and went through the legacy gate-paused / re-entry path, re-introducing the double-execution bug for already-completed sibling calls in the same batch. The signature of execute_action_calls switches leases from &LeaseManager to &Arc<LeaseManager> so the JoinSet tasks can clone an owned handle; all current callers are tests already constructing Arc::new(LeaseManager::new()), so this is a no-op call-site change. Refs: comments 3176333657, 3176333685 * test(engine): fix call_id_preserved_when_no_lease MockEffects inventory PR nearai#2868 added a callable-inventory check ahead of the lease lookup in execute_action_calls. The test was constructing MockEffects with an empty action list, so preflight short-circuited on "action is not callable in this execution context" before reaching the lease branch the test was actually trying to exercise. The assertion error.contains(\"no lease\") then failed against the not-callable message. Populate the mock with test_action(\"web_search\") so the inventory gate passes and the call reaches the lease check, matching the test's documented intent. * fix(engine): use gate-provided params in inline-await pause; drop Thread clone in parallel branches - orchestrator.rs::execute_single_action_with_inline_retry now sources GatePauseRequest.parameters from the gate-paused payload (which reflects safety-layer transformations/redactions) rather than the original caller params, matching structured.rs::execute_with_inline_gate_retry. - Both inline-retry helpers now take ThreadId + user_id instead of &Thread, so the parallel JoinSet branches no longer clone the full Thread (with message/event transcripts) per spawned task. The per-task ThreadExecutionContext clone is what the helpers actually need; in orchestrator.rs we build a base ctx once and override current_call_id per task instead of re-running thread_execution_context. Addresses Copilot PR nearai#3157 review comments 3176594866, 3176594886, 3176594898. * fix(engine): address serrrfirat review — audit event + stop-during-wait Three latest serrrfirat review threads on PR nearai#3157: * Medium: structured inline approval drops ApprovalRequested audit event (REAL — fixed). `execute_with_inline_gate_retry` previously swallowed `Err(GatePaused)` and returned only the post-retry outcome, so `classify_exec_result` never saw the gate and the `ApprovalRequested` event was lost. The orchestrator (Tier 1) path emits both events. Mirror that shape here: - `execute_with_inline_gate_retry` now returns `(Result<ActionResult, EngineError>, Vec<EventKind>)`. The Vec carries one `ApprovalRequested` per retry iteration that gated. - Slot type widens to `(ActionResult, EventKind, Vec<EventKind>)` so the merge phase flushes pre-terminal events before the terminal event in original-call order. - Two regression tests pin the contract (approved → Executed, denied → Failed); both assert ApprovalRequested precedes the terminal event. * High: stop/cancel does not unblock thread parked in inline gate await (REAL — fixed). `BridgeGateController::pause()` selected only on the resolution oneshot + 30-min expiry; `stop_thread()` sent `ThreadSignal::Stop` but the parked engine task wasn't polling the signal channel. - Adds `GateController::cancel_thread(thread_id)` to the engine-side trait (default no-op). - `ThreadManager::stop_thread()` calls `cancel_thread()` BEFORE sending `ThreadSignal::Stop` so any parked `pause()` future wakes promptly with `GateResolution::Cancelled`. - `BridgeGateController` tracks in-flight pauses in `active_pauses: HashMap<ThreadId, HashSet<Uuid>>` and walks that set on cancel, delivering `Cancelled` via the existing `GateResolutions::try_deliver` channel and discarding pending rows via `PendingGateStore::discard_for_thread`. - `pause()` always untrack on exit (idempotent — `cancel_thread` may have already removed the entry). - Two new bridge-tier unit tests: parked pause wakes within 2s of cancel_thread; cancel_thread on a thread with no parked pause is a no-op. * Medium: live inline gate waits are uncapped per user/process (PARTIAL — push back, follow-up). The reviewer's own comment notes this is a design-doc follow-up, not a blocker. The current bound is implicit: one pending gate per (user, thread) × the existing thread-creation budget × the 30-min expiry, which is enough to ship the inline-await substrate. Adding a per-user semaphore correctly requires designing the cap UX, fairness, and rejection error shape — out of scope for this PR. Add a TODO inside `pause()` pointing at the design doc and the follow-up issue, with the reasoning written down so the next contributor doesn't have to re-derive it. Pre-existing on this branch (NOT introduced by these changes): `runtime::manager::tests::stop_thread_works` flakes on the branch; the original PR commit message acknowledges "stop_thread_works reproduce on staging without these changes." Confirmed by stashing this commit and running the test: still fails. Out of scope here. Test totals after this commit: - executor::structured: 19 passing (incl. 2 new) - bridge::gate_controller: 6 passing (incl. 2 new) - cargo fmt + clippy --lib clean Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(engine): typed DenialOutcome + plug exhaustion-path lease leak Three review-driven fixes to the inline gate-await path. 1. CancellingGateController and bridge expiry/shutdown surfaced as `RuntimeError("user denied tool 'X': cancelled")` — misleading because the user never saw a prompt. Replace `Option<String>` helper with a typed `DenialOutcome { DeniedByUser, Unavailable }` so cancelled/expired/no-handler gates render as "approval for tool 'X' unavailable: …" while real user denials keep the "user denied" framing. Updated all six call sites in scripting.rs / structured.rs / orchestrator.rs through the typed surface so wording can't drift. 2. Both `execute_with_inline_gate_retry` and the orchestrator's `execute_single_action_with_inline_retry` consumed a fresh lease use on the final approved iteration, then exited the loop without ever calling execute_action. The exhaustion-path error didn't trip the caller's refund check, so a misbehaving tool that gates after every approval would slowly drain `max_uses`. Refund the unused lease before returning. 3. Drop the dead `parameters: serde_json::Value` field on `PendingFuture::Tool` and the matching `_parameters` arg on `resolve_tool_future`. The gate's own parameter snapshot is the source of truth on retry; threading the original through made the signature read like there was a second source. Quality gate: cargo fmt, cargo clippy --all --benches --tests --examples --all-features (clean), cargo test -p ironclaw_engine --lib (520 pass), cargo test --test engine_v2_gate_integration (27 pass), cargo test --lib bridge:: (456 pass). --------- Co-authored-by: Nikolay Pismenkov <nickpismenkov@gmail.com> Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(e2e): restore auth and approval coverage (#3430)
* test(e2e): avoid REPL auth retry race (#3437)
* feat: add pairing_approve tool for Slack binding via chat (#3396)
* feat: add pairing_approve tool for Slack binding via chat
Users can now paste their Slack pairing code in the IronClaw chat and
the LLM will call pairing_approve to bind their accounts. No need to
use the API directly.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: require approval before pairing + fix formatting
Address review comment: pairing_approve now requires UnlessAutoApproved
approval before executing, preventing accidental account binding.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* test: add regression test for pairing_approve tool
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: address review — Always approval, lock channel, add to protected list
1. Changed ApprovalRequirement to Always (not bypassable by auto-approve)
2. Locked channel to slack-relay constant (removed generic channel param)
3. Added pairing_approve to PROTECTED_TOOL_NAMES
4. Added tests: always-approval, protected-name, channel constant
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(web): isolate cross-tenant SSE/WS status events and thread access (#3390)
* fix(web): isolate cross-tenant SSE/WS status events and thread access
Plug a multi-tenant leak where unscoped `sse.broadcast(...)` calls
from `GatewayChannel::send_status`, sandbox `JobEvent` dispatch, the
WASM/Slack OAuth completion handlers, and any producer that lost
`metadata.user_id` along the way fan out to every connected
subscriber — exposing another tenant's tool calls, tool output,
onboarding state, and job lifecycle to anyone with an open SSE/WS.
Changes
- Extract `dispatch_status_event(sse, multi_tenant_mode, user_id, ev)`
from `Channel::send_status`. In multi-tenant mode an unscoped event
is dropped (with a WARN naming the producer to fix); single-tenant
keeps the global broadcast since there is one subscriber population.
- `IncomingMessage::new` now defaults `metadata` to `{"user_id": ...}`,
and `with_metadata` preserves the key so downstream `send_status`
consumers always have an owner to scope by.
- WASM/Slack OAuth completion broadcasts route through
`broadcast_for_user(&owner_id, ...)`. Sandbox `JobEvent` dispatch
in `main.rs` respects `multi_tenant_mode` for the empty-`user_id`
fallback.
- New pre-commit check #10 (`MULTITENANT`) flags unscoped
`sse.broadcast(...)` lines without a `// multi-tenant-safe: <reason>`
marker or a transport-only exemption. Marker regex accepts the marker
anywhere in a `//` comment so compound annotations on a single line
work.
Tests
- `src/channels/web/tests/status_event_isolation.rs` — 5 unit tests
covering both modes and the per-variant drop invariant.
- `src/channels/web/platform/sse.rs` — 2 quadrant tests for the
`subscribe_raw` filter (scoped/unscoped × matching/mismatched).
- `tests/thread_isolation_integration.rs` — 9 HTTP-level checks that
Bob cannot reach Alice's chat history (paginated and not), threads
list, engine v2 detail/steps/events, or Responses GET, plus an
unauthenticated-rejection guard.
- 8 new self-test cases for the `MULTITENANT` script check, including
the compound projection-exempt + multi-tenant-safe annotation case.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(web): pin cross-tenant boundaries on jobs, files, routines
Audit of the protected route surface for the same bug shape #3390 fixed
(handler that takes a user-controlled id and reads without an ownership
predicate) found that the implementations were correct but four
boundaries had no integration test. Lock them in before they regress.
- Sandbox job persisted-events history (`/api/jobs/{id}/events`):
Bob → 404 on Alice's job; Alice → 200 on her own.
- Sandbox job workspace listing (`/api/jobs/{id}/files/list`):
Bob → 404 on Alice's job.
- Sandbox job file read (`/api/jobs/{id}/files/read`): Bob → 404 on
Alice's job; Alice → 200 on her own; Alice → 403/404 on
`?path=../outside.txt` (path-traversal pin against the
`canonicalize() + starts_with(base_canonical)` guard).
- Routine run history (`/api/routines/{id}/runs`): Bob → 404; Alice
→ 200 with at least one seeded run.
The OAuth-state and NEAR-nonce stores were also flagged in the audit
but neither is a real cross-tenant bug: both are pre-auth, single-use,
and the token IS the secret. Documenting here so a future audit
doesn't re-flag them.
Project-static (`/projects/{id}/...`) is left for a follow-up — it
relies on `ironclaw_base_dir()` which is a process-wide `LazyLock`,
making per-test override fragile in the integration runner.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(web): address PR #3390 review — forge-resistant metadata, OAuth toast routing
Addresses six review comments from gemini-code-assist, copilot, and
serrrfirat on PR #3390. False-positives and the perf nit on
`with_metadata` are explained in the reply thread, not changed in code.
- (HIGH, serrrfirat) `IncomingMessage::with_metadata` now ALWAYS sets
`metadata.user_id` from `self.user_id`, dropping any caller-supplied
value. A WASM channel emitting `{"user_id":"victim"}` via
`apply_emitted_metadata` can no longer reroute downstream
`ToolStarted` / `ToolResult` SSE events into another tenant's
stream. New unit tests pin the forgery-resistance invariant.
- (MEDIUM, copilot + serrrfirat) Slack relay OAuth callback now
broadcasts the completion toast to the resolved `oauth_user`
(the IronClaw user who initiated the flow) rather than
`state.owner_id`. In multi-tenant deployments those differ and the
previous routing delivered the toast to the wrong browser tab.
Extracted the lookup into `resolve_relay_oauth_user`; two unit
tests cover the secret-present and secret-missing cases.
- (MEDIUM, gemini) `dispatch_status_event` treats empty-string
`user_id` the same as `None` so producers that lost the field
along the way fail-closed instead of falling through to a global
broadcast in multi-tenant mode.
- (MEDIUM, gemini) `main.rs` sandbox JobEvent dispatch now reuses
`dispatch_status_event` instead of duplicating the drop / WARN /
broadcast policy. `dispatch_status_event` is bumped from
`pub(crate)` to `pub` so the binary crate can call it.
- (LOW, copilot) Pre-commit `MULTITENANT` self-tests gain three
cases (`state.sse.broadcast(`, `gw_state.sse.broadcast(`,
annotated receiver-prefixed) to lock the existing boundary regex
behaviour against future tightening.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(channels): preserve i64 metadata.user_id from Telegram in with_metadata
PR #3390's forge-resistance fix made `IncomingMessage::with_metadata`
*always* overwrite `metadata.user_id` with `self.user_id` as a String.
That broke the Telegram WASM channel: it persists Telegram's chat user
ID as `metadata.user_id: i64` and re-deserializes it into
`TelegramMessageMetadata { user_id: i64, ... }` in `on_respond` /
`on_status`. After the fix, `respond` blew up with
`invalid type: string "999", expected i64 at line 1 column 87`,
failing 3 Telegram integration tests in CI.
Narrow the carve-out: overwrite only when the existing `user_id` is a
String (or missing). Non-string values are channel-private and the
SSE routing layer reads via `as_str()` — non-strings already fail
closed in multi-tenant mode, so the forge threat (WASM emits
`{"user_id":"victim"}` as a string) is still mitigated, while
Telegram's i64 use case survives.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(web): redact WARN payload + harden dotdot traversal test (PR #3390)
Two follow-up fixes from the second pass of review on #3390:
1. `dispatch_status_event`'s WARN log used `?event`, which on
`AppEvent::Response` / `Thinking` / `ToolResult` carries
user-authored content into operator logs in a multi-tenant
deployment. Replace with `event_kind = event.event_type()`
(the wire-stable variant name) — enough to identify the
misbehaving producer without leaking tenant data. Picked up via
Copilot's review on `src/channels/web/mod.rs:666`.
2. `alice_job_file_read_rejects_dotdot_traversal` planted
`outside.txt` under `outer.path()` (the `start_server_with_db`
fixture's tempdir holding `test.db`) but probed
`?path=../outside.txt` relative to `alice_proj` — a separate
`tempfile::tempdir()` rooted at the OS temp directory. The two
paths were unrelated, so the test could pass even if `..`
traversal was permitted (probe just hit empty space). Build the
directory tree by hand instead: `parent/alice_proj/` with the
planted file at `parent/outside.txt`, so the probe deterministically
resolves to the planted bytes. Add a body-content assertion that
fails loudly if those bytes leak. Picked up via Copilot's review on
`tests/cross_tenant_resource_isolation.rs:355`.
Plus a `cargo fmt` fix for `src/channels/channel.rs:1202` that was
breaking the Formatting CI check.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(web): address PR #3390 follow-ups — multi-tenant fallback WARN, exhaustive variant pin
- `resolve_relay_oauth_user`: take `multi_tenant_mode`; emit WARN when the
`relay:{ext}:oauth_user` secret is missing in multi-tenant mode so the
unrecoverable-initiator case surfaces in operator logs. Single-tenant
fallback stays silent (owner == only user).
- `dispatch_status_event`: doc note clarifying the function is `pub` only
for the sandbox JobEvent rx loop in `main.rs`; not part of a stable
public API.
- `_compile_time_appevent_variant_check`: exhaustive-match helper paired
with `unscoped_drop_holds_for_every_status_variant_in_multi_tenant`.
Adding a new `AppEvent` variant now fails the test build, prompting an
update to both the helper and the runtime leak-candidate list.
- `tests/thread_isolation_integration.rs`: honest scope note on the
engine-v2 thread tests — they pin handler shape (404/empty for
unknown id), not the cross-tenant ownership branch. Cross-tenant
engine-v2 coverage requires an `ENGINE_STATE` test fixture; tracked
as a follow-up in the comment block.
Tests: 8 unit (5 status_event_isolation + 3 resolve_relay_oauth_user)
and 9 thread_isolation_integration pass; clippy clean; pre-commit
safety scripts pass (regression suite 27 cases).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(web): unconditionally consume relay:{ext}:oauth_user secret in OAuth callback
The previous cleanup site lived inside the `if let Some(pairing_store)`
branch of the result block, which was unreachable on three failure
paths:
1. `pairing_store` is None (no identity pairing wired up)
2. The result block `?`-short-circuits before reaching the `if let`
(e.g. `set_setting` fails, `activate_stored_relay` fails, an inner
`relay_config()` / `list_connections()` errors)
3. The `if let` body itself errors before reaching the delete (e.g.
`list_connections` returns no matching team)
Leaving the secret behind lets a subsequent OAuth callback for the
same extension read a stale initiating user and misroute the
completion toast — Copilot review on PR #3390 (comment id 3211833864).
Move the delete to right after `resolve_relay_oauth_user` returns,
where it always runs once the value has been captured, regardless of
downstream failure mode. Updated the inner comment to document that
the secret is already gone by the time the pairing branch reads
`oauth_user`.
Regression test: `test_relay_oauth_callback_consumes_oauth_user_secret_on_failure_path`
seeds the secret, fires the callback against a fixture with no real
relay backend (so the result block deterministically errors), and
asserts the secret is gone afterward. Pre-fix this would have left
the secret behind on the failure path.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(auth): tighten Telegram pairing UX and OAuth-failure recovery (#3317, #3319, #3320) (#3381)
* fix(auth): tighten Telegram pairing UX and OAuth-failure recovery (#3317, #3319, #3320)
Three Bug Bash P1 issues from the same user journey: setup → use → fail.
The unifying root cause was per-channel auth tested in isolation; cross-channel
flows (Telegram → Gmail OAuth → resume) had no coverage and three small leaks
combined into a stuck conversation.
#3317 — Telegram pairing reply now names every IronClaw surface explicitly
(web settings, agent chat, terminal). The agent submission parser learns
`approve <channel> <code>`, dispatched through a new bridge handler that
mirrors `POST /api/pairing/{channel}/approve`.
#3319 — OAuth callback failures now log a category + correlation ID so a
user-reported "I saw 400" maps to one log line. Adds the
`OauthCallbackFailure` enum and `oauth_failure_correlation_id` helper.
#3320 — Two cleanup gaps fixed: (a) `/clear` now drains
`pending_oauth_flows` for the user (otherwise stale flows linger 5min and
mask new auth attempts); (b) OAuth provider-error and exchange-failure paths
now auto-cancel the engine pending auth gate via `clear_engine_pending_auth`,
so the conversation isn't blocked waiting for a resume that will never arrive.
Tests: 5 new submission-parser tests, 2 new bridge-handler tests, and one
new OAuth callback test verifying the pending-flow drain on provider error.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(auth): cross-channel pairing claim coverage + canary lane (#3317)
Adds the structural coverage that was missing when #3317 shipped:
- E2E (`tests/e2e/scenarios/test_telegram_pairing_chat_claim.py`):
three scenarios that drive the full Telegram pairing flow through
the gateway. Asserts the bot reply names every IronClaw surface
(web Settings, agent chat, terminal CLI), drives `approve telegram
CODE` through `/api/chat/send` and verifies the paired user
exchanges messages without re-prompting, and confirms invalid
codes get a clear rejection instead of an LLM-improvised reply.
- Rust integration (`tests/telegram_pairing_chat_claim_integration.rs`):
drives `Submission::PairingClaim` through a real `Agent` →
`bridge::handle_pairing_claim` → `PairingStore::approve` chain
using `TestRig` with engine v2 enabled. Covers the happy path
(mints a code, claims it via chat, asserts `Pairing approved`)
and the invalid-code rejection. The unit tests in `bridge/router`
cover only the no-extension-manager and invalid-channel branches —
this test exercises the wiring between submission parser, agent
loop dispatch, and bridge handler that #3317 specifically broke.
- Canary (`scripts/live_canary/auth_registry.py`): adds the two
user-visible scenarios to `AUTH_CHANNEL_TESTS` so the auth-channels
lane (scheduled every 6h) catches the regression class in CI
before any real user encounters it.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* docs: align pairing-claim and oauth-correlation comments with code (#3381)
Three Copilot review comments on PR #3381 flagged docstring/code drift in
already-merged PR #3317/#3319/#3320 changes. No behavior change — only
the doc strings move:
- `Submission::PairingClaim.code` and the inline `approve <channel> <code>`
parser comment claimed the user's casing was preserved, but the parser
builds the code from `lower` and the regression tests already lock in
the lowercased shape (`code == "abc12345"`). Update both comments to
describe the actual normalize-then-store contract.
- `oauth_failure_correlation_id` claimed the correlation appeared in the
user-facing error subtitle, but the failure path renders
`landing_html(label, false)` whose subtitle is fixed and never receives
the correlation. Mark the helper as logs-only and note that plumbing
the ID through the HTML is a follow-up.
[skip-regression-check] doc-only, behavior already covered by existing
pairing-claim parser tests in submission.rs.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* chore(telegram): bump channel registry to 0.2.11
The pairing-reply wording was updated in channels-src/telegram/, which
the version-check CI requires be matched by a registry version bump.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(telegram): fix import path in pairing chat claim e2e test
The scenarios/ folder is a Python package (has __init__.py), so a
flat `from test_telegram_e2e import …` fails with
ModuleNotFoundError during pytest collection. Switch to a relative
import that matches the package layout, and drop the unused
OWNER_USER_ID symbol while we're here.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(auth): address Copilot review on /clear OAuth drain + correlation doc
Two follow-ups on PR #3381's Copilot pass:
1. `Agent::process_clear` (engine v1 path) now drains in-flight OAuth
flows for the clearing user, mirroring the engine-v2 cleanup added in
`bridge::router::clear_engine_conversation`. Without this, `/clear`
was a clean slate on v2 but v1 left ghost flows in
`extension_manager.pending_oauth_flows()` until the 5-minute
`OAUTH_FLOW_EXPIRY` ticked over — same regression class #3320 fixed
on v2.
2. `oauth_failure_correlation_id`'s docstring previously said "redacted
state fingerprint", but callers seed it with the raw `state` query
value (or `flow.extension_name` for post-resolution failures).
Updated the doc to describe the actual behaviour: an arbitrary seed
that is hashed before any hex output, with a pointer to
`redact_oauth_state_for_logs` for the log-safe fingerprint.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(e2e): make Telegram pairing chat-claim suite actually run
The scenario landed in PR #3317 was orphaned — never wired into any CI
lane and could not pass even when run by hand. Three structural
issues, all fixed here:
1. `install_telegram` now overlays the locally-built WASM (and matching
capabilities file) on top of the registry-downloaded artifact when
present. The pairing-reply wording lives inside the WASM binary, so
without this overlay the test was asserting source-tree text against
the previous release's bytes. The overlay is best-effort: when the
local WASM is absent (CI groups that don't build the channel), the
test that depends on it skips with a clear message and the canary
lane in `scripts/live_canary/auth_registry.py` still covers the
wording end-to-end against the deployed binary.
2. `Submission::PairingClaim` is handled out-of-band by the bridge
layer; the response is delivered via `WebChannel::respond` →
`AppEvent::Response` over SSE only — no `Turn` is persisted, so
polling `/api/chat/history` could never see it. Refactored
`test_chat_surface_approves_pairing_code` and
`test_chat_surface_rejects_invalid_pairing_code` onto a
`_send_and_collect_response` helper that opens the SSE stream first
(so the broadcast doesn't fan out to zero subscribers) and matches
on the `response` event for the test thread.
3. Wired the file into `e2e.yml`'s `extensions` group so the suite
actually runs on every PR.
Verified locally: all three scenarios pass, plus the existing 22
Telegram e2e tests still green with the install-overlay change.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(bridge): bound and sanitize invalid-channel echo in pairing claim
Address Copilot review on PR #3381: `handle_pairing_claim`'s
invalid-channel branch was rendering the raw `channel` token (and the
underlying `IdentityError`, which itself echoes the offending input)
back to the user. Both routes are unbounded and could carry control
characters or markup since `channel` comes from chat input — a
hostile prompt could blow up the SSE / Telegram / TUI reply or smuggle
backticks/escape sequences through.
Cap the echo at 32 ASCII-alphanumeric (or `-`/`_`) characters and
replace the verbatim error with a fixed category description, so the
reply size and shape are bounded by what we render explicitly. Add a
regression test that drives a 200-char hostile blob (control chars +
backticks) through the handler and asserts the rendered reply stays
clean and short.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(bridge): include hyphens in invalid-channel error copy
Address Copilot review on PR #3381: the invalid-channel reply
listed "lowercase letters, digits, or underscores" as the valid
character set, but `ExtensionName::new` (and `web::features::pairing::
parse_channel`) intentionally accept hyphens too — they're folded to
underscores during canonicalization. A user typing `slack-relay` would
otherwise get an "invalid name" reply listing rules that contradict
the actual validator.
Updated the message to include hyphens with `telegram` and
`slack-relay` as concrete examples, and tightened the regression test
to assert against the user-controlled preview region between the
delimiter backticks rather than a global backtick count (which was
fragile to copy that includes example slugs in backticks).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(auth): address PR #3381 review on credential-scoped gate cleanup and Telegram surface promise
Three reviewer findings, one commit:
- OAuth provider-error and exchange-failure paths used
`clear_engine_pending_auth(user, None)`, which discards every
Authentication gate for the user. A failed Gmail callback could
silently wipe an unrelated Slack/MCP gate waiting on a different
thread. New `clear_engine_pending_auth_for_credential(user, credential)`
helper in bridge::router scopes cleanup to the failed flow.
Provider-error path tracks `removed_secret_name` alongside
`removed_user_id` so the scoped variant is callable.
- Expired-flow branch in the OAuth callback handler had two bugs: it
never cleared the engine pending auth gate (so the conversation sat
blocked forever, same #3320 class the provider-error fix addresses),
and the broader `clear_auth_mode` it called would re-discard via
the unscoped helper anyway. Now calls the credential-scoped helper
and the legacy-v1-only `clear_session_auth_mode_for_thread`.
- Telegram pairing reply advertised `approve telegram CODE` as
usable "in any IronClaw chat (TUI / web / Telegram)", but an
unpaired Telegram DM is intercepted by the allowlist gate before
the agent parser sees the command — the user would just get
another pairing reply. Reply now lists only the surfaces that
actually work (web / TUI / CLI) and a comment explains why.
Regression coverage:
- `clear_engine_pending_auth_for_credential_only_clears_matching_credential`
locks in helper scoping (Gmail/Slack two-gate scenario).
- `oauth_callback_expired_flow_clears_credential_scoped_engine_gate`
drives the full callback through axum oneshot with engine state
seeded; asserts the matching gate clears and the unrelated gate
survives.
- E2E `test_telegram_dm_approve_command_is_intercepted_by_allowlist_gate`
exercises the Telegram webhook path (not /api/chat/send) to lock in
the channel-layer interception, wired into the auth canary lane.
- Existing E2E pairing-reply test gains an assertion that
"TUI / web / Telegram" is *not* in the reply.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(oauth): mirror failure-cleanup contract on provider-error and reconcile stale comment
Copilot review on PR #3381 caught two real issues in the credential-scoped cleanup
landed in d45cf1bd8:
- Provider-error branch (`?error=access_denied`) returned the error page
without broadcasting `OnboardingState::Failed` or clearing the legacy v1
session `pending_auth`. The exchange-failure and expiry branches do both.
Net effect: the auth card stayed spinning and the next user message was
intercepted as a token. Now mirrors the other failure paths — keep the
full `flow`, emit Failed SSE, clear v1 session, clear credential-scoped
engine gate, then return the error page.
- Post-exchange comment said "failed callbacks should leave the gate
visible for retry" — that was the pre-#3320 contract. Rewrote it to
describe the new shape: each failure mode clears its own gate at the
failure site; this section only handles legacy-v1 session cleanup that
runs regardless of outcome. Also explains why we use
`clear_session_auth_mode_for_thread` here instead of `clear_auth_mode`
(the latter would re-clear the engine gate on the *success* path and
break the `ExternalCallback` resume).
Regression: `test_oauth_callback_provider_error_broadcasts_onboarding_failed`
in the oauth tests module — drives an `?error=access_denied` callback with
a flow whose `sse_manager` is attached, asserts the receiver gets
`OnboardingState::Failed` with the provider's `error_description` as the
message body.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(common): describe paths and platform helpers in crate description (#3498)
Align the package `description` and the lib.rs crate-level doc with
the modules now exposed from `ironclaw_common` (paths, platform,
env_helpers, attachment), which #3387 lifted out of `src/`. The
previous wording predates that extraction and only mentioned "types
and utilities".
This is also the release-plumbing trigger for v0.28.1: release-plz
proposes a leaf bump on source-path changes, and once `ironclaw_common`
crosses to a new patch, the root `ironclaw` bump can be added on top
of the release-plz branch (same approach as commit 070cbede1 for
v0.28.0). See PR #3372 for the equivalent v0.28.0 trigger.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* chore: release
* chore(release): bump ironclaw to 0.28.1
cargo-semver-checks did not detect API-breaking changes in the
ironclaw_common 0.4.1 -> 0.4.2 leaf bump, so release-plz did not
cascade a bump into the root ironclaw package. Add the root version
bump and CHANGELOG entry manually so this release-plz PR produces
an ironclaw-v0.28.1 tag and triggers cargo-dist.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor(llm): hide provider-specific auth, model fetch, and embeddings config behind facades (#3416)
* refactor(llm): hide provider-specific auth, model fetch, and embeddings config behind facades
External callers were reaching into provider-specific modules of
`ironclaw_llm` (`gemini_oauth::CredentialManager`,
`github_copilot_auth::*`, `OpenAiCodexSessionManager`,
`codex_auth::*`, `BedrockConfig`, etc.). Closes those leaks behind a
small set of verb-based public surfaces while keeping per-provider
behaviour inside the LLM crate.
Changes:
1. Extract `oauth_helpers.rs` into a new `ironclaw_oauth` crate. The
loopback OAuth callback listener (port 9876, landing pages,
`OAUTH_CALLBACK_HOST` rules) is shared by every IronClaw OAuth flow
(NEAR AI session login, WASM tool auth, MCP) and never depended on
`ironclaw_llm`. `src/auth/oauth.rs` now `pub use ironclaw_oauth::*`
directly. `ironclaw_llm` no longer depends on `ironclaw_oauth` —
the helper had zero internal callers.
2. Add `ironclaw_llm::auth` facade (`start_login`, `validate_token`,
`default_headers`, `load_persisted_credentials`,
`default_credentials_path`) with backend-agnostic types
(`AuthPrompt`, `LoginRequest`, `AuthOutcome`, `PersistedCredentials`,
`OpenAiCodexLoginOptions`, `AuthBackend`, `CredentialSource`).
Privatize `gemini_oauth`, `github_copilot_auth`, `openai_codex_session`,
`codex_auth` (`pub(crate) mod`). Migrate the wizard, the
`ironclaw login --openai-codex` CLI subcommand, and the LLM config
loader to the facade. Wizard introduces a single `WizardAuthPrompt`
that handles device-code prompts + browser launch for all backends.
3. Add `ironclaw_llm::models::fetch_models_for(provider_id, &opts)`
facade. Privatize `fetch_anthropic_models`, `fetch_openai_models`,
`fetch_ollama_models`, `fetch_openai_compatible_models`,
`is_openai_chat_model`, `openai_model_priority`, `sort_openai_models`.
Wizard's per-backend match collapses to one call. Move classifier
unit tests into `crates/ironclaw_llm/src/models.rs`; rewrite the
two wizard fallback tests through the public API.
4. Decouple embeddings from `ironclaw_llm::BedrockConfig`. New
`crate::workspace::BedrockEmbeddingSetup { region, profile }` carries
only what `BedrockEmbeddings` actually needs. `EmbeddingsConfig::create_provider`
and `BedrockEmbeddings::new` take the new type; callers translate from
`LlmConfig.bedrock` at the boundary (`src/app.rs`, `src/cli/mod.rs`).
5. Add `ironclaw_llm::testing::nearai_test_config(model)` helper for
tests that need a minimal `LlmConfig` shape (no retries, no caching,
NEAR AI backend). Replaces two duplicated 30-line struct literals
in the gateway settings hot-reload tests.
Boundary cleanup is behaviour-preserving: 4,932 main-binary unit tests,
729 ironclaw_llm unit tests, 4 ironclaw_oauth tests, 3 architecture
boundary tests all pass; `cargo clippy --all --benches --tests
--examples --all-features` is clean.
Three `pub` methods on `gemini_oauth::CredentialManager` /
`GeminiOauthProvider` (`get_valid_access_token`, `last_response_meta`,
`count_tokens`) and the `GeminiResponseMeta` struct are now reachable
only crate-internally and have no callers; marked `#[allow(dead_code)]`
with a comment rather than deleted to keep this PR purely a boundary
move (delete in a follow-up if no caller emerges).
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* refactor(llm): promote dedicated backends into the registry; absorb config validation, defaults, and per-provider overrides into ironclaw_llm
Continues the LLM boundary cleanup from 0addf3ac2. After that commit
provider-specific auth, model fetch, and embeddings config lived behind
facades inside `ironclaw_llm`, but four backend-specific knowledge
sources still leaked out:
1. Validation rules and default values for the dedicated-config
backends (Bedrock cross-region prefixes, OpenAI Codex endpoints
and client_id, Gemini OAuth credentials path defaults) lived
inline in `src/config/llm.rs::resolve`.
2. The dispatcher in `create_llm_provider` matched on backend strings
("nearai", "bedrock", ...) instead of a typed protocol value. The
same booleans (`is_nearai`, `is_bedrock`, `is_gemini_oauth`,
`is_openai_codex`) recurred across `src/config/llm.rs`,
`src/app.rs`, `src/cli/models.rs`, and the wizard.
3. The setup wizard had per-backend specialization in
`step_inference_provider` and `run_provider_setup` (manual menu
pushes for nearai/bedrock/codex/gemini_oauth, four dedicated
`setup_*` entry points dispatched on string compares).
4. `Settings` carried named `bedrock_region`, `bedrock_cross_region`,
`bedrock_profile` columns even though no other dedicated backend
had named columns and adding a new one would mean schema churn.
Layers A-D address each in turn:
* Layer A — `BedrockConfig::build`, `OpenAiCodexConfig::build`, and
`GeminiOauthConfig::build` own validation + defaults inside the
crate. `LlmConfigError` (`MissingRequired` / `InvalidValue`) carries
the failures across the boundary, with a `From` impl into the
binary's `ConfigError`. `src/config/llm.rs` calls the builders;
named-string defaults are gone from the binary. The orphaned
`tests/gemini_oauth_regression.rs` husk is deleted.
* Layer B — `ProviderProtocol` gains four new variants
(`Bedrock`, `OpenAiCodex`, `GeminiOauth`, `NearAi`) plus a
`has_dedicated_config()` predicate. The four dedicated-config
backends (with all aliases) become first-class registry entries in
`providers.json`, so `is_known()` / `model_env_var()` / the wizard /
the gateway handler iterate the registry uniformly. The
`is_nearai`/`is_bedrock`/`is_gemini_oauth`/`is_openai_codex` boolean
spaghetti collapses to protocol comparisons. `OpenAiCodex` and
`NearAi` carry explicit `#[serde(rename = "openai_codex" / "nearai",
alias = ...)]` so the wire-stable adapter strings the gateway and
frontend already use keep working. `LlmConfig::active_model_name()`
is now consumed by `cli/doctor.rs` instead of an inlined partial
dispatch.
* Layer C — `SetupHint` gains four credential-collection variants
(`AwsCredentials`, `OAuthDeviceCode`, `FileBasedCredentials`,
`SessionToken`). The wizard's `step_inference_provider` builds its
menu from a single `registry.selectable()` iteration with generic
env-detection (declared `api_key_env`, plus an Anthropic-specific
OAuth fallback). `run_provider_setup` dispatches on the SetupHint
variant; the remaining `def.id == "..."` checks live inside the
`ApiKey` arm only because Anthropic and GitHub Copilot present a
hybrid choice (API key OR OAuth) the simple `ApiKey` hint doesn't
capture. The synthetic bedrock + nearai entries in
`handlers/llm.rs::build_llm_providers` are deleted; a single
registry-driven loop covers both. ADAPTER_LABELS in
`static/js/surfaces/config.js` gains entries for the new protocols.
* Layer D — `LlmBuiltinOverride` gains a generic
`extras: HashMap<String, String>` bag with `extra(key)` /
`set_extra(key, value)` accessors. The bedrock resolver and wizard
read/write through this bag; `Settings::migrate_legacy_provider_fields()`
drains the named `bedrock_*` columns into `extras` on
`Settings::load_from()` so existing `settings.json` files migrate
losslessly. The named columns are kept (deprecated, marked with
`#[serde(skip_serializing_if = "Option::is_none")]`) for one
release; tracked for deletion in #3443.
`strip_admin_only_llm_keys` and `llm_setting_requires_reload` now
match dotted-path subkeys under `llm_builtin_overrides.*` so a
write to e.g. `llm_builtin_overrides.bedrock.extras.region`
triggers the right gating + chain reload.
Boundary cleanup is behaviour-preserving: 4,933 main-binary unit tests,
739 ironclaw_llm unit tests pass; `cargo clippy --all --benches --tests
--examples --all-features` is clean. New regression tests:
`crates/ironclaw_llm/src/config.rs` (6 builder tests),
`crates/ironclaw_llm/src/registry.rs::dedicated_config_backends_are_in_registry_and_selectable`,
and `src/setup/wizard.rs::legacy_bedrock_fields_migrate_into_extras_on_load`.
Three follow-ups tracked in #3443: delete the deprecated `bedrock_*`
named columns, move `BedrockEmbeddings` out of `src/workspace/` into
the LLM crate (last cargo-feature leak), and drive
`LlmConfig::active_model_name()` off `ProviderProtocol` instead of
backend strings.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test: add bug-bash regression-snapshot harness
Bug-bash fixtures pin specific open bugs to a deterministic snapshot.
When a bug is fixed, the snapshot diff is the reviewable proof; when
someone reintroduces the bug, the snapshot drifts and CI blocks the
merge.
This commit lands the harness plus the first recorded fixture for
issue #2541 (agent must call a tool, not answer from training data):
tests/e2e_bug_bash_snapshots.rs
`snapshot_summarization_uses_tools` replays the fixture, captures
`ReplayOutcome`, and asserts the YAML snapshot. Gated on
`feature = "libsql"`, same as other replay-snapshot tests.
tests/fixtures/llm_traces/bug_bash/summarization_uses_tools.json
Two-step recorded LLM trace (tool_call -> text) keyed off the
user prompt via `request_hint.last_user_message_contains`.
tests/fixtures/llm_traces/bug_bash/README.md
Coverage map for #2540-#2546 (one recorded, six TODO) plus the
`IRONCLAW_RECORD_TRACE` recording workflow.
tests/snapshots/replay__bug_bash_summarization_uses_tools.snap
Insta YAML snapshot pinning `tool_calls: [echo]`, 2 LLM calls,
and the event-kind histogram. Drift = regression.
* fix(settings): preserve pre-existing extras during legacy bedrock migration
`migrate_legacy_provider_fields` claimed to be idempotent and to drain
named `bedrock_*` columns into `llm_builtin_overrides["bedrock"].extras`
once on load. The previous implementation drained correctly but used
`HashMap::insert` unconditionally, which means a settings file
carrying BOTH a legacy `bedrock_region` column AND an already-populated
`extras["region"]` (manual hand-edit, or a future writer emitting both
shapes during a transition) would silently downgrade to the legacy
value.
Guard each `set_extra` call with `entry.extra(key).is_none()` so the
new-shape value always wins. Clarify the docstring to state this
explicitly.
Add three regression tests in `settings::tests`:
- `legacy_bedrock_migration_round_trips_through_save` — legacy JSON ->
load_from -> serialize -> reload, asserts the deprecated columns are
not re-emitted and extras survive the round trip.
- `legacy_bedrock_migration_preserves_existing_extras` — file with both
shapes; asserts the pre-existing extras value is kept and absent
extras are still backfilled from legacy fields.
- `legacy_bedrock_migration_is_idempotent_in_memory` — calling the
migration twice on the same Settings is a no-op (compares serialized
shape, since LlmBuiltinOverride does not derive PartialEq).
* fix(pr-3416): address PR review — migration on DB/TOML, admin-key gate, codex login, credential_kind/has_credentials
Addresses comments from gemini-code-assist, Copilot, and serrrfirat on PR #3416.
## Bugs
**Legacy bedrock fields not migrated on DB/TOML loads** (serrrfirat, High).
`Settings::load_from` (JSON) ran `migrate_legacy_provider_fields`, but
`from_db_map` and `load_toml` did not. Existing operators with
`bedrock_*` settings persisted in the DB or `config.toml` would silently
lose their AWS region/profile/cross-region after upgrade because the
resolver now reads only from `llm_builtin_overrides["bedrock"].extras`.
Both loaders now call the migration; added round-trip tests for each.
**Admin-only key write gate had narrower matching than read gate**
(Copilot, High). `strip_admin_only_llm_keys` matches both exact keys
and dotted subpaths under admin-only roots; `is_admin_only_setting_key`
in the web settings handler used `.contains(&key)` only. A non-admin
could write `llm_builtin_overrides.bedrock.extras.region` directly,
bypassing the gate. Promoted `is_admin_only_llm_key` to `pub(crate)`,
made the web write-side gate call it, added regression tests covering
dotted subpaths.
**`ironclaw login --openai-codex` dropped TOML/DB config** (Copilot,
High). The pre-refactor code resolved `Config::from_env` and used
`config.llm.openai_codex` so endpoint / client-id / session-path
overrides committed via TOML or DB stuck. The post-refactor code only
read env vars via `OpenAiCodexLoginOptions::from_env`. Added
`OpenAiCodexLoginOptions::from_resolved_config(&OpenAiCodexConfig)`;
the login command now prefers the resolved config when present and
falls back to env-only when `Config::from_env` itself fails (fresh
machine, no DB).
**Dedicated-auth backends marked configured without credentials**
(serrrfirat, Medium). `nearai` / `gemini_oauth` / `openai_codex` ship
`api_key_required: false` because they don't authenticate via a bearer
API key. The frontend `isProviderConfigured` treated that as "no
credentials needed" and rendered the Use button on a fresh install,
where clicking could trigger an interactive device-code OAuth from
inside a settings request.
Added `credential_kind` (wire-stable snake_case discriminator matching
`SetupHint::kind()`, e.g. `session_token`, `o_auth_device_code`,
`file_based_credentials`, `aws_credentials`) and `has_credentials`
(backend-authoritative; checks AWS env vars for Bedrock, codex session
file existence, file-based credential path expansion + existence) to
the web LLM providers payload. Frontend `isProviderConfigured` /
`providerMissingReason` now gate non-api-key kinds on `has_credentials`.
## Nits
**`fetch_models_for` doc overclaimed "Always returns something"**
(Copilot). The generic openai-compatible branch returns `vec![]` when
`base_url` is empty. Updated the docstring to call this out so callers
know to handle the empty case.
**`AuthError::Other` used for "validation not applicable"** (Gemini
bot). Added a dedicated `AuthError::TokenValidationNotSupported { backend }`
variant; `validate_token` now returns it for Gemini / OpenAiCodex
instead of stringly-formatted `Other`.
**Bug-bash regression-harness URLs pointed at `near/ironclaw`**
(Copilot, x2). The canonical tracker is `nearai/ironclaw`. Rewrote
all seven URLs in `tests/fixtures/llm_traces/bug_bash/README.md` and
the one in `tests/e2e_bug_bash_snapshots.rs`.
## Declined
The Gemini bot's MalformedConfig suggestion at
`crates/ironclaw_llm/src/models.rs:46` was not adopted: the call site is
the openai-compatible model-listing path, not a security-sensitive
request. The fetcher early-returns `vec![]` on empty `base_url` — no
URL parsing happens — and the docstring tightening above covers the
observable surprise. Promoting it to a typed error would change the
public-facing `fetch_models_for` signature for no behavioural gain.
## Tests
- `cargo fmt --check` clean
- `cargo clippy --all --benches --tests --examples --all-features` zero warnings
- `cargo test --lib` 4,941 / 4,941 pass
- `cargo test --features libsql --test e2e_bug_bash_snapshots` 1 / 1 pass
- New regression tests:
- `settings::tests::legacy_bedrock_fields_migrate_into_extras_on_db_load`
- `settings::tests::legacy_bedrock_fields_migrate_into_extras_on_toml_load`
- `channels::web::features::settings::tests::test_admin_only_setting_keys_cover_dotted_subpaths`
- `channels::web::handlers::llm::tests::test_llm_providers_expose_credential_kind_and_has_credentials`
- `channels::web::handlers::llm::tests::test_nearai_has_credentials_true_when_session_token_loaded`
* fix(pr-3416): tighten Bedrock/Codex has_credentials probes; collapse set_extra into one .into()
- `backend_has_credentials` for AWS now requires `AWS_PROFILE` OR
(`AWS_ACCESS_KEY_ID` AND `AWS_SECRET_ACCESS_KEY`). The lone
`AWS_ACCESS_KEY_ID` / `AWS_SESSION_TOKEN` arms previously flipped
has_credentials true even though the AWS SDK can't sign without the
secret key, so the UI was rendering Bedrock as configured on hosts
that would fail at first call.
- `backend_has_credentials` for OpenAI Codex now honours
`OPENAI_CODEX_SESSION_PATH` via `read_env` before falling back to
the default session path under `~/.ironclaw/`. Users with a custom
session location were seeing "not configured" despite a valid login.
- New regression tests `test_bedrock_partial_aws_env_reports_not_configured`
and `test_openai_codex_honours_session_path_env` drive the
`build_llm_providers` call site (not just the helper) so both gaps
stay closed.
- Tidied `LlmBuiltinOverride::set_extra` to convert the key once and
reuse it across the remove/insert branches; behaviour identical.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* chore(providers): default nearai model to "auto"
Switch the nearai registry entry's `default_model` from
`claude-sonnet-4-5` to `auto`, NEAR AI's server-side routing alias.
New installs without `NEARAI_MODEL` set now get auto-routed instead
of being pinned to a specific Anthropic model.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* Make Skills E2E lifecycle deterministic (#3309)
* test(e2e): make skills lifecycle deterministic
* test(e2e): address skills review comments (#3309)
* test(e2e): unxfail two auth-matrix tests now that contracts match (#3589)
Both xfails in tests/e2e/scenarios/test_v2_auth_oauth_matrix.py are
stale and pass against current code:
- test_wasm_tool_first_chat_auth_attempt_emits_auth_url
Marked xfail in #3235 because the engine-v2 callable-only contract
(#2868) stopped emitting an auth gate on direct LLM-driven tool
calls. PR #3157 (auth-preflight + inline-await) restored the
behavior the test asserts: when the LLM emits a direct call to a
not-yet-authed extension, the bridge raises an Authentication gate
with auth_url populated (src/bridge/effect_adapter.rs:1356-1392).
Marker removed; test passes.
- test_settings_first_custom_mcp_auth_then_chat_runs
The xfail reason claimed post-auth tool-output propagation was
broken. Real cause: engine-v2 gates the first MCP tool call on
`approval` and the browser fixture has no auto-approve UI, so the
chat sat in pending_gate forever. Same shape as the bugs fixed in
#3235 for test_wasm_tool_oauth_refresh_on_demand and
test_mcp_same_server_multi_user_via_browser. Inserted
_wait_for_tool_call between _send_chat and _wait_for_response_contains
to drive approval through the API; test passes.
Verified locally: both tests pass back-to-back in 27s on a fresh
auth_matrix_server.
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(extensions): restore chat-driven tool_install + fix double-invoke + auto-approve footgun (#3533) (#3559)
* fix(extensions): restore chat-driven tool_install + fix double-invoke + auto-approve footgun (#3533)
"Connect my telegram" was giving the user two options and not actually
installing anything because three layered issues had accumulated since
engine v2:
1. **`tool_install` was hidden from the agent** (#2868). The unified
`tool_activate` it was meant to be subsumed by was later removed in
#3166, but the hidden-from-callable-surface gate stayed. Restored
by dropping `hidden_from_model_callable_surface` from
`bridge::action_projector`. User consent is mediated by the tool's
own `ApprovalRequirement::UnlessAutoApproved` and the seeded
`AskEachTime` permission.
2. **Two competing Telegram registry entries** (`telegram` channel and
`telegram_mtproto` tool) both surfaced in the agent prompt's
`Activatable Integrations` section. The LLM correctly enumerated
them as "Option 1" and "Option 2" instead of installing the
canonical bot channel. Added a `hidden: bool` field to
`ExtensionManifest` / `RegistryEntry`, set `telegram_mtproto` to
`hidden: true`, and filter hidden entries out of the
"available-but-not-installed" appendix in `ExtensionManager::list`.
Hidden entries remain installable by explicit name.
3. **Updated the agent prompt** so `Activatable Integrations` instructs
the model to call `tool_install(name="<name>")` directly rather than
describing manual UI steps.
Fixes the double-`tool_install` invocation that surfaced once the agent
could install from chat:
- **`InlineGate` discarded cached output.** The bridge raised an
Authentication gate after `tool_install` succeeded, and the
inline-await retry re-executed the action (re-downloading the WASM
bundle) instead of returning the already-computed output. Added
`resume_output: Option<serde_json::Value>` to `InlineGate`; on
approval, return the cached output if present. Mirror fix in the
orchestrator's `execute_single_action_with_inline_retry` (reading
`result_json["resume_output"]`) and the structured-batch retry path.
- **`effect_adapter::auth_gate_from_extension_result`** now passes
`Some(output_value.clone())` as the gate's `resume_output` so the
retry has cached state to short-circuit on.
- **OAuth callback double-fired.** `oauth_callback_handler` now skips
the `ExternalCallback` re-entry when the inline-await path already
woke a parked waiter — eliminates the "thread already running" race.
- **`resolve_inline_gates_for_credential`** now also discards matching
Authentication rows from `pending_gates` so the row doesn't linger
in `HistoryResponse.pending_gate` after inline resolution.
Fixes the auto-approve footgun:
- **`ToolPermissionSnapshot::resolve_permission`** now collapses DB
values that match the seeded default to `explicit = None`. Before
this, the boot-time `seed_tool_permissions` write of `tool_install ->
AskEachTime` was indistinguishable from a user-explicit override,
causing `effect_adapter::enforce_tool_permission`'s `is_explicit_ask`
check to refuse `AGENT_AUTO_APPROVE_TOOLS=true`. Real overrides
(`AlwaysAllow`, `Disabled`) still surface as `Some(...)`.
Tests
- Unit: 4980/4980 pass (host) + 525/525 pass (engine).
- Unit: new tests in `bridge::tool_permissions::tests` lock seeded-vs-
explicit collapse; new test in `bridge::action_projector::tests`
asserts `tool_install` is callable; new manifest hidden-flag tests
in `registry::manifest::tests` and `extensions::manager::tests`.
- E2E: removed `@pytest.mark.xfail` on
`test_chat_first_gmail_installs_prompts_and_retries` (now passes
end-to-end via the chat-driven install path). Added
`test_chat_install_approval_then_auth_card` driving the
explicit-approval variant with a single Approve click (no Always
workaround needed) — wired into the `auth-full` canary lane.
- Mock LLM: extended the gmail-install-then-retry pattern to recognize
both the legacy "Extension not installed:" and the post-#3533 "is
not callable in this execution context" error strings, and to retry
`gmail(action="list_messages")` after a successful `tool_install`.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* fix(permissions): address #3559 review (permission bypass, lease accounting, hidden search filter)
Five fixes from the #3559 review (4× Copilot doc nits + 3× serrrfirat
security/correctness findings):
1. **Permission bypass (High).** Pre-#3559's `resolve_permission`
collapsed any DB row whose value matched the seeded default to
`explicit = None`, so a user who deliberately set `tool_install =
AskEachTime` had their explicit choice silently dropped and
`AGENT_AUTO_APPROVE_TOOLS=true` bypassed the gate. Provenance is now
handled at write time: `seed_tool_permissions` is gone and a
one-shot, sentinel-gated migration (`cleanup_ghost_seeded_tool_permissions`)
deletes existing ghost-seeded rows at startup. With no ghost rows,
the resolver treats every DB row as user-explicit and honors it.
2. **Lease/event accounting on `resume_output` replay (Medium).**
Inline-gate handlers in `structured.rs`, `scripting.rs`
(`resolve_tool_future` + `drive_inline_gate` retry loop), and
`orchestrator.rs` refunded the lease use the action just consumed,
then returned the cached `resume_output` on approval without
re-consuming — netting successful side-effecting actions to zero
lease uses. Skip the refund when the gate carries cached output.
3. **Hidden registry filter on `tool_search` (Medium).**
`RegistryCatalog::search` did not filter `hidden: true` entries,
so `telegram_mtproto` could resurface through the search path and
reintroduce the "two Telegram options" outcome that #3533 fixes
for the default-list path. Added the filter and a regression test.
4-7. Copilot doc nits: outdated `_set_tool_permission` docstring;
misleading "bridge-side auto-install implemented" comment in
`mock_llm.py`; `tool_install` described as "non-agent surface" in
`src/bridge/CLAUDE.md` while a paragraph below says the model
calls it directly; dangling `issue #3533 / PR —` placeholders
in both CLAUDE.md docs.
Regression tests:
- `bridge::tool_permissions::user_explicit_value_matching_seeded_default_stays_explicit` —
the original Copilot/serrrfirat bug case.
- `app::cleanup_ghost_seeded_tool_permissions_removes_seed_matching_rows` —
idempotent migration + sentinel.
- `extensions::registry::test_search_skips_hidden_entries` — hidden
entries excluded from search but still installable by exact name.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* test(#3559): caller-level regression coverage for review findings 1 & 2
Two follow-up regression tests for the #3559 security review, plus a
real bug surfaced by the first one.
1. `executor::structured::resume_output_replay_consumes_exactly_one_lease_use`
exercises the post-execution Authentication gate inline-retry path
with `max_uses=1` and asserts:
- Cached output is returned as a successful `ActionResult`.
- Exactly one `ActionExecuted` event is emitted.
- The lease budget is exhausted after one execution (refund-skip
keeps the consumption from being undone).
Writing this test surfaced a real bug: the structured cached-output
branch pushed `ActionExecuted` into `emitted_events`, and the
caller's `classify_exec_result` emitted ANOTHER terminal
`ActionExecuted` for the same Ok result — double-emit for one
action. Tier 1 (`scripting::drive_inline_gate`) and Tier 1 alt
(`orchestrator::execute_action_with_inline_gate`) emit themselves
because their callers don't run an Ok-branch classifier; structured
was the outlier. Dropped the redundant push; the classifier emits
the single canonical event.
2. `bridge::effect_adapter::explicit_ask_each_time_for_seeded_default_tool_still_gates`
drives `execute_action` end-to-end (the side-effecting caller) with
a tool whose `name()` matches a seeded-`AskEachTime` baseline
(`tool_install`) and an explicit `AskEachTime` user override. The
resolver collapse-to-implicit bug would have shown up here — not
just in the helper-level test that already exists in
`bridge::tool_permissions::tests`. Per `.claude/rules/testing.md`
"Test Through the Caller, Not Just the Helper".
Added `SeededAskEachTimeTestTool` as a `tool_install`-named test
fixture with `requires_approval: UnlessAutoApproved` to mirror the
real tool's contract.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* chore: release
* feat(engine): IRONCLAW_DISABLE_CODEACT flag to disable v2 CodeAct (#3665)
* add flag to disable codeact on engine v2
* fmt
* fix(engine): keep compact actions reachable when CodeAct is disabled
With IRONCLAW_DISABLE_CODEACT=true the structured-tools prompt told
the model to use the provider's tool_calls interface for every action,
but the bridge filtered the provider tool list down to
emits_full_schema_tool(). Most tools default to CompactToolInfo
(mission_create, gmail_send, notion_search, ...), so they appeared in
the prompt as "available" while being absent from the provider tool
list — i.e. unreachable. Addresses serrrfirat's review on PR #3665.
Fix coordinates both halves of the surface:
- src/bridge/llm_adapter.rs: in disabled-CodeAct mode, drop the
emits_full_schema_tool() filter and emit every action into the
provider tool list with its full schema.
- crates/ironclaw_engine/src/executor/prompt.rs: in disabled-CodeAct
mode, skip the "## Enabled Tools" section. The compact-form listing
with the tool_info(detail="schema") instruction is meaningless when
the provider already sends full schemas, and would just duplicate
the surface. "## Activatable Integrations" stays — the model still
needs to know what tool_install can target.
Test seam: build_codeact_system_prompt_inner now takes disable_codeact
as an explicit parameter, called once at the public entry points. This
lets prompt tests exercise both branches without process-global env
mutation.
Tests:
- executor::prompt::tests::disabled_codeact_omits_enabled_tools_section_and_keeps_activatable
- bridge::llm_adapter::tests::complete_emits_compact_actions_when_codeact_disabled
- existing complete_with_tools_only_emits_full_schema_provider_tools
now serialized via lock_env() so env mutation in the new test
can't leak across parallel runs.
cargo test -p ironclaw_engine --lib: 527 passed
cargo test --lib bridge::: 469 passed
cargo clippy -p ironclaw_engine --all-targets -- -D warnings: clean
cargo clippy --lib --tests -- -D warnings: clean
cargo fmt --check: clean
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Emil Bogomolov <emil.bogomolov@near.ai>
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* Fix markdown_to_mrkdwn to avoid converting emphasis inside generated <… (#3532)
* agent: Fix markdown_to_mrkdwn to avoid converting emphasis inside g…
* agent: Fix rustfmt/clippy CI failure by removing extra blank line b…
* agent: slack: fix markdown_to_mrkdwn replacement order to satisfy p…
* slack: protect generated links and sanitize sentinels in markdown_to_mrkdwn
Two issues raised on PR #3532 review:
1. Emphasis inside generated `<url|text>` was still rewritten because the
global `**`/`~~` → `*`/`~` substitution ran after link materialization.
Push the generated link span into the same protected arena used for
Slack-native `<...>` constructs so subsequent global replacements can't
reach inside it. Matches the PR's stated goal.
2. Untrusted input containing the private-use sentinel chars
(U+E000 / U+E001) could forge a protected-span reference and pull in
another span's content. Strip those chars from input up front.
Adds regression tests for both. Bumps registry/channels/slack.json
0.3.2 → 0.3.3 to satisfy the channel-source version-bump check.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* slack: escape link labels, drop pipe/gt URLs, expand nested sentinels
Addresses two follow-up review concerns on PR #3532:
Copilot review: `<url|text>` was built by string concatenation, so
`|` or `>` inside the URL would corrupt the entity, and `<` / `>`
inside the label would open/close a Slack span and break the link.
The label now escapes `<` → `<` and `>` → `>` (Slack's documented
literal-character form); a URL containing `<`, `>`, or `|` falls back
to leaving the original markdown form intact (those chars are not
valid URL characters per RFC 3986 anyway).
Latent nested-sentinel bug introduced by the previous fix: a markdown
link whose label contained a Slack-native `<...>` span (e.g.
`[<@U1> hi](url)`) ended up with the inner sentinel buried inside the
arena entry for the outer link span. The final restore pass advances
past the outer sentinel without rescanning what it just emitted, so
the raw U+E000/U+E001 characters would leak into the output. URL and
label are now pre-expanded before the link span is pushed.
While here, factor the duplicated restore loop into
`expand_protected_spans`, reused by both the pre-link expansion and
the final restore, and lift the sentinel constants to file scope.
Adds three regression tests covering label-bracket escaping, the URL
pipe/gt fallback, and the nested-span case. Bumps
registry/channels/slack.json 0.3.3 → 0.3.4.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
* feat(gateway): add logs download button (#3588)
* feat(web): support externally-provided tools in Responses API (#3122)
* feat(web): support externally-provided tools in Responses API
Lets callers of `/v1/responses` (and `/api/v1/responses`) declare their
own `function`-typed tools and feed back results via
`function_call_output` items, matching the OpenAI Responses wire shape.
Since IronClaw's engine has no per-request tool surface, integration
happens at the prompt level: the catalog is rendered as
`<external-tools>` in the user message and the agent signals a call by
ending its response with a fenced ```` ```tool_call ```` block. When
that fence is recognised, the reply is split into a leading `Message`
plus a `function_call` `ResponseOutputItem`.
Validation rejects unsupported tool types (`web_search`, `file_search`,
`code_interpreter`) and tools missing `name` with 400, with two new
integration tests covering both paths.
* refactor(responses-api): switch external tools to engine v2 native path
Replace the prompt-level fence protocol from PR #3122 with engine v2
native tool calls: caller-supplied `tools[]` are surfaced as real
LLM-callable actions, the engine pauses with `ResumeKind::External`
when one is invoked, and the bridge router projects the pause to a
new `AppEvent::ExternalToolCall` carrying the OpenAI-shaped
`function_call` wire fields.
The integration is small because v2 already has the right primitives:
- `ResumeKind::External { callback_id }` and
`GateResolution::ExternalCallback { payload }` already existed for
OAuth-style callbacks.
- `agent_loop.rs:1480` already routes Responses API messages to
`handle_with_engine` when `ENGINE_V2=true`, so no v2 migration of
the endpoint itself is needed.
- `EffectBridgeAdapter::execute_action` is the single chokepoint
where caller tools can be detected before they reach the dispatch
pipeline.
Changes:
- New `src/bridge/external_tools.rs` (`ExternalToolCatalog`) — per-thread
registry of caller-supplied `ActionDef`s, plus the `ext_tool:`
callback-id helpers used to disambiguate external-tool pauses from
OAuth/pairing pauses (which also use `ResumeKind::External`).
- `EffectBridgeAdapter` consults the catalog: any name in it is
short-circuited to a `GatePaused { resume_kind: External {
callback_id: ext_tool:<call_id> } }` before any registry dispatch,
and `available_action_inventory` merges the catalog into the
LLM-visible action surface (internal beats external on collision).
- `Submission::ExternalCallback` gains an optional `payload` field;
`bridge::handle_external_callback` plumbs it into
`GateResolution::ExternalCallback { payload }`. Fallback predicate
`gate_resume_is_external` lets non-auth External pauses (i.e.
caller-tool resumes) resolve through the same handler.
- New `AppEvent::ExternalToolCall` projected by `notify_pending_gate`
when a paused gate carries an `ext_tool:` callback id; OAuth/
pairing flows keep flowing through the existing `GateRequired`
channel.
- `responses_api.rs` is gutted of the prompt rendering and fence
parsing (`render_external_tools_preamble`, `extract_trailing_tool_call`,
`parse_external_tool_call`, `ParsedToolCall`, `external_tool_names`
accumulator field, and the `TOOL_CALL_FENCE` constants). The handler
now: rejects `tools[]` when `ENGINE_V2=false`, registers caller
tools in the catalog under the resolved thread id, detects resume
requests (`previous_response_id` + `function_call_output` items in
`input`) and submits them as `Submission::ExternalCallback` with
the outputs as the resolution payload, and surfaces
`AppEvent::ExternalToolCall` as a `function_call` `ResponseOutputItem`
in both streaming (`output_item.added`+`done`) and non-streaming.
- All existing OAuth/pairing `ExternalCallback` constructors updated
to pass `payload: None` (no behaviour change).
- Fence-protocol unit tests removed; replaced with coverage for the
new `responses_tools_to_action_defs` converter and the accumulator's
`ExternalToolCall` arm.
Existing 9 integration tests in `tests/responses_api_path_prefix.rs`
still pass.
Note for reviewers:
- The accumulator-side text response no longer tries to split the
reply on a fenced `tool_call` block. The wire shape that callers
receive for caller-tool invocations is purely event-driven now.
- Internal vs external collision is handled silently by the dedup in
`available_action_inventory` (internal wins). A request-time
rejection for shadowing names is a follow-up — the current behavior
is safe (the LLM only sees the internal version) but could surprise
a caller who expects their tool to run.
* test(responses-api): cover ENGINE_V2-off and resume-without-pending-gate
Two new integration tests for behaviours added by the engine-native
external-tool refactor:
- `external_tools_rejected_when_engine_v2_disabled`: a request with
caller-supplied `tools[]` while `ENGINE_V2` is off must 400 with a
message naming the flag, not silently fall through.
- `resume_without_pending_gate_returns_400`: a request with
`function_call_output` items and a `previous_response_id` that
doesn't correspond to a live external-tool gate must 400, not start
a fresh turn against the (unrelated) thread.
Both tests drive the full router (`start_test_server` + bearer auth)
per `.claude/rules/testing.md` "Test Through the Caller".
* test(responses-api): integration tests + drop unsafe env mutation
Three groups of changes:
1. **Drop unsafe env-var mutation in tests.** `responses_api.rs` no
longer reads `ENGINE_V2` directly: it keys off the presence of the
live `ExternalToolCatalog` (initialized by `init_engine`) as the
"engine v2 is up" signal. The path-prefix test that exercises the
no-engine branch no longer needs `unsafe { std::env::remove_var }`
— the absence of `init_engine` in `TestGatewayBuilder` is what
makes the catalog absent, which is what makes the request reject.
2.…
…arai#3589) Both xfails in tests/e2e/scenarios/test_v2_auth_oauth_matrix.py are stale and pass against current code: - test_wasm_tool_first_chat_auth_attempt_emits_auth_url Marked xfail in nearai#3235 because the engine-v2 callable-only contract (nearai#2868) stopped emitting an auth gate on direct LLM-driven tool calls. PR nearai#3157 (auth-preflight + inline-await) restored the behavior the test asserts: when the LLM emits a direct call to a not-yet-authed extension, the bridge raises an Authentication gate with auth_url populated (src/bridge/effect_adapter.rs:1356-1392). Marker removed; test passes. - test_settings_first_custom_mcp_auth_then_chat_runs The xfail reason claimed post-auth tool-output propagation was broken. Real cause: engine-v2 gates the first MCP tool call on `approval` and the browser fixture has no auto-approve UI, so the chat sat in pending_gate forever. Same shape as the bugs fixed in nearai#3235 for test_wasm_tool_oauth_refresh_on_demand and test_mcp_same_server_multi_user_via_browser. Inserted _wait_for_tool_call between _send_chat and _wait_for_response_contains to drive approval through the API; test passes. Verified locally: both tests pass back-to-back in 27s on a fresh auth_matrix_server. Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
… + auto-approve footgun (nearai#3533) (nearai#3559) * fix(extensions): restore chat-driven tool_install + fix double-invoke + auto-approve footgun (nearai#3533) "Connect my telegram" was giving the user two options and not actually installing anything because three layered issues had accumulated since engine v2: 1. **`tool_install` was hidden from the agent** (nearai#2868). The unified `tool_activate` it was meant to be subsumed by was later removed in nearai#3166, but the hidden-from-callable-surface gate stayed. Restored by dropping `hidden_from_model_callable_surface` from `bridge::action_projector`. User consent is mediated by the tool's own `ApprovalRequirement::UnlessAutoApproved` and the seeded `AskEachTime` permission. 2. **Two competing Telegram registry entries** (`telegram` channel and `telegram_mtproto` tool) both surfaced in the agent prompt's `Activatable Integrations` section. The LLM correctly enumerated them as "Option 1" and "Option 2" instead of installing the canonical bot channel. Added a `hidden: bool` field to `ExtensionManifest` / `RegistryEntry`, set `telegram_mtproto` to `hidden: true`, and filter hidden entries out of the "available-but-not-installed" appendix in `ExtensionManager::list`. Hidden entries remain installable by explicit name. 3. **Updated the agent prompt** so `Activatable Integrations` instructs the model to call `tool_install(name="<name>")` directly rather than describing manual UI steps. Fixes the double-`tool_install` invocation that surfaced once the agent could install from chat: - **`InlineGate` discarded cached output.** The bridge raised an Authentication gate after `tool_install` succeeded, and the inline-await retry re-executed the action (re-downloading the WASM bundle) instead of returning the already-computed output. Added `resume_output: Option<serde_json::Value>` to `InlineGate`; on approval, return the cached output if present. Mirror fix in the orchestrator's `execute_single_action_with_inline_retry` (reading `result_json["resume_output"]`) and the structured-batch retry path. - **`effect_adapter::auth_gate_from_extension_result`** now passes `Some(output_value.clone())` as the gate's `resume_output` so the retry has cached state to short-circuit on. - **OAuth callback double-fired.** `oauth_callback_handler` now skips the `ExternalCallback` re-entry when the inline-await path already woke a parked waiter — eliminates the "thread already running" race. - **`resolve_inline_gates_for_credential`** now also discards matching Authentication rows from `pending_gates` so the row doesn't linger in `HistoryResponse.pending_gate` after inline resolution. Fixes the auto-approve footgun: - **`ToolPermissionSnapshot::resolve_permission`** now collapses DB values that match the seeded default to `explicit = None`. Before this, the boot-time `seed_tool_permissions` write of `tool_install -> AskEachTime` was indistinguishable from a user-explicit override, causing `effect_adapter::enforce_tool_permission`'s `is_explicit_ask` check to refuse `AGENT_AUTO_APPROVE_TOOLS=true`. Real overrides (`AlwaysAllow`, `Disabled`) still surface as `Some(...)`. Tests - Unit: 4980/4980 pass (host) + 525/525 pass (engine). - Unit: new tests in `bridge::tool_permissions::tests` lock seeded-vs- explicit collapse; new test in `bridge::action_projector::tests` asserts `tool_install` is callable; new manifest hidden-flag tests in `registry::manifest::tests` and `extensions::manager::tests`. - E2E: removed `@pytest.mark.xfail` on `test_chat_first_gmail_installs_prompts_and_retries` (now passes end-to-end via the chat-driven install path). Added `test_chat_install_approval_then_auth_card` driving the explicit-approval variant with a single Approve click (no Always workaround needed) — wired into the `auth-full` canary lane. - Mock LLM: extended the gmail-install-then-retry pattern to recognize both the legacy "Extension not installed:" and the post-nearai#3533 "is not callable in this execution context" error strings, and to retry `gmail(action="list_messages")` after a successful `tool_install`. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(permissions): address nearai#3559 review (permission bypass, lease accounting, hidden search filter) Five fixes from the nearai#3559 review (4× Copilot doc nits + 3× serrrfirat security/correctness findings): 1. **Permission bypass (High).** Pre-nearai#3559's `resolve_permission` collapsed any DB row whose value matched the seeded default to `explicit = None`, so a user who deliberately set `tool_install = AskEachTime` had their explicit choice silently dropped and `AGENT_AUTO_APPROVE_TOOLS=true` bypassed the gate. Provenance is now handled at write time: `seed_tool_permissions` is gone and a one-shot, sentinel-gated migration (`cleanup_ghost_seeded_tool_permissions`) deletes existing ghost-seeded rows at startup. With no ghost rows, the resolver treats every DB row as user-explicit and honors it. 2. **Lease/event accounting on `resume_output` replay (Medium).** Inline-gate handlers in `structured.rs`, `scripting.rs` (`resolve_tool_future` + `drive_inline_gate` retry loop), and `orchestrator.rs` refunded the lease use the action just consumed, then returned the cached `resume_output` on approval without re-consuming — netting successful side-effecting actions to zero lease uses. Skip the refund when the gate carries cached output. 3. **Hidden registry filter on `tool_search` (Medium).** `RegistryCatalog::search` did not filter `hidden: true` entries, so `telegram_mtproto` could resurface through the search path and reintroduce the "two Telegram options" outcome that nearai#3533 fixes for the default-list path. Added the filter and a regression test. 4-7. Copilot doc nits: outdated `_set_tool_permission` docstring; misleading "bridge-side auto-install implemented" comment in `mock_llm.py`; `tool_install` described as "non-agent surface" in `src/bridge/CLAUDE.md` while a paragraph below says the model calls it directly; dangling `issue nearai#3533 / PR —` placeholders in both CLAUDE.md docs. Regression tests: - `bridge::tool_permissions::user_explicit_value_matching_seeded_default_stays_explicit` — the original Copilot/serrrfirat bug case. - `app::cleanup_ghost_seeded_tool_permissions_removes_seed_matching_rows` — idempotent migration + sentinel. - `extensions::registry::test_search_skips_hidden_entries` — hidden entries excluded from search but still installable by exact name. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(nearai#3559): caller-level regression coverage for review findings 1 & 2 Two follow-up regression tests for the nearai#3559 security review, plus a real bug surfaced by the first one. 1. `executor::structured::resume_output_replay_consumes_exactly_one_lease_use` exercises the post-execution Authentication gate inline-retry path with `max_uses=1` and asserts: - Cached output is returned as a successful `ActionResult`. - Exactly one `ActionExecuted` event is emitted. - The lease budget is exhausted after one execution (refund-skip keeps the consumption from being undone). Writing this test surfaced a real bug: the structured cached-output branch pushed `ActionExecuted` into `emitted_events`, and the caller's `classify_exec_result` emitted ANOTHER terminal `ActionExecuted` for the same Ok result — double-emit for one action. Tier 1 (`scripting::drive_inline_gate`) and Tier 1 alt (`orchestrator::execute_action_with_inline_gate`) emit themselves because their callers don't run an Ok-branch classifier; structured was the outlier. Dropped the redundant push; the classifier emits the single canonical event. 2. `bridge::effect_adapter::explicit_ask_each_time_for_seeded_default_tool_still_gates` drives `execute_action` end-to-end (the side-effecting caller) with a tool whose `name()` matches a seeded-`AskEachTime` baseline (`tool_install`) and an explicit `AskEachTime` user override. The resolver collapse-to-implicit bug would have shown up here — not just in the helper-level test that already exists in `bridge::tool_permissions::tests`. Per `.claude/rules/testing.md` "Test Through the Caller, Not Just the Helper". Added `SeededAskEachTimeTestTool` as a `tool_install`-named test fixture with `requires_approval: UnlessAutoApproved` to mirror the real tool's contract. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Context
This PR now includes the follow-up work from #2869, #2876, and #2889, so it should be read against the engine-v2 epic comment in #2767:
#2767 (comment)
The main alignment is:
available_actions()callable-only cleanuptool_activate(name=...)Old Experience
Before this stack:
available_actions()mixed together the wrong provider states in engine v2. Some blocked provider actions (needs_auth,needs_setup, inactive, not-installed, routed-only) could still show up as normal callable tools even though they were not directly invokable on that turn, while installed-but-not-yet-activated latent WASM provider actions could disappear entirely and fail to trigger auth-on-first-call.tool_infowas not consistently reading the same callable snapshot the executor was using, and hyphen/underscore aliases were easy to handle inconsistently at different call sites.tool_install,tool_auth, and latent synthetic actions all leaked overlapping lifecycle detail into the surfaced experience.tool_info-> promotion flow. That was leftover prototype behavior, not the contract we want to preserve.New Experience
After this stack:
Capabilitiesfor contextual/background runtime informationActivatable Integrationsfor blocked managed integrationstool_activate(name=...)is the single model-facing enablement path in engine v2 and CodeAct. The harness owns whether that means install, auth, activation, or a surfaced manual-setup failure.tool_infonow reads the current callable snapshot / inventory, resolves canonical names and hyphen/underscore aliases consistently, fails cleanly for names that are not callable in the current execution context, and does not gate or promote a new tool into the next LLM step's callable set.What Changed
ActionDiscoveryMetadata,ActionDiscoverySummary,ActionInventory, andActionDef::matches_name()so callable action metadata is canonical and alias-safeActionProjectoronly surfaces callable-now actions andCapabilityProjectorowns blocked provider/channel background plus action previewssrc/bridge/engine_actions.rstool_infoall read the same callable setCapabilitiesandActivatable Integrationssections, and refresh engine-owned system prompts across resume / compaction without duplicating appended sectionstool_activatethe normal surfaced enablement operation, while preserving install approval semantics when enablement would auto-install a not-yet-installed integrationopenai_compatibleprovider instead of relying on env-vs-DB precedencetool_activate/Activatable IntegrationscontractKey Coverage Added In This Stack
available_actions_omit_installed_needs_auth_provider_actionavailable_actions_omit_installed_needs_setup_provider_actionavailable_actions_omit_installed_inactive_provider_actionavailable_actions_omit_latent_inactive_provider_actionsavailable_capabilities_include_latent_provider_backgroundavailable_capabilities_include_latent_provider_activation_entrytool_activate_requires_approval_before_auto_installing_integrationskill_prompt_context_survives_pause_and_resumeskill_prompt_context_survives_compaction_and_resumeresume_refreshes_checkpointed_system_prompt_metadatatool_info_does_not_gate_callable_tool_into_next_llm_callable_setexecute_code_propagates_snapshot_to_tool_execution_contexttest_v2_tool_activate_surface.pyTesting
All of the following were run on this branch:
cargo fmt --all cargo clippy --all-targets -- -D warnings cargo test effect_adapter:: --lib tests/e2e/.venv/bin/pytest \ tests/e2e/scenarios/test_v2_auth_oauth_matrix.py::test_settings_first_gmail_auth_then_chat_runs \ tests/e2e/scenarios/test_v2_auth_oauth_matrix.py::test_mcp_oauth_roundtrip_via_browser \ tests/e2e/scenarios/test_telegram_hot_activation.py::test_telegram_auth_required_shows_configure_modal_and_can_cancel \ tests/e2e/scenarios/test_telegram_hot_activation.py::test_telegram_hot_activation_transitions_installed_to_pairing \ tests/e2e/scenarios/test_v2_tool_activate_surface.py -vResult: all passed.
https://github.com/nearai/ironclaw/actions/runs/24864084782