Skip to content

chore: promote staging to staging-promote/4c9a985b-23931806540 (2026-04-03 05:32 UTC) - #1953

Merged
henrypark133 merged 138 commits into
staging-promote/9c6d8cb3-23875318792from
staging-promote/aa59ca09-23935258795
Apr 10, 2026
Merged

henrypark133 merged 138 commits into
staging-promote/9c6d8cb3-23875318792from
staging-promote/aa59ca09-23935258795

Conversation

@ironclaw-ci

@ironclaw-ci ironclaw-ci Bot commented Apr 3, 2026 •

Copy link
Copy Markdown
Contributor

Auto-promotion from staging CI

Batch range: a55aff980a4e235590c3af57ded2542512e2f9f6..aa59ca09a7db58af593e94f718d79b0bb5be49c9
Promotion branch: staging-promote/aa59ca09-23935258795
Base: staging-promote/4c9a985b-23931806540
Triggered by: Staging CI batch at 2026-04-03 05:32 UTC

Commits in this batch (11):

Current commits in this promotion (85)

Current base: staging-promote/9c6d8cb3-23875318792
Current head: staging-promote/aa59ca09-23935258795
Current range: origin/staging-promote/9c6d8cb3-23875318792..origin/staging-promote/aa59ca09-23935258795

Auto-updated by staging promotion metadata workflow

Waiting for gates:

  • Tests: pending
  • E2E: pending
  • Claude Code review: pending (will post comments on this PR)

Auto-created by staging-ci workflow

@github-actions github-actions Bot added scope: llm LLM integration scope: setup Onboarding / setup scope: docs Documentation size: M 50-199 changed lines risk: high Safety, secrets, auth, or critical infrastructure contributor: core 20+ merged PRs labels Apr 3, 2026
@claude

claude Bot commented Apr 3, 2026

Copy link
Copy Markdown

Code review

Found 2 issues:

  1. [MEDIUM:50] Log suppression with Once may indicate config is resolved multiple times

    src/config/llm.rs wraps debug logging in static LOG_LLM_BACKEND_RESOLUTION: Once = Once::new() to prevent duplicate logs. This is a reasonable workaround but suggests create_llm_provider() or config resolution is being called multiple times during startup. Worth investigating if this is a design smell that could be refactored to avoid re-resolving.

    https://github.com/anthropics/ironclaw/blob/ff3ba94b1b8f8490ca3b51f366107593e03a7e18/src/config/llm.rs#L13-L85

  2. [LOW:60] Two independent tokio::spawn() tasks in auth token persistence

    src/main.rs lines 748-757 and 761-776 spawn two independent async tasks: one to write the token to bootstrap .env, and one to delete the legacy DB copy. While they have no runtime dependencies and shouldn't cause issues, they could theoretically complete out of order. Consider consolidating into a single spawned task if order matters for observability or logging.

    https://github.com/anthropics/ironclaw/blob/ff3ba94b1b8f8490ca3b51f366107593e03a7e18/src/main.rs#L748-L778

Other observations:

  • Logging level changes (info → debug) correctly follow CLAUDE.md guidance
  • HTTP bind address default change (0.0.0.0 → 127.0.0.1) is a security improvement
  • No panics, logic errors, or performance issues detected

ilblackdragon and others added 21 commits April 2, 2026 23:25
* Move safety benches into ironclaw_safety crate

* Annotate benchmark JSON unwraps for panic check
)

The import-feature-gated help snapshots were stale after the --auto-approve
flag and ACP subcommand were added, causing CI snapshot test failures.

https://claude.ai/code/session_01Vfo7aFbVDpibtH7EsDAUSE

Co-authored-by: Claude <noreply@anthropic.com>
The OpenAI Codex provider was missing the sanitize_tool_messages() call
that all other providers use to rewrite orphaned tool results as user
messages. This caused HTTP 400 errors when conversation history contained
tool result messages whose corresponding tool calls were dropped during
context compaction or thread resume.

Adds the call to both complete() and complete_with_tools(), matching the
pattern used by rig_adapter, nearai_chat, and bedrock providers.

Fixes #1969

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Add a three-state permission model (AlwaysAllow / AskEachTime / Disabled)
backed by per-user DB settings so that safe built-in tools run without
approval prompts by default, while destructive tools still require
explicit user confirmation.

Key changes:
- PermissionState enum + TOOL_RISK_DEFAULTS tier map (permissions.rs)
- Security: tools returning UnlessAutoApproved (http, create_job,
  event_emit, routine_create, routine_update) default to AskEachTime
  to prevent SSRF, resource abuse, and privilege escalation
- Dispatcher filters Disabled tools and pre-approves AlwaysAllow tools;
  clears and re-populates session auto-approvals each iteration so
  permission downgrades take effect immediately; caches permissions at
  iteration 0 to avoid repeated DB round-trips
- Graceful DB failure handling: keeps existing session approvals on
  transient DB errors instead of clearing them
- "Always Approve" persists to DB across sessions (thread_ops.rs);
  defense-in-depth skips persist for ApprovalRequirement::Always tools
- tool_permission_set LLM tool (always requires approval to change);
  returns error when no settings store configured
- tool_list extended with builtin tools + permission state fields;
  filters by builtin_tool_names() to avoid duplicating WASM/MCP tools
- Web UI: Tools settings tab with 3-way toggle + lock icons;
  JSON error body for locked tool rejection
- Dot-separated tool name validation at registration time
- Playwright E2E: 6 test scenarios for permissions lifecycle
- seed_tool_permissions writes tier defaults to DB at startup with
  idempotency test

ApprovalRequirement::Always is an unbypassable hard floor — even
AlwaysAllow cannot bypass it. All DB operations scoped to user_id via
TenantScope for multi-tenant correctness.

Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
* feat(docker): publish ironclaw-worker image alongside ironclaw

Build and push nearaidev/ironclaw-worker from Dockerfile.worker in the
same workflow. Both images share the same version/sha/tag scheme.

This lets ironclaw-dind pull the pre-built worker image for sandbox
baking instead of cloning the repo and building from source.

[skip-regression-check]

* feat(docker): daily scheduled build of :staging from staging branch

* perf(docker): worker image copies binary from ironclaw image instead of rebuilding

* revert Dockerfile.worker changes, keep it building from source
…B-backed pairing, and OwnershipCache (#1898)

* feat(ownership): add OwnerId, Identity, UserRole, can_act_on types

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>

* fix(ownership): private OwnerId field, ResourceScope serde derives, fix doc comment

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>

* refactor(tenant): replace SystemScope::db() escape hatch with typed workspace_for_user(), fix stale variable names

- Add SystemScope::workspace_for_user() that wraps Workspace::new_with_db
- Remove SystemScope::db() which exposed the raw Arc<dyn Database>
- Update 3 callers (routine_engine.rs x2, heartbeat.rs x1) to use the new method
- Fix stale comment: "admin context" -> "system context" in SystemScope
- Rename `admin` bindings to `system` in agent_loop.rs for clarity

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>

* fix(tenant): rename stale admin binding to system_store in heartbeat.rs

* refactor(tenant): TenantScope/TenantCtx carry Identity, add with_identity() constructor and bridge new()

- TenantScope: replace `user_id: String` field with `identity: Identity`; add `with_identity()` preferred constructor; keep `new(user_id, db)` as Member-role bridge; add `identity()` accessor; all internal method bodies use `identity.owner_id.as_str()` in place of `&self.user_id`
- TenantCtx: replace `user_id: String` field with `identity: Identity`; update constructor signature; add `identity()` accessor; `user_id()` delegates to `identity.owner_id.as_str()`; cost/rate methods updated accordingly
- agent_loop: split `tenant_ctx(&str)` into bridge + new `tenant_ctx_with_identity(Identity)` which holds the full body; bridge delegates to avoid duplication

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>

* feat(db): add V16 tool scope, V17 channel_identities, V18 pairing_requests migrations

- PostgreSQL: V16__tool_scope.sql adds scope column to wasm_tools/dynamic_tools
- PostgreSQL: V17__channel_identities.sql creates channel identity resolution table
- PostgreSQL: V18__pairing_requests.sql creates pairing request table replacing file-based store
- libSQL SCHEMA: adds scope column to wasm_tools/dynamic_tools, channel_identities, pairing_requests tables
- libSQL INCREMENTAL_MIGRATIONS: versions 17-19 for existing databases
- IDEMPOTENT_ADD_COLUMN_MIGRATIONS: handles fresh-install/upgrade dual path for scope columns
- Runner updated to check ALL idempotent columns per version before skipping SQL
- Test: test_ownership_model_tables_created verifies all new tables/columns exist after migrations

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>

* fix(db): use correct RFC3339 timestamp default in libSQL, document version sequence offset

Replace datetime('now') with strftime('%Y-%m-%dT%H:%M:%fZ', 'now') in the
channel_identities and pairing_requests table definitions (both in SCHEMA and
INCREMENTAL_MIGRATIONS) to match the project-standard RFC 3339 timestamp format
with millisecond precision. Also add a comment clarifying that libSQL incremental
migration version numbers are independent from PostgreSQL VN migration numbers.

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>

* feat(ownership): bootstrap_ownership(), migrate_default_owner, V19 FK migration, replace hardcoded 'default' user IDs

- Add V19__ownership_fk.sql (programmatic-only, not in auto-migration sweep)
- Add `migrate_default_owner` to Database trait + both PgBackend and LibSqlBackend
- Add `get_or_create_user` default method to UserStore trait
- Add `bootstrap_ownership()` to app.rs, called in init_database() after connect_with_handles
- Replace hardcoded "default" owner_id in cli/config.rs, cli/mcp.rs, cli/mod.rs, orchestrator/mod.rs
- Add TODO(ownership) comments in llm/session.rs and tools/mcp/client.rs for deferred constructors

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>

* fix(ownership): atomic get_or_create_user, transactional migrate_default_owner, V19 FK inline constant, fix remaining 'default' user IDs

- Delete migrations/V19__ownership_fk.sql so refinery no longer auto-applies FK constraints before bootstrap_ownership runs; add OWNERSHIP_FK_SQL constant with TODO for future programmatic application
- Remove racy SELECT+INSERT default in UserStore::get_or_create_user; both PostgreSQL (ON CONFLICT DO NOTHING) and libSQL (INSERT OR IGNORE) now use atomic upserts
- Wrap migrate_default_owner in explicit transactions on both backends for atomicity
- Make bootstrap_ownership failure fatal (propagate error instead of warn-and-continue)
- Fix mcp auth/test --user: change from default_value="default" to Option<String> resolved from configured owner_id
- Replace hardcoded "default" user IDs in channels/wasm/setup.rs with config.owner_id
- Replace "default" sentinel in OrchestratorState test helper with "<unset>" to make the test-only nature explicit

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>

* fix(ownership): remove default user_id from create_job(), change sentinel strings to <unset>

- Gate ContextManager::create_job() behind #[cfg(test)]; production code must
  use create_job_for_user() with an explicit user_id to prevent DB rows with
  user_id = 'default' being silently created on the production write path.
- Change the placeholder user_id in McpClient::new(), new_with_name(), and
  new_with_config() from "default" to "<unset>" so accidental secrets/settings
  lookups surface immediately rather than silently touching the wrong DB partition.
- Same sentinel change for SessionManager::new() and new_async() in session.rs;
  these are overwritten by attach_store() at startup with the real owner_id.
- Update tests that asserted the old "default" sentinel to expect "<unset>", and
  switch test_list_jobs_tool / test_job_status_tool to create_job_for_user("default")
  to keep ownership alignment with JobContext::default().

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>

* feat(db): add ChannelPairingStore sub-trait with resolve_channel_identity, upsert/approve pairing, PostgreSQL + libSQL implementations

Adds PairingRequestRecord, ChannelPairingStore trait (5 methods), and
generate_pairing_code() to src/db/mod.rs; implements for PgBackend in
postgres.rs and LibSqlBackend in libsql/pairing.rs; wires ChannelPairingStore
into the Database supertrait bound; all 6 libSQL unit tests pass.

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>

* fix(db): atomic libSQL approve_pairing with BEGIN IMMEDIATE, add case-insensitive/expired/double-approve tests

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>

* feat(ownership): add OwnershipCache for zero-DB-read identity resolution on warm path

Converts src/ownership.rs to src/ownership/ module directory and adds
src/ownership/cache.rs with a write-through in-process cache mapping
(channel, external_id) -> Identity. Wired as Arc<OwnershipCache> on
AppComponents for Task 8 pairing integration. All 7 cache unit tests pass.

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>

* test(e2e): add ownership model E2E tests and extend pairing tests for DB-backed store

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>

* fix(e2e): remove unused asyncio import, add fallback assertion in test_pairing_response_structure

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>

* test(tenant): unit tests for TenantScope::with_identity and AdminScope construction

Adds 5 focused unit tests verifying TenantScope::with_identity stores the
full Identity (owner_id + role), TenantScope::new creates a Member-role
identity, and AdminScope::new returns Some for Admin and None for Member.
Uses LibSqlBackend::new_memory() as the test DB stub.

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>

* fix(ownership): recover from RwLock poison instead of expect() in OwnershipCache

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>

* test(ownership): integration tests for bootstrap, tenant isolation, and ChannelPairingStore

Adds tests/ownership_integration.rs covering migrate_default_owner idempotency,
TenantScope per-user setting isolation (including Admin role bypass check),
and the full ChannelPairingStore lifecycle (upsert, approve, remove, multi-channel isolation).

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>

* fix(test): remove duplicate pairing tests and flaky random-code assertion from integration suite

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>

* feat(pairing): rewrite PairingStore to DB-backed async with OwnershipCache

Replaces the file-based pairing store (~/.ironclaw/*-pairing.json,
*-allowFrom.json) with a DB-backed async implementation that delegates
to ChannelPairingStore and writes through to OwnershipCache on reads.

- PairingStore::new(db, cache) uses the DB; new_noop() for test/no-DB
- resolve_identity() cache-first lookup via OwnershipCache
- approve(code, owner_id) removes channel arg (DB looks up by code)
- All WASM host functions updated: pairing_upsert_request uses block_in_place,
  pairing-is-allowed renamed to pairing-resolve-identity returning Option<String>,
  pairing-read-allow-from deprecated (returns empty list)
- Signal channel receives PairingStore via new(config, db) constructor
- Web gateway pairing handlers read from state.store (DB) directly
- extensions.rs derive_activation_status drops PairingStore dependency;
  derives status from extension.active and owner_binding flag instead
- All test call sites updated to use new_noop()

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>

* fix(pairing): add missing pairing_store field to all GatewayState initializers, fix disk-full post-edit compile

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>

* feat(channels): remove owner_id from IncomingMessage, user_id is the canonical resolved OwnerId

`owner_id` on `IncomingMessage` was always a duplicate of `user_id` —
both fields held the same value at every call site. Remove the field and
`with_owner_id()` builder, update the four WASM-wrapper and HTTP test
assertions to use `user_id`, and drop the redundant struct literal field
in the routine_engine test helper.

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>

* fix(channels): remove stale owner_id param from make_message test helper

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>

* test(e2e): add browser/Playwright tests for ownership model — auth screen, chat UI, owner login

Adds five Playwright-based browser tests to the ownership model E2E suite
verifying the web UI experience: authenticated owner sees chat input, unauthenticated
browser sees auth screen, owner can send a message and receive a response, settings
tab renders without errors, and basic page structure is correct after login.

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>

* feat(settings): migrate channel credentials from plaintext settings to encrypted secrets store

Moves nearai.session_token from the plaintext DB settings table to the
AES-256-GCM encrypted secrets store (key: nearai_session_token).

- SessionManager gains an `attach_secrets()` method that wires in the
  secrets store; `save_session` writes to it when available and
  `load_session_from_secrets` is called preferentially over settings
- `migrate_session_credential()` runs idempotently on each startup in
  `init_secrets()`, reading the JSON session from settings, writing it
  to secrets, then deleting the plaintext copy
- Wizard's `persist_session_to_db` now writes to secrets first, falling
  back to plaintext settings only when secrets store is unavailable
- Plaintext settings path is preserved as fallback for installs without
  a secrets store (no master key configured)

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>

* fix(settings): settings fallback only when no secrets store, verify decryption before deleting plaintext

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>

* fix(ownership): ROLLBACK in libSQL migrate_default_owner, shared OwnershipCache across channels, add dynamic_tools to migration, fix doc comment

- libSQL migrate_default_owner: wrap UPDATE loop in async closure + match to emit ROLLBACK on any mid-transaction failure (mirroring approve_pairing pattern)
- Both backends: add dynamic_tools to the migrate_default_owner table list so agent-built tools are migrated on first pairing
- setup_wasm_channels: accept Arc<OwnershipCache> parameter instead of allocating a fresh cache, share the AppComponents cache
- SignalChannel::new: accept Arc<OwnershipCache> parameter and pass it to PairingStore instead of allocating a new cache
- PairingStore: fix module-level and struct-level doc comments to accurately describe lazy cache population after approve()

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>

* fix(web): use can_act_on for authorization in job/routine handlers instead of raw string comparisons

Replace 12 raw `user_id != user.user_id` / `user_id == user.user_id` string comparisons
in jobs.rs and 4 in routines.rs with calls through the canonical `can_act_on` function
from `crate::ownership`, which is the spec-mandated authorization mechanism.

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>

* chore: include remaining modified files in ownership model branch

* fix: add pairing_store field to test GatewayState initializers, update PairingStore API calls in integration tests

Add missing `pairing_store: None` to all GatewayState struct initializers
in test files. Migrate old file-based PairingStore API calls
(PairingStore::new(), PairingStore::with_base_dir()) to the new DB-backed
API (PairingStore::new_noop()). Rewrite pairing_integration.rs to use
LibSqlBackend with the new async DB-backed PairingStore API.

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>

* chore: cargo fmt

* fix(pairing): truly no-op PairingStore noop mode, ensure owner user in CLI, fix signal safety comments

- PairingStore::upsert_request now returns a dummy record in noop mode instead of
  erroring, and approve silently succeeds (matching the doc promise of "writes
  are silently discarded").
- PairingStore::approve now accepts a channel parameter, matching the updated
  DB trait signature and propagated to all call sites (CLI, web server, tests).
- CLI run_pairing_command ensures the owner user row exists before approval to
  satisfy the FK constraint on channel_identities.owner_id.
- Signal channel block_in_place safety comments corrected from "WASM channel
  callbacks" to "Signal channel message processing".

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>

* fix(pairing): thread channel through approve_pairing, add created flag, retry on code collision, remove redundant indexes

Addresses PR review comments:
- approve_pairing validates code belongs to the given channel
- PairingRequestRecord.created replaces timing heuristic
- upsert retries on UNIQUE violation (up to 3 attempts)
- redundant indexes removed (UNIQUE creates implicit index)

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>

* fix(ownership): migrate api_tokens, serialize PG approvals, propagate resolved owner_id

Addresses PR review P1/P2 regressions:

- api_tokens included in migrate_default_owner (both backends)
- PostgreSQL approve_pairing uses FOR UPDATE to prevent concurrent approvals
- Signal resolve_sender_identity returns owner_id, set as IncomingMessage.user_id
  with raw phone number preserved as sender_id for reply routing
- Feishu uses resolved owner_id from pairing_resolve_identity in emitted message
- PairingStore noop mode logs warning when pairing admission is impossible

[skip-regression-check]

Co-Authored-By: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>

* fix(pr-review): sanitize DB errors in pairing handlers, fix doc comments, add TODO for derive_activation_status

- Pairing list/approve handlers no longer leak DB error details to clients
- NotFound errors return user-friendly 'Invalid or expired pairing code' message
- Module doc in pairing/store.rs corrected (remove -> evict, no insert method)
- wit_compat.rs stub comment corrected to match actual Val shape
- TODO added for derive_activation_status has_paired approximation

* fix(pr-review): propagate libSQL query errors in approve_pairing, round-trip validate session credential migration, fix test doc comment

- libSQL approve_pairing: .ok().flatten() replaced with .map_err() to propagate DB errors
- migrate_session_credential: round-trip compares decrypted secret against plaintext before deleting
- ownership_integration.rs: doc comment corrected to match actual test coverage

* fix(pairing): store meta, wrap upserts in transactions, case-insensitive role/channel, log Signal DB errors, use auth role in handlers

- Store meta JSONB/TEXT column in pairing_requests (PG migration V18, libSQL schema + incremental migration 19)
- Wrap upsert_pairing_request in transactions (PG: client.transaction(), libSQL: BEGIN IMMEDIATE/COMMIT/ROLLBACK)
- Case-insensitive role parsing: eq_ignore_ascii_case("admin") in both backends
- Case-insensitive channel matching in approve_pairing: LOWER(channel) = LOWER($2)
- Log DB errors in Signal resolve_sender_identity instead of silently discarding
- Use auth role from UserIdentity in web handlers (jobs.rs, routines.rs) via identity_from_auth helper
- Fix variable shadowing: rename `let channel` to `let req_channel` in libsql approve_pairing

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(security): add auth to pairing list, cache eviction on deactivate, runtime assert in Signal, remove default fallback, warn on noop pairing codes

Addresses zmanian's review:
- #1: pairing_list_handler requires AuthenticatedUser
- #2: OwnershipCache.evict_user() evicts all entries for a user on suspension
- #3: debug_assert! for multi-thread runtime in Signal block_in_place
- #9: Noop PairingStore warns when generating unredeemable codes
- #10: cli/mcp.rs default fallback replaced with <unset>

* fix(pairing): consistent LOWER() channel matching in resolve_channel_identity, fix wizard doc comment, fix E2E test assertion for ActionResponse convention

* fix(pairing): apply LOWER() consistently across all ChannelPairingStore queries (upsert, list_pending, remove)

All channel matching now uses LOWER() in both PostgreSQL and libSQL backends:
- upsert_pairing_request: WHERE LOWER(channel) = LOWER($1)
- list_pending_pairings: WHERE LOWER(channel) = LOWER($1)
- remove_channel_identity: WHERE LOWER(channel) = LOWER($1)

Previously only resolve_channel_identity and approve_pairing used LOWER(),
causing inconsistent matching when channel names differed by case.

* fix(pairing): unify code challenge flow and harden web pairing

* test: harden pairing review follow-ups

* fix: guard wasm pairing callbacks by runtime flavor

* fix(pairing): normalize channel keys and serialize pg upserts

* chore(web): clean up ownership review follow-ups

* Preserve WASM pairing allowlist compatibility

---------

Co-authored-by: Claude Sonnet 4.6 (1M context) <noreply@anthropic.com>
* Fix turn cost footer and per-turn usage accounting

* Avoid panicking on poisoned turn usage mutex

* Tighten SSE usage regression coverage

* Report usage for interrupted turns
…tags (#1952)

* fix(llm): invert reasoning default — unknown models skip <think>/<final> injection

When NEAR AI model="auto" resolves server-side to Qwen 3.5, the system
prompt injected <think>/<final> tags because "auto" didn't match any
known native-thinking pattern. This caused empty responses:

1. Qwen 3.5's native thinking puts reasoning in a `reasoning` field
   (not `reasoning_content`) — silently dropped due to field name mismatch
2. Content contained only <think> tags or <tool_call> XML, which
   clean_response() stripped to empty → "I'm not sure how to respond"

Three fixes:
- Invert the default: new requires_think_final_tags() with empty allowlist
  means unknown/alias models get the safe direct-answer prompt
- Add #[serde(alias = "reasoning")] so vLLM's field name is accepted
- Update active_model from API response.model so capability checks
  use the resolved model name after the first call

Confirmed via direct API testing against NEAR AI staging with
Qwen/Qwen3.5-122B-A10B.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* remove model alias resolution from nearai_chat

auto should stay as the active model name — no reason to overwrite it
with the resolved model since requires_think_final_tags() returns false
for both "auto" and the resolved name.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix wording: remove native-thinking assumption from direct-answer prompt

The direct-answer prompt is now the default for all models, not just
native-thinking ones. Remove misleading "handled natively" language.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* Fix bootstrap ownership migration for dynamic tools

* Add bootstrap ownership regression coverage
* feat(embeddings): add bedrock provider

* refactor(embeddings): address review feedback

* fix: CI failures and review feedback for Bedrock embeddings

- Replace ENV_MUTEX.lock() with lock_env() to match staging's test pattern
- Use db_first_or_default() for non-bedrock model resolution (staging API)
- Validate returned embedding dimension matches configured dimension

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Illia Polosukhin <ilblackdragon@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(security): run safety checks on truncated tool output

Previously, `sanitize_tool_output()` returned immediately after
truncating oversized output, skipping leak detection, policy
enforcement, and injection scanning entirely. This allowed an attacker
to embed malicious payloads in the first N bytes of oversized tool
output and have them delivered unsanitized to the LLM.

Restructure the truncation path so it feeds into the same safety
pipeline as non-truncated content: leak detection, policy checks,
and Aho-Corasick injection scanning all run on the (possibly
truncated) content before it is returned.

Adds regression tests to verify truncated output is still scanned.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* style: apply rustfmt to fix CI formatting check

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Wui <wui@Wui-Work-2.local>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Illia Polosukhin <ilblackdragon@gmail.com>
…ding (#1652) (#1875)

- Verify primary user_id switches correctly and original workspace is unchanged
  - Verify private memory layers rescope to new user while shared layers stay intact
  - Verify secondary read scopes are preserved exactly and old primary is removed with no duplicates
  - Verify identity reads via both read_primary() and read() pin to new primary, not old or shared scopes
  - Verify non-identity reads still span preserved shared scopes after old primary removal
  - Verify bootstrap flags (pending + completed) preserve on same-user rebind and reset on different-user rebind
  - Establish precondition probes to ensure bootstrap reset tests validate true→false, not default false
…pair (#1991)

* fix(self-repair): skip built-in tools in broken tool detection and repair

Built-in tools (http, shell, json, etc.) are part of the ironclaw binary.
Errors on them are caller-side issues (bad LLM parameters), not tool
defects. The self-repair system was attempting to rebuild these via
SoftwareBuilder, wasting LLM tokens and spamming users with notifications.

- Add PROTECTED_TOOL_NAMES list covering all 47+ built-in tool names
- Export is_protected_tool_name() for use outside the registry module
- Filter built-in tools in detect_broken_tools() (primary guard)
- Add defense-in-depth guard in repair_broken_tool() (rejects builtins
  even if detection failed to filter them)
- Document SelfRepair trait contract about built-in tool exclusion
- Add regression tests: detect_broken_tools_filters_out_builtins (with
  real libSQL store), repair_broken_tool_skips_builtin (with mock builder),
  is_protected_tool_name_covers_common_builtins (spot-check)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: remove duplicate tool_info from PROTECTED_TOOL_NAMES

Addressed Gemini review: tool_info was listed twice (extension management
and incorrectly under image tools after merge conflict resolution).
Moved tool_permission_set under its own "Permission tools" category.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: add plan_update to PROTECTED_TOOL_NAMES

The plan_update builtin was missing from the protected list, allowing
the self-repair system to incorrectly attempt rebuilding it on errors.

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: use SystemScope in test after ownership model rebase

AdminScope::new now takes (Identity, Database) and returns Option.
The test needs SystemScope which has the simpler single-arg constructor.

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: j-bloggs <j-bloggs@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat(test): add dual-mode live/replay test harness with LLM judge

Add a general-purpose test infrastructure for running E2E tests in two modes:
- Live mode (IRONCLAW_LIVE_TEST=1): real LLM calls with real tools, records
  traces to disk for future replay
- Replay mode (default): loads saved trace fixtures, deterministic, no API keys

The harness uses Config::from_env() in live mode so the test agent mirrors
the real binary's behavior (engine_v2, allow_local_tools, approval gates).
Includes an LLM judge for semantic verification of non-deterministic output,
and saves human-readable session logs alongside trace fixtures for inspection
and diffing between live and replay runs.

First test case: zizmor security scanner against ironclaw's own workflows.

New files:
- tests/support/live_harness.rs — LiveTestHarness, builder, LLM judge
- tests/e2e_live.rs — zizmor_scan test
- tests/fixtures/llm_traces/live/ — recorded trace + session log

TestRigBuilder additions:
- with_http_interceptor() for injecting RecordingHttpInterceptor
- with_config() for real-binary config parity (respects allow_local_tools,
  engine_v2 from env instead of forcing test defaults)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* test(live): add engine v2 zizmor scan test, add with_engine_v2 to harness

Add zizmor_scan_v2 test that exercises the same scenario through engine v2.
Documents the current v2 limitation: auto_approve_tools config flag is not
honored by EffectBridgeAdapter — it only checks the per-session "always"
set, so shell calls pause at the approval gate.

Also:
- Add with_engine_v2() to LiveTestHarnessBuilder for config override
- Refactor v1 test to use shared run_zizmor_scan() helper
- V2 test has relaxed assertions matching current v2 behavior

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(test): address PR review feedback

- Fix UTF-8 unsafe string truncation in session log (use char_indices
  to find safe boundary instead of byte-index slicing)
- Remove forced auto_approve_tools(true) from LiveTestHarness build_live;
  let Config::from_env() drive it, with per-test override via new
  with_auto_approve_tools() builder method
- Apply engine_v2 builder override in TestRig's config-override branch
  so with_engine_v2() is not silently ignored when with_config() is used
- Remove unused timeout field and with_timeout() from LiveTestHarnessBuilder
- Tighten judge_response parsing to require strict PASS:/FAIL: prefix;
  anything else is treated as a failure with diagnostic message

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(agent): prevent self-repair notification spam for stuck jobs

When a stuck job exceeds max repair attempts, self-repair returns
ManualRequired but never transitions the job to a terminal state.
detect_stuck_jobs() re-finds it every cycle (~60s), sending a
Telegram notification each time — infinite spam.

Two-layer fix:
1. repair_stuck_job: transition to Failed before returning
   ManualRequired, so detect_stuck_jobs stops finding the job
2. Agent loop: HashSet dedup prevents duplicate ManualRequired
   notifications per job (defense-in-depth if transition fails)

Also adds Pending → Failed to the state machine — stuck Pending
jobs (dispatched but never started) could not be terminated.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(agent): handle transition failure in ManualRequired and adjust message

Log error if Failed transition fails, and adjust the ManualRequired
message to accurately reflect whether the job was marked failed or not.

Addresses gemini-code-assist feedback on PR #1867.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(agent): flatten nested Result, dedup Failed notifications, fix test state path

Adversarial review findings:
1. CRITICAL: update_context returns Result<Result<()>>. Using .is_ok()
   on the outer only checks job existence, not transition success.
   Fixed: matches!(result, Ok(Ok(()))).
2. IMPORTANT: RepairResult::Failed arm had same spam potential as
   ManualRequired. Applied same dedup pattern.
3. IMPORTANT: Test exercised Pending→Failed (bypassing production
   path). Now transitions through InProgress→Stuck→Failed.
4. Dropped unrelated Cargo.lock version bump.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* docs: update state machine diagram to include Pending -> Failed

Adversarial review finding: CLAUDE.md state diagram was the canonical
reference but didn't show the new Pending -> Failed transition.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: j-bloggs <j-bloggs@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(bridge): sanitize orphaned tool results in v2 adapter

* test(bridge): cover no-tool sanitizer path
* fix: harden WASM channel HTTP SSRF protections

* fix: clean up wasm http security helpers for clippy
* Add Telegram local regression test harness

* Add local Telegram smoke test runner

* test: add high-priority Telegram regression tests

Cover 6 previously untested API-level flows using fake axum Telegram
servers: photo attachment download, voice attachment download, long
message splitting (>4096 chars), Markdown parse error fallback to
plain text, sendChatAction typing indicator, and polling mode
(getUpdates with offset tracking).

Test count: 13 → 19. All use real WASM channel execution with
env-var URL override.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* test: add full-process Telegram E2E tests

Add 4 end-to-end tests that boot IronClaw, activate the Telegram WASM
channel via the setup API, POST webhook updates, and verify the
sendMessage round-trip through the mock LLM to a fake Telegram API.

Tests cover:
- DM round-trip (setup → webhook → LLM → sendMessage)
- Edited message handling
- Unauthorized user rejection (dm_policy = pairing)
- Invalid webhook secret rejection (401)

New files:
- fake_telegram_api.py: aiohttp server faking the Telegram Bot API
- test_telegram_e2e.py: the 4 test scenarios

conftest.py changes:
- Add fake_telegram_server and telegram_e2e_server fixtures
- Extend _wasm_build_symlinks to also cover channels-src/

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* test: expand Telegram E2E coverage with 8 new regression tests

Add 8 new tests covering core functionality gaps and high-priority
error resilience scenarios for the Telegram WASM channel:

Round 1 (functionality):
- Group mention filtering (ignore without @bot, reply with @bot)
- Long message chunking (>4096 chars split correctly)
- Polling mode roundtrip (getUpdates picks up queued messages)
- Markdown fallback (400 parse error triggers plain-text retry)

Round 2 (resilience):
- Missing webhook secret header (401 rejection)
- 429 rate limit resilience (system survives, recovers)
- Document download failure (getFile 500, text still processed)
- Malformed payload resilience (invalid JSON handled, bot continues)

Also extends fake_telegram_api.py with reject_markdown, rate_limit,
and fail_downloads simulation flags plus control endpoints, and adds
a "long response" canned pattern to mock_llm.py.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix: resolve CI failures in Telegram test suite

- Add #[cfg(feature = "integration")] gate to
  test_bot_mention_detection_case_insensitive and
  build_telegram_update_value (fixes compilation on default/libsql)
- Run cargo fmt on telegram_auth_integration.rs
- Fix race condition in fake_telegram_api.py get_updates
- Increase rate_limit_count from 5 to 20 for retry resilience
- Move helper functions to proper section in test_telegram_e2e.py

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
…2580

chore: promote staging to staging-promote/9399fccc-24203835221 (2026-04-09 18:19 UTC)
…5221

chore: promote staging to staging-promote/b819d704-24201344542 (2026-04-09 17:24 UTC)
…4542

chore: promote staging to staging-promote/af9b59a2-24198673070 (2026-04-09 16:27 UTC)
…6101

chore: promote staging to staging-promote/6895cdad-24185214226 (2026-04-09 13:31 UTC)
…4226

chore: promote staging to staging-promote/63a48e4e-24182836482 (2026-04-09 10:24 UTC)
…6482

chore: promote staging to staging-promote/288fe49a-24110798843 (2026-04-09 09:26 UTC)
…8843

chore: promote staging to staging-promote/79c1b0fd-24108317021 (2026-04-08 00:16 UTC)
…7021

chore: promote staging to staging-promote/86c15903-24100112892 (2026-04-07 22:53 UTC)
…2892

chore: promote staging to staging-promote/00fd2e88-24092158668 (2026-04-07 19:23 UTC)
…8668

chore: promote staging to staging-promote/f765958f-24078644272 (2026-04-07 16:22 UTC)
…4272

chore: promote staging to staging-promote/13774cc0-24076446124 (2026-04-07 11:19 UTC)
…3070

chore: promote staging to staging-promote/13c458e3-24192916101 (2026-04-09 15:30 UTC)
…6124

chore: promote staging to staging-promote/6fa2d0ec-24071779175 (2026-04-07 10:20 UTC)
…9175

chore: promote staging to staging-promote/0ab1a474-24064762647 (2026-04-07 08:24 UTC)
…2647

chore: promote staging to staging-promote/5d6a247d-24057733730 (2026-04-07 04:46 UTC)
…3730

chore: promote staging to staging-promote/8b629851-24041939370 (2026-04-07 00:16 UTC)
…9370

chore: promote staging to staging-promote/9cf37364-24039632441 (2026-04-06 17:15 UTC)
…2441

chore: promote staging to staging-promote/d0096dfc-24035427523 (2026-04-06 16:13 UTC)
…7523

chore: promote staging to staging-promote/f9ed8152-24023233420 (2026-04-06 14:19 UTC)
…3420

chore: promote staging to staging-promote/13852ff5-24021660555 (2026-04-06 07:31 UTC)
…0555

chore: promote staging to staging-promote/5083aed4-24002418644 (2026-04-06 06:33 UTC)
…8644

chore: promote staging to staging-promote/733678dd-23996777140 (2026-04-05 13:23 UTC)
…7140

chore: promote staging to staging-promote/e1695914-23995878265 (2026-04-05 07:24 UTC)
…8265

chore: promote staging to staging-promote/f3036388-23995135957 (2026-04-05 06:24 UTC)
@github-actions github-actions Bot added scope: tool/wasm WASM tool sandbox scope: tool/builder Dynamic tool builder scope: secrets Secrets management labels Apr 10, 2026
@henrypark133
henrypark133 merged commit 0facf16 into staging-promote/9c6d8cb3-23875318792 Apr 10, 2026
13 of 14 checks passed
@henrypark133
henrypark133 deleted the staging-promote/aa59ca09-23935258795 branch April 10, 2026 23:10
theredspoon pushed a commit to theredspoon/ironclaw that referenced this pull request Jun 21, 2026
…3935258795

chore: promote staging to staging-promote/fa7d94d4-23931806540 (2026-04-03 05:32 UTC)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: high Safety, secrets, auth, or critical infrastructure scope: agent Agent core (agent loop, router, scheduler) scope: channel/cli TUI / CLI channel scope: channel/wasm WASM channel runtime scope: channel/web Web gateway channel scope: channel Channel infrastructure scope: ci CI/CD workflows scope: config Configuration scope: db/libsql libSQL / Turso backend scope: db/postgres PostgreSQL backend scope: db Database trait / abstraction scope: dependencies Dependency updates scope: docs Documentation scope: extensions Extension management scope: llm LLM integration scope: orchestrator Container orchestrator scope: pairing Pairing mode scope: sandbox Docker sandbox scope: secrets Secrets management scope: setup Onboarding / setup scope: tool/builder Dynamic tool builder scope: tool/builtin Built-in tools scope: tool/mcp MCP client scope: tool/wasm WASM tool sandbox scope: tool Tool infrastructure scope: worker Container worker scope: workspace Persistent memory / workspace size: XL 500+ changed lines staging-promotion

Projects

None yet

Development

Successfully merging this pull request may close these issues.