Skip to content

chore: promote staging to staging-promote/9ba10eac-23686921981 (2026-03-28 15:07 UTC) - #1727

Merged
henrypark133 merged 36 commits into
staging-promote/9ba10eac-23686921981from
staging-promote/9bb19a98-23687925861
Mar 30, 2026
Merged

henrypark133 merged 36 commits into
staging-promote/9ba10eac-23686921981from
staging-promote/9bb19a98-23687925861

Conversation

@ironclaw-ci

@ironclaw-ci ironclaw-ci Bot commented Mar 28, 2026 •

Copy link
Copy Markdown
Contributor

Auto-promotion from staging CI

Batch range: 2f4eb08613cefff1af8b7b1a475fda00c84dd855..9bb19a98f767f2e4db0d4f3dda0e495a356a726a
Promotion branch: staging-promote/9bb19a98-23687925861
Base: staging-promote/9ba10eac-23686921981
Triggered by: Staging CI batch at 2026-03-28 15:07 UTC

Commits in this batch (6):

Current commits in this promotion (2)

Current base: staging-promote/9ba10eac-23686921981
Current head: staging-promote/9bb19a98-23687925861
Current range: origin/staging-promote/9ba10eac-23686921981..origin/staging-promote/9bb19a98-23687925861

Auto-updated by staging promotion metadata workflow

Waiting for gates:

  • Tests: pending
  • E2E: pending
  • Claude Code review: pending (will post comments on this PR)

Auto-created by staging-ci workflow

Achieve added 2 commits March 28, 2026 15:10
)

* fix(oauth): tighten legacy state validation and fallback handling

* style: fix formatting

* refactor: separate validation checks for clearer error messages
@github-actions github-actions Bot added scope: channel/cli TUI / CLI channel scope: channel/web Web gateway channel size: M 50-199 changed lines risk: medium Business logic, config, or moderate-risk modules contributor: core 20+ merged PRs labels Mar 28, 2026
Achieve and others added 22 commits March 28, 2026 16:31
- Implement broadcast_dm() that creates a DM channel with the target
  user (POST /users/@me/channels, cached by Discord) and sends the
  message to it
- Extract DISCORD_API_BASE constant for all Discord REST API URLs
- Extract send_channel_message() shared helper to deduplicate message
  posting between on_respond and broadcast_dm
- Add snowflake validation on user_id before API calls
- Fix pre-existing clippy redundant_closure warning
- Use typed DmChannelResponse struct instead of serde_json::Value

Closes no specific issue — completes the previously stubbed on_broadcast.

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…PTY (#1678)

- Add pty-process crate (MIT, tokio async support) for PTY allocation
- Spawn claude CLI with pty-process::Command::arg() chaining instead of
  building a shell string for script -qfc
- Eliminates all shell injection surfaces: prompt, model, session_id
  are passed via execve, never interpreted by a shell
- Keep stderr on separate pipe to prevent NDJSON parse breakage
  (pty-process attaches PTY to all fds by default)
- Gate PTY behind #[cfg(unix)] with direct-spawn fallback for Windows CI
- Read stdout from PTY master (implements tokio::io::AsyncRead)
- Add regression tests: arg vector construction + PTY allocation

Addresses review feedback from zmanian and gemini-code-assist.

Co-authored-by: j-bloggs <j-bloggs@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…#1677)

* fix(worker): treat empty LLM response after text output as completion

When a job's LLM produces a substantive text response (e.g., formatted
results from a routine) and the next LLM call returns empty or errors,
the worker now treats this as successful completion instead of
continuing the loop until failure.

Previously, empty responses always triggered TextAction::Continue,
causing the loop to re-call the LLM. The LLM had nothing more to say,
so the provider returned "Response contained no message or tool call
(empty)". This made routine jobs that successfully produced results
report as "failed".

The fix adds a `has_text_response` flag to JobDelegate:
- After any non-empty text response: flag is set
- Empty text after flag is set: treated as completion
- LLM errors (select_tools/respond_with_tools) after flag: treated
  as completion instead of propagating
- Empty text before any output: still retries (rate-limit backoff)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(worker): restrict error swallowing to EmptyResponse variant only

- Add LlmError::EmptyResponse variant for when LLM returns no content
- Update nearai_chat and github_copilot providers to emit EmptyResponse
  instead of InvalidResponse for empty/no-choice responses
- try_complete_on_error now only swallows EmptyResponse (not AuthFailed,
  ContextLengthExceeded, Http, Io, etc.)
- Extract is_completion_eligible_error as testable pure function
- Log mark_completed errors at warn level instead of silently dropping
- Add EmptyResponse to retry and circuit breaker transient classifications
- Rewrite test to exercise real classification logic against all variants

Addresses review feedback from zmanian and gemini-code-assist.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* refactor(worker): extract mark_completed_or_warn helper to DRY completion logic

Extract shared mark-completed + warn-on-failure pattern into a single
helper method used by both try_complete_on_error and handle_text_response.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: j-bloggs <j-bloggs@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(routines): persist full LLM transcript and remove sandbox gate for full_job

Routine execution output was invisible — routine_fire returned a one-liner,
routine_history had no actual output, and the conversation thread contained
only a summary. Full-job routines also hard-failed without Docker.

Three fixes:

1. **Full transcript persistence**: execute_lightweight now persists every
   message (prompt, LLM responses, tool calls with params, tool results) to
   the routine's conversation thread as it executes, not just a summary
   after the fact.

2. **Routine output visibility**: routine_history includes conversation_id
   and recent_output messages. routine_fire tells the user to check
   routine_history. Web detail page has a "View Execution Thread" button
   that navigates to the chat tab. ROUTINE_OK stores "No issues found"
   instead of None. Full-job summary pulls actual job output instead of
   generic "Job X finished".

3. **Remove SandboxReadiness gate**: full_job routines dispatch through the
   scheduler like regular /job commands — no Docker required. The
   SandboxReadiness enum is removed entirely.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* style: apply cargo fmt

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(worker): treat AutonomousUnavailable tool errors as recoverable

The job worker crashed the entire job when a tool was denied for
autonomous execution (e.g. secret_list). The error was already recorded
in reason_ctx for the LLM to see, but process_tool_result_job returned
Err which propagated through the agentic loop and terminated the job.

Now all tool errors (including AutonomousUnavailable) return Ok,
letting the LLM see the denial and try a different approach.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(llm): sanitize tool names for OpenAI Codex Responses API

The Codex API requires tool names to match `^[a-zA-Z0-9_-]+$` but
MCP/extension tools can have dots in their names (e.g. `mcp.server.tool`).
This caused HTTP 400 errors when the job worker sent tool calls back
to the LLM.

Sanitize tool names in both `convert_tool_definition` and
`convert_message` (function_call items) by replacing invalid characters
with underscores.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(routines): inject execution context into full_job description [skip-regression-check]

When a full_job routine dispatches a job, the LLM had no context that
it was already executing inside a routine. It wasted iterations on
infrastructure (discovering tools, creating routines, setting up auth)
instead of doing the actual work.

Prepend a clear directive to the job description telling the LLM that
tools and the routine are already configured, and to execute the task
directly.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(mcp): auto-refresh expired OAuth tokens on access [skip-regression-check]

When IronClaw restarts, MCP servers fail with "Secret has expired"
because get_access_token() checks token expiry locally and returns an
error before any HTTP request is made — so the existing 401-retry
refresh logic never triggers.

Now get_access_token() catches SecretError::Expired and automatically
calls refresh_access_token() using the stored refresh token. If the
refresh succeeds, the new token is returned transparently. If it fails,
the error message includes both the expiry and the refresh failure.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(mcp): align refresh token naming and set expiry on stored tokens

Two bugs prevented MCP OAuth token auto-refresh on restart:

1. Naming mismatch: the hosted OAuth flow stored the refresh token as
   `{token_secret_name}_refresh_token` (e.g. `mcp_notion_access_token_refresh_token`)
   but `McpServerConfig::refresh_token_secret_name()` returned
   `mcp_notion_refresh_token`. The refresh token was there but unfindable.

2. Missing expiry: `store_tokens` in auth.rs never called `with_expiry()`
   even though `AccessToken::expires_in` was available. Combined with the
   fix from the previous commit (auto-refresh on Expired), tokens stored
   via the MCP auth flow will now also trigger refresh correctly.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(web): show activity and transitions for agent jobs in job detail [skip-regression-check]

The job events endpoint only checked sandbox jobs for ownership,
returning 404 for agent jobs dispatched from routines. The detail
handler also returned empty transitions for agent jobs.

- events handler: fall back to agent job ownership check
- detail handler: populate transitions from job's state history

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(routines): expose max_iterations for full_job routines (default 25)

The max_iterations parameter was hardcoded to 10 and not configurable
via routine_create or routine_update, causing complex tasks to hit the
iteration cap.

- Add max_iterations to full_job execution schema (1-200, default 25)
- Thread it through parse → build → RoutineAction
- Support updating via routine_update
- Raise default from 10 to 25

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(routines): break self-dialogue loop after full_job plan execution

After plan execution, the completion-check Q&A ("Is the job complete?" /
"No, not complete...") was left in the message context, causing the
agentic loop to repeat the same analysis instead of calling tools.

Replace the stale dialogue with an action-oriented continuation prompt
that instructs the LLM to use tools for remaining work. Also strip
<suggestions> tags from all job output since they're only meaningful
for interactive chat sessions.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(repl): prevent test hang in single-message mode

In single-message mode, start() stored a clone of the mpsc sender in
self.msg_tx for approval injection. After the thread sent /quit and
exited, the stored clone kept the stream alive, so stream.next()
blocked forever in the test assertion that the stream ends.

Skip storing the sender in single-message mode since interactive
approval is not needed.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(jobs): treat text responses as final answer in agentic loop

When the LLM produces a non-empty text response with no tool intent
(already filtered by the nudge mechanism), it is the job's final
answer. Previously, handle_text_response only exited the loop if the
text matched rigid completion phrases like "job is complete". Natural
summaries like "Weekly review completed and saved to Notion" were
added to context and the loop continued, causing the LLM to restate
the same summary until max_iterations was hit.

Now any non-empty text response marks the job complete and stops the
loop, matching the chat dispatcher behavior.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* perf(tests): reduce skills catalog network failure test from 10s to 1s

The test_search_returns_error_on_network_failure test connects to an
unreachable RFC 5737 TEST-NET IP and waited for the full 10s production
REQUEST_TIMEOUT. Add with_url_and_timeout test helper and use a 1s
timeout instead. [skip-regression-check]

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(tools): accept 'message' as alias for 'content' in message tool

LLMs frequently call the message tool with {"message": "..."} instead
of {"content": "..."}. Fall back to the 'message' key when 'content'
is missing to avoid InvalidParameters errors during autonomous job
execution.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(tools): attach thread_id for gateway broadcast in message tool

When the message tool broadcasts to all channels (channel=null), it
sent an OutgoingResponse without a thread_id. The gateway silently
dropped these messages (returned Ok but never sent the SSE event),
so they appeared in repl but not in the web UI.

The thread_id was only populated when channel was explicitly "gateway".
Now it is always populated from notify_thread_id metadata, so
broadcast_all delivers to the gateway correctly.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(gateway): return error instead of silently dropping messages

Gateway broadcast() and respond() previously returned Ok(()) when
thread_id was missing, silently swallowing the message. Callers
(message tool, agent loop) believed delivery succeeded when it didn't.

Now returns ChannelError::MissingRoutingTarget so callers can detect
and report the failure. Four regression tests verify the contract:
respond/broadcast with and without thread_id.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: resolve rebase conflicts with staging

Restore sandbox_readiness field removed by pre-rebase commits (staging
still uses it). Update repl test to match staging's single-message
behavior (no longer sends /quit). Add missing reasoning field to
ToolCall in codex test.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(tools): log error when routine conversation lookup fails

The routine_history tool silently swallowed errors from
get_or_create_routine_conversation, returning empty output without
any diagnostic logging. Add tracing::warn so failures are visible
in logs. [skip-regression-check]

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: address PR #1650 review comments

- E2E test: accept submitted/accepted as success states in job assertion
- TimeTool: remove operation from required schema (defaults to "now")
- jobs handler: log DB errors server-side, return generic message to client
- routines handler: use read-only find_routine_conversation on GET
- codex provider: reverse-map sanitized tool names so MCP tools resolve

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: address zmanian review feedback on PR #1650

- MCP refresh token: fall back to legacy secret name (mcp_{name}_refresh_token)
  so existing users don't need to re-authenticate after the naming fix
- Job worker: replace fragile messages.pop() with truncate-to-saved-count
  to avoid maintenance hazard if message flow changes
- Document cost implications of max_iterations 10->25 default bump
- Revert Cargo.toml dist profile change (thin LTO comment, codegen-units=16)
  as it's unrelated to this PR

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: resolve rebase conflicts and address new Copilot comments

- Fix no_silent_drop tests for updated GatewayConfig (user_id moved to
  GatewayChannel::new second arg, user_tokens removed)
- Fix handle_text_response param name (_reason_ctx -> reason_ctx)
- Fix missing has_text_response field in test JobDelegate
- Propagate row.get errors in find_routine_conversation instead of
  unwrap_or_default
- Only fall back to legacy refresh token name on NotFound/Expired,
  propagate real errors (DB, decryption)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat(discord): restore gateway channel flow in wasm

* chore(discord): bump channel version to 0.2.1

* fix(discord): address review feedback on gateway channel PR

- Add #[serde(default)] to DiscordMessageMetadata for backward compat
  with old Option<String> serialized metadata
- Restore mention polling alongside Gateway (on_poll processes gateway
  events first, then runs poll_for_mentions if configured)
- Update on_respond to handle source_message_id with message_reference
  for mention-poll reply threading
- Implement Gateway presence status: dnd before pairing, online after
- Implement Gateway resume (OP 6) with session_id tracking, falling
  back to fresh identify on Invalid Session (OP 9)
- Extract WebsocketSessionState and spawn_websocket_poll to reduce
  nesting in start_websocket_runtime
- Simplify should_apply_dm_pairing tautology
- Remove completed plan docs
- Fix clippy items_after_test_module in extensions handler

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(discord): address review findings in gateway channel PR

- Fix gateway presence always showing "online" by filtering empty
  owner_id strings from workspace store reads
- Fix interaction followup using POST instead of PATCH to
  /messages/@original, which left deferred "thinking" state unresolved
- Restore mention-poll pagination (up to 5 pages of 100 messages)
- Remove dead ed25519-dalek and hex dependencies from WASM crate
- Remove unused _channel_id parameter from remember_processed_id
- Clean up redundant let binding in send_pairing_reply

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(discord): address second-round review findings

- Log warning when gateway event queue JSON fails to deserialize
  instead of silently returning empty (zmanian review item 1)
- Defer presence update from OP 10 Hello to after OP 0 READY, per
  Discord gateway protocol which requires READY before non-Identify
  commands (zmanian review item 2)
- Add 0-25% random jitter to websocket reconnect backoff per Discord's
  reconnection recommendations (zmanian suggestion)
- Extract WebsocketPollContext struct to replace 19-parameter
  spawn_websocket_poll function (zmanian suggestion)
- Document intent bitmask 4609 = GUILDS + GUILD_MESSAGES +
  DIRECT_MESSAGES in capabilities JSON (zmanian suggestion)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: zhyaoyu <zhyaoyu@aliyun.com>
Co-authored-by: ilblackdragon@gmail.com <ilblackdragon@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…#1667)

Add support for bundle layouts where directories without SKILL.md are
recursed into to find nested skills (e.g., skills/my-org/skill-a/SKILL.md).

- Add configurable max_scan_depth (SKILLS_MAX_SCAN_DEPTH env, default 3)
- Recurse into subdirectories lacking SKILL.md up to depth limit
- Share remaining discovery cap across recursive levels
- Replace try_exists + read_dir with single read_dir (eliminates TOCTOU)
- Box::pin recursive async calls for correct future sizing

Closes #1664

Co-authored-by: Rajul Bhatnagar <brajul@amazon.com>
* Clarify message tool and channel setup guidance

* Add target format hints to proactive messaging prompt

* Clarify search and message tool edge cases

* Fix stale tool_search e2e assertion

* Update src/llm/reasoning.rs

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>

* Tighten prompt guidance for message replies

* Format prompt guidance assertions

---------

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* fix: pin staging ci jobs to a single tested sha

* chore(ci): retrigger regression gate

[skip-regression-check]

---------

Co-authored-by: Firat Sertgoz <f@nuff.tech>
* fix: prevent UTF-8 panics in byte-index string truncation

Replace unsafe `&s[..n]` patterns with `floor_char_boundary(s, n)` at 3
production code sites where the truncation index could land mid-multibyte
character, panicking on non-ASCII input:

- src/llm/nearai_chat.rs: API response truncation in error message
- src/cli/memory.rs: memory content display truncation
- src/cli/config.rs: config value display truncation

All 3 sites operate on external or user-supplied strings that may contain
non-ASCII characters. The existing `crate::util::floor_char_boundary`
utility (used at 18 other call sites) walks back to the nearest char
boundary, preventing the panic.

Adds regression test with multi-byte characters (combining accents and
4-byte emoji) for truncate_content.

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: clarify test comment and use exact assertions

Address Gemini review feedback:
- Fix misleading comment: \u{00e9} is precomposed e-acute, not combining accent
- Replace weak assertions (ends_with/is_empty) with exact assert_eq!

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…nt (#1630)

* fix(bedrock): strip tool blocks from messages when toolConfig is absent

Bedrock's Converse API requires `toolConfig` whenever messages contain
`toolUse` or `toolResult` content blocks. When the agentic loop reaches
its force_text iteration (e.g. lightweight routine at max_iterations),
it switches from `complete_with_tools()` to `complete()` — but the
message history still carries tool blocks from prior iterations.

`convert_messages()` faithfully converts these into Bedrock content
blocks, and without `toolConfig` Bedrock rejects the request:

  "The toolConfig field must be defined when using toolUse and
   toolResult content blocks."

Add `strip_tool_blocks()` that converts tool interaction data to text:
- Assistant `tool_calls` → dropped (text content preserved)
- `Role::Tool` → `Role::User` with `[Tool ... returned: ...]` text

Wire it into:
- `complete()`: unconditionally, since it never sends toolConfig
- `complete_with_tools()`: when `build_tool_config()` returns None
  (empty tools or tool_choice="none")

Closes #1629

* fix(bedrock): address review feedback on strip_tool_blocks

- Add tracing::debug\! when tool blocks are stripped (zmanian suggestion)
- Add test for tool_choice="none" path (zmanian suggestion)
- Add inline comment on empty-content assistant behavior

---------

Co-authored-by: Rajul Bhatnagar <brajul@amazon.com>
* feat: support custom LLM provider configuration via web UI

Users can now define custom LLM providers through the web UI and have
them take effect without modifying environment variables or config files.

- Add `CustomLlmProviderSettings` struct and `llm_custom_providers`
  field to `Settings` so custom provider definitions are persisted and
  loaded from the DB settings table
- Add `LlmConfig::resolve_custom_provider()` to build a
  `RegistryProviderConfig` from user-defined provider data (base_url,
  adapter, model, api_key)
- Flip resolution priority to `db > env > default` so active provider
  set through the UI takes precedence over deployment env vars
- Warn when a custom provider is missing base_url or model
- Add startup info logs for backend source and provider creation
- Add regression tests for custom provider resolution and DB priority

* feat: add test connection for custom LLM providers

- Add POST /api/llm/test_connection endpoint that validates
  connectivity and auth for OpenAI-compatible, Anthropic, and
  Ollama adapters (10s timeout, per-adapter request logic)
- Add "Test" button next to Save/Cancel in the add-provider form;
  result shown inline with green/red styling
- Hide delete button for the active provider instead of showing
  an error toast
- Sort the active provider to the top of the provider list
- Clear selected_model when switching providers to avoid
  model-not-supported errors on the new provider
- Add i18n keys for test/testing states (en + zh-CN)

* feat: add built-in provider API key and model configuration

- Add Configure button on built-in provider cards (openai, anthropic,
  gemini, ollama, etc.) to set API key and default model via web UI
- Store overrides as `llm_builtin_overrides` setting (per-provider
  key/model map) using the existing generic settings k/v API
- Add LlmBuiltinOverride struct in settings.rs; resolve in
  resolve_registry_provider() with priority:
  env var > selected_model > llm_builtin_overrides[id] > default
- Restore provider's configured model to selected_model on provider
  switch, so /model command always takes precedence at runtime
- Fix fetch-models button in built-in configure mode: use hardcoded
  base_url from BUILTIN_PROVIDERS instead of the hidden form field
- Add edit support for custom providers with pre-filled dialog
- Show current model on active and configured provider cards
- Convert add/edit provider form to a modal dialog
- Sync selected_model when editing or deleting an active custom provider

* feat: move Config tab into Settings as Providers subtab

* feat(web): merge Providers into Inference tab with UX improvements

* chore: resolve conflicts

* fix(llm): address security and correctness issues in custom LLM provider

* fix(llm): address security and correctness issues in custom LLM provider

* feat(web): fall back to env vars for LLM provider config in UI

* fix(llm): enforce db > env > default config priority for provider setting

* fix: address review feedback on provider config priority

* feat: extract BUILTIN_PROVIDERS into providers.js

* fix(security): store LLM API keys in encrypted secrets store instead of plaintext

* fix(security): harden LLM API key handling across settings and LLM endpoints

* fix: test_connection sends actual chat completion

* refactor(web): derive LLM Provider display from active Model Provider

* fix(settings): language switch not working for llm provider

* feat(web): add restart notice to LLM Provider settings

* fix: review fixes for custom LLM provider PR

- Add server-side validation of custom provider ID format (lowercase
  alphanumeric + hyphens, 1-64 chars) to match frontend regex
- Tighten is_nearai_private_endpoint to exact-match private.near.ai
  or *.private.near.ai, rejecting lookalikes like private-evil.near.ai
- Fix misleading priority doc comments in config/mod.rs and settings.rs
  to reflect the split model: LLM uses DB > env, others use env > DB
- Clean up #1581 artifacts: remove TOML file creation from
  persist_selected_model (DB is sufficient), update stale priority
  comments in commands.rs, fix contradictory test assertions
- Add 18 new tests for provider ID validation, adapter validation,
  and nearai private endpoint matching

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: address PR review comments for custom LLM provider

- Move LLM handlers (test_connection, list_models, env_defaults) from
  server.rs to handlers/llm.rs for consistency with other handler modules
- Merge validate_custom_providers into single pass (ID + adapter check)
- Allow underscores in custom provider IDs to match builtin naming
- Add missing i18n key config.fetchingModels (en + zh-CN)
- Fix optional_env().ok().flatten() error swallowing in config/llm.rs;
  propagate ConfigError with ? instead of silently discarding
- Narrow settings.rs module docs to scope DB>env precedence to LLM
- Add unit tests for hydrate_llm_keys_from_secrets

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* refactor: replace static providers.js with API endpoint from registry

- Delete providers.js; serve provider list from /api/llm/providers
  endpoint that reads from the embedded ProviderRegistry (providers.json)
- Centralize secret naming (builtin_secret_name, custom_secret_name)
  into settings.rs; replace 8 duplicated format! calls across 4 files
- Extract JS API_KEY_UNCHANGED constant; replace 6 magic string literals
- Replace hard-coded API key placeholder strings with i18n keys
  (config.apiKeyConfigured, config.apiKeyFromEnv, config.apiKeyEnter)
- Simplify apiFetchVoid to delegate to apiFetch
- Remove unnecessary Vec clones in guard_active_provider_not_removed

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Robert Yan <46699230+think-in-universe@users.noreply.github.com>
Co-authored-by: Illia Polosukhin <ilblackdragon@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* Handle empty tool completions in autonomous jobs

* Address malformed tool recovery review comments

* style: apply rustfmt to reasoning tests

---------

Co-authored-by: Firat Sertgoz <f@nuff.tech>
…1463)

* feat(gateway): add OIDC JWT authentication for reverse-proxy deployments

Add an optional OIDC JWT auth path to the web gateway, enabling
deployments behind identity-aware proxies like AWS ALB with Okta/Cognito.
When GATEWAY_OIDC_ENABLED=true, the gateway reads a signed JWT from a
configurable HTTP header (default: x-amzn-oidc-data), fetches the
signing key from a JWKS endpoint, and verifies the signature + claims.

Auth flow: Bearer token → OIDC JWT → query-string token → 401.

Key design decisions:
- Split signature verification from claim extraction to handle AWS ALB's
  non-standard base64 padding (ALB includes '=' padding in JWT segments,
  but jsonwebtoken's decode() strips it, changing the signing input).
  We verify against the original token text, then extract claims from a
  normalized copy.
- JWKS keys cached for 1 hour with per-kid granularity.
- Supports both ALB-style per-key PEM URLs ({kid} placeholder) and
  standard JWKS endpoints.
- DER-to-raw ECDSA signature conversion for IdPs that use DER encoding.
- Frontend auto-detects proxy auth via /api/gateway/status probe,
  skipping the login screen when OIDC is active.

Configuration (env vars):
  GATEWAY_OIDC_ENABLED=true
  GATEWAY_OIDC_HEADER=x-amzn-oidc-data  (default)
  GATEWAY_OIDC_JWKS_URL=https://public-keys.auth.elb.us-east-1.amazonaws.com/{kid}
  GATEWAY_OIDC_ISSUER=https://example.okta.com  (optional)
  GATEWAY_OIDC_AUDIENCE=my-client-id  (optional)

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* Address code review feedback on OIDC auth PR

- Add EdDSA PEM key parsing support (was falling through to RSA)
- Fix issuer validation: remove set_issuer(&[]) else branch that
  rejected all tokens when GATEWAY_OIDC_ISSUER is unset
- Make missing `sub` claim a validation error instead of silently
  defaulting to "unknown"
- Extract initApp() in app.js so OIDC auto-auth actually initializes
  the UI (was calling undefined function)
- Add regression tests for sub claim and issuer validation fixes

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* Harden OIDC auth: address claude[bot] security review

- SSRF: URL-encode kid before substituting into JWKS URL template
- Cache bounds: cap key cache at 64 entries, evict expired + oldest
- DER parsing: support long-form length encoding (>= 128 bytes),
  validate component lengths against expected curve size
- Production safety: replace .expect() with Result in OidcState::from_config
- Fetch backoff: cache failed JWKS fetches for 10s to prevent retry storms
- Body limit: cap JWKS responses at 256 KB to prevent OOM from rogue endpoint
- Add regression tests for DER long-form, kid encoding, cache bounds

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* test(auth): add regression test for OIDC identity resolution

Add two integration tests that exercise the full OIDC middleware path
through to AuthenticatedUser extraction:

- test_oidc_auth_inserts_user_identity_for_handler: sends a valid OIDC
  JWT through the middleware and verifies the handler receives the sub
  claim as user_id. Returns 401 if identity insertion is missing —
  verified by temporarily removing the insert and confirming failure.

- test_oidc_auth_user_gets_member_role: confirms OIDC-authenticated
  users receive role=member (not admin).

Uses a seed_key() test helper on OidcState to pre-populate the key
cache with an HS256 secret, avoiding the need for an HTTP JWKS mock.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* test(auth): comprehensive OIDC test coverage for edge cases

Add 17 new OIDC tests covering middleware integration, auth priority,
invalid JWTs, issuer/audience validation, and key cache behavior:

Middleware auth priority & fallthrough:
- Bearer works when OIDC configured but header absent
- Bearer takes priority when both Bearer and OIDC header present
- Bad OIDC signature returns 401 (not 500)
- Invalid OIDC doesn't block valid bearer auth
- No auth at all with OIDC configured → 401

Expired / invalid JWT edge cases:
- Expired JWT (exp in the past) rejected
- JWT without kid header rejected
- Malformed JWTs rejected (empty, 2-part, 4-part, garbage)
- Non-string sub claim (integer) rejected
- Empty-string sub passes auth (documented behavior)
- Missing sub rejected through full middleware path

Issuer / audience validation:
- Matching issuer accepted, wrong issuer rejected
- Matching audience accepted, wrong audience rejected
- Missing iss/aud when configured: passes (jsonwebtoken v9 behavior,
  documented with notes on potential hardening)

Key cache:
- Expired cache entries not served
- Fetch failure backoff blocks retry within 10s
- Backoff expiry allows retry
- Cache max entries constant verified

Also adds shared test helpers (encode_test_jwt, test_oidc_state,
oidc_auth_state, oidc_test_app) to reduce boilerplate.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(ci): resolve formatting and no-panics check failures

- Run cargo fmt to wrap long assert lines in OIDC tests
- Add // safety: test helper comments to suppress false positives
  from check_no_panics.py (unwraps in #[cfg(test)] helper fns)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: synner88 <29090601+synner88@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: ilblackdragon@gmail.com <ilblackdragon@gmail.com>
…ts (#1529)

* fix(wasm): inject Content-Length: 0 for bodyless mutating requests [skip-version-check]

The WASM host http_request now auto-injects Content-Length: 0 for
POST/PUT/PATCH/DELETE requests with no body, unless the tool already
provides the header. This fixes Gmail returning 411 on trash_message
and proactively covers all other tools (Google Calendar DELETE,
Google Drive DELETE, etc.).

Extracted needs_content_length_zero() with 8 regression tests
covering all HTTP methods and case-insensitive header detection.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(wasm): use eq_ignore_ascii_case to avoid allocation [skip-version-check]

Replace matches!(method.to_uppercase().as_str(), ...) with
eq_ignore_ascii_case() to avoid a per-request String allocation.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: ilblackdragon@gmail.com <ilblackdragon@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(auth): make shared Google tool status scope-aware

* fix(auth): simplify google docs auth status test

* fix(auth): skip scope expansion for env-var tokens and add dual-source test

Env-var-provided tokens are externally managed, so the scope-expansion
check must not apply — otherwise tools regress to NeedsAuth when no
scopes record exists in the secrets store. Split the token detection
into managed vs env-var paths and only run scope checks for managed
tokens.

Also adds tests verifying: (1) env-var-only tokens return Ready without
scope checks, and (2) when both a managed token and env var are present,
the managed path with scope checks takes priority.

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: ilblackdragon@gmail.com <ilblackdragon@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
… gate regressions (#1746)

* fix: resolve 11 test failures from multi-tenant bootstrap and sandbox gate regressions

Three root causes fixed:

1. Per-user bootstrap greeting in tests: After the multi-tenant isolation
   PR, `tenant_ctx("test-user")` creates a per-user workspace that seeds
   BOOTSTRAP.md and triggers an unwanted bootstrap greeting. This threw
   off response counting and caused message-drain races in 7 e2e tests.
   Fix: pre-seed the "test-user" workspace in the test rig DB so the
   first tenant_ctx call finds existing documents.

2. Sandbox gate blocking full_job routines: The full_job reliability
   overhaul (#1650) intended to remove the SandboxReadiness gate from
   execute_full_job (since full_job routines dispatch through the
   scheduler, not Docker). The gate was accidentally re-added during
   rebase, breaking 4 routine tests. Fix: remove the gate and clean up
   the unused sandbox_readiness field from EngineContext.

3. Owner-gate tests expecting old failure path: Two tests expected
   RunStatus::Failed from the sandbox gate. With the gate removed, the
   tool is now blocked at execution time by the approval context and the
   job completes normally. Fix: update traces and assertions to match the
   new behavior (RunStatus::Ok, owner_gate_count == 0).

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: keep DockerUnavailable gate for full_job routines

Only remove the DisabledByConfig gate — when sandbox is enabled but
Docker is unavailable, full_job routines should still fail rather than
silently running without the expected sandbox isolation.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* refactor: address PR review feedback

- Remove unused `include_completion` param from `owner_gate_trace()`
  and update all 5 call sites
- Use `.expect()` instead of `let _ =` on `seed_if_empty()` in test rig
  to surface seeding failures early
- Rename owner-gate tests from `_blocks_` to `_denies_tool_` to clarify
  the denial-with-success semantics

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* refactor: let per-user bootstrap fire naturally, filter in TestRig

Instead of pre-seeding the "test-user" workspace to prevent the
per-user bootstrap greeting, let it happen naturally and make the
TestRig resilient to it. `wait_for_responses` now transparently
filters bootstrap greetings from the response stream:

- Normal tests: all greetings filtered (bootstrap_greetings_to_keep=0)
- `.with_bootstrap()` tests: 1 greeting kept (the startup greeting),
  additional per-user duplicates filtered

Also updates owner-gate test section headers to match the
denial-with-success semantics.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: assert tool denial event in owner-gate tests

The owner-gate denial tests previously only checked RunStatus::Ok +
owner_gate_count == 0, which could pass if the tool was never called
at all. Now both tests also verify that a tool_result event with
success=false exists for "owner_gate" in the job's event log,
confirming the tool was attempted and blocked by the approval context.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* style: apply rustfmt to collapsed function signature

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: use async lock in bootstrap filter loop, add DisabledByConfig unit test

- Switch TestRig polling loop from `captured_responses()` (try_lock,
  panics on contention) to `captured_responses_async()` (async lock,
  safe under concurrent response pushing)
- Add unit test asserting DisabledByConfig does NOT match the
  DockerUnavailable gate (verifies the intended behavior change)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(slack): respond to thread replies in channels without requiring @mention

Two fixes:

1. Host bug: `on_respond` callback never committed workspace writes or
   injected workspace reader, unlike all other WASM callbacks. Any WASM
   channel persisting state during on_respond silently lost data.

2. Slack WASM channel: track threads where the bot has participated via
   workspace storage. When a message event arrives in a channel thread
   the bot previously replied to, process it without requiring @mention.

Closes #1404

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(slack): log workspace_write error instead of silently discarding

Address code review feedback: handle the Result from workspace_write
when tracking thread participation, logging a warning on failure
instead of using `let _ =` which would silently swallow errors.

Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>

* fix(slack): harden thread reply tracking

---------

Co-authored-by: synner88 <29090601+synner88@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com>
Co-authored-by: Firat Sertgoz <f@nuff.tech>
Co-authored-by: firat.sertgoz <firat.sertgoz@near.ai>
…esh (#1756)

* fix(routines): clone Arc before await in web handler event cache refresh (#1076)

Address review: drop superseded ticker changes, keep only the .cloned()
fix that prevents holding RwLockReadGuard across .await in toggle/delete
handlers. Add regression test for web toggle disabling a system_event
routine.

Closes #1076

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: use explicit block to drop RwLockReadGuard before await

Address review feedback: in Rust 2024, `if let` scrutinee temporaries
live through the body, so the `.cloned()` approach still held the
RwLockReadGuard across `refresh_event_cache().await`. Extract into an
explicit block to ensure the guard is dropped, matching the existing
pattern in `routines_trigger_handler`.

Also add retry loop for `routine_by_name` in the integration test to
avoid flakiness from potential race conditions.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…1762)

* test(e2e): align wasm reinstall expectations with uninstall cleanup

* test(e2e): clarify wasm reinstall fixture semantics
…2238

chore: promote staging to staging-promote/d567d94c-23755250735 (2026-03-30 18:33 UTC)
…5295

chore: promote staging to staging-promote/70214c4a-23719079615 (2026-03-29 22:05 UTC)
…9615

chore: promote staging to staging-promote/86389dab-23706696435 (2026-03-29 21:07 UTC)
…6435

chore: promote staging to staging-promote/e0e530e6-23703082447 (2026-03-29 10:07 UTC)
…2447

chore: promote staging to staging-promote/a8e83210-23702343584 (2026-03-29 06:21 UTC)
…3584

chore: promote staging to staging-promote/8a320ae9-23693265249 (2026-03-29 05:32 UTC)
…5249

chore: promote staging to staging-promote/fd41bdf4-23691145719 (2026-03-28 20:05 UTC)
…5719

chore: promote staging to staging-promote/de5a1c7b-23688974037 (2026-03-28 18:06 UTC)
…4037

chore: promote staging to staging-promote/9bb19a98-23687925861 (2026-03-28 16:06 UTC)
@henrypark133
henrypark133 merged commit e24be45 into staging-promote/9ba10eac-23686921981 Mar 30, 2026
12 of 13 checks passed
@henrypark133
henrypark133 deleted the staging-promote/9bb19a98-23687925861 branch March 30, 2026 19:33
@github-actions github-actions Bot added scope: agent Agent core (agent loop, router, scheduler) scope: channel/wasm WASM channel runtime scope: tool/builtin Built-in tools scope: tool/wasm WASM tool sandbox scope: tool/mcp MCP client scope: db Database trait / abstraction scope: db/postgres PostgreSQL backend scope: llm LLM integration scope: worker Container worker scope: config Configuration scope: extensions Extension management scope: pairing Pairing mode scope: ci CI/CD workflows scope: docs Documentation scope: dependencies Dependency updates size: XL 500+ changed lines risk: high Safety, secrets, auth, or critical infrastructure and removed size: M 50-199 changed lines risk: medium Business logic, config, or moderate-risk modules labels Mar 30, 2026
drchirag1991 pushed a commit to drchirag1991/ironclaw that referenced this pull request Apr 8, 2026
…3687925861

chore: promote staging to staging-promote/851fbb54-23686921981 (2026-03-28 15:07 UTC)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: high Safety, secrets, auth, or critical infrastructure scope: agent Agent core (agent loop, router, scheduler) scope: channel/cli TUI / CLI channel scope: channel/wasm WASM channel runtime scope: channel/web Web gateway channel scope: ci CI/CD workflows scope: config Configuration scope: db/postgres PostgreSQL backend scope: db Database trait / abstraction scope: dependencies Dependency updates scope: docs Documentation scope: extensions Extension management scope: llm LLM integration scope: pairing Pairing mode scope: tool/builtin Built-in tools scope: tool/mcp MCP client scope: tool/wasm WASM tool sandbox scope: worker Container worker size: XL 500+ changed lines staging-promotion

Projects

None yet

Development

Successfully merging this pull request may close these issues.