chore: promote staging to staging-promote/9ba10eac-23686921981 (2026-03-28 15:07 UTC) - #1727
Merged
henrypark133 merged 36 commits intoMar 30, 2026
Conversation
added 2 commits
March 28, 2026 15:10
- Implement broadcast_dm() that creates a DM channel with the target user (POST /users/@me/channels, cached by Discord) and sends the message to it - Extract DISCORD_API_BASE constant for all Discord REST API URLs - Extract send_channel_message() shared helper to deduplicate message posting between on_respond and broadcast_dm - Add snowflake validation on user_id before API calls - Fix pre-existing clippy redundant_closure warning - Use typed DmChannelResponse struct instead of serde_json::Value Closes no specific issue — completes the previously stubbed on_broadcast. Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…PTY (#1678) - Add pty-process crate (MIT, tokio async support) for PTY allocation - Spawn claude CLI with pty-process::Command::arg() chaining instead of building a shell string for script -qfc - Eliminates all shell injection surfaces: prompt, model, session_id are passed via execve, never interpreted by a shell - Keep stderr on separate pipe to prevent NDJSON parse breakage (pty-process attaches PTY to all fds by default) - Gate PTY behind #[cfg(unix)] with direct-spawn fallback for Windows CI - Read stdout from PTY master (implements tokio::io::AsyncRead) - Add regression tests: arg vector construction + PTY allocation Addresses review feedback from zmanian and gemini-code-assist. Co-authored-by: j-bloggs <j-bloggs@users.noreply.github.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…#1677) * fix(worker): treat empty LLM response after text output as completion When a job's LLM produces a substantive text response (e.g., formatted results from a routine) and the next LLM call returns empty or errors, the worker now treats this as successful completion instead of continuing the loop until failure. Previously, empty responses always triggered TextAction::Continue, causing the loop to re-call the LLM. The LLM had nothing more to say, so the provider returned "Response contained no message or tool call (empty)". This made routine jobs that successfully produced results report as "failed". The fix adds a `has_text_response` flag to JobDelegate: - After any non-empty text response: flag is set - Empty text after flag is set: treated as completion - LLM errors (select_tools/respond_with_tools) after flag: treated as completion instead of propagating - Empty text before any output: still retries (rate-limit backoff) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(worker): restrict error swallowing to EmptyResponse variant only - Add LlmError::EmptyResponse variant for when LLM returns no content - Update nearai_chat and github_copilot providers to emit EmptyResponse instead of InvalidResponse for empty/no-choice responses - try_complete_on_error now only swallows EmptyResponse (not AuthFailed, ContextLengthExceeded, Http, Io, etc.) - Extract is_completion_eligible_error as testable pure function - Log mark_completed errors at warn level instead of silently dropping - Add EmptyResponse to retry and circuit breaker transient classifications - Rewrite test to exercise real classification logic against all variants Addresses review feedback from zmanian and gemini-code-assist. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * refactor(worker): extract mark_completed_or_warn helper to DRY completion logic Extract shared mark-completed + warn-on-failure pattern into a single helper method used by both try_complete_on_error and handle_text_response. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: j-bloggs <j-bloggs@users.noreply.github.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(routines): persist full LLM transcript and remove sandbox gate for full_job
Routine execution output was invisible — routine_fire returned a one-liner,
routine_history had no actual output, and the conversation thread contained
only a summary. Full-job routines also hard-failed without Docker.
Three fixes:
1. **Full transcript persistence**: execute_lightweight now persists every
message (prompt, LLM responses, tool calls with params, tool results) to
the routine's conversation thread as it executes, not just a summary
after the fact.
2. **Routine output visibility**: routine_history includes conversation_id
and recent_output messages. routine_fire tells the user to check
routine_history. Web detail page has a "View Execution Thread" button
that navigates to the chat tab. ROUTINE_OK stores "No issues found"
instead of None. Full-job summary pulls actual job output instead of
generic "Job X finished".
3. **Remove SandboxReadiness gate**: full_job routines dispatch through the
scheduler like regular /job commands — no Docker required. The
SandboxReadiness enum is removed entirely.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* style: apply cargo fmt
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(worker): treat AutonomousUnavailable tool errors as recoverable
The job worker crashed the entire job when a tool was denied for
autonomous execution (e.g. secret_list). The error was already recorded
in reason_ctx for the LLM to see, but process_tool_result_job returned
Err which propagated through the agentic loop and terminated the job.
Now all tool errors (including AutonomousUnavailable) return Ok,
letting the LLM see the denial and try a different approach.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(llm): sanitize tool names for OpenAI Codex Responses API
The Codex API requires tool names to match `^[a-zA-Z0-9_-]+$` but
MCP/extension tools can have dots in their names (e.g. `mcp.server.tool`).
This caused HTTP 400 errors when the job worker sent tool calls back
to the LLM.
Sanitize tool names in both `convert_tool_definition` and
`convert_message` (function_call items) by replacing invalid characters
with underscores.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(routines): inject execution context into full_job description [skip-regression-check]
When a full_job routine dispatches a job, the LLM had no context that
it was already executing inside a routine. It wasted iterations on
infrastructure (discovering tools, creating routines, setting up auth)
instead of doing the actual work.
Prepend a clear directive to the job description telling the LLM that
tools and the routine are already configured, and to execute the task
directly.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(mcp): auto-refresh expired OAuth tokens on access [skip-regression-check]
When IronClaw restarts, MCP servers fail with "Secret has expired"
because get_access_token() checks token expiry locally and returns an
error before any HTTP request is made — so the existing 401-retry
refresh logic never triggers.
Now get_access_token() catches SecretError::Expired and automatically
calls refresh_access_token() using the stored refresh token. If the
refresh succeeds, the new token is returned transparently. If it fails,
the error message includes both the expiry and the refresh failure.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(mcp): align refresh token naming and set expiry on stored tokens
Two bugs prevented MCP OAuth token auto-refresh on restart:
1. Naming mismatch: the hosted OAuth flow stored the refresh token as
`{token_secret_name}_refresh_token` (e.g. `mcp_notion_access_token_refresh_token`)
but `McpServerConfig::refresh_token_secret_name()` returned
`mcp_notion_refresh_token`. The refresh token was there but unfindable.
2. Missing expiry: `store_tokens` in auth.rs never called `with_expiry()`
even though `AccessToken::expires_in` was available. Combined with the
fix from the previous commit (auto-refresh on Expired), tokens stored
via the MCP auth flow will now also trigger refresh correctly.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(web): show activity and transitions for agent jobs in job detail [skip-regression-check]
The job events endpoint only checked sandbox jobs for ownership,
returning 404 for agent jobs dispatched from routines. The detail
handler also returned empty transitions for agent jobs.
- events handler: fall back to agent job ownership check
- detail handler: populate transitions from job's state history
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat(routines): expose max_iterations for full_job routines (default 25)
The max_iterations parameter was hardcoded to 10 and not configurable
via routine_create or routine_update, causing complex tasks to hit the
iteration cap.
- Add max_iterations to full_job execution schema (1-200, default 25)
- Thread it through parse → build → RoutineAction
- Support updating via routine_update
- Raise default from 10 to 25
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(routines): break self-dialogue loop after full_job plan execution
After plan execution, the completion-check Q&A ("Is the job complete?" /
"No, not complete...") was left in the message context, causing the
agentic loop to repeat the same analysis instead of calling tools.
Replace the stale dialogue with an action-oriented continuation prompt
that instructs the LLM to use tools for remaining work. Also strip
<suggestions> tags from all job output since they're only meaningful
for interactive chat sessions.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(repl): prevent test hang in single-message mode
In single-message mode, start() stored a clone of the mpsc sender in
self.msg_tx for approval injection. After the thread sent /quit and
exited, the stored clone kept the stream alive, so stream.next()
blocked forever in the test assertion that the stream ends.
Skip storing the sender in single-message mode since interactive
approval is not needed.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(jobs): treat text responses as final answer in agentic loop
When the LLM produces a non-empty text response with no tool intent
(already filtered by the nudge mechanism), it is the job's final
answer. Previously, handle_text_response only exited the loop if the
text matched rigid completion phrases like "job is complete". Natural
summaries like "Weekly review completed and saved to Notion" were
added to context and the loop continued, causing the LLM to restate
the same summary until max_iterations was hit.
Now any non-empty text response marks the job complete and stops the
loop, matching the chat dispatcher behavior.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* perf(tests): reduce skills catalog network failure test from 10s to 1s
The test_search_returns_error_on_network_failure test connects to an
unreachable RFC 5737 TEST-NET IP and waited for the full 10s production
REQUEST_TIMEOUT. Add with_url_and_timeout test helper and use a 1s
timeout instead. [skip-regression-check]
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(tools): accept 'message' as alias for 'content' in message tool
LLMs frequently call the message tool with {"message": "..."} instead
of {"content": "..."}. Fall back to the 'message' key when 'content'
is missing to avoid InvalidParameters errors during autonomous job
execution.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(tools): attach thread_id for gateway broadcast in message tool
When the message tool broadcasts to all channels (channel=null), it
sent an OutgoingResponse without a thread_id. The gateway silently
dropped these messages (returned Ok but never sent the SSE event),
so they appeared in repl but not in the web UI.
The thread_id was only populated when channel was explicitly "gateway".
Now it is always populated from notify_thread_id metadata, so
broadcast_all delivers to the gateway correctly.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(gateway): return error instead of silently dropping messages
Gateway broadcast() and respond() previously returned Ok(()) when
thread_id was missing, silently swallowing the message. Callers
(message tool, agent loop) believed delivery succeeded when it didn't.
Now returns ChannelError::MissingRoutingTarget so callers can detect
and report the failure. Four regression tests verify the contract:
respond/broadcast with and without thread_id.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: resolve rebase conflicts with staging
Restore sandbox_readiness field removed by pre-rebase commits (staging
still uses it). Update repl test to match staging's single-message
behavior (no longer sends /quit). Add missing reasoning field to
ToolCall in codex test.
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(tools): log error when routine conversation lookup fails
The routine_history tool silently swallowed errors from
get_or_create_routine_conversation, returning empty output without
any diagnostic logging. Add tracing::warn so failures are visible
in logs. [skip-regression-check]
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: address PR #1650 review comments
- E2E test: accept submitted/accepted as success states in job assertion
- TimeTool: remove operation from required schema (defaults to "now")
- jobs handler: log DB errors server-side, return generic message to client
- routines handler: use read-only find_routine_conversation on GET
- codex provider: reverse-map sanitized tool names so MCP tools resolve
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: address zmanian review feedback on PR #1650
- MCP refresh token: fall back to legacy secret name (mcp_{name}_refresh_token)
so existing users don't need to re-authenticate after the naming fix
- Job worker: replace fragile messages.pop() with truncate-to-saved-count
to avoid maintenance hazard if message flow changes
- Document cost implications of max_iterations 10->25 default bump
- Revert Cargo.toml dist profile change (thin LTO comment, codegen-units=16)
as it's unrelated to this PR
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: resolve rebase conflicts and address new Copilot comments
- Fix no_silent_drop tests for updated GatewayConfig (user_id moved to
GatewayChannel::new second arg, user_tokens removed)
- Fix handle_text_response param name (_reason_ctx -> reason_ctx)
- Fix missing has_text_response field in test JobDelegate
- Propagate row.get errors in find_routine_conversation instead of
unwrap_or_default
- Only fall back to legacy refresh token name on NotFound/Expired,
propagate real errors (DB, decryption)
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat(discord): restore gateway channel flow in wasm * chore(discord): bump channel version to 0.2.1 * fix(discord): address review feedback on gateway channel PR - Add #[serde(default)] to DiscordMessageMetadata for backward compat with old Option<String> serialized metadata - Restore mention polling alongside Gateway (on_poll processes gateway events first, then runs poll_for_mentions if configured) - Update on_respond to handle source_message_id with message_reference for mention-poll reply threading - Implement Gateway presence status: dnd before pairing, online after - Implement Gateway resume (OP 6) with session_id tracking, falling back to fresh identify on Invalid Session (OP 9) - Extract WebsocketSessionState and spawn_websocket_poll to reduce nesting in start_websocket_runtime - Simplify should_apply_dm_pairing tautology - Remove completed plan docs - Fix clippy items_after_test_module in extensions handler Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(discord): address review findings in gateway channel PR - Fix gateway presence always showing "online" by filtering empty owner_id strings from workspace store reads - Fix interaction followup using POST instead of PATCH to /messages/@original, which left deferred "thinking" state unresolved - Restore mention-poll pagination (up to 5 pages of 100 messages) - Remove dead ed25519-dalek and hex dependencies from WASM crate - Remove unused _channel_id parameter from remember_processed_id - Clean up redundant let binding in send_pairing_reply Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(discord): address second-round review findings - Log warning when gateway event queue JSON fails to deserialize instead of silently returning empty (zmanian review item 1) - Defer presence update from OP 10 Hello to after OP 0 READY, per Discord gateway protocol which requires READY before non-Identify commands (zmanian review item 2) - Add 0-25% random jitter to websocket reconnect backoff per Discord's reconnection recommendations (zmanian suggestion) - Extract WebsocketPollContext struct to replace 19-parameter spawn_websocket_poll function (zmanian suggestion) - Document intent bitmask 4609 = GUILDS + GUILD_MESSAGES + DIRECT_MESSAGES in capabilities JSON (zmanian suggestion) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: zhyaoyu <zhyaoyu@aliyun.com> Co-authored-by: ilblackdragon@gmail.com <ilblackdragon@gmail.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…#1667) Add support for bundle layouts where directories without SKILL.md are recursed into to find nested skills (e.g., skills/my-org/skill-a/SKILL.md). - Add configurable max_scan_depth (SKILLS_MAX_SCAN_DEPTH env, default 3) - Recurse into subdirectories lacking SKILL.md up to depth limit - Share remaining discovery cap across recursive levels - Replace try_exists + read_dir with single read_dir (eliminates TOCTOU) - Box::pin recursive async calls for correct future sizing Closes #1664 Co-authored-by: Rajul Bhatnagar <brajul@amazon.com>
* Clarify message tool and channel setup guidance * Add target format hints to proactive messaging prompt * Clarify search and message tool edge cases * Fix stale tool_search e2e assertion * Update src/llm/reasoning.rs Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com> * Tighten prompt guidance for message replies * Format prompt guidance assertions --------- Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* fix: pin staging ci jobs to a single tested sha * chore(ci): retrigger regression gate [skip-regression-check] --------- Co-authored-by: Firat Sertgoz <f@nuff.tech>
* fix: prevent UTF-8 panics in byte-index string truncation
Replace unsafe `&s[..n]` patterns with `floor_char_boundary(s, n)` at 3
production code sites where the truncation index could land mid-multibyte
character, panicking on non-ASCII input:
- src/llm/nearai_chat.rs: API response truncation in error message
- src/cli/memory.rs: memory content display truncation
- src/cli/config.rs: config value display truncation
All 3 sites operate on external or user-supplied strings that may contain
non-ASCII characters. The existing `crate::util::floor_char_boundary`
utility (used at 18 other call sites) walks back to the nearest char
boundary, preventing the panic.
Adds regression test with multi-byte characters (combining accents and
4-byte emoji) for truncate_content.
[skip-regression-check]
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix: clarify test comment and use exact assertions
Address Gemini review feedback:
- Fix misleading comment: \u{00e9} is precomposed e-acute, not combining accent
- Replace weak assertions (ends_with/is_empty) with exact assert_eq!
Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
---------
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…nt (#1630) * fix(bedrock): strip tool blocks from messages when toolConfig is absent Bedrock's Converse API requires `toolConfig` whenever messages contain `toolUse` or `toolResult` content blocks. When the agentic loop reaches its force_text iteration (e.g. lightweight routine at max_iterations), it switches from `complete_with_tools()` to `complete()` — but the message history still carries tool blocks from prior iterations. `convert_messages()` faithfully converts these into Bedrock content blocks, and without `toolConfig` Bedrock rejects the request: "The toolConfig field must be defined when using toolUse and toolResult content blocks." Add `strip_tool_blocks()` that converts tool interaction data to text: - Assistant `tool_calls` → dropped (text content preserved) - `Role::Tool` → `Role::User` with `[Tool ... returned: ...]` text Wire it into: - `complete()`: unconditionally, since it never sends toolConfig - `complete_with_tools()`: when `build_tool_config()` returns None (empty tools or tool_choice="none") Closes #1629 * fix(bedrock): address review feedback on strip_tool_blocks - Add tracing::debug\! when tool blocks are stripped (zmanian suggestion) - Add test for tool_choice="none" path (zmanian suggestion) - Add inline comment on empty-content assistant behavior --------- Co-authored-by: Rajul Bhatnagar <brajul@amazon.com>
* feat: support custom LLM provider configuration via web UI Users can now define custom LLM providers through the web UI and have them take effect without modifying environment variables or config files. - Add `CustomLlmProviderSettings` struct and `llm_custom_providers` field to `Settings` so custom provider definitions are persisted and loaded from the DB settings table - Add `LlmConfig::resolve_custom_provider()` to build a `RegistryProviderConfig` from user-defined provider data (base_url, adapter, model, api_key) - Flip resolution priority to `db > env > default` so active provider set through the UI takes precedence over deployment env vars - Warn when a custom provider is missing base_url or model - Add startup info logs for backend source and provider creation - Add regression tests for custom provider resolution and DB priority * feat: add test connection for custom LLM providers - Add POST /api/llm/test_connection endpoint that validates connectivity and auth for OpenAI-compatible, Anthropic, and Ollama adapters (10s timeout, per-adapter request logic) - Add "Test" button next to Save/Cancel in the add-provider form; result shown inline with green/red styling - Hide delete button for the active provider instead of showing an error toast - Sort the active provider to the top of the provider list - Clear selected_model when switching providers to avoid model-not-supported errors on the new provider - Add i18n keys for test/testing states (en + zh-CN) * feat: add built-in provider API key and model configuration - Add Configure button on built-in provider cards (openai, anthropic, gemini, ollama, etc.) to set API key and default model via web UI - Store overrides as `llm_builtin_overrides` setting (per-provider key/model map) using the existing generic settings k/v API - Add LlmBuiltinOverride struct in settings.rs; resolve in resolve_registry_provider() with priority: env var > selected_model > llm_builtin_overrides[id] > default - Restore provider's configured model to selected_model on provider switch, so /model command always takes precedence at runtime - Fix fetch-models button in built-in configure mode: use hardcoded base_url from BUILTIN_PROVIDERS instead of the hidden form field - Add edit support for custom providers with pre-filled dialog - Show current model on active and configured provider cards - Convert add/edit provider form to a modal dialog - Sync selected_model when editing or deleting an active custom provider * feat: move Config tab into Settings as Providers subtab * feat(web): merge Providers into Inference tab with UX improvements * chore: resolve conflicts * fix(llm): address security and correctness issues in custom LLM provider * fix(llm): address security and correctness issues in custom LLM provider * feat(web): fall back to env vars for LLM provider config in UI * fix(llm): enforce db > env > default config priority for provider setting * fix: address review feedback on provider config priority * feat: extract BUILTIN_PROVIDERS into providers.js * fix(security): store LLM API keys in encrypted secrets store instead of plaintext * fix(security): harden LLM API key handling across settings and LLM endpoints * fix: test_connection sends actual chat completion * refactor(web): derive LLM Provider display from active Model Provider * fix(settings): language switch not working for llm provider * feat(web): add restart notice to LLM Provider settings * fix: review fixes for custom LLM provider PR - Add server-side validation of custom provider ID format (lowercase alphanumeric + hyphens, 1-64 chars) to match frontend regex - Tighten is_nearai_private_endpoint to exact-match private.near.ai or *.private.near.ai, rejecting lookalikes like private-evil.near.ai - Fix misleading priority doc comments in config/mod.rs and settings.rs to reflect the split model: LLM uses DB > env, others use env > DB - Clean up #1581 artifacts: remove TOML file creation from persist_selected_model (DB is sufficient), update stale priority comments in commands.rs, fix contradictory test assertions - Add 18 new tests for provider ID validation, adapter validation, and nearai private endpoint matching Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: address PR review comments for custom LLM provider - Move LLM handlers (test_connection, list_models, env_defaults) from server.rs to handlers/llm.rs for consistency with other handler modules - Merge validate_custom_providers into single pass (ID + adapter check) - Allow underscores in custom provider IDs to match builtin naming - Add missing i18n key config.fetchingModels (en + zh-CN) - Fix optional_env().ok().flatten() error swallowing in config/llm.rs; propagate ConfigError with ? instead of silently discarding - Narrow settings.rs module docs to scope DB>env precedence to LLM - Add unit tests for hydrate_llm_keys_from_secrets Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * refactor: replace static providers.js with API endpoint from registry - Delete providers.js; serve provider list from /api/llm/providers endpoint that reads from the embedded ProviderRegistry (providers.json) - Centralize secret naming (builtin_secret_name, custom_secret_name) into settings.rs; replace 8 duplicated format! calls across 4 files - Extract JS API_KEY_UNCHANGED constant; replace 6 magic string literals - Replace hard-coded API key placeholder strings with i18n keys (config.apiKeyConfigured, config.apiKeyFromEnv, config.apiKeyEnter) - Simplify apiFetchVoid to delegate to apiFetch - Remove unnecessary Vec clones in guard_active_provider_not_removed Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Robert Yan <46699230+think-in-universe@users.noreply.github.com> Co-authored-by: Illia Polosukhin <ilblackdragon@gmail.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* Handle empty tool completions in autonomous jobs * Address malformed tool recovery review comments * style: apply rustfmt to reasoning tests --------- Co-authored-by: Firat Sertgoz <f@nuff.tech>
…1463) * feat(gateway): add OIDC JWT authentication for reverse-proxy deployments Add an optional OIDC JWT auth path to the web gateway, enabling deployments behind identity-aware proxies like AWS ALB with Okta/Cognito. When GATEWAY_OIDC_ENABLED=true, the gateway reads a signed JWT from a configurable HTTP header (default: x-amzn-oidc-data), fetches the signing key from a JWKS endpoint, and verifies the signature + claims. Auth flow: Bearer token → OIDC JWT → query-string token → 401. Key design decisions: - Split signature verification from claim extraction to handle AWS ALB's non-standard base64 padding (ALB includes '=' padding in JWT segments, but jsonwebtoken's decode() strips it, changing the signing input). We verify against the original token text, then extract claims from a normalized copy. - JWKS keys cached for 1 hour with per-kid granularity. - Supports both ALB-style per-key PEM URLs ({kid} placeholder) and standard JWKS endpoints. - DER-to-raw ECDSA signature conversion for IdPs that use DER encoding. - Frontend auto-detects proxy auth via /api/gateway/status probe, skipping the login screen when OIDC is active. Configuration (env vars): GATEWAY_OIDC_ENABLED=true GATEWAY_OIDC_HEADER=x-amzn-oidc-data (default) GATEWAY_OIDC_JWKS_URL=https://public-keys.auth.elb.us-east-1.amazonaws.com/{kid} GATEWAY_OIDC_ISSUER=https://example.okta.com (optional) GATEWAY_OIDC_AUDIENCE=my-client-id (optional) Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * Address code review feedback on OIDC auth PR - Add EdDSA PEM key parsing support (was falling through to RSA) - Fix issuer validation: remove set_issuer(&[]) else branch that rejected all tokens when GATEWAY_OIDC_ISSUER is unset - Make missing `sub` claim a validation error instead of silently defaulting to "unknown" - Extract initApp() in app.js so OIDC auto-auth actually initializes the UI (was calling undefined function) - Add regression tests for sub claim and issuer validation fixes Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * Harden OIDC auth: address claude[bot] security review - SSRF: URL-encode kid before substituting into JWKS URL template - Cache bounds: cap key cache at 64 entries, evict expired + oldest - DER parsing: support long-form length encoding (>= 128 bytes), validate component lengths against expected curve size - Production safety: replace .expect() with Result in OidcState::from_config - Fetch backoff: cache failed JWKS fetches for 10s to prevent retry storms - Body limit: cap JWKS responses at 256 KB to prevent OOM from rogue endpoint - Add regression tests for DER long-form, kid encoding, cache bounds Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * test(auth): add regression test for OIDC identity resolution Add two integration tests that exercise the full OIDC middleware path through to AuthenticatedUser extraction: - test_oidc_auth_inserts_user_identity_for_handler: sends a valid OIDC JWT through the middleware and verifies the handler receives the sub claim as user_id. Returns 401 if identity insertion is missing — verified by temporarily removing the insert and confirming failure. - test_oidc_auth_user_gets_member_role: confirms OIDC-authenticated users receive role=member (not admin). Uses a seed_key() test helper on OidcState to pre-populate the key cache with an HS256 secret, avoiding the need for an HTTP JWKS mock. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * test(auth): comprehensive OIDC test coverage for edge cases Add 17 new OIDC tests covering middleware integration, auth priority, invalid JWTs, issuer/audience validation, and key cache behavior: Middleware auth priority & fallthrough: - Bearer works when OIDC configured but header absent - Bearer takes priority when both Bearer and OIDC header present - Bad OIDC signature returns 401 (not 500) - Invalid OIDC doesn't block valid bearer auth - No auth at all with OIDC configured → 401 Expired / invalid JWT edge cases: - Expired JWT (exp in the past) rejected - JWT without kid header rejected - Malformed JWTs rejected (empty, 2-part, 4-part, garbage) - Non-string sub claim (integer) rejected - Empty-string sub passes auth (documented behavior) - Missing sub rejected through full middleware path Issuer / audience validation: - Matching issuer accepted, wrong issuer rejected - Matching audience accepted, wrong audience rejected - Missing iss/aud when configured: passes (jsonwebtoken v9 behavior, documented with notes on potential hardening) Key cache: - Expired cache entries not served - Fetch failure backoff blocks retry within 10s - Backoff expiry allows retry - Cache max entries constant verified Also adds shared test helpers (encode_test_jwt, test_oidc_state, oidc_auth_state, oidc_test_app) to reduce boilerplate. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(ci): resolve formatting and no-panics check failures - Run cargo fmt to wrap long assert lines in OIDC tests - Add // safety: test helper comments to suppress false positives from check_no_panics.py (unwraps in #[cfg(test)] helper fns) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: synner88 <29090601+synner88@users.noreply.github.com> Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> Co-authored-by: ilblackdragon@gmail.com <ilblackdragon@gmail.com>
…ts (#1529) * fix(wasm): inject Content-Length: 0 for bodyless mutating requests [skip-version-check] The WASM host http_request now auto-injects Content-Length: 0 for POST/PUT/PATCH/DELETE requests with no body, unless the tool already provides the header. This fixes Gmail returning 411 on trash_message and proactively covers all other tools (Google Calendar DELETE, Google Drive DELETE, etc.). Extracted needs_content_length_zero() with 8 regression tests covering all HTTP methods and case-insensitive header detection. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(wasm): use eq_ignore_ascii_case to avoid allocation [skip-version-check] Replace matches!(method.to_uppercase().as_str(), ...) with eq_ignore_ascii_case() to avoid a per-request String allocation. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: ilblackdragon@gmail.com <ilblackdragon@gmail.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(auth): make shared Google tool status scope-aware * fix(auth): simplify google docs auth status test * fix(auth): skip scope expansion for env-var tokens and add dual-source test Env-var-provided tokens are externally managed, so the scope-expansion check must not apply — otherwise tools regress to NeedsAuth when no scopes record exists in the secrets store. Split the token detection into managed vs env-var paths and only run scope checks for managed tokens. Also adds tests verifying: (1) env-var-only tokens return Ready without scope checks, and (2) when both a managed token and env var are present, the managed path with scope checks takes priority. [skip-regression-check] Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: ilblackdragon@gmail.com <ilblackdragon@gmail.com> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
… gate regressions (#1746) * fix: resolve 11 test failures from multi-tenant bootstrap and sandbox gate regressions Three root causes fixed: 1. Per-user bootstrap greeting in tests: After the multi-tenant isolation PR, `tenant_ctx("test-user")` creates a per-user workspace that seeds BOOTSTRAP.md and triggers an unwanted bootstrap greeting. This threw off response counting and caused message-drain races in 7 e2e tests. Fix: pre-seed the "test-user" workspace in the test rig DB so the first tenant_ctx call finds existing documents. 2. Sandbox gate blocking full_job routines: The full_job reliability overhaul (#1650) intended to remove the SandboxReadiness gate from execute_full_job (since full_job routines dispatch through the scheduler, not Docker). The gate was accidentally re-added during rebase, breaking 4 routine tests. Fix: remove the gate and clean up the unused sandbox_readiness field from EngineContext. 3. Owner-gate tests expecting old failure path: Two tests expected RunStatus::Failed from the sandbox gate. With the gate removed, the tool is now blocked at execution time by the approval context and the job completes normally. Fix: update traces and assertions to match the new behavior (RunStatus::Ok, owner_gate_count == 0). Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: keep DockerUnavailable gate for full_job routines Only remove the DisabledByConfig gate — when sandbox is enabled but Docker is unavailable, full_job routines should still fail rather than silently running without the expected sandbox isolation. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * refactor: address PR review feedback - Remove unused `include_completion` param from `owner_gate_trace()` and update all 5 call sites - Use `.expect()` instead of `let _ =` on `seed_if_empty()` in test rig to surface seeding failures early - Rename owner-gate tests from `_blocks_` to `_denies_tool_` to clarify the denial-with-success semantics Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * refactor: let per-user bootstrap fire naturally, filter in TestRig Instead of pre-seeding the "test-user" workspace to prevent the per-user bootstrap greeting, let it happen naturally and make the TestRig resilient to it. `wait_for_responses` now transparently filters bootstrap greetings from the response stream: - Normal tests: all greetings filtered (bootstrap_greetings_to_keep=0) - `.with_bootstrap()` tests: 1 greeting kept (the startup greeting), additional per-user duplicates filtered Also updates owner-gate test section headers to match the denial-with-success semantics. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: assert tool denial event in owner-gate tests The owner-gate denial tests previously only checked RunStatus::Ok + owner_gate_count == 0, which could pass if the tool was never called at all. Now both tests also verify that a tool_result event with success=false exists for "owner_gate" in the job's event log, confirming the tool was attempted and blocked by the approval context. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * style: apply rustfmt to collapsed function signature Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: use async lock in bootstrap filter loop, add DisabledByConfig unit test - Switch TestRig polling loop from `captured_responses()` (try_lock, panics on contention) to `captured_responses_async()` (async lock, safe under concurrent response pushing) - Add unit test asserting DisabledByConfig does NOT match the DockerUnavailable gate (verifies the intended behavior change) Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(slack): respond to thread replies in channels without requiring @mention Two fixes: 1. Host bug: `on_respond` callback never committed workspace writes or injected workspace reader, unlike all other WASM callbacks. Any WASM channel persisting state during on_respond silently lost data. 2. Slack WASM channel: track threads where the bot has participated via workspace storage. When a message event arrives in a channel thread the bot previously replied to, process it without requiring @mention. Closes #1404 Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(slack): log workspace_write error instead of silently discarding Address code review feedback: handle the Result from workspace_write when tracking thread participation, logging a warning on failure instead of using `let _ =` which would silently swallow errors. Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com> * fix(slack): harden thread reply tracking --------- Co-authored-by: synner88 <29090601+synner88@users.noreply.github.com> Co-authored-by: Claude Opus 4.6 <noreply@anthropic.com> Co-authored-by: Firat Sertgoz <f@nuff.tech> Co-authored-by: firat.sertgoz <firat.sertgoz@near.ai>
…esh (#1756) * fix(routines): clone Arc before await in web handler event cache refresh (#1076) Address review: drop superseded ticker changes, keep only the .cloned() fix that prevents holding RwLockReadGuard across .await in toggle/delete handlers. Add regression test for web toggle disabling a system_event routine. Closes #1076 Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: use explicit block to drop RwLockReadGuard before await Address review feedback: in Rust 2024, `if let` scrutinee temporaries live through the body, so the `.cloned()` approach still held the RwLockReadGuard across `refresh_event_cache().await`. Extract into an explicit block to ensure the guard is dropped, matching the existing pattern in `routines_trigger_handler`. Also add retry loop for `routine_by_name` in the integration test to avoid flakiness from potential race conditions. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…1762) * test(e2e): align wasm reinstall expectations with uninstall cleanup * test(e2e): clarify wasm reinstall fixture semantics
…2238 chore: promote staging to staging-promote/d567d94c-23755250735 (2026-03-30 18:33 UTC)
…5295 chore: promote staging to staging-promote/70214c4a-23719079615 (2026-03-29 22:05 UTC)
…9615 chore: promote staging to staging-promote/86389dab-23706696435 (2026-03-29 21:07 UTC)
…6435 chore: promote staging to staging-promote/e0e530e6-23703082447 (2026-03-29 10:07 UTC)
…2447 chore: promote staging to staging-promote/a8e83210-23702343584 (2026-03-29 06:21 UTC)
…3584 chore: promote staging to staging-promote/8a320ae9-23693265249 (2026-03-29 05:32 UTC)
…5249 chore: promote staging to staging-promote/fd41bdf4-23691145719 (2026-03-28 20:05 UTC)
…5719 chore: promote staging to staging-promote/de5a1c7b-23688974037 (2026-03-28 18:06 UTC)
…4037 chore: promote staging to staging-promote/9bb19a98-23687925861 (2026-03-28 16:06 UTC)
henrypark133
merged commit Mar 30, 2026
e24be45
into
staging-promote/9ba10eac-23686921981
12 of 13 checks passed
drchirag1991
pushed a commit
to drchirag1991/ironclaw
that referenced
this pull request
Apr 8, 2026
…3687925861 chore: promote staging to staging-promote/851fbb54-23686921981 (2026-03-28 15:07 UTC)
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Auto-promotion from staging CI
Batch range:
2f4eb08613cefff1af8b7b1a475fda00c84dd855..9bb19a98f767f2e4db0d4f3dda0e495a356a726aPromotion branch:
staging-promote/9bb19a98-23687925861Base:
staging-promote/9ba10eac-23686921981Triggered by: Staging CI batch at 2026-03-28 15:07 UTC
Commits in this batch (6):
Current commits in this promotion (2)
Current base:
staging-promote/9ba10eac-23686921981Current head:
staging-promote/9bb19a98-23687925861Current range:
origin/staging-promote/9ba10eac-23686921981..origin/staging-promote/9bb19a98-23687925861Auto-updated by staging promotion metadata workflow
Waiting for gates:
Auto-created by staging-ci workflow