Skip to content

chore: promote staging to staging-promote/8f8cb7f7-23680994633 (2026-03-28 14:09 UTC) - #1726

Merged
henrypark133 merged 40 commits into
staging-promote/8f8cb7f7-23680994633from
staging-promote/9ba10eac-23686921981
Mar 30, 2026
Merged

henrypark133 merged 40 commits into
staging-promote/8f8cb7f7-23680994633from
staging-promote/9ba10eac-23686921981

Conversation

@ironclaw-ci

@ironclaw-ci ironclaw-ci Bot commented Mar 28, 2026 •

Copy link
Copy Markdown
Contributor

Auto-promotion from staging CI

Batch range: 2f4eb08613cefff1af8b7b1a475fda00c84dd855..9ba10eac35debd6b77a6b4acbaf1b65a96e52a86
Promotion branch: staging-promote/9ba10eac-23686921981
Base: staging-promote/8f8cb7f7-23680994633
Triggered by: Staging CI batch at 2026-03-28 14:09 UTC

Commits in this batch (4):

Current commits in this promotion (3)

Current base: staging-promote/8f8cb7f7-23680994633
Current head: staging-promote/9ba10eac-23686921981
Current range: origin/staging-promote/8f8cb7f7-23680994633..origin/staging-promote/9ba10eac-23686921981

Auto-updated by staging promotion metadata workflow

Waiting for gates:

  • Tests: pending
  • E2E: pending
  • Claude Code review: pending (will post comments on this PR)

Auto-created by staging-ci workflow

henrypark133 and others added 3 commits March 28, 2026 14:46
* Clean up extension credentials on uninstall

* Address PR review feedback

* Cover channel webhook secrets on uninstall

* Harden tool secret cleanup detection
…rse_timestamp tests (#1700)

* fix(db): add tracing warn for naive timestamp fallback and improve parse_timestamp tests

* style: fix formatting
@github-actions github-actions Bot added scope: tool/wasm WASM tool sandbox scope: extensions Extension management scope: docs Documentation size: XL 500+ changed lines risk: medium Business logic, config, or moderate-risk modules contributor: core 20+ merged PRs labels Mar 28, 2026
)

* fix(oauth): tighten legacy state validation and fallback handling

* style: fix formatting

* refactor: separate validation checks for clearer error messages
@claude

claude Bot commented Mar 28, 2026

Copy link
Copy Markdown

Code review

No issues found.

@claude

claude Bot commented Mar 28, 2026

Copy link
Copy Markdown

Code review

Found 3 performance/production considerations:

  1. [HIGH:85] Sequential async I/O in extension cleanup

    • collect_referenced_secret_names() iterates tools/channels and awaits capability loads sequentially
    • Impact: Uninstall performance degrades with many installed extensions
    • Location: Extension cleanup logic in src/extensions/manager.rs
    • Optimization: Use futures::join_all() for concurrent capability discovery
  2. [MEDIUM:70] Full discovery re-run on every removal

    • Every remove() call performs complete discovery of all tools/channels
    • Impact: Scales linearly with extension count
    • Optimization opportunity: Consider caching or lazy discovery
  3. [MEDIUM:75] Silent failure in best-effort secret deletion

    • Individual secret delete failures are logged but silently continue
    • Could mask systemic secrets backend issues during uninstall
    • Note: Proper early return if reference check fails; only individual deletes fail gracefully

Summary: No blocking bugs found. The 3 findings are optimization opportunities rather than correctness issues. Sequential I/O would only impact users with dozens of extensions during uninstall (non-critical path). Code properly handles errors with early return on reference check failure.

@claude

claude Bot commented Mar 28, 2026

Copy link
Copy Markdown

Code review findings from architecture review agent:

Finding 1 (HIGH:85) - TOCTOU race condition: Agent flagged timing between cleanup_plan collection and reference check. However, checking AFTER unregistration (not before) is the correct approach. Logic appears sound - requires verification.

Finding 2 (HIGH:80) - Stringly-typed secrets: Agent flagged string-based secret names. However, secret names ARE fundamentally strings. Strong types would over-engineer. Design is appropriate.

Finding 3 (MEDIUM:75) - DRY violation: Three functions extract secrets from different capability types independently. Minor code duplication, acceptable given different structures.

Finding 4 (MEDIUM:70) - Case normalization: to_lowercase() consistently applied. Minor concern about secrets store consistency.

Finding 5 (MEDIUM:75) - Error type mismatch: collect_referenced_secret_names returns String error instead of ExtensionError. VALID - should use proper error type per CLAUDE.md thiserror pattern.

Status: One valid finding (error type), others are design choices or false positives. No blocking issues.

@claude

claude Bot commented Mar 28, 2026

Copy link
Copy Markdown

CRITICAL BUG FOUND - Bug Scan Review

[CRITICAL:95] Missing companion secrets in reference tracking

The code has a logic error in the secret cleanup process:

PROBLEM:
When collecting referenced secrets from other extensions, the code only includes BASE secret names (e.g., 'github_token') but NOT the companion secrets (e.g., 'github_token_refresh_token', 'github_token_scopes').

EXAMPLE:

  • Tool A and Tool B both use OAuth token 'github_token'
  • Tool A is uninstalled
  • cleanup_plan includes: [github_token, github_token_refresh_token, github_token_scopes]
  • collect_referenced_secret_names() finds only 'github_token' from Tool B (missing companions)
  • Companion secrets are deleted because they don't appear referenced!

AFFECTED CODE:

  • tool_secret_names() line ~3240: only extracts auth.secret_name, not companions
  • mcp_server_secret_names() line ~3260: only extracts base tokens, missing OAuth companions
  • channel_secret_names() line ~3267: missing webhook companion secrets

FIX NEEDED:
Update tool_secret_names, mcp_server_secret_names, channel_secret_names to include companion secrets:

  • For OAuth: add oauth_refresh_secret_name() and oauth_scopes_secret_name() companions
  • For webhook: add signature_key and hmac companion secrets

IMPACT: Orphaned secrets, credential leaks, broken OAuth flows after uninstall

@claude

claude Bot commented Mar 28, 2026

Copy link
Copy Markdown

Security Review

Found 1 potential security concern:

[HIGH:75] TOCTOU race condition in secret cleanup

  • Between collect_referenced_secret_names() and delete operations, concurrent extension install could reference deleted secrets
  • Scenario: Tool B installed WHILE Tool A is being uninstalled after reference check
  • Impact: Secret deletion during concurrent install could break Tool B's auth
  • Mitigation: Code uses best-effort deletion (logs but doesn't fail), reducing impact

Other security aspects reviewed as safe:

  • SQL queries properly parameterized
  • No production unwrap/expect calls
  • Secrets not logged (key names only)
  • Path construction safe
  • Proper error handling

NOTE: This TOCTOU concern is secondary to the CRITICAL bug found in companion secret tracking (see separate comment).

@claude

claude Bot commented Mar 28, 2026

Copy link
Copy Markdown

CODE REVIEW SUMMARY - PR 1726

CRITICAL ISSUES FOUND

  1. CRITICAL [95] Missing companion secrets in reference tracking
    Cause: OAuth refresh/scopes and webhook secrets not included when collecting referenced secrets
    Impact: Companion secrets deleted while main secret still referenced
    Fix: Update tool_secret_names, mcp_server_secret_names, channel_secret_names to include companions

  2. HIGH [85] Sequential async I/O performance bottleneck
    Location: collect_referenced_secret_names() uses awaits in loops
    Impact: Timeout risk with many extensions
    Fix: Use futures::join_all() for concurrent discovery

  3. HIGH [75] TOCTOU race condition during concurrent install/uninstall
    Risk: Extension installed while another uninstalls same secret
    Mitigation: best-effort deletion reduces impact

  4. HIGH [80] String-based error types violate CLAUDE.md pattern
    Issue: Result<_, String> instead of ExtensionError
    Fix: Return proper ExtensionError for consistency

  5. MEDIUM [75] DRY violation: 3 similar secret extraction functions

POSITIVE FINDINGS

  • No unwrap/expect in production code
  • Comprehensive E2E tests
  • Well-designed SecretCleanupPlan
  • Proper SQL parameter binding
  • Best-effort error handling

RECOMMENDATION: Fix CRITICAL companion secrets bug before merge. Address HIGH-severity issues for stability.

Achieve and others added 12 commits March 28, 2026 16:31
- Implement broadcast_dm() that creates a DM channel with the target
  user (POST /users/@me/channels, cached by Discord) and sends the
  message to it
- Extract DISCORD_API_BASE constant for all Discord REST API URLs
- Extract send_channel_message() shared helper to deduplicate message
  posting between on_respond and broadcast_dm
- Add snowflake validation on user_id before API calls
- Fix pre-existing clippy redundant_closure warning
- Use typed DmChannelResponse struct instead of serde_json::Value

Closes no specific issue — completes the previously stubbed on_broadcast.

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…PTY (#1678)

- Add pty-process crate (MIT, tokio async support) for PTY allocation
- Spawn claude CLI with pty-process::Command::arg() chaining instead of
  building a shell string for script -qfc
- Eliminates all shell injection surfaces: prompt, model, session_id
  are passed via execve, never interpreted by a shell
- Keep stderr on separate pipe to prevent NDJSON parse breakage
  (pty-process attaches PTY to all fds by default)
- Gate PTY behind #[cfg(unix)] with direct-spawn fallback for Windows CI
- Read stdout from PTY master (implements tokio::io::AsyncRead)
- Add regression tests: arg vector construction + PTY allocation

Addresses review feedback from zmanian and gemini-code-assist.

Co-authored-by: j-bloggs <j-bloggs@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…#1677)

* fix(worker): treat empty LLM response after text output as completion

When a job's LLM produces a substantive text response (e.g., formatted
results from a routine) and the next LLM call returns empty or errors,
the worker now treats this as successful completion instead of
continuing the loop until failure.

Previously, empty responses always triggered TextAction::Continue,
causing the loop to re-call the LLM. The LLM had nothing more to say,
so the provider returned "Response contained no message or tool call
(empty)". This made routine jobs that successfully produced results
report as "failed".

The fix adds a `has_text_response` flag to JobDelegate:
- After any non-empty text response: flag is set
- Empty text after flag is set: treated as completion
- LLM errors (select_tools/respond_with_tools) after flag: treated
  as completion instead of propagating
- Empty text before any output: still retries (rate-limit backoff)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(worker): restrict error swallowing to EmptyResponse variant only

- Add LlmError::EmptyResponse variant for when LLM returns no content
- Update nearai_chat and github_copilot providers to emit EmptyResponse
  instead of InvalidResponse for empty/no-choice responses
- try_complete_on_error now only swallows EmptyResponse (not AuthFailed,
  ContextLengthExceeded, Http, Io, etc.)
- Extract is_completion_eligible_error as testable pure function
- Log mark_completed errors at warn level instead of silently dropping
- Add EmptyResponse to retry and circuit breaker transient classifications
- Rewrite test to exercise real classification logic against all variants

Addresses review feedback from zmanian and gemini-code-assist.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* refactor(worker): extract mark_completed_or_warn helper to DRY completion logic

Extract shared mark-completed + warn-on-failure pattern into a single
helper method used by both try_complete_on_error and handle_text_response.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: j-bloggs <j-bloggs@users.noreply.github.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* fix(routines): persist full LLM transcript and remove sandbox gate for full_job

Routine execution output was invisible — routine_fire returned a one-liner,
routine_history had no actual output, and the conversation thread contained
only a summary. Full-job routines also hard-failed without Docker.

Three fixes:

1. **Full transcript persistence**: execute_lightweight now persists every
   message (prompt, LLM responses, tool calls with params, tool results) to
   the routine's conversation thread as it executes, not just a summary
   after the fact.

2. **Routine output visibility**: routine_history includes conversation_id
   and recent_output messages. routine_fire tells the user to check
   routine_history. Web detail page has a "View Execution Thread" button
   that navigates to the chat tab. ROUTINE_OK stores "No issues found"
   instead of None. Full-job summary pulls actual job output instead of
   generic "Job X finished".

3. **Remove SandboxReadiness gate**: full_job routines dispatch through the
   scheduler like regular /job commands — no Docker required. The
   SandboxReadiness enum is removed entirely.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* style: apply cargo fmt

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(worker): treat AutonomousUnavailable tool errors as recoverable

The job worker crashed the entire job when a tool was denied for
autonomous execution (e.g. secret_list). The error was already recorded
in reason_ctx for the LLM to see, but process_tool_result_job returned
Err which propagated through the agentic loop and terminated the job.

Now all tool errors (including AutonomousUnavailable) return Ok,
letting the LLM see the denial and try a different approach.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(llm): sanitize tool names for OpenAI Codex Responses API

The Codex API requires tool names to match `^[a-zA-Z0-9_-]+$` but
MCP/extension tools can have dots in their names (e.g. `mcp.server.tool`).
This caused HTTP 400 errors when the job worker sent tool calls back
to the LLM.

Sanitize tool names in both `convert_tool_definition` and
`convert_message` (function_call items) by replacing invalid characters
with underscores.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(routines): inject execution context into full_job description [skip-regression-check]

When a full_job routine dispatches a job, the LLM had no context that
it was already executing inside a routine. It wasted iterations on
infrastructure (discovering tools, creating routines, setting up auth)
instead of doing the actual work.

Prepend a clear directive to the job description telling the LLM that
tools and the routine are already configured, and to execute the task
directly.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(mcp): auto-refresh expired OAuth tokens on access [skip-regression-check]

When IronClaw restarts, MCP servers fail with "Secret has expired"
because get_access_token() checks token expiry locally and returns an
error before any HTTP request is made — so the existing 401-retry
refresh logic never triggers.

Now get_access_token() catches SecretError::Expired and automatically
calls refresh_access_token() using the stored refresh token. If the
refresh succeeds, the new token is returned transparently. If it fails,
the error message includes both the expiry and the refresh failure.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(mcp): align refresh token naming and set expiry on stored tokens

Two bugs prevented MCP OAuth token auto-refresh on restart:

1. Naming mismatch: the hosted OAuth flow stored the refresh token as
   `{token_secret_name}_refresh_token` (e.g. `mcp_notion_access_token_refresh_token`)
   but `McpServerConfig::refresh_token_secret_name()` returned
   `mcp_notion_refresh_token`. The refresh token was there but unfindable.

2. Missing expiry: `store_tokens` in auth.rs never called `with_expiry()`
   even though `AccessToken::expires_in` was available. Combined with the
   fix from the previous commit (auto-refresh on Expired), tokens stored
   via the MCP auth flow will now also trigger refresh correctly.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(web): show activity and transitions for agent jobs in job detail [skip-regression-check]

The job events endpoint only checked sandbox jobs for ownership,
returning 404 for agent jobs dispatched from routines. The detail
handler also returned empty transitions for agent jobs.

- events handler: fall back to agent job ownership check
- detail handler: populate transitions from job's state history

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* feat(routines): expose max_iterations for full_job routines (default 25)

The max_iterations parameter was hardcoded to 10 and not configurable
via routine_create or routine_update, causing complex tasks to hit the
iteration cap.

- Add max_iterations to full_job execution schema (1-200, default 25)
- Thread it through parse → build → RoutineAction
- Support updating via routine_update
- Raise default from 10 to 25

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(routines): break self-dialogue loop after full_job plan execution

After plan execution, the completion-check Q&A ("Is the job complete?" /
"No, not complete...") was left in the message context, causing the
agentic loop to repeat the same analysis instead of calling tools.

Replace the stale dialogue with an action-oriented continuation prompt
that instructs the LLM to use tools for remaining work. Also strip
<suggestions> tags from all job output since they're only meaningful
for interactive chat sessions.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(repl): prevent test hang in single-message mode

In single-message mode, start() stored a clone of the mpsc sender in
self.msg_tx for approval injection. After the thread sent /quit and
exited, the stored clone kept the stream alive, so stream.next()
blocked forever in the test assertion that the stream ends.

Skip storing the sender in single-message mode since interactive
approval is not needed.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(jobs): treat text responses as final answer in agentic loop

When the LLM produces a non-empty text response with no tool intent
(already filtered by the nudge mechanism), it is the job's final
answer. Previously, handle_text_response only exited the loop if the
text matched rigid completion phrases like "job is complete". Natural
summaries like "Weekly review completed and saved to Notion" were
added to context and the loop continued, causing the LLM to restate
the same summary until max_iterations was hit.

Now any non-empty text response marks the job complete and stops the
loop, matching the chat dispatcher behavior.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* perf(tests): reduce skills catalog network failure test from 10s to 1s

The test_search_returns_error_on_network_failure test connects to an
unreachable RFC 5737 TEST-NET IP and waited for the full 10s production
REQUEST_TIMEOUT. Add with_url_and_timeout test helper and use a 1s
timeout instead. [skip-regression-check]

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(tools): accept 'message' as alias for 'content' in message tool

LLMs frequently call the message tool with {"message": "..."} instead
of {"content": "..."}. Fall back to the 'message' key when 'content'
is missing to avoid InvalidParameters errors during autonomous job
execution.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(tools): attach thread_id for gateway broadcast in message tool

When the message tool broadcasts to all channels (channel=null), it
sent an OutgoingResponse without a thread_id. The gateway silently
dropped these messages (returned Ok but never sent the SSE event),
so they appeared in repl but not in the web UI.

The thread_id was only populated when channel was explicitly "gateway".
Now it is always populated from notify_thread_id metadata, so
broadcast_all delivers to the gateway correctly.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(gateway): return error instead of silently dropping messages

Gateway broadcast() and respond() previously returned Ok(()) when
thread_id was missing, silently swallowing the message. Callers
(message tool, agent loop) believed delivery succeeded when it didn't.

Now returns ChannelError::MissingRoutingTarget so callers can detect
and report the failure. Four regression tests verify the contract:
respond/broadcast with and without thread_id.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: resolve rebase conflicts with staging

Restore sandbox_readiness field removed by pre-rebase commits (staging
still uses it). Update repl test to match staging's single-message
behavior (no longer sends /quit). Add missing reasoning field to
ToolCall in codex test.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(tools): log error when routine conversation lookup fails

The routine_history tool silently swallowed errors from
get_or_create_routine_conversation, returning empty output without
any diagnostic logging. Add tracing::warn so failures are visible
in logs. [skip-regression-check]

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: address PR #1650 review comments

- E2E test: accept submitted/accepted as success states in job assertion
- TimeTool: remove operation from required schema (defaults to "now")
- jobs handler: log DB errors server-side, return generic message to client
- routines handler: use read-only find_routine_conversation on GET
- codex provider: reverse-map sanitized tool names so MCP tools resolve

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: address zmanian review feedback on PR #1650

- MCP refresh token: fall back to legacy secret name (mcp_{name}_refresh_token)
  so existing users don't need to re-authenticate after the naming fix
- Job worker: replace fragile messages.pop() with truncate-to-saved-count
  to avoid maintenance hazard if message flow changes
- Document cost implications of max_iterations 10->25 default bump
- Revert Cargo.toml dist profile change (thin LTO comment, codegen-units=16)
  as it's unrelated to this PR

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: resolve rebase conflicts and address new Copilot comments

- Fix no_silent_drop tests for updated GatewayConfig (user_id moved to
  GatewayChannel::new second arg, user_tokens removed)
- Fix handle_text_response param name (_reason_ctx -> reason_ctx)
- Fix missing has_text_response field in test JobDelegate
- Propagate row.get errors in find_routine_conversation instead of
  unwrap_or_default
- Only fall back to legacy refresh token name on NotFound/Expired,
  propagate real errors (DB, decryption)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
* feat(discord): restore gateway channel flow in wasm

* chore(discord): bump channel version to 0.2.1

* fix(discord): address review feedback on gateway channel PR

- Add #[serde(default)] to DiscordMessageMetadata for backward compat
  with old Option<String> serialized metadata
- Restore mention polling alongside Gateway (on_poll processes gateway
  events first, then runs poll_for_mentions if configured)
- Update on_respond to handle source_message_id with message_reference
  for mention-poll reply threading
- Implement Gateway presence status: dnd before pairing, online after
- Implement Gateway resume (OP 6) with session_id tracking, falling
  back to fresh identify on Invalid Session (OP 9)
- Extract WebsocketSessionState and spawn_websocket_poll to reduce
  nesting in start_websocket_runtime
- Simplify should_apply_dm_pairing tautology
- Remove completed plan docs
- Fix clippy items_after_test_module in extensions handler

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(discord): address review findings in gateway channel PR

- Fix gateway presence always showing "online" by filtering empty
  owner_id strings from workspace store reads
- Fix interaction followup using POST instead of PATCH to
  /messages/@original, which left deferred "thinking" state unresolved
- Restore mention-poll pagination (up to 5 pages of 100 messages)
- Remove dead ed25519-dalek and hex dependencies from WASM crate
- Remove unused _channel_id parameter from remember_processed_id
- Clean up redundant let binding in send_pairing_reply

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(discord): address second-round review findings

- Log warning when gateway event queue JSON fails to deserialize
  instead of silently returning empty (zmanian review item 1)
- Defer presence update from OP 10 Hello to after OP 0 READY, per
  Discord gateway protocol which requires READY before non-Identify
  commands (zmanian review item 2)
- Add 0-25% random jitter to websocket reconnect backoff per Discord's
  reconnection recommendations (zmanian suggestion)
- Extract WebsocketPollContext struct to replace 19-parameter
  spawn_websocket_poll function (zmanian suggestion)
- Document intent bitmask 4609 = GUILDS + GUILD_MESSAGES +
  DIRECT_MESSAGES in capabilities JSON (zmanian suggestion)

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: zhyaoyu <zhyaoyu@aliyun.com>
Co-authored-by: ilblackdragon@gmail.com <ilblackdragon@gmail.com>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…#1667)

Add support for bundle layouts where directories without SKILL.md are
recursed into to find nested skills (e.g., skills/my-org/skill-a/SKILL.md).

- Add configurable max_scan_depth (SKILLS_MAX_SCAN_DEPTH env, default 3)
- Recurse into subdirectories lacking SKILL.md up to depth limit
- Share remaining discovery cap across recursive levels
- Replace try_exists + read_dir with single read_dir (eliminates TOCTOU)
- Box::pin recursive async calls for correct future sizing

Closes #1664

Co-authored-by: Rajul Bhatnagar <brajul@amazon.com>
* Clarify message tool and channel setup guidance

* Add target format hints to proactive messaging prompt

* Clarify search and message tool edge cases

* Fix stale tool_search e2e assertion

* Update src/llm/reasoning.rs

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>

* Tighten prompt guidance for message replies

* Format prompt guidance assertions

---------

Co-authored-by: Copilot <175728472+Copilot@users.noreply.github.com>
* fix: pin staging ci jobs to a single tested sha

* chore(ci): retrigger regression gate

[skip-regression-check]

---------

Co-authored-by: Firat Sertgoz <f@nuff.tech>
* fix: prevent UTF-8 panics in byte-index string truncation

Replace unsafe `&s[..n]` patterns with `floor_char_boundary(s, n)` at 3
production code sites where the truncation index could land mid-multibyte
character, panicking on non-ASCII input:

- src/llm/nearai_chat.rs: API response truncation in error message
- src/cli/memory.rs: memory content display truncation
- src/cli/config.rs: config value display truncation

All 3 sites operate on external or user-supplied strings that may contain
non-ASCII characters. The existing `crate::util::floor_char_boundary`
utility (used at 18 other call sites) walks back to the nearest char
boundary, preventing the panic.

Adds regression test with multi-byte characters (combining accents and
4-byte emoji) for truncate_content.

[skip-regression-check]

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: clarify test comment and use exact assertions

Address Gemini review feedback:
- Fix misleading comment: \u{00e9} is precomposed e-acute, not combining accent
- Replace weak assertions (ends_with/is_empty) with exact assert_eq!

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…nt (#1630)

* fix(bedrock): strip tool blocks from messages when toolConfig is absent

Bedrock's Converse API requires `toolConfig` whenever messages contain
`toolUse` or `toolResult` content blocks. When the agentic loop reaches
its force_text iteration (e.g. lightweight routine at max_iterations),
it switches from `complete_with_tools()` to `complete()` — but the
message history still carries tool blocks from prior iterations.

`convert_messages()` faithfully converts these into Bedrock content
blocks, and without `toolConfig` Bedrock rejects the request:

  "The toolConfig field must be defined when using toolUse and
   toolResult content blocks."

Add `strip_tool_blocks()` that converts tool interaction data to text:
- Assistant `tool_calls` → dropped (text content preserved)
- `Role::Tool` → `Role::User` with `[Tool ... returned: ...]` text

Wire it into:
- `complete()`: unconditionally, since it never sends toolConfig
- `complete_with_tools()`: when `build_tool_config()` returns None
  (empty tools or tool_choice="none")

Closes #1629

* fix(bedrock): address review feedback on strip_tool_blocks

- Add tracing::debug\! when tool blocks are stripped (zmanian suggestion)
- Add test for tool_choice="none" path (zmanian suggestion)
- Add inline comment on empty-content assistant behavior

---------

Co-authored-by: Rajul Bhatnagar <brajul@amazon.com>
…5295

chore: promote staging to staging-promote/70214c4a-23719079615 (2026-03-29 22:05 UTC)
…9615

chore: promote staging to staging-promote/86389dab-23706696435 (2026-03-29 21:07 UTC)
…6435

chore: promote staging to staging-promote/e0e530e6-23703082447 (2026-03-29 10:07 UTC)
…2447

chore: promote staging to staging-promote/a8e83210-23702343584 (2026-03-29 06:21 UTC)
…3584

chore: promote staging to staging-promote/8a320ae9-23693265249 (2026-03-29 05:32 UTC)
…5249

chore: promote staging to staging-promote/fd41bdf4-23691145719 (2026-03-28 20:05 UTC)
…5719

chore: promote staging to staging-promote/de5a1c7b-23688974037 (2026-03-28 18:06 UTC)
…4037

chore: promote staging to staging-promote/9bb19a98-23687925861 (2026-03-28 16:06 UTC)
…5861

chore: promote staging to staging-promote/9ba10eac-23686921981 (2026-03-28 15:07 UTC)
@henrypark133
henrypark133 merged commit 6c2c47b into staging-promote/8f8cb7f7-23680994633 Mar 30, 2026
12 of 13 checks passed
@henrypark133
henrypark133 deleted the staging-promote/9ba10eac-23686921981 branch March 30, 2026 19:33
@github-actions github-actions Bot added scope: agent Agent core (agent loop, router, scheduler) scope: channel/cli TUI / CLI channel scope: channel/web Web gateway channel scope: channel/wasm WASM channel runtime scope: tool/builtin Built-in tools scope: tool/mcp MCP client scope: db Database trait / abstraction scope: db/postgres PostgreSQL backend scope: llm LLM integration scope: worker Container worker scope: config Configuration scope: pairing Pairing mode scope: ci CI/CD workflows scope: dependencies Dependency updates risk: high Safety, secrets, auth, or critical infrastructure and removed risk: medium Business logic, config, or moderate-risk modules labels Mar 30, 2026
drchirag1991 pushed a commit to drchirag1991/ironclaw that referenced this pull request Apr 8, 2026
…3686921981

chore: promote staging to staging-promote/b365bb77-23680994633 (2026-03-28 14:09 UTC)
errol-t3 added a commit to Terminal-3/t3-claw that referenced this pull request Jun 30, 2026
The sidecar Dockerfile filtered `@terminal-3/t3n-mcp` (hyphenated), but
trinity nearai#1726 (dual-channel release) renamed the MCP package to
`@terminal3/t3n-mcp` (no hyphen). Once trinity-sdk-ref was bumped past
nearai#1726, the install filter `--filter @terminal-3/t3n-mcp...` matched no
project, so the t3n-sdk dev-dependencies (cross-env, rollup, …) were
never installed and the SDK build failed with `sh: cross-env: not found`.

Point both filters at the current package name. Verified
`pnpm --filter "@terminal3/t3n-mcp..."` now selects t3n-mcp + t3n-sdk,
restoring the dependency closure the build needs.

Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: high Safety, secrets, auth, or critical infrastructure scope: agent Agent core (agent loop, router, scheduler) scope: channel/cli TUI / CLI channel scope: channel/wasm WASM channel runtime scope: channel/web Web gateway channel scope: ci CI/CD workflows scope: config Configuration scope: db/postgres PostgreSQL backend scope: db Database trait / abstraction scope: dependencies Dependency updates scope: docs Documentation scope: extensions Extension management scope: llm LLM integration scope: pairing Pairing mode scope: tool/builtin Built-in tools scope: tool/mcp MCP client scope: tool/wasm WASM tool sandbox scope: worker Container worker size: XL 500+ changed lines staging-promotion

Projects

None yet

Development

Successfully merging this pull request may close these issues.