Skip to content

chore: promote staging to staging-promote/ecd37e10-24483216739 (2026-04-16 12:18 UTC) - #2525

Merged
henrypark133 merged 3 commits into
mainfrom
staging-promote/017218d0-24509761033
Apr 18, 2026
Merged

henrypark133 merged 3 commits into
mainfrom
staging-promote/017218d0-24509761033

Conversation

@ironclaw-ci

@ironclaw-ci ironclaw-ci Bot commented Apr 16, 2026 •

Copy link
Copy Markdown
Contributor

Auto-promotion from staging CI

Batch range: a53eac5c2dec6b6cd5c08189086093fde64aa9cb..017218d0a342ba6bda4c9b9f755db3c352a0e65b
Promotion branch: staging-promote/017218d0-24509761033
Base: staging-promote/ecd37e10-24483216739
Triggered by: Staging CI batch at 2026-04-16 12:18 UTC

Commits in this batch (55):

Current commits in this promotion (0)

Current base: main
Current head: staging-promote/017218d0-24509761033
Current range: origin/main..origin/staging-promote/017218d0-24509761033

  • (no non-merge commits in range)

Auto-updated by staging promotion metadata workflow

Waiting for gates:

  • Tests: pending
  • E2E: pending
  • Claude Code review: pending (will post comments on this PR)

Auto-created by staging-ci workflow

italic-jinxin and others added 3 commits April 16, 2026 15:07
…#2518)

* fix(gateway): extend settings search to card, tool, and user sections

Settings search only filtered .settings-row elements, leaving Channels,
Extensions, MCP, Skills, Tools, and User Management sections unsearchable.
Add filtering for .ext-card, .tool-permission-row, and #users-tbody tr
elements, and update CSS to hide matched elements.

* fix(gateway): extend settings search to card, tool, and user sections

Settings search only filtered .settings-row elements, leaving Channels,
Extensions, MCP, Skills, Tools, and User Management sections unsearchable.
Add filtering for .ext-card, .tool-permission-row, and #users-tbody tr
elements, update CSS to hide matched elements, and reorder logic so
container visibility checks run after all items are filtered.

Add E2E tests covering search across tool rows and extension cards.

* chore: minor

---------

Co-authored-by: Robert Yan <46699230+think-in-universe@users.noreply.github.com>
…2458)

* fix(engine): normalize granted action aliases across lease checks

Keep lease preflight, policy, and consumption consistent for hyphen/underscore action aliases so installed tools do not fail mid-turn after being allowed.

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>

* fix(web): preserve pending gate call ids on auth resume

Resolve or synthesize the original action call id when resuming auth or external callback gates so resumed ActionResult messages remain correctly paired with the waiting assistant call.

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>

* test(engine): cover install resume followed by aliased tool use

Add a higher-fidelity v2 gate integration regression that proves an install auth resume can flow directly into an aliased follow-up tool call and still complete the thread instead of stalling.

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>

* fix(engine): apply policy checks to aliased action names

Resolve structured preflight action definitions with the same hyphen/underscore alias semantics as lease matching so aliased calls cannot bypass approval or deny policies.

Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent)

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>

* style: cargo fmt

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(bridge): scan internal_messages in legacy call_id fallback

The `resolved_call_id_for_pending_action` legacy fallback scanned only
`thread.messages`, but in production the orchestrator writes ActionResult
and assistant messages to `thread.internal_messages` via
`sync_runtime_state`. This meant the `resolved_ids` set was always empty
and the fallback never found a match, silently falling through to a
synthetic id.

Scan both `messages` and `internal_messages` so the legacy path works
correctly for orchestrator-driven threads.

Addresses review feedback from @standardtoaster on #2458.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* style: format bridge router

---------

Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai>
Co-authored-by: Zaki <zaki@iqlusion.io>
Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…on (#2483)

* feat(engine): add code execution failure categorization instrumentation

Add structured error classification to the v2 engine's CodeAct execution
path so we can measure whether REPL failures come from Monty VM
limitations, LLM logic errors, tool dispatch issues, or resource limits.

- Add CodeExecutionFailure enum (8 categories: SyntaxError, RuntimeError,
  NameLookup, VmPanic, ResourceLimit, ToolError, GatePause, OsDenied)
- Add CodeExecutionFailed event kind to EventKind for event sourcing
- Tag every error return path in scripting.rs with the correct category
- Emit CodeExecutionFailed events from the orchestrator on code errors
- Enhance trace analyzer to use structured events (with fallback to
  message-level pattern matching for pre-instrumentation threads)
- Expand fallback error patterns from 4 to 10 Python exception types
- Add 13 regression tests covering classification and trace detection

This enables aggregate queries like "what % of code failures are Monty
VM panics vs LLM generating bad Python" to inform runtime decisions.

https://claude.ai/code/session_018jFKVTjv1pkzwJobw43HwP

* fix(engine): address ilblackdragon + gemini review — failure instrumentation correctness (#2483)

- Replace `had_error: bool` + `failure_category: Option<_>` with single
  `failure: Option<CodeExecutionFailure>` field, making invalid states
  unrepresentable
- Remove `GatePause` variant (gate pauses are suspensions, not failures)
- Convert all 9 catch_unwind Err(_) paths to emit VmPanic instead of
  propagating EngineError, so panics get proper instrumentation events
- Fix error_text truncation: take last 500 chars (where tracebacks are),
  not first 500 chars (where print output is)
- Thread real `Instant::now()` timing through `duration_ms` instead of
  hardcoded 0
- Tighten `classify_runtime_error`: "syntax" → "syntaxerror", remove
  loose `"os" && "denied"` substring match, reorder checks
- Replace `DefaultHasher` with FNV-1a for stable cross-version hashing
- Add `#[serde(rename_all = "snake_case")]` so Serialize matches Display
- Add `#[serde(other)] Unknown` to EventKind for forward-compat
- Fix misleading CodeExecutionFailed docstring
- Add orchestrator caller test for CodeExecutionFailed event emission
- Use `to_ascii_lowercase()` per project convention

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix: resolve clippy warnings — simplify boolean expressions and remove tautological assert

- Replace `!result.failure.is_some()` with `result.failure.is_none()` in scripting tests
- Remove tautological `duration_ms >= 0` assertion on unsigned type in orchestrator test
- Add missing V24 migration checksum to checksums.lock

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(engine): address ilblackdragon review nits — tail_chars helper, tighten fuel match, clippy (#2483)

- Extract `tail_chars(s, n)` helper to deduplicate last-N-chars logic in
  `handle_execute_code_step` (used by both ActionFailed and
  CodeExecutionFailed event emission)
- Tighten `classify_runtime_error` fuel match from `contains("fuel")` to
  `contains("out of fuel") || contains("fuel exhausted")` to avoid
  miscategorizing runtime errors that mention the word "fuel"
- Fix test message typo: "should set had_error" → "should set failure"
- Fix clippy warnings: `!x.is_some()` → `x.is_none()`, remove
  tautological `u64 >= 0` assertion

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(engine): drop stray V24 checksums.lock entry (#2483)

Rebase artifact — no V24 migration exists in this PR.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude <noreply@anthropic.com>
@github-actions github-actions Bot added size: XL 500+ changed lines risk: low Changes to docs, tests, or low-risk modules contributor: core 20+ merged PRs labels Apr 16, 2026
Base automatically changed from staging-promote/ecd37e10-24483216739 to main April 18, 2026 01:00
@henrypark133
henrypark133 merged commit 017218d into main Apr 18, 2026
38 of 47 checks passed
@henrypark133
henrypark133 deleted the staging-promote/017218d0-24509761033 branch April 18, 2026 01:00

This branch had an error being deployed

1 failed and 5 inactive deployments
venice-ironclaw / production — 017218d0 Deployed Apr 16, 2026 by railway-app[bot]
cosmose-ironclaw / production — 017218d0 Deployed Apr 16, 2026 by railway-app[bot]
Ironclaw-QA / production — 017218d0 Deployed Apr 16, 2026 by railway-app[bot]
humble-cat / staging-cameron — 017218d0 Deployed Apr 16, 2026 by railway-app[bot]
Near Foundation Ironclaw / production — 017218d0 Deployed Apr 16, 2026 by railway-app[bot]
ironclaw-nearai / production — 017218d0 Deployed Apr 16, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: low Changes to docs, tests, or low-risk modules size: XL 500+ changed lines staging-promotion

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants