chore: promote staging to staging-promote/ecd37e10-24483216739 (2026-04-16 12:18 UTC) - #2525
Merged
Merged
Conversation
…#2518) * fix(gateway): extend settings search to card, tool, and user sections Settings search only filtered .settings-row elements, leaving Channels, Extensions, MCP, Skills, Tools, and User Management sections unsearchable. Add filtering for .ext-card, .tool-permission-row, and #users-tbody tr elements, and update CSS to hide matched elements. * fix(gateway): extend settings search to card, tool, and user sections Settings search only filtered .settings-row elements, leaving Channels, Extensions, MCP, Skills, Tools, and User Management sections unsearchable. Add filtering for .ext-card, .tool-permission-row, and #users-tbody tr elements, update CSS to hide matched elements, and reorder logic so container visibility checks run after all items are filtered. Add E2E tests covering search across tool rows and extension cards. * chore: minor --------- Co-authored-by: Robert Yan <46699230+think-in-universe@users.noreply.github.com>
…2458) * fix(engine): normalize granted action aliases across lease checks Keep lease preflight, policy, and consumption consistent for hyphen/underscore action aliases so installed tools do not fail mid-turn after being allowed. Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent) Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai> * fix(web): preserve pending gate call ids on auth resume Resolve or synthesize the original action call id when resuming auth or external callback gates so resumed ActionResult messages remain correctly paired with the waiting assistant call. Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent) Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai> * test(engine): cover install resume followed by aliased tool use Add a higher-fidelity v2 gate integration regression that proves an install auth resume can flow directly into an aliased follow-up tool call and still complete the thread instead of stalling. Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent) Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai> * fix(engine): apply policy checks to aliased action names Resolve structured preflight action definitions with the same hyphen/underscore alias semantics as lease matching so aliased calls cannot bypass approval or deny policies. Ultraworked with [Sisyphus](https://github.com/code-yeongyu/oh-my-openagent) Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai> * style: cargo fmt Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(bridge): scan internal_messages in legacy call_id fallback The `resolved_call_id_for_pending_action` legacy fallback scanned only `thread.messages`, but in production the orchestrator writes ActionResult and assistant messages to `thread.internal_messages` via `sync_runtime_state`. This meant the `resolved_ids` set was always empty and the fallback never found a match, silently falling through to a synthetic id. Scan both `messages` and `internal_messages` so the legacy path works correctly for orchestrator-driven threads. Addresses review feedback from @standardtoaster on #2458. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * style: format bridge router --------- Co-authored-by: Sisyphus <clio-agent@sisyphuslabs.ai> Co-authored-by: Zaki <zaki@iqlusion.io> Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
…on (#2483) * feat(engine): add code execution failure categorization instrumentation Add structured error classification to the v2 engine's CodeAct execution path so we can measure whether REPL failures come from Monty VM limitations, LLM logic errors, tool dispatch issues, or resource limits. - Add CodeExecutionFailure enum (8 categories: SyntaxError, RuntimeError, NameLookup, VmPanic, ResourceLimit, ToolError, GatePause, OsDenied) - Add CodeExecutionFailed event kind to EventKind for event sourcing - Tag every error return path in scripting.rs with the correct category - Emit CodeExecutionFailed events from the orchestrator on code errors - Enhance trace analyzer to use structured events (with fallback to message-level pattern matching for pre-instrumentation threads) - Expand fallback error patterns from 4 to 10 Python exception types - Add 13 regression tests covering classification and trace detection This enables aggregate queries like "what % of code failures are Monty VM panics vs LLM generating bad Python" to inform runtime decisions. https://claude.ai/code/session_018jFKVTjv1pkzwJobw43HwP * fix(engine): address ilblackdragon + gemini review — failure instrumentation correctness (#2483) - Replace `had_error: bool` + `failure_category: Option<_>` with single `failure: Option<CodeExecutionFailure>` field, making invalid states unrepresentable - Remove `GatePause` variant (gate pauses are suspensions, not failures) - Convert all 9 catch_unwind Err(_) paths to emit VmPanic instead of propagating EngineError, so panics get proper instrumentation events - Fix error_text truncation: take last 500 chars (where tracebacks are), not first 500 chars (where print output is) - Thread real `Instant::now()` timing through `duration_ms` instead of hardcoded 0 - Tighten `classify_runtime_error`: "syntax" → "syntaxerror", remove loose `"os" && "denied"` substring match, reorder checks - Replace `DefaultHasher` with FNV-1a for stable cross-version hashing - Add `#[serde(rename_all = "snake_case")]` so Serialize matches Display - Add `#[serde(other)] Unknown` to EventKind for forward-compat - Fix misleading CodeExecutionFailed docstring - Add orchestrator caller test for CodeExecutionFailed event emission - Use `to_ascii_lowercase()` per project convention Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix: resolve clippy warnings — simplify boolean expressions and remove tautological assert - Replace `!result.failure.is_some()` with `result.failure.is_none()` in scripting tests - Remove tautological `duration_ms >= 0` assertion on unsigned type in orchestrator test - Add missing V24 migration checksum to checksums.lock Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): address ilblackdragon review nits — tail_chars helper, tighten fuel match, clippy (#2483) - Extract `tail_chars(s, n)` helper to deduplicate last-N-chars logic in `handle_execute_code_step` (used by both ActionFailed and CodeExecutionFailed event emission) - Tighten `classify_runtime_error` fuel match from `contains("fuel")` to `contains("out of fuel") || contains("fuel exhausted")` to avoid miscategorizing runtime errors that mention the word "fuel" - Fix test message typo: "should set had_error" → "should set failure" - Fix clippy warnings: `!x.is_some()` → `x.is_none()`, remove tautological `u64 >= 0` assertion Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(engine): drop stray V24 checksums.lock entry (#2483) Rebase artifact — no V24 migration exists in this PR. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude <noreply@anthropic.com>
This branch had an error being deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Auto-promotion from staging CI
Batch range:
a53eac5c2dec6b6cd5c08189086093fde64aa9cb..017218d0a342ba6bda4c9b9f755db3c352a0e65bPromotion branch:
staging-promote/017218d0-24509761033Base:
staging-promote/ecd37e10-24483216739Triggered by: Staging CI batch at 2026-04-16 12:18 UTC
Commits in this batch (55):
ironclaw profile listsubcommand (feat(cli): addironclaw profile listsubcommand #2288)reasoning_contentfields in chat completions response (fix: duplicatereasoning_contentfields in chat completions response #2493)Current commits in this promotion (0)
Current base:
mainCurrent head:
staging-promote/017218d0-24509761033Current range:
origin/main..origin/staging-promote/017218d0-24509761033Auto-updated by staging promotion metadata workflow
Waiting for gates:
Auto-created by staging-ci workflow