Skip to content

chore: promote staging to staging-promote/e88236ab-24647433078 (2026-04-20 04:51 UTC) - #2708

Merged
henrypark133 merged 4 commits into
mainfrom
staging-promote/b5dde504-24649025467
Apr 21, 2026
Merged

henrypark133 merged 4 commits into
mainfrom
staging-promote/b5dde504-24649025467

Conversation

@ironclaw-ci

@ironclaw-ci ironclaw-ci Bot commented Apr 20, 2026 •

Copy link
Copy Markdown
Contributor

Auto-promotion from staging CI

Batch range: 7fb41555a9e55677d1aaea29ca567a5b369c2b05..b5dde5048671b8be51430c79ef1e08c963c5360b
Promotion branch: staging-promote/b5dde504-24649025467
Base: staging-promote/e88236ab-24647433078
Triggered by: Staging CI batch at 2026-04-20 04:51 UTC

Commits in this batch (29):

Current commits in this promotion (4)

Current base: staging-promote/e88236ab-24647433078
Current head: staging-promote/b5dde504-24649025467
Current range: origin/staging-promote/e88236ab-24647433078..origin/staging-promote/b5dde504-24649025467

Auto-updated by staging promotion metadata workflow

Waiting for gates:

  • Tests: pending
  • E2E: pending
  • Claude Code review: pending (will post comments on this PR)

Auto-created by staging-ci workflow

ilblackdragon and others added 4 commits April 20, 2026 12:46
…helpers — ironclaw#2599 stage-6 prereq (#2704)

Promotes three `GatewayState` builders out of `server.rs::tests` (where
they were private to that test module) into
`src/channels/web/test_helpers.rs` as `pub(crate)` functions:

- `test_gateway_state(ext_mgr)`
- `test_gateway_state_with_dependencies(ext_mgr, store, db_auth, pairing_store)`
- `test_gateway_state_with_store_and_session_manager(store, session_manager)`

Why this is a prerequisite for stage 6 (deleting `server.rs`): the
caller-level tests in `server.rs::tests` that exercise `features/chat`,
`features/extensions`, `features/oauth`, and `features/pairing` all
consume at least one of these three builders. Until the builders had a
reachable home, the tests couldn't migrate into their respective slice
`mod tests` blocks — and without that, `server.rs::tests` can't be
deleted, so the back-compat shim can't go either.

The promoted functions keep the exact positional signatures they had
when they lived in `server.rs::tests`, so the future test migration in
stage 6 becomes a pure `git mv` + import-path update with zero API
surface changes.

Visibility / compilation scope:
- `TestGatewayBuilder` stays `pub` and always-compiled (integration
  tests in `tests/` import the crate without `cfg(test)` set).
- The three cross-slice builders are individually `#[cfg(test)]`-gated
  because their only callers are in-crate unit tests; keeping them
  un-gated would produce dead-code warnings in release builds.
- The `DbAuthenticator` and `ActiveConfigSnapshot` imports that only
  the cross-slice builders need are also `#[cfg(test)]`-gated.

Also updates:
- `features/chat/mod.rs` pending-follow-up comment: the reference
  helpers are now in `test_helpers`, so the note points to stage 6 as
  the migration step rather than a "promote helpers first" prereq.
- `channels/web/CLAUDE.md` file map: adds a `test_helpers.rs` row
  describing both the public builder and the three `pub(crate)` fns.

Mechanical verification:
- `cargo check -p ironclaw --tests --all-features` — clean
- `cargo check -p ironclaw --all-features` — clean (no dead-code warnings)
- `cargo check -p ironclaw --no-default-features --features libsql --tests` — clean
- `cargo clippy -p ironclaw --tests --all-features` — zero warnings
- `cargo test -p ironclaw --lib channels::web` — 431 passed
- `cargo test -p ironclaw --test multi_tenant_integration` — 40 passed
- `cargo test -p ironclaw --test openai_compat_integration` — 16 passed

Regression coverage: this is a pure relocation with no behavior change.
The existing 64 tests in `server.rs::tests` that consume these three
helpers continue to pass unmodified, which is the regression evidence.
A "test that would have caught this" would necessarily be identical to
the existing tests — no new test adds coverage.
[skip-regression-check]

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…2532)

* feat(gateway): expose engine v2 threads in chat history and sidebar

Engine v2 threads weren't appearing in the gateway sidebar and
deep-linking to one by id (`#/chat/<engine-thread-id>`) returned an
empty history because the v1 `assistant` flow dual-writes into the
single assistant conversation id, not the engine thread id.

Three coordinated fixes:

- `chat_history_handler`: extend the ownership check with an engine v2
  lookup so an engine thread id is recognized, then fall back to
  loading messages via `bridge::get_engine_thread` when the v1
  conversation table has nothing.
- `chat_threads_handler`: merge engine threads from
  `bridge::list_engine_threads` into the sidebar, label them with
  their goal, and re-sort by `updated_at`. Bump the v1 conversation
  cap from 50 to 500 so older threads stop silently aging off the
  sidebar.
- Gateway frontend (`app.js`): when restoring from `#/chat/<id>` on
  load, switch even if the id is not in the loaded sidebar list — the
  history endpoint resolves it via the DB. Log a warning instead of
  silently dropping the URL.

Cherry-picked from 8df22ab (feat/skills-engine-fixes), adapted to
staging where chat_history_handler and chat_threads_handler still live
in both server.rs and handlers/chat.rs.

Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com>

* fix(gateway): address v2 history review comments

* test(gateway): caller-level coverage for v2 thread ownership + history

Adds tests exercising chat_history_handler through the three ownership
branches the PR introduces, and two HTTP-driven e2e scenarios against
an ENGINE_V2=true server. Fills the caller-level coverage gap flagged
in review.

- Rust caller-level tests on chat_history_handler:
  - v2-owned: engine thread owned by user returns synthesized history
  - cross-user: alice can't read bob's engine thread (404)
  - session-only: in-memory session-owned thread returns 200 without DB
- Playwright e2e under ENGINE_V2=true:
  - engine-only thread appears in sidebar with channel=engine
  - deep-link by engine thread id returns synthesized turns

Exposes a minimal bridge::test_support module (ThreadTestStore,
install_engine_state_with_threads, clear_engine_state, shared test
lock) so cross-module tests can seed ENGINE_STATE without the weight
of the bridge's own full-featured TestStore.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* fix(ci): align with staging EngineState fields and clippy rules

- `EngineState` on staging now has `extension_manager` and `project_root`
  fields (added in #2549 and the attachments flow). Update the test-only
  `install_engine_state_with_threads` helper to populate them.
- Rewrite the `let Some(...) else { return None }` in
  `engine_history_entry_to_message` as a `?` — clippy::question_mark is
  denied on staging's all-features CI.
- Add missing `AuthenticatedUser` + `Query` imports to the server.rs
  test module for the `history_request` helper.
- `cargo fmt` on the 500-error test.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com>
Co-authored-by: Illia Polosukhin <ilblackdragon@gmail.com>
* feat(web): add debug inspector panel for web gateway chat UI (#1493)

Add a debug inspector sidebar with three tabs (Prompt, Activity, Stats)
activated via ?debug=true URL parameter. Consolidate theme-init.js into
a new init.js for early initialization. The panel shows real-time SSE
event timeline, system prompt component breakdown with token estimates,
and session-wide statistics including per-model usage.

- New files: init.js, debug-panel.js, debug-panel.css
- New endpoint: /api/debug/debug/prompt for system prompt inspection
- i18n support (en + zh-CN) for all debug panel strings
- Responsive layout: sidebar on desktop, overlay on tablet, hidden on mobile

* chore: minor

* feat(web): add debug inspection endpoints and verbose SSE mode (#1492)

Add per-subscriber verbose filtering to SSE, new AppEvent variants
(ToolResultFull, TurnMetrics), and enhanced debug prompt endpoint.
Debug subscribers (?debug=true) receive full tool output, per-LLM-call
metrics with model/duration/cache tokens, and tool parameters on
success. Non-debug subscribers see no change (backward compatible).

- New AppEvent variants: tool_result_full, turn_metrics with is_verbose_only()
- SseManager.subscribe()/subscribe_raw() accept verbose flag
- Emit TurnMetrics from dispatcher after each LLM call
- Emit ToolResultFull with 50KB cap after tool execution
- tool_completed() always includes redacted parameters
- /api/debug/prompt returns system_prompt, model, context_limit
- Frontend: turn-based activity tracking, turn navigation, message click
- Frontend: prompt tab with model name, progress bar, full prompt view
- Unit test for verbose SSE filtering

* fix(debug-panel): start turn counter at 0 so first message shows turn 1

* chore: minor

* chore: minor

* chore: fix lint

* fix(i18n): add Korean debug panel translations and fix hardcoded string

* fix(i18n): add Korean debug panel translations and fix hardcoded string

* fix: fix lint

* fix(web): propagate call_id to SSE events, gate debug mode on admin role, and skip verbose broadcasts without subscribers

- Add call_id field to AppEvent::ToolStarted/ToolCompleted/ToolResult and
  propagate from StatusUpdate conversion instead of silently dropping it,
  fixing mismatched tool start/complete pairs during concurrent same-name
  tool calls in the debug panel
- Update debug-panel.js to key pending tools by call_id (flat map) instead
  of FIFO name-based queues
- Skip ToolResultFull/TurnMetrics allocation and broadcast when no
  SSE/WebSocket subscribers are connected (SseManager::has_receivers)
- Require admin role for verbose/debug SSE and WebSocket event streams,
  matching the existing AdminUser gate on /api/debug/prompt
- Add audit log (tracing::debug) on debug prompt endpoint access

* fix(gateway): add admin gate to WS debug mode, add call_id to ToolResultFull, fix debug panel i18n

- Require admin role for WebSocket debug mode (server.rs), matching the
  existing SSE handler check — prevents non-admin users from receiving
  verbose tool output via ?debug=true
- Add call_id: Option<String> to StatusUpdate::ToolResultFull and
  AppEvent::ToolResultFull for correct concurrent same-name tool
  matching; update dispatcher, web gateway conversion, and debug-panel.js
- Remove dead chat_ws_handler from handlers/chat.rs (superseded by
  server.rs local version)
- Fix debug panel overlay blocking page on viewport resize by switching
  from inline style to CSS class toggle with transparent background
- Internationalize hardcoded English strings in debug panel (In/Out/
  Cost/Model/Cache labels) with en/ko/zh-CN translations
- Fix activity entries not updating on language switch: store labelKey,
  resolve during render without mutating entry, rebuild activity DOM
  in refreshDynamicI18n
- Fix pre-existing subscribe_raw() test compilation errors (missing
  verbose parameter)

---------

Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com>
Co-authored-by: Illia Polosukhin <ilblackdragon@gmail.com>
* fix(gateway): address PR #2622 review feedback

Five small fixes flagged by reviewers on the restage commit:

1. Trim trailing punctuation from `result.message` before
   `format!("{}. Resuming...", ...)` so messages like
   "Configuration saved for 'telegram'." don't render as
   "...telegram'.. Resuming..." (Copilot review).

2. Rename the error-context strings in `fail_waiting_thread` from
   "reconcile waiting thread" / "save reconciled thread" to "fail
   waiting thread" / "save failed thread" — the helper is now used
   for the no-auth-backend failure path too, not just orphan
   reconciliation (Copilot review).

3. Add a `debug!` log when an auth credential is intentionally
   dropped on the SkippedNoBackend + resume_output bare-test path,
   so the silent drop is observable to operators tracing this case
   (serrrfirat). Uses `debug!` (not `info!`) per the REPL/TUI
   logging rule in CLAUDE.md.

4. Document `submit_target` vs `credential_name` asymmetry in
   `submit_pending_auth_credential`'s doc comment — steps 1-2 take
   the extension identity, step 3 takes the credential identity
   because the secrets store has no extension concept (serrrfirat).

5. Document the credential-name validation chain on the step-3
   `secrets_store` fallback — upstream typing (`CredentialName`
   newtype validated at construction + pending-gate insertion by
   the engine) is the trust source, not a per-call check
   (serrrfirat).

Quality gate:
- cargo fmt --all
- cargo clippy --all --benches --tests --examples --all-features (zero warnings)
- 5 router tests pass

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

* test(gateway): regression test for auth-completed message trim

Extracts the trailing-period trim from `resolve_gate`'s auth-completed
arm into `format_auth_completed_resuming(raw: &str) -> String` and adds
a unit test pinning the expected behavior:

- Strips trailing period(s) from upstream backend messages so
  "Configuration saved for 'telegram'." renders as
  "Configuration saved for 'telegram'. Resuming..." instead of the
  prior "...telegram'.. Resuming..." double period.
- Multiple trailing periods + whitespace collapse to a single period.
- Messages with no trailing punctuation get exactly one period.
- Non-period punctuation (`!`, `?`) is intentionally left intact —
  the spec is "trim periods only", matching the motivating bug.

Also clarifies the inline doc comment to say "trailing period(s) +
whitespace" instead of "trailing punctuation + whitespace", per
@Copilot review feedback on PR #2701 — the predicate only matches `.`
plus whitespace, so the comment now matches the code.

Closes the regression-test enforcement gap that failed CI on the
original follow-up commit.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@github-actions github-actions Bot added size: XL 500+ changed lines scope: agent Agent core (agent loop, router, scheduler) scope: channel Channel infrastructure scope: channel/web Web gateway channel scope: channel/wasm WASM channel runtime scope: docs Documentation risk: medium Business logic, config, or moderate-risk modules contributor: core 20+ merged PRs and removed size: XL 500+ changed lines labels Apr 20, 2026
@claude

claude Bot commented Apr 20, 2026

Copy link
Copy Markdown

Code review

Found 4 issues:

  1. [HIGH:75] Memory leak: statsTimer never cleared in debug panel

    In crates/ironclaw_gateway/static/debug-panel.js, the statsTimer setInterval is created in init() but never cleared when the panel is closed or the page unloads. This causes the fetchGatewayStats() polling to continue making requests every 30 seconds even after the debug panel is closed.

    https://github.com/anthropics/ironclaw/blob/b5dde5048671b8be51430c79ef1e08c963c5360b/crates/ironclaw_gateway/static/debug-panel.js#L61-L64

    statsTimer = setInterval(function () {
      fetchGatewayStats();
      updateSseHealthDisplay();
    }, STATS_POLL_INTERVAL);

    The closePanel() function (line 207) does not call clearInterval(statsTimer). Add cleanup to prevent unbounded timer accumulation:

    function closePanel() {
      if (statsTimer) clearInterval(statsTimer);
      // ... rest of cleanup
    }
  2. [MEDIUM:75] Unnecessary string allocations before short-circuit check

    In src/agent/dispatcher.rs (ToolResultFull emission), the tool output string is cloned and capped at 50KB before send_status() is called. The actual short-circuit check for verbose-only events happens inside send_status(), so these allocations occur unconditionally even when no debug subscribers are connected. This results in ~50KB allocations on every tool execution.

    https://github.com/anthropics/ironclaw/blob/b5dde5048671b8be51430c79ef1e08c963c5360b/src/agent/dispatcher.rs#L1214-L1237

    Recommend: Check has_verbose_receivers() earlier before cloning the output, or defer the string capping to inside the channel's send_status() implementation.

  3. [MEDIUM:50] O(N²) activity log eviction in JavaScript

    In crates/ironclaw_gateway/static/debug-panel.js, the activity log eviction logic contains a nested loop. When the activity log approaches its cap (1000 entries) and new events arrive rapidly, the eviction scan for empty turns becomes O(N²). In high-activity debug sessions, this could cause frame drops in the debug panel.

    https://github.com/anthropics/ironclaw/blob/b5dde5048671b8be51430c79ef1e08c963c5360b/crates/ironclaw_gateway/static/debug-panel.js#L500-L530

    Suggest: Track turn membership counts in a Map<turn, count> to avoid re-scanning the entire activity log on each eviction.

  4. [LOW:25] Unbounded string concatenation in appendActivityOutput()

    In crates/ironclaw_gateway/static/debug-panel.js, tool output is concatenated with entry.output + '\n' + text on each tool_result_full event. With the 50KB backend cap, single large tool outputs could still accumulate multiple string copies in the DOM. Consider streaming large outputs or storing them separately from the activity log to minimize string allocation overhead.

Base automatically changed from staging-promote/e88236ab-24647433078 to main April 21, 2026 03:18
@henrypark133
henrypark133 merged commit b5dde50 into main Apr 21, 2026
56 of 68 checks passed
@henrypark133
henrypark133 deleted the staging-promote/b5dde504-24649025467 branch April 21, 2026 03:18

This branch had an error being deployed

1 failed and 5 inactive deployments
cosmose-ironclaw / production — b5dde504 Deployed Apr 20, 2026 by railway-app[bot]
Ironclaw-QA / production — b5dde504 Deployed Apr 20, 2026 by railway-app[bot]
humble-cat / staging-cameron — b5dde504 Deployed Apr 20, 2026 by railway-app[bot]
Near Foundation Ironclaw / production — b5dde504 Deployed Apr 20, 2026 by railway-app[bot]
ironclaw-nearai / production — b5dde504 Deployed Apr 20, 2026 by railway-app[bot]
venice-ironclaw / production — b5dde504 Deployed Apr 20, 2026 by railway-app[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

contributor: core 20+ merged PRs risk: medium Business logic, config, or moderate-risk modules scope: agent Agent core (agent loop, router, scheduler) scope: channel/wasm WASM channel runtime scope: channel/web Web gateway channel scope: channel Channel infrastructure scope: docs Documentation staging-promotion

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants