chore: promote staging to staging-promote/e88236ab-24647433078 (2026-04-20 04:51 UTC) - #2708
Conversation
…helpers — ironclaw#2599 stage-6 prereq (#2704) Promotes three `GatewayState` builders out of `server.rs::tests` (where they were private to that test module) into `src/channels/web/test_helpers.rs` as `pub(crate)` functions: - `test_gateway_state(ext_mgr)` - `test_gateway_state_with_dependencies(ext_mgr, store, db_auth, pairing_store)` - `test_gateway_state_with_store_and_session_manager(store, session_manager)` Why this is a prerequisite for stage 6 (deleting `server.rs`): the caller-level tests in `server.rs::tests` that exercise `features/chat`, `features/extensions`, `features/oauth`, and `features/pairing` all consume at least one of these three builders. Until the builders had a reachable home, the tests couldn't migrate into their respective slice `mod tests` blocks — and without that, `server.rs::tests` can't be deleted, so the back-compat shim can't go either. The promoted functions keep the exact positional signatures they had when they lived in `server.rs::tests`, so the future test migration in stage 6 becomes a pure `git mv` + import-path update with zero API surface changes. Visibility / compilation scope: - `TestGatewayBuilder` stays `pub` and always-compiled (integration tests in `tests/` import the crate without `cfg(test)` set). - The three cross-slice builders are individually `#[cfg(test)]`-gated because their only callers are in-crate unit tests; keeping them un-gated would produce dead-code warnings in release builds. - The `DbAuthenticator` and `ActiveConfigSnapshot` imports that only the cross-slice builders need are also `#[cfg(test)]`-gated. Also updates: - `features/chat/mod.rs` pending-follow-up comment: the reference helpers are now in `test_helpers`, so the note points to stage 6 as the migration step rather than a "promote helpers first" prereq. - `channels/web/CLAUDE.md` file map: adds a `test_helpers.rs` row describing both the public builder and the three `pub(crate)` fns. Mechanical verification: - `cargo check -p ironclaw --tests --all-features` — clean - `cargo check -p ironclaw --all-features` — clean (no dead-code warnings) - `cargo check -p ironclaw --no-default-features --features libsql --tests` — clean - `cargo clippy -p ironclaw --tests --all-features` — zero warnings - `cargo test -p ironclaw --lib channels::web` — 431 passed - `cargo test -p ironclaw --test multi_tenant_integration` — 40 passed - `cargo test -p ironclaw --test openai_compat_integration` — 16 passed Regression coverage: this is a pure relocation with no behavior change. The existing 64 tests in `server.rs::tests` that consume these three helpers continue to pass unmodified, which is the regression evidence. A "test that would have caught this" would necessarily be identical to the existing tests — no new test adds coverage. [skip-regression-check] Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…2532) * feat(gateway): expose engine v2 threads in chat history and sidebar Engine v2 threads weren't appearing in the gateway sidebar and deep-linking to one by id (`#/chat/<engine-thread-id>`) returned an empty history because the v1 `assistant` flow dual-writes into the single assistant conversation id, not the engine thread id. Three coordinated fixes: - `chat_history_handler`: extend the ownership check with an engine v2 lookup so an engine thread id is recognized, then fall back to loading messages via `bridge::get_engine_thread` when the v1 conversation table has nothing. - `chat_threads_handler`: merge engine threads from `bridge::list_engine_threads` into the sidebar, label them with their goal, and re-sort by `updated_at`. Bump the v1 conversation cap from 50 to 500 so older threads stop silently aging off the sidebar. - Gateway frontend (`app.js`): when restoring from `#/chat/<id>` on load, switch even if the id is not in the loaded sidebar list — the history endpoint resolves it via the DB. Log a warning instead of silently dropping the URL. Cherry-picked from 8df22ab (feat/skills-engine-fixes), adapted to staging where chat_history_handler and chat_threads_handler still live in both server.rs and handlers/chat.rs. Co-Authored-By: Claude Opus 4.6 (1M context) <noreply@anthropic.com> * fix(gateway): address v2 history review comments * test(gateway): caller-level coverage for v2 thread ownership + history Adds tests exercising chat_history_handler through the three ownership branches the PR introduces, and two HTTP-driven e2e scenarios against an ENGINE_V2=true server. Fills the caller-level coverage gap flagged in review. - Rust caller-level tests on chat_history_handler: - v2-owned: engine thread owned by user returns synthesized history - cross-user: alice can't read bob's engine thread (404) - session-only: in-memory session-owned thread returns 200 without DB - Playwright e2e under ENGINE_V2=true: - engine-only thread appears in sidebar with channel=engine - deep-link by engine thread id returns synthesized turns Exposes a minimal bridge::test_support module (ThreadTestStore, install_engine_state_with_threads, clear_engine_state, shared test lock) so cross-module tests can seed ENGINE_STATE without the weight of the bridge's own full-featured TestStore. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * fix(ci): align with staging EngineState fields and clippy rules - `EngineState` on staging now has `extension_manager` and `project_root` fields (added in #2549 and the attachments flow). Update the test-only `install_engine_state_with_threads` helper to populate them. - Rewrite the `let Some(...) else { return None }` in `engine_history_entry_to_message` as a `?` — clippy::question_mark is denied on staging's all-features CI. - Add missing `AuthenticatedUser` + `Query` imports to the server.rs test module for the `history_request` helper. - `cargo fmt` on the 500-error test. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.6 (1M context) <noreply@anthropic.com> Co-authored-by: Illia Polosukhin <ilblackdragon@gmail.com>
* feat(web): add debug inspector panel for web gateway chat UI (#1493) Add a debug inspector sidebar with three tabs (Prompt, Activity, Stats) activated via ?debug=true URL parameter. Consolidate theme-init.js into a new init.js for early initialization. The panel shows real-time SSE event timeline, system prompt component breakdown with token estimates, and session-wide statistics including per-model usage. - New files: init.js, debug-panel.js, debug-panel.css - New endpoint: /api/debug/debug/prompt for system prompt inspection - i18n support (en + zh-CN) for all debug panel strings - Responsive layout: sidebar on desktop, overlay on tablet, hidden on mobile * chore: minor * feat(web): add debug inspection endpoints and verbose SSE mode (#1492) Add per-subscriber verbose filtering to SSE, new AppEvent variants (ToolResultFull, TurnMetrics), and enhanced debug prompt endpoint. Debug subscribers (?debug=true) receive full tool output, per-LLM-call metrics with model/duration/cache tokens, and tool parameters on success. Non-debug subscribers see no change (backward compatible). - New AppEvent variants: tool_result_full, turn_metrics with is_verbose_only() - SseManager.subscribe()/subscribe_raw() accept verbose flag - Emit TurnMetrics from dispatcher after each LLM call - Emit ToolResultFull with 50KB cap after tool execution - tool_completed() always includes redacted parameters - /api/debug/prompt returns system_prompt, model, context_limit - Frontend: turn-based activity tracking, turn navigation, message click - Frontend: prompt tab with model name, progress bar, full prompt view - Unit test for verbose SSE filtering * fix(debug-panel): start turn counter at 0 so first message shows turn 1 * chore: minor * chore: minor * chore: fix lint * fix(i18n): add Korean debug panel translations and fix hardcoded string * fix(i18n): add Korean debug panel translations and fix hardcoded string * fix: fix lint * fix(web): propagate call_id to SSE events, gate debug mode on admin role, and skip verbose broadcasts without subscribers - Add call_id field to AppEvent::ToolStarted/ToolCompleted/ToolResult and propagate from StatusUpdate conversion instead of silently dropping it, fixing mismatched tool start/complete pairs during concurrent same-name tool calls in the debug panel - Update debug-panel.js to key pending tools by call_id (flat map) instead of FIFO name-based queues - Skip ToolResultFull/TurnMetrics allocation and broadcast when no SSE/WebSocket subscribers are connected (SseManager::has_receivers) - Require admin role for verbose/debug SSE and WebSocket event streams, matching the existing AdminUser gate on /api/debug/prompt - Add audit log (tracing::debug) on debug prompt endpoint access * fix(gateway): add admin gate to WS debug mode, add call_id to ToolResultFull, fix debug panel i18n - Require admin role for WebSocket debug mode (server.rs), matching the existing SSE handler check — prevents non-admin users from receiving verbose tool output via ?debug=true - Add call_id: Option<String> to StatusUpdate::ToolResultFull and AppEvent::ToolResultFull for correct concurrent same-name tool matching; update dispatcher, web gateway conversion, and debug-panel.js - Remove dead chat_ws_handler from handlers/chat.rs (superseded by server.rs local version) - Fix debug panel overlay blocking page on viewport resize by switching from inline style to CSS class toggle with transparent background - Internationalize hardcoded English strings in debug panel (In/Out/ Cost/Model/Cache labels) with en/ko/zh-CN translations - Fix activity entries not updating on language switch: store labelKey, resolve during render without mutating entry, rebuild activity DOM in refreshDynamicI18n - Fix pre-existing subscribe_raw() test compilation errors (missing verbose parameter) --------- Co-authored-by: Claude Sonnet 4.6 <noreply@anthropic.com> Co-authored-by: Illia Polosukhin <ilblackdragon@gmail.com>
* fix(gateway): address PR #2622 review feedback Five small fixes flagged by reviewers on the restage commit: 1. Trim trailing punctuation from `result.message` before `format!("{}. Resuming...", ...)` so messages like "Configuration saved for 'telegram'." don't render as "...telegram'.. Resuming..." (Copilot review). 2. Rename the error-context strings in `fail_waiting_thread` from "reconcile waiting thread" / "save reconciled thread" to "fail waiting thread" / "save failed thread" — the helper is now used for the no-auth-backend failure path too, not just orphan reconciliation (Copilot review). 3. Add a `debug!` log when an auth credential is intentionally dropped on the SkippedNoBackend + resume_output bare-test path, so the silent drop is observable to operators tracing this case (serrrfirat). Uses `debug!` (not `info!`) per the REPL/TUI logging rule in CLAUDE.md. 4. Document `submit_target` vs `credential_name` asymmetry in `submit_pending_auth_credential`'s doc comment — steps 1-2 take the extension identity, step 3 takes the credential identity because the secrets store has no extension concept (serrrfirat). 5. Document the credential-name validation chain on the step-3 `secrets_store` fallback — upstream typing (`CredentialName` newtype validated at construction + pending-gate insertion by the engine) is the trust source, not a per-call check (serrrfirat). Quality gate: - cargo fmt --all - cargo clippy --all --benches --tests --examples --all-features (zero warnings) - 5 router tests pass Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> * test(gateway): regression test for auth-completed message trim Extracts the trailing-period trim from `resolve_gate`'s auth-completed arm into `format_auth_completed_resuming(raw: &str) -> String` and adds a unit test pinning the expected behavior: - Strips trailing period(s) from upstream backend messages so "Configuration saved for 'telegram'." renders as "Configuration saved for 'telegram'. Resuming..." instead of the prior "...telegram'.. Resuming..." double period. - Multiple trailing periods + whitespace collapse to a single period. - Messages with no trailing punctuation get exactly one period. - Non-period punctuation (`!`, `?`) is intentionally left intact — the spec is "trim periods only", matching the motivating bug. Also clarifies the inline doc comment to say "trailing period(s) + whitespace" instead of "trailing punctuation + whitespace", per @Copilot review feedback on PR #2701 — the predicate only matches `.` plus whitespace, so the comment now matches the code. Closes the regression-test enforcement gap that failed CI on the original follow-up commit. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Code reviewFound 4 issues:
|
Auto-promotion from staging CI
Batch range:
7fb41555a9e55677d1aaea29ca567a5b369c2b05..b5dde5048671b8be51430c79ef1e08c963c5360bPromotion branch:
staging-promote/b5dde504-24649025467Base:
staging-promote/e88236ab-24647433078Triggered by: Staging CI batch at 2026-04-20 04:51 UTC
Commits in this batch (29):
Current commits in this promotion (4)
Current base:
staging-promote/e88236ab-24647433078Current head:
staging-promote/b5dde504-24649025467Current range:
origin/staging-promote/e88236ab-24647433078..origin/staging-promote/b5dde504-24649025467Auto-updated by staging promotion metadata workflow
Waiting for gates:
Auto-created by staging-ci workflow