Repository navigation
feat(responses): virtualize previous_response_id continuation regardless of upstream support - #10262
Conversation
|
CI here is failing on pre-existing base-red on
|
…ess of upstream support OmniRoute now exposes OpenAI-compatible previous_response_id/store continuation to clients unconditionally, even when the selected upstream provider has no native Responses-API state support. Reconstruction happens server-side in handleChatImplementation, before any downstream validation or provider translation: OmniRoute resolves the response id back to the full input/output it previously produced, prepends it to the client's delta, and forwards the full reconstructed history upstream exactly as it does today. Client<->OmniRoute traffic shrinks to the new delta only; OmniRoute<->provider traffic is unchanged. Storage reuses the existing call-log pipeline artifact (already gated by call_log_pipeline_enabled, already retained/cleaned up by the existing call-log lifecycle) instead of duplicating conversation content into a second store -- only a lightweight call_logs.response_id index is new. Every lookup is scoped by api_key_id so one client can never resolve another client's stored conversation, and any unresolvable/missing/ size-limit-omitted state fails closed with OpenAI's own previous_response_not_found contract. Stacked on feat/openai-responses-store-toggle (diegosouzapw#10121).
check-db-rules requires every db/ module to be re-exported (or explicitly allowlisted as intentionally-internal) for discoverability. Missed this when the module was first added.
bdac2ec to
2287c1e
Compare
…gosouzapw#10262 153/154 (originally 154/155) were chosen back when this branch stacked on top of the previous_response_id migration (153_call_logs_response_id.sql). Decoupling removed that migration from this branch's history, leaving an unused 153 slot that check-migration-numbering.test.ts correctly flags as a gap.
|
Thanks for this — reusing the call-log artifact instead of a second store is the right call, and the fail-closed/tenant-scoped design in Two things I'd like resolved before merge:
Happy to take another look once those are addressed — the core mechanism (fail-closed reconstruction from the call-log artifact) looks good. |
…nses-previous-response-id-virtualization
The migration was numbered 153, but release/v3.8.50 already carries 153_radar_local_model_state.sql. The emngrating runner's collision guard throws on two live .sql files sharing a numeric prefix, so the refreshed merge would fail DB startup. Renumber to the next free slot (154). Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>
The responses-continuation store adds one migration, so the docs' migration count is now 149 (was 148). Update README/AGENTS/llm.txt and regenerate the i18n llm.txt mirrors to keep check:docs-all green. Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>
Babysit summaryThe failing checks on this PR were STALE — the branch was 60 commits behind What was done:
Verified locally: Gate: all non-advisory checks pass on head Remaining: no unresolved review threads. Ready for human review & merge (not auto-merged). |
- Un-export ResponsesContinuationState: it's never imported outside responsesContinuationStore.ts, its own defining file. Fixes the check:dead-code regression (410 > baseline 409). - Scope the previous_response_id virtualization interception in chat.ts to skip entirely when responsesPreviousResponseIdMode=preserve. The interception ran unconditionally before target/connection selection, ahead of applyResponsesPreviousResponseIdPolicy (chatCore.ts) -- the existing per-target enforcement point for this setting -- so "preserve" (the explicit, connection-independent contract for "let the upstream resolve previous_response_id natively") was silently unreachable: the field was already deleted and replaced with locally-reconstructed input by the time that policy ran. This also broke Codex's own executor, which relies on an untouched previous_response_id to delegate history resolution upstream (see stripOrphanedCodexFunctionCallOutputs in codex.ts). "auto" and "strip" modes are unaffected -- virtualization is a strict improvement over their old "drop the field, hope the client resent everything" behavior. - Add a regression test exercising the actual chat.ts handler (not just the policy helper in isolation): confirms mode=preserve now proceeds to normal routing instead of the virtualization's previous_response_not_found rejection, and that default/auto mode's existing virtualization behavior is unchanged. Verified the test fails for the right reason against pre-fix chat.ts. Addresses PR review feedback.
|
Addressed both items:
Verified: Pushed: |
|
Thank you @hartmark for the Responses continuation fix. The exact head passed the complete PR checks and the combined FAST train; no PR-specific red was observed. The release merge is being serialized after the prior verified base update. |
…gosouzapw#10262 153/154 (originally 154/155) were chosen back when this branch stacked on top of the previous_response_id migration (153_call_logs_response_id.sql). Decoupling removed that migration from this branch's history, leaving an unused 153 slot that check-migration-numbering.test.ts correctly flags as a gap.
…age-architecture concern resolved (#10263) * feat(responses): virtualize previous_response_id continuation regardless of upstream support OmniRoute now exposes OpenAI-compatible previous_response_id/store continuation to clients unconditionally, even when the selected upstream provider has no native Responses-API state support. Reconstruction happens server-side in handleChatImplementation, before any downstream validation or provider translation: OmniRoute resolves the response id back to the full input/output it previously produced, prepends it to the client's delta, and forwards the full reconstructed history upstream exactly as it does today. Client<->OmniRoute traffic shrinks to the new delta only; OmniRoute<->provider traffic is unchanged. Storage reuses the existing call-log pipeline artifact (already gated by call_log_pipeline_enabled, already retained/cleaned up by the existing call-log lifecycle) instead of duplicating conversation content into a second store -- only a lightweight call_logs.response_id index is new. Every lookup is scoped by api_key_id so one client can never resolve another client's stored conversation, and any unresolvable/missing/ size-limit-omitted state fails closed with OpenAI's own previous_response_not_found contract. Stacked on feat/openai-responses-store-toggle (#10121). * feat(dashboard): agentic conversation tracking with live transcript view Every agentic chat request now gets a conversation id (X-ConversationId response header). OmniRoute detects when a follow-up request continues the same conversation via fingerprint + bounded prefix-hash matching, with a strict-growth invariant to prevent false merges between independent single-shot requests that happen to share identical opening content. Continuation detection excludes the system message from the identity anchor, since real coding-agent CLIs commonly regenerate it every request with live context (timestamp, cwd, git status) — without this, that volatility alone broke every continuation check against real traffic. - `/dashboard/logs`: new toggleable Conversation column. - `/dashboard/logs/timeline`: requests sharing a conversation id share a timeline lane, connected by an arrow, with a configurable lane-reuse window. - Request detail panel: new Full Conversation transcript above the raw SSE event stream — Markdown rendering, per-turn timestamps, turn-relative view, click-any-turn navigation, live auto-refresh building the transcript in real time from the in-flight SSE chunk buffer while a request is still streaming, auto-scroll-to-bottom as the live turn grows. - New `/dashboard/conversations` page listing conversations with 2+ turns, no-forking model (an edited/duplicated mid-history turn mints its own independent conversation instead of merging), pagination, duplicate- anchor fix. - Configurable auto-refresh intervals on both the timeline and conversations list pages. - Responses API tool-call gap fix: turnsFromOpenAiMessages only handled role-based Chat Completions messages, so bare {type:"function_call"} / {type:"function_call_output"} / {type:"reasoning"} items (real Responses API traffic) silently vanished from the Conversation Context panel. - truncateForLog now counts input[] (Responses API), not just messages[] (Chat Completions), so a truncated /v1/responses request still shows a placeholder instead of nothing. - RequestTimeline.tsx now reads the same debugEnabled/emailsVisible settings RequestLoggerV2.tsx already used, instead of hardcoding both false — the timeline view never showed SSE/stream-chunk events or respected email-masking, regardless of the actual setting. Migrations 147/148 (agentic_conversations, conversation_turn_nodes) — 135 and 136 are now taken upstream; 143-145 are documented KNOWN_GAPS, so this uses the next free slot past upstream's current highest. Test plan: - npm run typecheck:core — clean - npm run lint — clean - node --import tsx/esm scripts/check/check-migration-numbering.mjs — OK, 0 collisions - 109 unit tests across the conversation-tracking, migration-renumber, and dashboard-wiring surface — 0 failures * refactor(dashboard): reuse call-log artifacts for conversation transcript content conversation_turn_nodes no longer stores turn text/tool-call content (text_preview/block_kind/tool_name) -- it's identity-only now (id/parent/ content_hash), matching agentic_conversations' existing lightweight-index shape. Every node's originating request is already fully captured by the call-log pipeline artifact its last_correlation_id points at, so the /dashboard/conversations tree view resolves each node's actual display content on demand from there (open-sse/services/conversationTurnContent.ts), re-running the same extractCanonicalTurns/hashTurnContent the write path used and matching by content_hash, instead of duplicating conversation content into a second store under a separate retention/gating policy. This also drops the old 8000-char text_preview truncation entirely -- resolved content is always full and untruncated. The frontend contract is unchanged (tree API still returns {textPreview, blockKind, toolName} per node), so the dashboard UI itself (page.tsx, RequestLoggerDetail/RequestTimeline, sidebar, i18n) needed no changes. Renumbered the cherry-picked 147/148 migrations to 153/154 -- 147 now collides with 147_api_keys_model_access_mode.sql, which landed on release/v3.8.50 after this work was originally built. Also includes a standalone, unrelated fix carried along from this rebase: close isProviderModelHidden's missing function-body brace in modelSelectModalHelpers.ts (separately landed as #10206). Stacked on feat/responses-previous-response-id-virtualization (#3), which is itself stacked on feat/openai-responses-store-toggle (#10121). * fix(dashboard): resync conversation list on open so the live-text poll starts immediately openConversation() seeded activeConversation (and therefore activeCallLogId, which gates the live-partial-text poll effect) from whatever row snapshot the list's own fixed-interval poll last produced. A conversation opened right after a reply started streaming -- after that tick, before the next -- had activeCallLogId still null, so the live-text poll never started; only a subsequent background list-poll resync (already existed) picked it up, which is why closing and reopening the same conversation "just worked". loadConversations() is now a shared callback so openConversation can force one immediately on open instead of waiting on pollSeconds. Live-verified against omniroute-dev: opening a conversation mid-stream now shows live reasoning on the first open. * style: prettier formatting for conversationTurnContent.test.ts * fix(db): close migration numbering gap left by decoupling from #3/#10262 153/154 (originally 154/155) were chosen back when this branch stacked on top of the previous_response_id migration (153_call_logs_response_id.sql). Decoupling removed that migration from this branch's history, leaving an unused 153 slot that check-migration-numbering.test.ts correctly flags as a gap. * refactor(dashboard): split RequestTimeline/RequestLoggerDetail under the 1000-line file-size cap Both files exceeded check-file-size's new-file cap after this PR's own additions (RequestTimeline 1048, RequestLoggerDetail 1163). Extracted pure non-component logic (types, constants, allocateLanes and its helpers) out of RequestTimeline.tsx into RequestTimeline.utils.ts, and the two self-contained presentational sub-components (PayloadSection, ConversationContextSection + its private helper) out of RequestLoggerDetail.tsx into RequestLoggerDetail.sections.tsx. No behavior change; existing external imports (default exports, allocateLanes, TimelineLog, CONVERSATION_LANE_REUSE_STORAGE_KEY) still resolve from the original file paths. * fix(db): renumber agentic-conversation migrations to clear 153 collision + sync migration-count docs The refresh-merge of release/v3.8.50 exposed that the feature's three migrations collided at slot 153 with the base's radar_local_model_state (153) and its own call_logs_response_id. Migration runner enforces unique numeric prefixes -> every DB init threw, red-ing Vitest, all Unit shards and the DB-backed quality gates. Renumber the feature's pair to 155_agentic_conversations / 156_conversation_turn_nodes and move call_logs_response_id to 154 (keeps 153_radar base-owned, preserves agentic-before-turn_nodes ordering). Update SQL headers and the 154/156 references in feature code + tests. Migration count is now 151 (was 148 stale in README/AGENTS/llm.txt) — sync the doc counts to clear the docs-accuracy gate. Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com> * fix(ui): drop unused CONVERSATION_LANE_REUSE_STORAGE_KEY re-export from RequestTimeline Knip 6.32 (baseline 415) flags the public re-export of CONVERSATION_LANE_REUSE_STORAGE_KEY from RequestTimeline.tsx as dead: no external consumer imports it through that re-export (it is imported and used directly from RequestTimeline.utils.ts inside the component). Removed the unused re-export; the internal import stays. DEAD_TOTAL 416 -> 415, back to the frozen baseline. Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com> * fix(agentic-conversations): guard resolveConversationId, drop dead whole-chain export - Wrap resolveConversationId() in try/catch in chat.ts, matching the defensive pattern used by every other best-effort side call nearby, so a DB hiccup in conversation tracking can't turn a working chat request into a hard failure. - Remove getConversationTurnTree: knip's project scope excludes tests/**, so an export used only by tests can never register as used there. Swap its 8 test call sites to the paginated getConversationTurnPage (already the dashboard's canonical query) with a generous limit, collapsing to one query path instead of keeping a second whole-chain export alive solely for test convenience. - Regenerate i18n llm.txt mirrors from root (pre-existing drift on this branch, unrelated to the above, caught by the docs-sync pre-commit gate). Addresses PR review feedback. * fix(i18n): close requestLogger conversation-column gap, fix domain-modules count drift - fr.json, vi.json were missing requestLogger.columns.conversation (added in the conversation-tracking feature), failing i18n-vi-completeness.test.ts. - docs/i18n/*/llm.txt mirrors still said 117 domain-specific files after an earlier rebase fixed the migration count but missed this companion number, failing check-docs-sync.mjs across all 42 locales. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * fix(docs): restore PROXY_LOG_INCLUDE_IPS env/doc entries (env-doc-sync red) .env.example and docs/reference/ENVIRONMENT.md were both missing the PROXY_LOG_INCLUDE_IPS entry that src/lib/proxyLogger.ts already reads (confirmed present at this branch's merge-base too, so this predates the conversation-tracking work and is unrelated to it) -- the entry was added on release/v3.8.50 after this branch's last sync and this branch never picked it up. That gap red-lines tests/unit/check-env-doc-sync.test.ts and tests/unit/issue-7793-env-doc-sync-repro.test.ts (Unit Tests fast-path 2/4 in CI). Restore both entries verbatim from the current release/v3.8.50 tip -- no feature-code change. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: hartmark <hartmark@users.noreply.github.com> Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
…age-architecture concern resolved (diegosouzapw#10263) * feat(responses): virtualize previous_response_id continuation regardless of upstream support OmniRoute now exposes OpenAI-compatible previous_response_id/store continuation to clients unconditionally, even when the selected upstream provider has no native Responses-API state support. Reconstruction happens server-side in handleChatImplementation, before any downstream validation or provider translation: OmniRoute resolves the response id back to the full input/output it previously produced, prepends it to the client's delta, and forwards the full reconstructed history upstream exactly as it does today. Client<->OmniRoute traffic shrinks to the new delta only; OmniRoute<->provider traffic is unchanged. Storage reuses the existing call-log pipeline artifact (already gated by call_log_pipeline_enabled, already retained/cleaned up by the existing call-log lifecycle) instead of duplicating conversation content into a second store -- only a lightweight call_logs.response_id index is new. Every lookup is scoped by api_key_id so one client can never resolve another client's stored conversation, and any unresolvable/missing/ size-limit-omitted state fails closed with OpenAI's own previous_response_not_found contract. Stacked on feat/openai-responses-store-toggle (diegosouzapw#10121). * feat(dashboard): agentic conversation tracking with live transcript view Every agentic chat request now gets a conversation id (X-ConversationId response header). OmniRoute detects when a follow-up request continues the same conversation via fingerprint + bounded prefix-hash matching, with a strict-growth invariant to prevent false merges between independent single-shot requests that happen to share identical opening content. Continuation detection excludes the system message from the identity anchor, since real coding-agent CLIs commonly regenerate it every request with live context (timestamp, cwd, git status) — without this, that volatility alone broke every continuation check against real traffic. - `/dashboard/logs`: new toggleable Conversation column. - `/dashboard/logs/timeline`: requests sharing a conversation id share a timeline lane, connected by an arrow, with a configurable lane-reuse window. - Request detail panel: new Full Conversation transcript above the raw SSE event stream — Markdown rendering, per-turn timestamps, turn-relative view, click-any-turn navigation, live auto-refresh building the transcript in real time from the in-flight SSE chunk buffer while a request is still streaming, auto-scroll-to-bottom as the live turn grows. - New `/dashboard/conversations` page listing conversations with 2+ turns, no-forking model (an edited/duplicated mid-history turn mints its own independent conversation instead of merging), pagination, duplicate- anchor fix. - Configurable auto-refresh intervals on both the timeline and conversations list pages. - Responses API tool-call gap fix: turnsFromOpenAiMessages only handled role-based Chat Completions messages, so bare {type:"function_call"} / {type:"function_call_output"} / {type:"reasoning"} items (real Responses API traffic) silently vanished from the Conversation Context panel. - truncateForLog now counts input[] (Responses API), not just messages[] (Chat Completions), so a truncated /v1/responses request still shows a placeholder instead of nothing. - RequestTimeline.tsx now reads the same debugEnabled/emailsVisible settings RequestLoggerV2.tsx already used, instead of hardcoding both false — the timeline view never showed SSE/stream-chunk events or respected email-masking, regardless of the actual setting. Migrations 147/148 (agentic_conversations, conversation_turn_nodes) — 135 and 136 are now taken upstream; 143-145 are documented KNOWN_GAPS, so this uses the next free slot past upstream's current highest. Test plan: - npm run typecheck:core — clean - npm run lint — clean - node --import tsx/esm scripts/check/check-migration-numbering.mjs — OK, 0 collisions - 109 unit tests across the conversation-tracking, migration-renumber, and dashboard-wiring surface — 0 failures * refactor(dashboard): reuse call-log artifacts for conversation transcript content conversation_turn_nodes no longer stores turn text/tool-call content (text_preview/block_kind/tool_name) -- it's identity-only now (id/parent/ content_hash), matching agentic_conversations' existing lightweight-index shape. Every node's originating request is already fully captured by the call-log pipeline artifact its last_correlation_id points at, so the /dashboard/conversations tree view resolves each node's actual display content on demand from there (open-sse/services/conversationTurnContent.ts), re-running the same extractCanonicalTurns/hashTurnContent the write path used and matching by content_hash, instead of duplicating conversation content into a second store under a separate retention/gating policy. This also drops the old 8000-char text_preview truncation entirely -- resolved content is always full and untruncated. The frontend contract is unchanged (tree API still returns {textPreview, blockKind, toolName} per node), so the dashboard UI itself (page.tsx, RequestLoggerDetail/RequestTimeline, sidebar, i18n) needed no changes. Renumbered the cherry-picked 147/148 migrations to 153/154 -- 147 now collides with 147_api_keys_model_access_mode.sql, which landed on release/v3.8.50 after this work was originally built. Also includes a standalone, unrelated fix carried along from this rebase: close isProviderModelHidden's missing function-body brace in modelSelectModalHelpers.ts (separately landed as diegosouzapw#10206). Stacked on feat/responses-previous-response-id-virtualization (#3), which is itself stacked on feat/openai-responses-store-toggle (diegosouzapw#10121). * fix(dashboard): resync conversation list on open so the live-text poll starts immediately openConversation() seeded activeConversation (and therefore activeCallLogId, which gates the live-partial-text poll effect) from whatever row snapshot the list's own fixed-interval poll last produced. A conversation opened right after a reply started streaming -- after that tick, before the next -- had activeCallLogId still null, so the live-text poll never started; only a subsequent background list-poll resync (already existed) picked it up, which is why closing and reopening the same conversation "just worked". loadConversations() is now a shared callback so openConversation can force one immediately on open instead of waiting on pollSeconds. Live-verified against omniroute-dev: opening a conversation mid-stream now shows live reasoning on the first open. * style: prettier formatting for conversationTurnContent.test.ts * fix(db): close migration numbering gap left by decoupling from #3/diegosouzapw#10262 153/154 (originally 154/155) were chosen back when this branch stacked on top of the previous_response_id migration (153_call_logs_response_id.sql). Decoupling removed that migration from this branch's history, leaving an unused 153 slot that check-migration-numbering.test.ts correctly flags as a gap. * refactor(dashboard): split RequestTimeline/RequestLoggerDetail under the 1000-line file-size cap Both files exceeded check-file-size's new-file cap after this PR's own additions (RequestTimeline 1048, RequestLoggerDetail 1163). Extracted pure non-component logic (types, constants, allocateLanes and its helpers) out of RequestTimeline.tsx into RequestTimeline.utils.ts, and the two self-contained presentational sub-components (PayloadSection, ConversationContextSection + its private helper) out of RequestLoggerDetail.tsx into RequestLoggerDetail.sections.tsx. No behavior change; existing external imports (default exports, allocateLanes, TimelineLog, CONVERSATION_LANE_REUSE_STORAGE_KEY) still resolve from the original file paths. * fix(db): renumber agentic-conversation migrations to clear 153 collision + sync migration-count docs The refresh-merge of release/v3.8.50 exposed that the feature's three migrations collided at slot 153 with the base's radar_local_model_state (153) and its own call_logs_response_id. Migration runner enforces unique numeric prefixes -> every DB init threw, red-ing Vitest, all Unit shards and the DB-backed quality gates. Renumber the feature's pair to 155_agentic_conversations / 156_conversation_turn_nodes and move call_logs_response_id to 154 (keeps 153_radar base-owned, preserves agentic-before-turn_nodes ordering). Update SQL headers and the 154/156 references in feature code + tests. Migration count is now 151 (was 148 stale in README/AGENTS/llm.txt) — sync the doc counts to clear the docs-accuracy gate. Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com> * fix(ui): drop unused CONVERSATION_LANE_REUSE_STORAGE_KEY re-export from RequestTimeline Knip 6.32 (baseline 415) flags the public re-export of CONVERSATION_LANE_REUSE_STORAGE_KEY from RequestTimeline.tsx as dead: no external consumer imports it through that re-export (it is imported and used directly from RequestTimeline.utils.ts inside the component). Removed the unused re-export; the internal import stays. DEAD_TOTAL 416 -> 415, back to the frozen baseline. Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com> * fix(agentic-conversations): guard resolveConversationId, drop dead whole-chain export - Wrap resolveConversationId() in try/catch in chat.ts, matching the defensive pattern used by every other best-effort side call nearby, so a DB hiccup in conversation tracking can't turn a working chat request into a hard failure. - Remove getConversationTurnTree: knip's project scope excludes tests/**, so an export used only by tests can never register as used there. Swap its 8 test call sites to the paginated getConversationTurnPage (already the dashboard's canonical query) with a generous limit, collapsing to one query path instead of keeping a second whole-chain export alive solely for test convenience. - Regenerate i18n llm.txt mirrors from root (pre-existing drift on this branch, unrelated to the above, caught by the docs-sync pre-commit gate). Addresses PR review feedback. * fix(i18n): close requestLogger conversation-column gap, fix domain-modules count drift - fr.json, vi.json were missing requestLogger.columns.conversation (added in the conversation-tracking feature), failing i18n-vi-completeness.test.ts. - docs/i18n/*/llm.txt mirrors still said 117 domain-specific files after an earlier rebase fixed the migration count but missed this companion number, failing check-docs-sync.mjs across all 42 locales. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * fix(docs): restore PROXY_LOG_INCLUDE_IPS env/doc entries (env-doc-sync red) .env.example and docs/reference/ENVIRONMENT.md were both missing the PROXY_LOG_INCLUDE_IPS entry that src/lib/proxyLogger.ts already reads (confirmed present at this branch's merge-base too, so this predates the conversation-tracking work and is unrelated to it) -- the entry was added on release/v3.8.50 after this branch's last sync and this branch never picked it up. That gap red-lines tests/unit/check-env-doc-sync.test.ts and tests/unit/issue-7793-env-doc-sync-repro.test.ts (Unit Tests fast-path 2/4 in CI). Restore both entries verbatim from the current release/v3.8.50 tip -- no feature-code change. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: hartmark <hartmark@users.noreply.github.com> Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
…age-architecture concern resolved (diegosouzapw#10263) * feat(responses): virtualize previous_response_id continuation regardless of upstream support OmniRoute now exposes OpenAI-compatible previous_response_id/store continuation to clients unconditionally, even when the selected upstream provider has no native Responses-API state support. Reconstruction happens server-side in handleChatImplementation, before any downstream validation or provider translation: OmniRoute resolves the response id back to the full input/output it previously produced, prepends it to the client's delta, and forwards the full reconstructed history upstream exactly as it does today. Client<->OmniRoute traffic shrinks to the new delta only; OmniRoute<->provider traffic is unchanged. Storage reuses the existing call-log pipeline artifact (already gated by call_log_pipeline_enabled, already retained/cleaned up by the existing call-log lifecycle) instead of duplicating conversation content into a second store -- only a lightweight call_logs.response_id index is new. Every lookup is scoped by api_key_id so one client can never resolve another client's stored conversation, and any unresolvable/missing/ size-limit-omitted state fails closed with OpenAI's own previous_response_not_found contract. Stacked on feat/openai-responses-store-toggle (diegosouzapw#10121). * feat(dashboard): agentic conversation tracking with live transcript view Every agentic chat request now gets a conversation id (X-ConversationId response header). OmniRoute detects when a follow-up request continues the same conversation via fingerprint + bounded prefix-hash matching, with a strict-growth invariant to prevent false merges between independent single-shot requests that happen to share identical opening content. Continuation detection excludes the system message from the identity anchor, since real coding-agent CLIs commonly regenerate it every request with live context (timestamp, cwd, git status) — without this, that volatility alone broke every continuation check against real traffic. - `/dashboard/logs`: new toggleable Conversation column. - `/dashboard/logs/timeline`: requests sharing a conversation id share a timeline lane, connected by an arrow, with a configurable lane-reuse window. - Request detail panel: new Full Conversation transcript above the raw SSE event stream — Markdown rendering, per-turn timestamps, turn-relative view, click-any-turn navigation, live auto-refresh building the transcript in real time from the in-flight SSE chunk buffer while a request is still streaming, auto-scroll-to-bottom as the live turn grows. - New `/dashboard/conversations` page listing conversations with 2+ turns, no-forking model (an edited/duplicated mid-history turn mints its own independent conversation instead of merging), pagination, duplicate- anchor fix. - Configurable auto-refresh intervals on both the timeline and conversations list pages. - Responses API tool-call gap fix: turnsFromOpenAiMessages only handled role-based Chat Completions messages, so bare {type:"function_call"} / {type:"function_call_output"} / {type:"reasoning"} items (real Responses API traffic) silently vanished from the Conversation Context panel. - truncateForLog now counts input[] (Responses API), not just messages[] (Chat Completions), so a truncated /v1/responses request still shows a placeholder instead of nothing. - RequestTimeline.tsx now reads the same debugEnabled/emailsVisible settings RequestLoggerV2.tsx already used, instead of hardcoding both false — the timeline view never showed SSE/stream-chunk events or respected email-masking, regardless of the actual setting. Migrations 147/148 (agentic_conversations, conversation_turn_nodes) — 135 and 136 are now taken upstream; 143-145 are documented KNOWN_GAPS, so this uses the next free slot past upstream's current highest. Test plan: - npm run typecheck:core — clean - npm run lint — clean - node --import tsx/esm scripts/check/check-migration-numbering.mjs — OK, 0 collisions - 109 unit tests across the conversation-tracking, migration-renumber, and dashboard-wiring surface — 0 failures * refactor(dashboard): reuse call-log artifacts for conversation transcript content conversation_turn_nodes no longer stores turn text/tool-call content (text_preview/block_kind/tool_name) -- it's identity-only now (id/parent/ content_hash), matching agentic_conversations' existing lightweight-index shape. Every node's originating request is already fully captured by the call-log pipeline artifact its last_correlation_id points at, so the /dashboard/conversations tree view resolves each node's actual display content on demand from there (open-sse/services/conversationTurnContent.ts), re-running the same extractCanonicalTurns/hashTurnContent the write path used and matching by content_hash, instead of duplicating conversation content into a second store under a separate retention/gating policy. This also drops the old 8000-char text_preview truncation entirely -- resolved content is always full and untruncated. The frontend contract is unchanged (tree API still returns {textPreview, blockKind, toolName} per node), so the dashboard UI itself (page.tsx, RequestLoggerDetail/RequestTimeline, sidebar, i18n) needed no changes. Renumbered the cherry-picked 147/148 migrations to 153/154 -- 147 now collides with 147_api_keys_model_access_mode.sql, which landed on release/v3.8.50 after this work was originally built. Also includes a standalone, unrelated fix carried along from this rebase: close isProviderModelHidden's missing function-body brace in modelSelectModalHelpers.ts (separately landed as diegosouzapw#10206). Stacked on feat/responses-previous-response-id-virtualization (#3), which is itself stacked on feat/openai-responses-store-toggle (diegosouzapw#10121). * fix(dashboard): resync conversation list on open so the live-text poll starts immediately openConversation() seeded activeConversation (and therefore activeCallLogId, which gates the live-partial-text poll effect) from whatever row snapshot the list's own fixed-interval poll last produced. A conversation opened right after a reply started streaming -- after that tick, before the next -- had activeCallLogId still null, so the live-text poll never started; only a subsequent background list-poll resync (already existed) picked it up, which is why closing and reopening the same conversation "just worked". loadConversations() is now a shared callback so openConversation can force one immediately on open instead of waiting on pollSeconds. Live-verified against omniroute-dev: opening a conversation mid-stream now shows live reasoning on the first open. * style: prettier formatting for conversationTurnContent.test.ts * fix(db): close migration numbering gap left by decoupling from #3/diegosouzapw#10262 153/154 (originally 154/155) were chosen back when this branch stacked on top of the previous_response_id migration (153_call_logs_response_id.sql). Decoupling removed that migration from this branch's history, leaving an unused 153 slot that check-migration-numbering.test.ts correctly flags as a gap. * refactor(dashboard): split RequestTimeline/RequestLoggerDetail under the 1000-line file-size cap Both files exceeded check-file-size's new-file cap after this PR's own additions (RequestTimeline 1048, RequestLoggerDetail 1163). Extracted pure non-component logic (types, constants, allocateLanes and its helpers) out of RequestTimeline.tsx into RequestTimeline.utils.ts, and the two self-contained presentational sub-components (PayloadSection, ConversationContextSection + its private helper) out of RequestLoggerDetail.tsx into RequestLoggerDetail.sections.tsx. No behavior change; existing external imports (default exports, allocateLanes, TimelineLog, CONVERSATION_LANE_REUSE_STORAGE_KEY) still resolve from the original file paths. * fix(db): renumber agentic-conversation migrations to clear 153 collision + sync migration-count docs The refresh-merge of release/v3.8.50 exposed that the feature's three migrations collided at slot 153 with the base's radar_local_model_state (153) and its own call_logs_response_id. Migration runner enforces unique numeric prefixes -> every DB init threw, red-ing Vitest, all Unit shards and the DB-backed quality gates. Renumber the feature's pair to 155_agentic_conversations / 156_conversation_turn_nodes and move call_logs_response_id to 154 (keeps 153_radar base-owned, preserves agentic-before-turn_nodes ordering). Update SQL headers and the 154/156 references in feature code + tests. Migration count is now 151 (was 148 stale in README/AGENTS/llm.txt) — sync the doc counts to clear the docs-accuracy gate. Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com> * fix(ui): drop unused CONVERSATION_LANE_REUSE_STORAGE_KEY re-export from RequestTimeline Knip 6.32 (baseline 415) flags the public re-export of CONVERSATION_LANE_REUSE_STORAGE_KEY from RequestTimeline.tsx as dead: no external consumer imports it through that re-export (it is imported and used directly from RequestTimeline.utils.ts inside the component). Removed the unused re-export; the internal import stays. DEAD_TOTAL 416 -> 415, back to the frozen baseline. Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com> * fix(agentic-conversations): guard resolveConversationId, drop dead whole-chain export - Wrap resolveConversationId() in try/catch in chat.ts, matching the defensive pattern used by every other best-effort side call nearby, so a DB hiccup in conversation tracking can't turn a working chat request into a hard failure. - Remove getConversationTurnTree: knip's project scope excludes tests/**, so an export used only by tests can never register as used there. Swap its 8 test call sites to the paginated getConversationTurnPage (already the dashboard's canonical query) with a generous limit, collapsing to one query path instead of keeping a second whole-chain export alive solely for test convenience. - Regenerate i18n llm.txt mirrors from root (pre-existing drift on this branch, unrelated to the above, caught by the docs-sync pre-commit gate). Addresses PR review feedback. * fix(i18n): close requestLogger conversation-column gap, fix domain-modules count drift - fr.json, vi.json were missing requestLogger.columns.conversation (added in the conversation-tracking feature), failing i18n-vi-completeness.test.ts. - docs/i18n/*/llm.txt mirrors still said 117 domain-specific files after an earlier rebase fixed the migration count but missed this companion number, failing check-docs-sync.mjs across all 42 locales. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * fix(docs): restore PROXY_LOG_INCLUDE_IPS env/doc entries (env-doc-sync red) .env.example and docs/reference/ENVIRONMENT.md were both missing the PROXY_LOG_INCLUDE_IPS entry that src/lib/proxyLogger.ts already reads (confirmed present at this branch's merge-base too, so this predates the conversation-tracking work and is unrelated to it) -- the entry was added on release/v3.8.50 after this branch's last sync and this branch never picked it up. That gap red-lines tests/unit/check-env-doc-sync.test.ts and tests/unit/issue-7793-env-doc-sync-repro.test.ts (Unit Tests fast-path 2/4 in CI). Restore both entries verbatim from the current release/v3.8.50 tip -- no feature-code change. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: hartmark <hartmark@users.noreply.github.com> Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
…e2e assertion and 10 integration reds Electron Package Smoke — a packaging defect that had been hidden behind another packaging defect for nine days. Once the loginHeaderCapture fix let the main process start, the server underneath died on 'Cannot find module next': resources/app/server.js shipped without resources/app/node_modules. electron-builder discards the ROOT node_modules in code, not by configuration — app-builder-lib/out/util/filter.js:42 has a hard-coded `if (relative === "node_modules") return false` that runs before any filter pattern. The second extraResources entry pointing INTO ../.build/electron-standalone/node_modules is what sidesteps it, because those relative paths are never equal to "node_modules". #10325 removed that entry as an apparent duplicate and flipped the test to assert "exactly once", freezing the regression as if it were the contract. Restored, and the unit guard now pins both entries — proven by mutation: reverting package.json to the post-#10325 shape fails the guard 3/4, restoring it passes 4/4. group-b-quota-plans-config — the assertion was impossible to satisfy on ANY route, and the page was never broken. layout.tsx hands the whole message catalogue to NextIntlClientProvider, React serialises that prop into the RSC payload, and en.json carries "Internal Server Error" twice, so page.content() always contains it: probing /dashboard, /dashboard/costs, /dashboard/settings and /login showed the string present with every page rendering fine, and a pageerror probe on the failing run captured zero client exceptions. This is the same trap that killed the sibling not.toContain("500") in fc77100 ("raw HTML is unreliable") — that one was removed, this one was kept. Now asserts on rendered text, which still catches a real error boundary. The pageerror capture stays: the CI failure carried no stack trace, which is why it was misread twice. Integration — 10 of the 14 shard-2 reds, all sibling-test gaps behind security fixes: monitoring health now takes a Request and requires management auth (GHSA-mvf8-qc78-5mxm); the OAuth import routes moved to requireManagementAuth (GHSA-mg76) — the test accepts both guard shapes and gained a stronger anchor that every exported handler awaits a guard on its own request, mutation-verified; skill tool names are derived from encodeSkillToolName() and the fake upstream now returns the encoded name so decodeSkillToolName() is exercised too; previous_response_id now fails closed (#10262); proxy_logs persist as an async batch (#11182) so the test flushes first; providerQuotaOverrides joined GET /api/resilience (#9871); the reasoning fixture used a model that stopped being thinking-incompatible, replaced and pinned with a premise assert so it cannot rot silently again. A vacuous assert.ok(true, "all 10 streams completed without hanging") was replaced with real anchors — content must arrive on every stream and the active Timeout count must not grow. Four are deliberately left red rather than aligned, each now tracked: #11551 (the /v1/models after() wiring is dead — the route passes a third argument to a two-parameter function and catalogCache never imports after, so the #8728 contract is unimplemented), #11552 (~27% of requests emit an extra discarded upstream call; the delivered distribution is exactly 0.70, so weighted routing is correct and the waste is the real finding), the fixed-account combo pin (aligning it would destroy the per-step attribution the test exists for), and the web_search fallback already tracked as #11524. Package Artifact — the provenance stamp I added last round used git rev-parse HEAD, which under pull_request is the ephemeral merge commit and therefore never an ancestor of the release branch. Now takes the PR head sha. Refs #10692
* fix(release): restore diegosouzapw#10534 quota recovery and validate the volcengine connect bodies Two base-reds on the v3.8.50 tip, found by the release pre-flight. 1. diegosouzapw#11355 regressed diegosouzapw#10534. It replaced the per-window recovery check with an unconditional `hasActiveCooldown()` stop, which is right for an upstream-derived cooldown but also blocks the case diegosouzapw#10534 exists for: a Claude-subscription 429 persists a SYNTHETIC 1h rateLimitedUntil because the upstream sends no parseable reset. When the later poll shows every governing window has really reset with quota left, holding that synthetic cooldown just deadlocks the connection for an hour. The orphaned `windowStillExhaustedAfterRealReset()` helper and the three unused claudeExtraUsage imports that ESLint flagged were the fingerprint of this regression, not dead code: they are the two halves of the original gate. Re-wired as `isQuotaExhaustedCooldownReleasable()`, deliberately narrow — only lastErrorType "quota_exhausted" is eligible, one still-exhausted or unknown-reset window keeps the lock, and an extra-usage POLICY block stays locked even though its quota windows do look recovered in the same fetch. diegosouzapw#11277/diegosouzapw#11355 semantics are untouched (both guards still pass). Regression guard: tests/unit/provider-limits-recovery.test.ts already pinned this contract and was red on the tip. 15/15 now. 2. The three volcengine-plan connect routes read `request.json()` and handed the raw fields to a headless-browser login service after ad-hoc typeof checks (`check:route-validation:t06`, Hard Rule #7). `String(body.code ?? "")` turned 123 into "123" and an absent code into "", both reaching the service as a plausible SMS code. Now parsed with Zod schemas, before the session lookup, so a malformed body answers 400 instead of a misleading 404. New: tests/unit/volcengine-plan-connect-validation.test.ts (8 cases, red before the fix). Gate: 687 route files scanned, PASS. Also drops a genuinely dead import (formatVideoTimestamp in videoBridge.ts — only used inside the helpers module that defines it). * fix(release): drain the v3.8.50 docs/golden/GLM base-reds and restore a masked assert Second base-red batch from the release pre-flight, measured on the .113 with a clean npm ci (the devbox tree resolves eslint-plugin-react-hooks 7.1.1 from a stray pnpm store instead of the lockfile 7.0.1 and reports 925 phantom errors). Provider count 350 -> 352, one root cause behind three reds. Two providers landed this cycle (volcengine-agent-plan, volcengine-coding-plan) without regenerating the artifacts that quote the count: - docs/reference/PROVIDER_REFERENCE.md regenerated (gen:provider-reference). - README / AGENTS / llm.txt (+42 mirrors) / package.json description / 4 SVG diagrams updated, including the section heading AND the anchor that links to it, so the link does not break. - tests/snapshots/provider/translate-path.json regenerated. The diff is purely additive: 46 insertions, 0 deletions, exactly the two new providers. GLM effort tiers. diegosouzapw#11415 added the explicit glm-5.3-max tier and left two sibling vitest specs pinning the old 16-model inventory and an empty tier list for it. Aligned to the shipped contract (inventory order matches glmProvider.ts; glm-5.3-max declares ["max"]). Test-masking. Four assert reductions surfaced once the deleted-file signal was resolved. Three are legitimate and are allowlisted with their reasoning: diegosouzapw#11355 inverted the startup-cooldown contract (preserve future quota cooldowns), diegosouzapw#11280 replaced two unrolled hops with a 3-hop loop that asserts MORE, and the Gemini 3.5 Flash retirement removed the models those capability asserts described. The fourth was real masking: diegosouzapw#10960 rewrote the oneproxy status test to install a stream mock, immediately overwrite it with a passthrough to the real fetch, and assert `calls.length >= 0` — always true. Restored to assert what the test name claims (the JSON-RPC tools/call carries omniroute_oneproxy_stats and its result reaches the caller), with a scope note that it pins the MCP client contract rather than the commander wiring. Also allowlists the Gemini 3.5 Flash test deletion as _deletedWithReplacement (the model was retired by 2764812; gemini-models-parser.test.ts pins the new "excluded from the parsed list" contract), and rebaselines bundleSize 8045 -> 8461 with per-entry measurements — every entrypoint stays far below its absolute budget. * test(a2a): call the agent-card route handlers with a NextRequest PR diegosouzapw#11418 (S2 topology sanitisation) removed the hardcoded localhost:20128 from both well-known agent-card routes and made them derive the base URL from `request.nextUrl.origin` via `getBaseUrl(request)` (src/lib/wellKnown.ts). That changed the handler contract: `GET` now requires the request Next.js always passes it. Three sibling test files were never aligned and still invoked the handler as a bare `GET()`, so every case blew up with `TypeError: Cannot read properties of undefined (reading nextUrl)` before reaching a single assertion — 8 base-reds from one moved contract, not from a skill-count drift. Align the callers to the shipped contract with a local `makeCardRequest()` helper mirroring tests/unit/security-s1-s2-s4.test.ts (a Request with a defined `nextUrl`). No assertion was removed, loosened or skipped; the assert counts are unchanged and the cases now actually execute. Refs diegosouzapw#11418 * test(router-eval): give spawned CLI children an explicit DATA_DIR The router-eval CLI test spawns the CLI with spawnSync and asserts stderr stays empty. NODE_TEST_CONTEXT is inherited by those children, so since diegosouzapw#10432 (guard diegosouzapw#10428) resolveWritableDataDir() detects a test context with no DATA_DIR and warns on stderr before falling back to a throwaway dir - 194 chars that broke three cases. Pass an isolated DATA_DIR in the child env (the resolution the guard message itself prescribes) instead of loosening the assertions. * test(quota): freeze the clock in the GLM absolute-ISO-reset regression test The diegosouzapw#11353 regression test pins an ABSOLUTE upstream reset instant (2026-08-29 21:01:21) in its production fixture body but measured the resulting cooldown against the real wall clock. The remaining window therefore shrank every day: from 2026-08-25 it dropped under the 5-day floor the two assertions use, and past 2026-08-29 it would parse to null and collapse onto the 24h WEEKLY_QUOTA_COOLDOWN_MS default - a guaranteed future red. The shipped parser (parseIsoDateTimeResetMs / parseDayGranularityResetMs / buildWeeklyQuotaFallback / checkFallbackError) is correct: it returned the real multi-day reset, just measured from today instead of the fixtures NOW. Freeze Date at NOW via node:test mock timers in the two time-dependent cases so they assert the parser rather than the calendar. No assertion weakened, no production code touched. * test(providers): use a non-reserved prefix in the CC-compatible node create cases 93da24c ("fix(providers): reject reserved provider prefixes on compatible-node create/update") made createProviderNodeSchema reject any prefix that is a built-in REGISTRY id or alias. "cc" is the alias of the built-in `claude` provider, so the two provider-nodes create cases in cc-compatible-provider.test.ts started getting a 400 schema rejection before the route ever reached its feature-flag gate (403) or the create path (201) — the guard PR updated its own tests but missed this sibling file, leaving a base-red on release/v3.8.50. The operator-chosen prefix is incidental to what these cases assert (the ENABLE_CC_COMPATIBLE_PROVIDER gate, the dedicated anthropic-compatible-cc- id prefix, baseUrl sanitization and the nulled modelsPath), so switch it to a non-reserved "cc-proxy". No assertion was removed or loosened. * fix(quality): drain the three v3.8.50 inventory/coverage base-reds All three guards were drifting behind legitimate cycle growth, not catching a defect. Nothing was weakened: no assertion removed, no floor lowered, no blanket-allow added. providers-constants-split: APIKEY_PROVIDERS 231 -> 233. The delta is exactly the two Volcano Ark plan providers (volcengine-agent-plan, volcengine-coding-plan) added to the regional family in d732cf6. The invariant the guard exists for still holds, measured on the tip: 233 merged keys, 233 unique, family sum 233 (gateways 92 + frontier-labs 25 + inference-hosts 29 + enterprise-cloud 17 + regional 43 + specialty-media 27) with an empty cross-family duplicate set and an empty symmetric difference between the merged object and the family union - so the six files are still a strict partition, no loss and no dup. openapi-coverage: the operation floor (34.6%) is untouched. The cycle grew the denominator 985 -> 1002 while covered only moved 343 -> 345 (34.4%). Fixed by DOCUMENTING five real public operations rather than moving the floor, taking it to 350/1002 = 34.9%: GET /api/health, GET /api/v1/voices, POST /api/v1/speech-to-text, POST /api/v1/text-to-speech/{voiceId} and GET /api/v1/explain/routing. Each entry was written from the route source (auth mode, path-param pattern, limit clamp, upstream relay behaviour and the 400 / 401 / 429 branches), not from memory. hard-session-lease-bypass-inventory: three new connection-query sites classified, none silenced. open-sse/services/combo.ts (readConnectionForCooldownGate) reads the row backing the pre-dispatch persisted-cooldown gate, so it sits on the routing path and joins the class-B list next to combo/providerWildcard.ts and autoComboCandidates.ts. src/lib/providers/volcenginePlanBinding.ts and src/lib/providers/volcPlanAutoSyncBackfill.ts are connection persistence, not dispatch - the first resolves update-vs-create during connect, the second is a one-shot boot backfill of a providerSpecificData flag with no upstream call - so both stay class C alongside oauth/connectionPersistence.ts. * fix(release): drain two v3.8.50 base-reds (volcengine vision metadata, antigravity BYOP contract) Two unrelated real reds on the release tip: * fix(providers): flag MiniMax M3 as multimodal on the Volcengine Ark plans. d732cf6 ("feat(volcengine): add Ark plan providers") added volcengine-agent-plan/minimax-m3 and volcengine-coding-plan/minimax-m3 without supportsVision, breaking the LEDGER-4 invariant that every minimax-m3 registry entry except PromptQL (text-only upstream) is flagged multimodal. Every other provider carrying the model (opencode-zen, opencode-go, bazaarlink, ollama-cloud, codebuddy-cn, trae) sets it. Registry metadata defect, not a stale test. * test(antigravity): align the empty-projectId onboarding test to the contract shipped by diegosouzapw#11284/diegosouzapw#11358 (6de542b). That change made an onboardUser 200 whose body carries NO cloudaicompanionProject mean Google BYOP — no project was created and none ever will be — so it short-circuits before the retry loadCodeAssist. The older test still mocked onboardUser with the bare { done: true } BYOP shape while asserting the retry path, so it pinned a contract that was deliberately moved. The mock now returns a real onboarding-success body; every assertion is kept, and the id in the onboard body deliberately differs from the expected one so the test still proves the projectId came from the retry discovery. Refs diegosouzapw#11284 * fix(cli): do not treat a tmpfs mount as proof a config path reaches the host hasBindMountAt() accepted ANY mount as evidence that a would-be CLI config write reaches the operator's host: it matched on the mount point alone and never looked at the filesystem type. An in-memory mount therefore cleared the ephemeral flag, so guardCliConfigWrite() let the write through and both POST /api/cli-tools/apply and the dashboard's guide-settings writer answered 200 instead of the safe 422 that diegosouzapw#10057 added. That is the exact case the guard exists to refuse, and the worst one: a container running with `--tmpfs /tmp` (or a home on tmpfs) loses the file even before the container is recreated, while the UI reports success. Parse the filesystem type from mountinfo (the field after the lone "-" separator) and skip mounts backed by RAM or kernel state. Real bind mounts (ext4/xfs/nfs/virtiofs/fuse.*) still count, including one nested under a tmpfs path, so the compose `host` profile is unaffected. A line carrying no separator proves nothing and is skipped too. Regression cover added to tests/unit/container-env-detect.test.ts; this also un-reds tests/unit/cli-tools-apply-container-422.test.ts and tests/unit/api/cli-tools/apply-container-guard.test.ts, which were failing on any box whose /tmp is a tmpfs. * chore(docs): commit the next-dev agent-rules block into AGENTS.md `next dev` writes and re-adds this block (see node_modules/next/dist/server/lib/generate-agent-files.js), so leaving it out of a diff only recreates the uncommitted change on the next dev run. Committing it keeps the working tree clean, which is what the block's own note prescribes. * docs(changelog): aggregate the 363 changelog.d fragments into [3.8.50] Release reconciliation (Phase 0a.1). `scripts/release/aggregate-changelog.mjs` folds each changelog.d/<section>/*.md fragment into its heading in the living [3.8.50] section and deletes the fragment, which is the whole point of the fragment convention: two PRs never touch the same file, so the CHANGELOG never conflicts mid-cycle and no bullet is eaten by a merge auto-resolve. Section bullets 731 -> 1041. The remaining uncovered commits (mostly merges from diegosouzapw#11397 onward, which landed without a fragment) are reconciled separately. * docs(changelog): reconcile the v3.8.50 section with the cycle's uncovered commits Adds 131 consolidated bullets (45 features, 71 fixes, 15 maintenance) covering the ~490 user-facing commits and the ~100 chore/ci/test/refactor/docs commits that landed in the cycle without a CHANGELOG entry, grouped by subsystem and citing their PR references. Uncovered report: 594 -> 175 (the remainder are commits carrying no #N in their subject, which the matcher can never resolve; they are covered in prose). * docs(changelog): date the 3.8.50 header, inject contributors and sync the 42 i18n mirrors * fix(ci): raise the unit-shard heap ceiling to 8192 MB to stop the SIGABRT OOM The 8 unit shards run under V8 coverage instrumentation, which retains far more memory than the bare suite. With the 4096 MB ceiling they began aborting with exit 134 ("Ineffective mark-compacts near heap limit") at ~4086 MB as the provider catalogue grew during the v3.8.50 cycle: every test in the shard passed and the process died at the end, which reads as a test failure without being one. Aligns test:unit:ci:shard and test:unit:serial with the 8192 MB the non-sharded variants already use. GitHub-hosted runners have 16 GB, so the headroom is real. Validated by the CI run on this commit — the shards are the gate. Refs diegosouzapw#10692 * fix(ci): set NODE_OPTIONS on the unit-shard step and prune stale eslint suppressions The previous heap bump only touched test:unit:ci:shard, i.e. the node the shard script spawns. The process that actually runs out of memory is the `c8` wrapper around it — it aggregates ~577 MB of raw V8 coverage JSON — so the ceiling stayed at the V8 default (~4 GB) and the shards kept aborting at ~4083 MB, byte for byte the same failure. Setting NODE_OPTIONS on the step covers c8 and every child, which is the pattern the coverage-merge job already uses. Also prunes three eslint suppression entries whose violations no longer exist: videoBridgeContactSheet.ts and videoBridgeRuntime.ts (no-unused-vars, fixed during this cycle) and cli-oneproxy-commands.test.ts (no-explicit-any 14 -> 13, a consequence of restoring the real mock in that test). Stale entries make `npm run lint` exit 2 with 'There are suppressions left that do not occur anymore'. Pruned and verified on an uncontaminated checkout, not the devbox. Refs diegosouzapw#10692 * fix(ci): drain four inherited base-reds in packaging, electron and integration tests All four predate this session — each reproduces identically on f95b03d (2026-08-24), so none is a cycle regression. Draining them here because the release pre-flight is where inherited reds get resolved. 1. Package Artifact: the job runs `build:cli`, which assembles dist/ but never writes dist/BUILD_SHA — only `build:release` does, via write-build-sha.mjs. The diegosouzapw#10427 provenance guard inside check:pack-artifact then rejects the artifact as untraceable, and rejects it even under OMNIROUTE_ALLOW_CANARY_BUILD. The job's build+validate pair was structurally incompatible and failed 100% of the time. Stamps the SHA between the two steps. 2. Electron Package Smoke: electron/package.json's build.files allowlist enumerates each lib/*.js by hand and never got lib/loginHeaderCapture.js, added alongside its require() in diegosouzapw#9984. The file therefore stayed out of app.asar and the packaged app died at startup on 'Cannot find module ./lib/loginHeaderCapture'. 3. proxy-pipeline: the breaker assertion grepped chat.ts for executeChatWithBreaker(, but that call moved behind the chatDispatch.ts seam. Rather than drop the check, it now pins both hops — chat.ts dispatches through the seam and the seam calls the breaker — so the extraction cannot silently take the breaker off the path. 4. skills-pipeline: diegosouzapw#9058 began encoding skill tool names as omr_skill_<base64url> because providers require ^[a-zA-Z0-9_-]+$, and these assertions still expected the raw name@version. They now derive the expected name from encodeSkillToolName(), the same helper production uses, so the test tracks the contract instead of duplicating it. Only the assertions about names on the wire were converted; the identifiers passed straight to skillExecutor.execute() stay raw, because those are not encoded. Integration suite for these two files: 54/55. The one still red — 'web_search fallback preserves Responses API output' — is a separate pre-existing defect, deliberately left failing rather than papered over: on the /v1/responses path resolveSearchCredentials() returns null for the seeded serper-search connection, so executeWebSearch.ts:185-200 falls through to the cheapest fallbackOnly provider (duckduckgo-free) and the results come back empty. The sibling chat-path test seeds identically and does resolve serper-search. Needs its own investigation. Refs diegosouzapw#10692 * test(vitest): realign five sibling-test contracts unmasked by the green mcp shard None of these are cycle regressions. The Vitest job runs test:vitest (mcp shard) then test:vitest:ui; the mcp shard was failing on a missing glm-5.3-max and aborted the job before the ui shard ever ran. Fixing that shard this cycle unmasked 34 ui failures that had been broken since 18-23 Aug — four separate PRs that moved a contract and updated their own tests but not their siblings. - ProviderCard gained useRouter() in diegosouzapw#10448; four test files render it without mocking next/navigation and died on 'invariant expected app router to be mounted'. The sibling created alongside diegosouzapw#10448 already had the mock — it just was not applied to the other four consumers. - SkillCoverage gained a required config category. The four fixtures in agent-skills-page still described only api/cli, so the component read config.have off undefined. Values were chosen per scenario rather than pasted: full coverage gets 2/2 so its bar stays emerald, the amber fixture gets 3/4 so it stays amber. CoverageBar renders api -> config -> cli, so the new bar lands in the MIDDLE and the cli assertions moved from index [1] to [2]; without that the cli checks would have passed while measuring the config bar. The aria test now pins all three bars. - CliAgentsPage hardcoded AGENT_IDS, which had already drifted once (6 -> 8 with omp/letta, per its own comment) and drifted again with prime-agent (diegosouzapw#11166). It is now derived from CLI_TOOLS. This is why an agent missing from that list is not cosmetic: it never enters the status map, defaults to not_installed, and adds a phantom card to the filter and count tests. Deriving keeps the fixture in sync by construction instead of waiting for the next agent. - claudeTlsClient asserted proxyUrl was undefined inside a test literally named 'falls back to env var when per-call proxyUrl not provided' — it pinned the old behaviour where testOverride bypassed proxy resolution. diegosouzapw#10910 moved resolution ahead of the override on purpose ('so test overrides and the real path both see it'), so the assertion now checks the fallback the test name promises. test:vitest:ui goes from 34 failures to 14. The remaining 14 sit in six files none of this commit touches (AutoComboCatalog, CoolingConnectionsPanel, ProxyRegistryManager x2, connectionsSearchFilter) plus one claudeTlsClient case that passes in isolation and only fails in the full run — i.e. cross-file pollution. They need a clean environment to judge: this devbox resolves part of its tree through a stray pnpm store and has already produced one phantom failure count this cycle. Refs diegosouzapw#10692 * fix(i18n): unstale trafficInspectorSubtitle across 32 locales and regenerate the omni-inference skill Two gates in the Lint job, both inherited — each was hidden behind the one before it. i18n value drift: diegosouzapw#11283 rewrote sidebar.trafficInspectorSubtitle in en.json without touching the 42 translations, so 32 locales kept serving a sentence the English no longer says. Most take the documented __MISSING__: placeholder, which makes the runtime serve the corrected English until the translation pipeline catches up. Three do not: - vi cannot take a placeholder at all — tests/unit/i18n-vi-completeness.test.ts bans any __MISSING__/__TODO__ value outright, so it needs a real translation. - pt and pt-BR are translated for real rather than placeheld, because a placeholder there means this project's own maintainer reads the sidebar in English. Each file changed by exactly one line; the JSON was not reserialised wholesale. agent-skills-sync: skills/omni-inference/SKILL.md was missing the ElevenLabs voices and speech-to-text routes added by diegosouzapw#11312, so the generator reported one file out of date and the gate exited 2. Regenerated — purely additive, 48 lines, no deletions. Verified: check-ui-value-drift PASS, i18n:check-ui-coverage PASS (42 locales), i18n-vi-completeness 5/5, check:agent-skills-sync 46 unchanged. Refs diegosouzapw#10692 * test(vitest): pay heavy module imports at collection time, not out of the per-test budget The remaining ui-shard reds were one class, not six bugs: every one of them did `await import(<heavy component>)` INSIDE an `it()`, so Vite's transform of the dependency graph was billed to that test's timeout. Measured costs against the budgets they had to fit in: ProxyRegistryManager 86s import vs 30s / 60s / 5s budgets (render itself: 567ms) claudeTlsClient ~12s import vs 5s default useProviderConnections 1050-line hook, whole dashboard graph, vs 5s default That is why they looked like cross-file pollution: on an idle box the import squeaked under the limit, and under the ui suite's 20 parallel workers it did not. Running claudeTlsClient ALONE on a loaded box reproduces it — the trigger is CPU contention, not a neighbouring file. The sibling chatgptTlsClient/grokTlsClient tests import the same graph and never fail, because they import statically at module scope, where the cost falls on the collection phase which has no per-test budget. Every fix here does the same: static import or a beforeAll with its own budget. AutoComboCatalog also explains its own blast radius: the timeout aborted inside an open act(), leaking an unbalanced act scope that then failed the file's three remaining tests in ~20ms with 'overlapping act() calls'. One slow import, four reds. CoolingConnectionsPanel is the one production change. It imported providerText from the ../providerPageHelpers barrel, but that symbol is DEFINED in the ../providerCredentialText leaf and only re-exported by the barrel — which drags providerRegistry (352 providers) and the rest of the provider-page graph into a "use client" component for one string helper. Verified before accepting: the component used nothing else from the barrel, the barrel has no top-level side-effect to lose (the empty-registry hazard this repo has hit before does not apply), typecheck:core is clean, and the panel's first test drops from ~4s to 95ms. The import was suboptimal, never broken — the screen was not failing for users. No assertion was weakened anywhere. expect() counts are unchanged (25/25, 4/4) or up by one (AutoComboCatalog 11 -> 12); the diegosouzapw#8855 autofill sentinels, the data-1p-ignore / data-lpignore guards and the dead-status round-trip are intact. The diegosouzapw#5918 TDZ guard was proven still live by mutation, not by absence of red: moving useProxyBatchOperations(load) above its const reproduced 'ReferenceError: Cannot access load before initialization' in 207ms, then the production file was restored (diff empty). tests/unit/ui under load: 17 failed files / 45 failed tests -> 4 failed files / 4 failed tests, none of them these. The four left are compression-guidance-7530, compressionPanel, compressionUltraTier and lobe-provider-icons-stepfun, untouched and uninvestigated. Refs diegosouzapw#10692 * test(integration): realign three suites to security and version contracts that moved All three are the sibling-test gap again: a PR moved a contract, updated its own tests, and left these behind. None is a production defect — in two of the three the production side is a deliberate security fix. v1-contracts-behavior (4 failures, one cause): the job env sets INITIAL_PASSWORD, which makes isAuthRequired() true, and diegosouzapw#9320 (b07182c) made the /v1 catalogue gate on-by-default instead of opt-in via settings.requireAuthForModels. The four contract reads were calling the catalogue routes with no credential and correctly getting 401. Bisected the job's four env vars to confirm INITIAL_PASSWORD alone reproduces it (5 pass / 4 fail with it, 9 / 0 without). The tests now send a Bearer token; the shape assertions are untouched, and the auth contract itself stays owned by tests/unit/v1-models-auth-leak-9320.test.ts rather than being duplicated here. opencode-config-startup: two independent drifts. OPENCODE_VERSION was pinned to 1.18.8 while the installed opencode-ai is 1.18.18 (Dependabot 7f69589, diegosouzapw#10626) — now read from require("opencode-ai/package.json").version, which is exactly as strict but cannot drift on the next bump. And the no-limit-metadata case asserted limit === undefined, but diegosouzapw#11054 made the generator always emit a limit; it now pins the actual fallback {context: 128_000, output: 8_192} instead of an absence. memory-pipeline: diegosouzapw#11040 (GHSA-cpv3-xr7r-xf8q) made the resolved caller principal always win over a caller-supplied apiKeyId, so a spoofed id can no longer write into another principal's store. That PR updated the unit sibling but not this one. The test now asserts the stronger property — and deliberately not just the absence: the spoofed principal's store is empty AND the caller can still read the entry, which proves the write was redirected rather than dropped and keeps the emptiness check from passing vacuously with a disabled store. (The old assertion was count === 0, which a switched-off memory store would satisfy.) Assertion counts: 43 -> 43, 13 -> 14, 76 -> 81. Nothing weakened or removed. Verified: 24/24 pass, with and without the CI env vars. Refs diegosouzapw#10692 * fix(ci): drain the electron packaging regression, the models-catalog e2e assertion and 10 integration reds Electron Package Smoke — a packaging defect that had been hidden behind another packaging defect for nine days. Once the loginHeaderCapture fix let the main process start, the server underneath died on 'Cannot find module next': resources/app/server.js shipped without resources/app/node_modules. electron-builder discards the ROOT node_modules in code, not by configuration — app-builder-lib/out/util/filter.js:42 has a hard-coded `if (relative === "node_modules") return false` that runs before any filter pattern. The second extraResources entry pointing INTO ../.build/electron-standalone/node_modules is what sidesteps it, because those relative paths are never equal to "node_modules". diegosouzapw#10325 removed that entry as an apparent duplicate and flipped the test to assert "exactly once", freezing the regression as if it were the contract. Restored, and the unit guard now pins both entries — proven by mutation: reverting package.json to the post-diegosouzapw#10325 shape fails the guard 3/4, restoring it passes 4/4. group-b-quota-plans-config — the assertion was impossible to satisfy on ANY route, and the page was never broken. layout.tsx hands the whole message catalogue to NextIntlClientProvider, React serialises that prop into the RSC payload, and en.json carries "Internal Server Error" twice, so page.content() always contains it: probing /dashboard, /dashboard/costs, /dashboard/settings and /login showed the string present with every page rendering fine, and a pageerror probe on the failing run captured zero client exceptions. This is the same trap that killed the sibling not.toContain("500") in fc77100 ("raw HTML is unreliable") — that one was removed, this one was kept. Now asserts on rendered text, which still catches a real error boundary. The pageerror capture stays: the CI failure carried no stack trace, which is why it was misread twice. Integration — 10 of the 14 shard-2 reds, all sibling-test gaps behind security fixes: monitoring health now takes a Request and requires management auth (GHSA-mvf8-qc78-5mxm); the OAuth import routes moved to requireManagementAuth (GHSA-mg76) — the test accepts both guard shapes and gained a stronger anchor that every exported handler awaits a guard on its own request, mutation-verified; skill tool names are derived from encodeSkillToolName() and the fake upstream now returns the encoded name so decodeSkillToolName() is exercised too; previous_response_id now fails closed (diegosouzapw#10262); proxy_logs persist as an async batch (diegosouzapw#11182) so the test flushes first; providerQuotaOverrides joined GET /api/resilience (diegosouzapw#9871); the reasoning fixture used a model that stopped being thinking-incompatible, replaced and pinned with a premise assert so it cannot rot silently again. A vacuous assert.ok(true, "all 10 streams completed without hanging") was replaced with real anchors — content must arrive on every stream and the active Timeout count must not grow. Four are deliberately left red rather than aligned, each now tracked: diegosouzapw#11551 (the /v1/models after() wiring is dead — the route passes a third argument to a two-parameter function and catalogCache never imports after, so the diegosouzapw#8728 contract is unimplemented), diegosouzapw#11552 (~27% of requests emit an extra discarded upstream call; the delivered distribution is exactly 0.70, so weighted routing is correct and the waste is the real finding), the fixed-account combo pin (aligning it would destroy the per-step attribution the test exists for), and the web_search fallback already tracked as diegosouzapw#11524. Package Artifact — the provenance stamp I added last round used git rev-parse HEAD, which under pull_request is the ephemeral merge commit and therefore never an ancestor of the release branch. Now takes the PR head sha. Refs diegosouzapw#10692 * test(ui): unmount before asserting so the auto-sync timer cannot outlive the test Vitest went red on 'synchronizes upstream models only when autoFetchModels is explicitly true' with new URL throwing inside a fetch dispatched from Timeout._onTimeout (useProviderModels.ts:69). It is intermittent: green in the two previous CI runs, green every time in isolation, red only under the ui suite's 20 parallel workers. The hook schedules its auto-sync in a setTimeout whose callback only checks the flag on entry — and that flag stays false while the component is mounted. Both tests asserted first and unmounted last, so under contention the timer escaped the test window, fired after afterEach had already run vi.unstubAllGlobals(), and reached the REAL fetch with a relative URL. Unmounting before the assertions closes the window: cleanup flips , the callback returns early, and the calls already recorded on fetchMock are still there to assert against. No assertion changed. Not a regression from this cycle — the file's last change is fd76271 (diegosouzapw#10603). Fixed rather than tracked because an intermittent red in a blocking job is worse than a permanent one: it teaches people to re-run instead of to look. Refs diegosouzapw#10692 * fix(search): prefer the configured search connection over duckduckgo-free When no explicit provider is requested and the auto-selected cheapest provider has no credentials, executeWebSearch ran the fallbackOnly loop first. duckduckgo-free (costPerQuery 0, authType none) always won there with an empty credentials object, so a configured paid connection such as serper-search was silently ignored and the caller got success:true with zero results. Move the sweep for other credentialed regular providers ahead of the fallbackOnly loop (and exclude fallbackOnly ids from it, so a free last-resort provider never outranks a configured one on cost). The fallbackOnly loop stays as the true last resort. The chat path only appeared correct because duckduckgo-free happened to fail there and handleSearch retried the alternate provider; on /v1/responses it "succeeded" with no results. Closes diegosouzapw#11524 * fix(api): restore the after() injection point for the /v1/models SWR refresh `/v1/models` has passed a third argument to `getUnifiedModelsResponse()` (`{ scheduleBackgroundRefresh: (task) => after(task) }`) ever since diegosouzapw#10198, but diegosouzapw#9199 had already removed the parameter: the function takes two, so the object was silently dropped and `catalogCache` kept scheduling the stale-while- revalidate rebuild with `setTimeout(..., 0)`. The builder is overwhelmingly synchronous under the single-threaded App Router, so it pinned the event loop before the stale response was flushed — the diegosouzapw#8728 guarantee did not exist. - `catalogCache` now imports `after` from `next/server` and exposes `defaultBackgroundRefreshScheduler`, which defers to `after()` and falls back to a macrotask outside a Next request scope (instrumentation warm-up, tests). - `resolveCachedCatalogResponse` takes `scheduleBackgroundRefresh` and `getStaleWhileRevalidateMs` on its existing options object. - `getUnifiedModelsResponse` accepts the options object the route already passes and propagates the scheduler. The excess argument was invisible to CI: `tsconfig.typecheck-core.json` is a curated 27-file allowlist, `check:dashboard-typecheck` only covers `src/app/(dashboard)`, and `next.config.mjs` sets `ignoreBuildErrors: true`. Closes diegosouzapw#11551 * fix(sse): stop re-summarizing a universal handoff that never parses A universal handoff whose summary comes back unparseable persists nothing, so shouldGenerateUniversalHandoff keeps answering "generate" and the very next model switch in the same session re-issues the same full-history summarization call and discards the answer again — forever. With a switch-heavy combo strategy (weighted, random, round-robin, p2c) the models alternate on almost every turn, so that background call lands on a large fraction of requests: an upstream call whose response nobody reads is real money on a paid provider and real quota on a metered one. Measured on the weighted 70/30 matrix, that inflated the observed openai share to 0.895 where routing actually delivered 0.70, and at the unit level 199 of 200 model switches issued a fresh discarded summarization call. Back off per (session, combo) after an answer that is not a usable handoff: exponential 5min -> 1h, cleared on the first successful generation, capped at 500 tracked keys. Deliberately narrow — a transient upstream failure (!response.ok) is NOT tracked, so it still retries on the next switch, which is the behavior the context-relay path already depends on. After the fix, same harness at n=200: 201 upstream calls for 200 requests (1 extra, 0.5%), and the measured openai share equals the delivered share (0.725). The unit guard drops 199 discarded calls to 1. Refs diegosouzapw#11552 * fix(sse): honor per-step connection pins instead of rotating accounts on fallback A combo step pinned with an explicit `connectionId` (or a request pinned via `x-omniroute-connection`) is an operator instruction, not a hint. The generic account-fallback branch in handleSingleModelChat excluded the pinned connection after an upstream failure and re-selected a sibling account of the same provider, so a priority combo repeating one provider/model with two different fixed accounts ran both attempts under the FIRST step: the second step, with its own pin, never executed and per-step attribution (comboStepId / comboExecutionKey) was wrong. Gate the rotation on `!hasForcedConnection`, matching the antigravity stream-readiness, pre-response-timeout and account-semaphore branches that already let pinned steps fall through to combo orchestration. Cooldown recording via markAccountUnavailable is unchanged, and unpinned selection still excludes burned connections. Refs tests/integration/combo-routing-e2e.test.ts * fix(ci): point the pack-artifact provenance gate at the branch under test The guard added for diegosouzapw#10427 checks ancestry against origin/main by default. That is the right ref at publish time (npm-publish.yml runs on main), but in a pull_request context it can never hold: while the PR is open its head is by construction not an ancestor of main, and the shallow checkout does not even bring origin/main into the local graph, so the probe answers false regardless. The job therefore failed 100% of the time and only became visible now that it stopped being cancelled behind Build. Pre-merge the one checkable invariant is that the stamp matches the branch under test, so resolve the ref from refs/pull/<N>/head — which exists on origin even for fork PRs, unlike head.ref, which only exists on the author's repository. * test(dashboard): open the proxy toolbar overflow menu before bulk assign PR diegosouzapw#9870 moved the proxy-registry bulk-assign action into the toolbar's "More actions" (…) overflow menu, which renders its items only while open. The e2e smoke flow still clicked the testid directly, so the locator never resolved and the test burned its full 180s budget. Open the menu first and assert it is visible before clicking; every existing assertion is unchanged. * fix(dashboard): restore expert-mode manual model entry in the combo builder PR diegosouzapw#8285 (global model search) replaced the expert-mode "Manual model" block positionally with the new GlobalModelSearchPanel, dropping the only way to type a provider/model pair by hand in expert mode. The supporting state and handlers (manualModelInput, manualModelError, manualModelHasDuplicate, handleAddManualModel) survived as dead code, so neither typecheck nor lint flagged the loss. Re-render the block above the search panel, unchanged from its pre-diegosouzapw#8285 form. Regression guard: tests/e2e/combos-flow.spec.ts "expert mode shows a single-page combo form with manual model entry", which had been failing with a 180s timeout on locator.fill for combo-manual-model-input. * test(api): accept the post-diegosouzapw#9320 auth gate on the /v1/models e2e check b07182c (diegosouzapw#9320) inverted the catalog auth rule: /v1/models now requires auth whenever management auth is configured, unless requireAuthForModels is explicitly false. The e2e harness boots with INITIAL_PASSWORD set, so the endpoint has been answering 401 since 2026-08-04 and this check has been red ever since — invisible only because the job kept being cancelled behind Build. Mirror the sibling /api/providers check in this same file: assert the catalog shape whenever the catalog is actually served, and otherwise pin the auth gate by status AND error type, so a 401 from an unrelated misroute cannot pass for the deliberate one. --------- Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> Co-authored-by: Xiangzhe <diegosouza.pw@gmail.com> Co-authored-by: Xiangzhe <bakryun0718@proton.me> Co-authored-by: TheDemonTuan <nguyenviettuanbp@gmail.com>
…ess of upstream support (diegosouzapw#10262) * feat(responses): virtualize previous_response_id continuation regardless of upstream support OmniRoute now exposes OpenAI-compatible previous_response_id/store continuation to clients unconditionally, even when the selected upstream provider has no native Responses-API state support. Reconstruction happens server-side in handleChatImplementation, before any downstream validation or provider translation: OmniRoute resolves the response id back to the full input/output it previously produced, prepends it to the client's delta, and forwards the full reconstructed history upstream exactly as it does today. Client<->OmniRoute traffic shrinks to the new delta only; OmniRoute<->provider traffic is unchanged. Storage reuses the existing call-log pipeline artifact (already gated by call_log_pipeline_enabled, already retained/cleaned up by the existing call-log lifecycle) instead of duplicating conversation content into a second store -- only a lightweight call_logs.response_id index is new. Every lookup is scoped by api_key_id so one client can never resolve another client's stored conversation, and any unresolvable/missing/ size-limit-omitted state fails closed with OpenAI's own previous_response_not_found contract. Stacked on feat/openai-responses-store-toggle (diegosouzapw#10121). * fix(db): re-export responsesContinuationStore from the localDb barrel check-db-rules requires every db/ module to be re-exported (or explicitly allowlisted as intentionally-internal) for discoverability. Missed this when the module was first added. * fix(db): renumber previous_response_id index migration to 154 The migration was numbered 153, but release/v3.8.50 already carries 153_radar_local_model_state.sql. The emngrating runner's collision guard throws on two live .sql files sharing a numeric prefix, so the refreshed merge would fail DB startup. Renumber to the next free slot (154). Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com> * docs(db): sync migration count to 149 across llm.txt mirrors The responses-continuation store adds one migration, so the docs' migration count is now 149 (was 148). Update README/AGENTS/llm.txt and regenerate the i18n llm.txt mirrors to keep check:docs-all green. Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com> * fix(responses-continuation): respect preserve mode, drop dead export - Un-export ResponsesContinuationState: it's never imported outside responsesContinuationStore.ts, its own defining file. Fixes the check:dead-code regression (410 > baseline 409). - Scope the previous_response_id virtualization interception in chat.ts to skip entirely when responsesPreviousResponseIdMode=preserve. The interception ran unconditionally before target/connection selection, ahead of applyResponsesPreviousResponseIdPolicy (chatCore.ts) -- the existing per-target enforcement point for this setting -- so "preserve" (the explicit, connection-independent contract for "let the upstream resolve previous_response_id natively") was silently unreachable: the field was already deleted and replaced with locally-reconstructed input by the time that policy ran. This also broke Codex's own executor, which relies on an untouched previous_response_id to delegate history resolution upstream (see stripOrphanedCodexFunctionCallOutputs in codex.ts). "auto" and "strip" modes are unaffected -- virtualization is a strict improvement over their old "drop the field, hope the client resent everything" behavior. - Add a regression test exercising the actual chat.ts handler (not just the policy helper in isolation): confirms mode=preserve now proceeds to normal routing instead of the virtualization's previous_response_not_found rejection, and that default/auto mode's existing virtualization behavior is unchanged. Verified the test fails for the right reason against pre-fix chat.ts. Addresses PR review feedback. --------- Co-authored-by: adevwithpurpose <adevwithpurpose@users.noreply.github.com> Co-authored-by: hartmark <hartmark@users.noreply.github.com> Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>
…age-architecture concern resolved (diegosouzapw#10263) * feat(responses): virtualize previous_response_id continuation regardless of upstream support OmniRoute now exposes OpenAI-compatible previous_response_id/store continuation to clients unconditionally, even when the selected upstream provider has no native Responses-API state support. Reconstruction happens server-side in handleChatImplementation, before any downstream validation or provider translation: OmniRoute resolves the response id back to the full input/output it previously produced, prepends it to the client's delta, and forwards the full reconstructed history upstream exactly as it does today. Client<->OmniRoute traffic shrinks to the new delta only; OmniRoute<->provider traffic is unchanged. Storage reuses the existing call-log pipeline artifact (already gated by call_log_pipeline_enabled, already retained/cleaned up by the existing call-log lifecycle) instead of duplicating conversation content into a second store -- only a lightweight call_logs.response_id index is new. Every lookup is scoped by api_key_id so one client can never resolve another client's stored conversation, and any unresolvable/missing/ size-limit-omitted state fails closed with OpenAI's own previous_response_not_found contract. Stacked on feat/openai-responses-store-toggle (diegosouzapw#10121). * feat(dashboard): agentic conversation tracking with live transcript view Every agentic chat request now gets a conversation id (X-ConversationId response header). OmniRoute detects when a follow-up request continues the same conversation via fingerprint + bounded prefix-hash matching, with a strict-growth invariant to prevent false merges between independent single-shot requests that happen to share identical opening content. Continuation detection excludes the system message from the identity anchor, since real coding-agent CLIs commonly regenerate it every request with live context (timestamp, cwd, git status) — without this, that volatility alone broke every continuation check against real traffic. - `/dashboard/logs`: new toggleable Conversation column. - `/dashboard/logs/timeline`: requests sharing a conversation id share a timeline lane, connected by an arrow, with a configurable lane-reuse window. - Request detail panel: new Full Conversation transcript above the raw SSE event stream — Markdown rendering, per-turn timestamps, turn-relative view, click-any-turn navigation, live auto-refresh building the transcript in real time from the in-flight SSE chunk buffer while a request is still streaming, auto-scroll-to-bottom as the live turn grows. - New `/dashboard/conversations` page listing conversations with 2+ turns, no-forking model (an edited/duplicated mid-history turn mints its own independent conversation instead of merging), pagination, duplicate- anchor fix. - Configurable auto-refresh intervals on both the timeline and conversations list pages. - Responses API tool-call gap fix: turnsFromOpenAiMessages only handled role-based Chat Completions messages, so bare {type:"function_call"} / {type:"function_call_output"} / {type:"reasoning"} items (real Responses API traffic) silently vanished from the Conversation Context panel. - truncateForLog now counts input[] (Responses API), not just messages[] (Chat Completions), so a truncated /v1/responses request still shows a placeholder instead of nothing. - RequestTimeline.tsx now reads the same debugEnabled/emailsVisible settings RequestLoggerV2.tsx already used, instead of hardcoding both false — the timeline view never showed SSE/stream-chunk events or respected email-masking, regardless of the actual setting. Migrations 147/148 (agentic_conversations, conversation_turn_nodes) — 135 and 136 are now taken upstream; 143-145 are documented KNOWN_GAPS, so this uses the next free slot past upstream's current highest. Test plan: - npm run typecheck:core — clean - npm run lint — clean - node --import tsx/esm scripts/check/check-migration-numbering.mjs — OK, 0 collisions - 109 unit tests across the conversation-tracking, migration-renumber, and dashboard-wiring surface — 0 failures * refactor(dashboard): reuse call-log artifacts for conversation transcript content conversation_turn_nodes no longer stores turn text/tool-call content (text_preview/block_kind/tool_name) -- it's identity-only now (id/parent/ content_hash), matching agentic_conversations' existing lightweight-index shape. Every node's originating request is already fully captured by the call-log pipeline artifact its last_correlation_id points at, so the /dashboard/conversations tree view resolves each node's actual display content on demand from there (open-sse/services/conversationTurnContent.ts), re-running the same extractCanonicalTurns/hashTurnContent the write path used and matching by content_hash, instead of duplicating conversation content into a second store under a separate retention/gating policy. This also drops the old 8000-char text_preview truncation entirely -- resolved content is always full and untruncated. The frontend contract is unchanged (tree API still returns {textPreview, blockKind, toolName} per node), so the dashboard UI itself (page.tsx, RequestLoggerDetail/RequestTimeline, sidebar, i18n) needed no changes. Renumbered the cherry-picked 147/148 migrations to 153/154 -- 147 now collides with 147_api_keys_model_access_mode.sql, which landed on release/v3.8.50 after this work was originally built. Also includes a standalone, unrelated fix carried along from this rebase: close isProviderModelHidden's missing function-body brace in modelSelectModalHelpers.ts (separately landed as diegosouzapw#10206). Stacked on feat/responses-previous-response-id-virtualization (diegosouzapw#3), which is itself stacked on feat/openai-responses-store-toggle (diegosouzapw#10121). * fix(dashboard): resync conversation list on open so the live-text poll starts immediately openConversation() seeded activeConversation (and therefore activeCallLogId, which gates the live-partial-text poll effect) from whatever row snapshot the list's own fixed-interval poll last produced. A conversation opened right after a reply started streaming -- after that tick, before the next -- had activeCallLogId still null, so the live-text poll never started; only a subsequent background list-poll resync (already existed) picked it up, which is why closing and reopening the same conversation "just worked". loadConversations() is now a shared callback so openConversation can force one immediately on open instead of waiting on pollSeconds. Live-verified against omniroute-dev: opening a conversation mid-stream now shows live reasoning on the first open. * style: prettier formatting for conversationTurnContent.test.ts * fix(db): close migration numbering gap left by decoupling from diegosouzapw#3/diegosouzapw#10262 153/154 (originally 154/155) were chosen back when this branch stacked on top of the previous_response_id migration (153_call_logs_response_id.sql). Decoupling removed that migration from this branch's history, leaving an unused 153 slot that check-migration-numbering.test.ts correctly flags as a gap. * refactor(dashboard): split RequestTimeline/RequestLoggerDetail under the 1000-line file-size cap Both files exceeded check-file-size's new-file cap after this PR's own additions (RequestTimeline 1048, RequestLoggerDetail 1163). Extracted pure non-component logic (types, constants, allocateLanes and its helpers) out of RequestTimeline.tsx into RequestTimeline.utils.ts, and the two self-contained presentational sub-components (PayloadSection, ConversationContextSection + its private helper) out of RequestLoggerDetail.tsx into RequestLoggerDetail.sections.tsx. No behavior change; existing external imports (default exports, allocateLanes, TimelineLog, CONVERSATION_LANE_REUSE_STORAGE_KEY) still resolve from the original file paths. * fix(db): renumber agentic-conversation migrations to clear 153 collision + sync migration-count docs The refresh-merge of release/v3.8.50 exposed that the feature's three migrations collided at slot 153 with the base's radar_local_model_state (153) and its own call_logs_response_id. Migration runner enforces unique numeric prefixes -> every DB init threw, red-ing Vitest, all Unit shards and the DB-backed quality gates. Renumber the feature's pair to 155_agentic_conversations / 156_conversation_turn_nodes and move call_logs_response_id to 154 (keeps 153_radar base-owned, preserves agentic-before-turn_nodes ordering). Update SQL headers and the 154/156 references in feature code + tests. Migration count is now 151 (was 148 stale in README/AGENTS/llm.txt) — sync the doc counts to clear the docs-accuracy gate. Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com> * fix(ui): drop unused CONVERSATION_LANE_REUSE_STORAGE_KEY re-export from RequestTimeline Knip 6.32 (baseline 415) flags the public re-export of CONVERSATION_LANE_REUSE_STORAGE_KEY from RequestTimeline.tsx as dead: no external consumer imports it through that re-export (it is imported and used directly from RequestTimeline.utils.ts inside the component). Removed the unused re-export; the internal import stays. DEAD_TOTAL 416 -> 415, back to the frozen baseline. Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com> * fix(agentic-conversations): guard resolveConversationId, drop dead whole-chain export - Wrap resolveConversationId() in try/catch in chat.ts, matching the defensive pattern used by every other best-effort side call nearby, so a DB hiccup in conversation tracking can't turn a working chat request into a hard failure. - Remove getConversationTurnTree: knip's project scope excludes tests/**, so an export used only by tests can never register as used there. Swap its 8 test call sites to the paginated getConversationTurnPage (already the dashboard's canonical query) with a generous limit, collapsing to one query path instead of keeping a second whole-chain export alive solely for test convenience. - Regenerate i18n llm.txt mirrors from root (pre-existing drift on this branch, unrelated to the above, caught by the docs-sync pre-commit gate). Addresses PR review feedback. * fix(i18n): close requestLogger conversation-column gap, fix domain-modules count drift - fr.json, vi.json were missing requestLogger.columns.conversation (added in the conversation-tracking feature), failing i18n-vi-completeness.test.ts. - docs/i18n/*/llm.txt mirrors still said 117 domain-specific files after an earlier rebase fixed the migration count but missed this companion number, failing check-docs-sync.mjs across all 42 locales. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> * fix(docs): restore PROXY_LOG_INCLUDE_IPS env/doc entries (env-doc-sync red) .env.example and docs/reference/ENVIRONMENT.md were both missing the PROXY_LOG_INCLUDE_IPS entry that src/lib/proxyLogger.ts already reads (confirmed present at this branch's merge-base too, so this predates the conversation-tracking work and is unrelated to it) -- the entry was added on release/v3.8.50 after this branch's last sync and this branch never picked it up. That gap red-lines tests/unit/check-env-doc-sync.test.ts and tests/unit/issue-7793-env-doc-sync-repro.test.ts (Unit Tests fast-path 2/4 in CI). Restore both entries verbatim from the current release/v3.8.50 tip -- no feature-code change. Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com> --------- Co-authored-by: hartmark <hartmark@users.noreply.github.com> Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com> Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
…e2e assertion and 10 integration reds Electron Package Smoke — a packaging defect that had been hidden behind another packaging defect for nine days. Once the loginHeaderCapture fix let the main process start, the server underneath died on 'Cannot find module next': resources/app/server.js shipped without resources/app/node_modules. electron-builder discards the ROOT node_modules in code, not by configuration — app-builder-lib/out/util/filter.js:42 has a hard-coded `if (relative === "node_modules") return false` that runs before any filter pattern. The second extraResources entry pointing INTO ../.build/electron-standalone/node_modules is what sidesteps it, because those relative paths are never equal to "node_modules". diegosouzapw#10325 removed that entry as an apparent duplicate and flipped the test to assert "exactly once", freezing the regression as if it were the contract. Restored, and the unit guard now pins both entries — proven by mutation: reverting package.json to the post-diegosouzapw#10325 shape fails the guard 3/4, restoring it passes 4/4. group-b-quota-plans-config — the assertion was impossible to satisfy on ANY route, and the page was never broken. layout.tsx hands the whole message catalogue to NextIntlClientProvider, React serialises that prop into the RSC payload, and en.json carries "Internal Server Error" twice, so page.content() always contains it: probing /dashboard, /dashboard/costs, /dashboard/settings and /login showed the string present with every page rendering fine, and a pageerror probe on the failing run captured zero client exceptions. This is the same trap that killed the sibling not.toContain("500") in feb0709 ("raw HTML is unreliable") — that one was removed, this one was kept. Now asserts on rendered text, which still catches a real error boundary. The pageerror capture stays: the CI failure carried no stack trace, which is why it was misread twice. Integration — 10 of the 14 shard-2 reds, all sibling-test gaps behind security fixes: monitoring health now takes a Request and requires management auth (GHSA-mvf8-qc78-5mxm); the OAuth import routes moved to requireManagementAuth (GHSA-mg76) — the test accepts both guard shapes and gained a stronger anchor that every exported handler awaits a guard on its own request, mutation-verified; skill tool names are derived from encodeSkillToolName() and the fake upstream now returns the encoded name so decodeSkillToolName() is exercised too; previous_response_id now fails closed (diegosouzapw#10262); proxy_logs persist as an async batch (diegosouzapw#11182) so the test flushes first; providerQuotaOverrides joined GET /api/resilience (diegosouzapw#9871); the reasoning fixture used a model that stopped being thinking-incompatible, replaced and pinned with a premise assert so it cannot rot silently again. A vacuous assert.ok(true, "all 10 streams completed without hanging") was replaced with real anchors — content must arrive on every stream and the active Timeout count must not grow. Four are deliberately left red rather than aligned, each now tracked: diegosouzapw#11551 (the /v1/models after() wiring is dead — the route passes a third argument to a two-parameter function and catalogCache never imports after, so the diegosouzapw#8728 contract is unimplemented), diegosouzapw#11552 (~27% of requests emit an extra discarded upstream call; the delivered distribution is exactly 0.70, so weighted routing is correct and the waste is the real finding), the fixed-account combo pin (aligning it would destroy the per-step attribution the test exists for), and the web_search fallback already tracked as diegosouzapw#11524. Package Artifact — the provenance stamp I added last round used git rev-parse HEAD, which under pull_request is the ephemeral merge commit and therefore never an ancestor of the release branch. Now takes the PR head sha. Refs diegosouzapw#10692
Summary
previous_response_id/storecontinuation to clients unconditionally, even when the selected upstream provider has no native Responses-API state support. Reconstruction happens server-side inhandleChatImplementation, before any downstream validation or provider translation: OmniRoute resolves the response id back to the full input/output it previously produced, prepends it to the client's delta, and forwards the full reconstructed history upstream exactly as it does today. Client<->OmniRoute traffic shrinks to the new delta only; OmniRoute<->provider traffic is unchanged.call_log_pipeline_enabled, already retained/cleaned up by the existing call-log lifecycle) instead of duplicating conversation content into a second store — only a lightweightcall_logs.response_idindex is new.api_key_idso one client can never resolve another client's stored conversation, and any unresolvable/missing/size-limit-omitted state fails closed with OpenAI's ownprevious_response_not_foundcontract.Originally stacked on
feat/openai-responses-store-toggle, which has since merged as #10121 — this PR is now rebased directly onto currentrelease/v3.8.50and stands alone.Related Issues
Validation
node --import tsx --test tests/unit/responses-continuation-store.test.ts(6 pass)npm run lint— cleannpx tsc -p tsconfig.typecheck-core.json— cleanTests Added Or Updated
tests/unit/responses-continuation-store.test.ts(new, 6 tests: reconstruction from artifact, unknown response id, cross-tenant isolation, missing artifact on disk, size-limit-omitted payload fails closed, never-captured detail logging)Coverage Notes
src/lib/db/responsesContinuationStore.ts(new),src/sse/handlers/chat.ts,src/lib/usage/callLogs.ts,open-sse/handlers/chatCore/attemptLogging.ts— all covered by the new test file plus existing chat-handler test coverage for the surrounding request path.Reviewer Notes
153_call_logs_response_id.sql(ALTER TABLE call_logs ADD COLUMN response_id TEXT DEFAULT NULL+ index) is purely additive, no backfill needed.previous_response_idreturns OpenAI's ownprevious_response_not_founderror rather than proceeding without the prior context.