Skip to content

fix: stabilize reasoning streams and request logs - #3879

Merged
diegosouzapw merged 17 commits into
diegosouzapw:release/v3.8.26from
rdself:coder/stream-log-root-cause-main
Jun 15, 2026
Merged

diegosouzapw merged 17 commits into
diegosouzapw:release/v3.8.26from
rdself:coder/stream-log-root-cause-main

Conversation

@rdself

@rdself rdself commented Jun 15, 2026 •

Copy link
Copy Markdown
Contributor

Summary

  • Keep successful streaming requests pending until the chat completion callback finalizes the exact request id, so active request rows do not disappear before durable call-log persistence catches up.
  • Re-serialize passthrough OpenAI-compatible stream chunks after splitting mixed reasoning_content plus content deltas, so downstream clients receive a reasoning-only event followed by a content-only event.
  • Bind stream chunk debug payloads and completion finalization to the exact pending request id before falling back to the legacy model/provider/connection lookup.
  • Stop sending redact-thinking-2026-02-12 by default on Claude Code-compatible relays; expose it as a per-connection CC Compatible toggle for upstreams that require redacted thinking streams.
  • Rebase onto release/v3.8.26 and align the quality ratchet baselines to values measured on that release base, keeping this PR neutral against the target branch.

Root cause

There were three independent streaming/logging issues that could combine into the same user-visible failure.

First, createSSEStream() split mixed reasoning/content OpenAI-compatible deltas, but the passthrough branch could still reuse the original serialized SSE payload unless a later mutation forced reserialization. Clients that process reasoning before content, including SillyTavern-style OpenAI-compatible parsers, could then receive a mixed event and treat later reasoning/content boundaries incorrectly.

Second, the successful stream flush() path cleared the pending request before onComplete ran. The upper layer now finalizes by exact pending request id, but the request had already been removed from the pending map, so the in-memory completion fallback could be lost before the durable call_logs row was visible. That matched the observed behavior where the activity row spun for a while and then disappeared even though the upstream provider completed and billed normally.

Third, anthropic-compatible-cc-* relays sent redact-thinking-2026-02-12 unconditionally. With extended thinking enabled, Claude can spend most or all of the output budget on thinking. If the upstream redacts that thinking, OmniRoute receives little or no visible stream content even though the provider reports a successful request.

Changes

  • Leave successful stream lifecycle cleanup to onComplete when completion finalization is available; only the legacy no-callback success path clears pending requests directly.
  • Keep failure, idle-timeout, and empty-Claude-stream cleanup explicit and idempotent so errored requests still leave the active set.
  • Move chatCore stream finalization before usage/cost/persistence side effects, so UI lifecycle state is stabilized even if later accounting work fails.
  • Add regression coverage proving successful passthrough streams remain pending until onComplete finalizes the exact request id.
  • Preserve provider-delimited reasoning by reserializing split SSE events, without inventing reasoning from unmarked content.
  • Make CC Compatible thinking redaction opt-in through connection-level providerSpecificData.requestDefaults.redactThinking and Dashboard add/edit toggles.

Validation

  • node --import tsx/esm --test tests/unit/provider-request-failure-pipeline.test.ts
  • node --import tsx/esm --test tests/unit/claude-code-compatible-helpers.test.ts tests/unit/request-defaults-store-session.test.ts tests/unit/provider-specific-data-schema.test.ts tests/unit/executor-default-base.test.ts tests/unit/provider-page-helpers-3501.test.ts
  • node --import tsx/esm --test tests/unit/request-logger-endpoints.test.ts tests/unit/stream-utils.test.ts tests/unit/active-request-stream-chunks-lifecycle.test.ts
  • npx vitest run open-sse/mcp-server/__tests__/glmCodingProviderConfig.test.ts
  • npx vitest run 'src/app/(dashboard)/dashboard/providers/[id]/components/modals/__tests__/connModals.test.tsx'
  • npm run typecheck:core
  • npm run check:file-size
  • npm run quality:collect
  • npm run quality:ratchet -- --allow-missing
  • npm run check:cognitive-complexity
  • npm run check:type-coverage
  • npm run check:dead-code
  • git diff --check

Release-base quality comparison

  • origin/release/v3.8.26: npm run quality:collect measured eslintWarnings=3760, eslintErrors=0, openapiCoverage.pct=38, i18nUiCoverage.pct=79.7.
  • This PR after rebase: npm run quality:collect measured the same values.
  • origin/release/v3.8.26: dedicated gates measured cognitiveComplexity=753 and deadExports=339.
  • This PR after rebase: dedicated gates measured the same values.

Live stream checks

  • openrouter/deepseek/deepseek-v4-flash: 59.8s stream, 840 SSE events, 379 reasoning_content chunks followed by 456 content chunks, first reasoning at 6.14s, first content at 30.79s, mixedChunks=0, completed with [DONE].
  • anthropic-compatible-cc-sp-anthropic/claude-opus-4-6 before disabling default redaction: 46.9s stream, 4 SSE events, 0 reasoning chunks, 0 content chunks, output_tokens=2048, thinking_tokens=2047, stop_reason=max_tokens.
  • anthropic-compatible-cc-sp-anthropic/claude-opus-4-6 after disabling default redaction: 38.7s stream, 62 SSE events, 26 reasoning_content chunks, 32 content chunks, first reasoning at 1.73s, first content at 23.86s, mixedChunks=0, completed with [DONE].
  • openrouter/deepseek/deepseek-v4-pro@preset/prefer through the local production build: 70.6s stream, 6,434 SSE events, first reasoning at 1.59s, first content at 48.2s, mixedChunks=0, completed with [DONE] and durable call_logs.detail_state=ready.
  • anthropic-compatible-cc-sp-anthropic/claude-opus-4-6 through the local production build: 71.8s stream, 107 SSE events, 101 reasoning_content chunks, mixedChunks=0, completed with [DONE] and durable call_logs.detail_state=ready.

@rdself
rdself requested a review from diegosouzapw as a code owner June 15, 2026 09:04
Copilot AI review requested due to automatic review settings June 15, 2026 09:04

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces exact pending request tracking and finalization using a unique request ID, preventing mismatches when multiple requests for the same model overlap. It also updates the call logs API to include completed in-memory fallback rows, extends the completed details TTL to 120 seconds, and ensures proper reserialization when splitting mixed reasoning and content deltas in SSE streams. Feedback on the changes highlights a potential TypeError in open-sse/utils/requestLogger.ts due to inconsistent optional chaining on entry when accessing entry.provider directly.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment thread open-sse/utils/requestLogger.ts

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Note

Copilot was unable to run its full agentic suite in this review.

Adds stronger in-memory request lifecycle tracking so stream chunks and “just completed” requests reliably show up in usage logs and can be finalized by exact request id.

Changes:

  • Introduces finalizePendingRequestById and extends the completed-details cache TTL to support slower UI polling.
  • Updates request logging to bind streamChunks to the exact pending request via requestId (with safe fallback matching).
  • Extends /api/usage/call-logs to include “completed in-memory” rows and adds test coverage around the new behaviors.

Reviewed changes

Copilot reviewed 8 out of 8 changed files in this pull request and generated 7 comments.

Show a summary per file
File Description
tests/unit/stream-utils.test.ts Adds a regression test ensuring passthrough SSE output is consumable by SillyTavern-style reasoning parsing.
tests/unit/request-logger-endpoints.test.ts Adds tests for completed in-memory rows and requestId-based stream chunk binding.
tests/unit/active-request-stream-chunks-lifecycle.test.ts Updates lifecycle coverage for overlap finalization and longer completed-details visibility.
src/lib/usage/usageHistory.ts Adds completed-details TTL, exact-id finalization, and refactors finalization/cleanup logic.
src/app/api/usage/call-logs/route.ts Prepends pending + completed in-memory rows to persisted call logs with deduping.
open-sse/utils/stream.ts Ensures mixed reasoning/content split triggers reserialization reliably.
open-sse/utils/requestLogger.ts Adds requestId option and tightens matching to avoid cross-connection streamChunk attachment.
open-sse/handlers/chatCore.ts Passes pendingRequestId into logger and finalizes by id first, falling back to legacy behavior.

Comment thread src/app/api/usage/call-logs/route.ts Outdated
Comment thread src/app/api/usage/call-logs/route.ts Outdated
Comment thread src/lib/usage/usageHistory.ts Outdated
Comment thread src/lib/usage/usageHistory.ts Outdated
Comment thread src/lib/usage/usageHistory.ts Outdated
Comment thread tests/unit/stream-utils.test.ts
Comment thread tests/unit/request-logger-endpoints.test.ts Outdated
@rdself
rdself marked this pull request as draft June 15, 2026 09:11
@rdself
rdself marked this pull request as ready for review June 15, 2026 19:38
@rdself
rdself changed the base branch from main to release/v3.8.26 June 15, 2026 22:09
@rdself
rdself marked this pull request as draft June 15, 2026 22:44
diegosouzapw and others added 2 commits June 15, 2026 19:54
…am-log-root-cause-main

# Conflicts:
#	src/app/(dashboard)/dashboard/providers/[id]/components/modals/AddApiKeyModal.tsx
#	src/app/(dashboard)/dashboard/providers/[id]/components/modals/EditConnectionModal.tsx
…reconcile

PR #3879 reconciled with #3921: AddApiKeyModal.tsx 843->845 (+2 = the
CcCompatibleRequestDefaultsFields context1m+redact-thinking component sits
inside the cc-compatible settings block alongside the OpenRouter preset input).

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>
@diegosouzapw

Copy link
Copy Markdown
Owner

Thanks, @rdself! 🙏 Three real, observed fixes in one well-tested PR: (1) pending-request lifecycle — keeping the request pending until onComplete() finalizes by exact id (finalizePendingRequestById) so durable call_logs persist (activity rows stopped vanishing); (2) mixed reasoning/content SSE re-serialization for strict parsers (SillyTavern-style); (3) CC-compatible redact-thinking now opt-in per-connection (it was unconditional and could eliminate output under extended thinking). Guarded fallbacks, behavioral tests, live-verified.

I reconciled the modal conflict with #3921 (just merged): your CcCompatibleRequestDefaultsFields (context1m + redact-thinking) now lives inside the cc-compatible settings block next to the OpenRouter preset input. Validated: typecheck:core clean, 115 behavioral tests green (lifecycle/stream-utils/request-logger/cc-helpers), conn-modals vitest 8/8, lint clean. Merging into release/v3.8.26. 🚀

@diegosouzapw
diegosouzapw marked this pull request as ready for review June 15, 2026 22:57
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Warning

You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again!

@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@diegosouzapw
diegosouzapw merged commit cdefc98 into diegosouzapw:release/v3.8.26 Jun 15, 2026
1 check passed
diegosouzapw pushed a commit that referenced this pull request Jun 15, 2026
Integrated into release/v3.8.26 — mid-stream failure persistence (follow-up to #3879). Validated locally: typecheck:core clean, 75/75 stream/usage tests, eslint 0 errors, file-size + any-budget OK.
diegosouzapw added a commit to felipesartori/OmniRoute that referenced this pull request Jun 18, 2026
…unner collects them

The 3 React-component tests landed at top-level tests/unit/*.test.tsx, which
no runner collects: test:unit only globs tests/unit/*.test.ts, and test:vitest:ui
filters to tests/unit/ui. check:test-discovery flagged them as NEW orphans (never
run). Moved them under tests/unit/ui/ (collected by vitest.config.ts + the ui path
filter) and bumped the relative import one level (../../ -> ../../../). Also bumped
the file-size baseline for the two connection modals (cohesive free-models toggle,
mirroring diegosouzapw#3879/diegosouzapw#2997). 12/12 vitest + 16/16 node tests pass; test-discovery clean.

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>
diegosouzapw pushed a commit that referenced this pull request Jun 18, 2026
…4176)

Integrated into release/v3.8.29 — import-only-free-models connection option + free/paid list filters and free-first sort. Thanks @felipesartori! On review we relocated the 3 React-component tests from tests/unit/ to tests/unit/ui/ so a runner actually collects them (check:test-discovery had flagged them as orphans), bumped the file-size baseline for the two connection modals (cohesive toggle, mirroring #3879/#2997), and synced the branch with the current release. 12/12 vitest + 16/16 node tests green.
@rdself
rdself deleted the coder/stream-log-root-cause-main branch June 24, 2026 03:25
tkgo11 pushed a commit to tkgo11/OmniRoute that referenced this pull request Sep 23, 2026
…iegosouzapw#4176)

Integrated into release/v3.8.29 — import-only-free-models connection option + free/paid list filters and free-first sort. Thanks @felipesartori! On review we relocated the 3 React-component tests from tests/unit/ to tests/unit/ui/ so a runner actually collects them (check:test-discovery had flagged them as orphans), bumped the file-size baseline for the two connection modals (cohesive toggle, mirroring diegosouzapw#3879/diegosouzapw#2997), and synced the branch with the current release. 12/12 vitest + 16/16 node tests green.
tkgo11 pushed a commit to tkgo11/OmniRoute that referenced this pull request Sep 23, 2026
Integrated into release/v3.8.26 (reconciled cc-defaults UI with diegosouzapw#3921)
tkgo11 pushed a commit to tkgo11/OmniRoute that referenced this pull request Sep 23, 2026
Integrated into release/v3.8.26 — mid-stream failure persistence (follow-up to diegosouzapw#3879). Validated locally: typecheck:core clean, 75/75 stream/usage tests, eslint 0 errors, file-size + any-budget OK.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

3 participants