Skip to content

fix(sse): non-SSE JSON upstream on streaming path + SSE-wrap cache hits (#3089, #2952) - #3108

Merged
diegosouzapw merged 1 commit into
release/v3.8.9from
fix/3089-stream-json
Jun 3, 2026
Merged

diegosouzapw merged 1 commit into
release/v3.8.9from
fix/3089-stream-json

Conversation

@diegosouzapw

Copy link
Copy Markdown
Owner

Closes #3089
Closes #2952

#3089 — STREAM_EARLY_EOF on reasoning openai-compatible upstreams

Reproduced on a live instance: a reasoning openai-compatible upstream that ignores stream:true and returns a complete application/json body makes a streaming request 502 with STREAM_EARLY_EOF (stream omitted → 502; stream:false → 200). The readiness check only recognizes SSE data: frames.

Fix: chatCore detects a non-SSE JSON upstream body on the streaming path and synthesizes an equivalent OpenAI SSE stream via the new synthesizeOpenAiSseFromJson util, preserving content + reasoning_content. Normal SSE upstreams (content-type text/event-stream) are untouched.

#2952 — semantic-cache hits drop reasoning_content for streaming clients

The cache-hit path returned application/json regardless of the stream flag, so OpenAI-compatible streaming clients got a non-stream body and lost reasoning_content. Stream requests now SSE-wrap the cached completion via the same helper (non-OpenAI shapes fall back to JSON unchanged).

Tests

  • tests/unit/json-to-sse-3089.test.ts — 4 cases (reasoning, content-only, tool_calls, non-completion/invalid → "").
  • End-to-end validated on the local VM (mock reasoning upstream returning JSON; see PR discussion).

ESLint clean.

…rap cache hits (#3089, #2952)

#3089: reasoning openai-compatible upstreams that ignore stream:true and return application/json produced STREAM_EARLY_EOF because readiness only scans SSE data: frames. chatCore now detects a non-SSE JSON upstream body on the streaming path and synthesizes an equivalent OpenAI SSE stream (new synthesizeOpenAiSseFromJson util), preserving content + reasoning_content. #2952: semantic-cache hits returned application/json regardless of stream flag, so streaming clients lost reasoning_content; stream requests now SSE-wrap the cached completion via the same helper. Unit tests for the converter (4).
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@diegosouzapw
diegosouzapw merged commit b57afb5 into release/v3.8.9 Jun 3, 2026
2 checks passed
@diegosouzapw
diegosouzapw deleted the fix/3089-stream-json branch June 3, 2026 13:23

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request introduces changes to handle OpenAI-compatible upstreams that ignore the stream: true flag and return a complete JSON response instead of an SSE stream, as well as serving semantic-cache hits as SSE streams for streaming clients. It adds a utility synthesizeOpenAiSseFromJson to convert a complete chat-completion JSON body into an equivalent SSE stream, along with corresponding unit tests. The review feedback suggests optimizing cache hit handling by avoiding duplicate JSON serialization, checking if the provider response is successful (providerResponse.ok) before attempting to parse it, and ensuring that the usage field is only appended to the final chunk of the last choice in multi-choice responses to prevent duplicate usage fields.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

Comment on lines +2300 to +2311
const cachedSse = stream ? synthesizeOpenAiSseFromJson(JSON.stringify(cached)) : "";
const cacheHitMetaHeaders = buildOmniRouteResponseMetaHeaders({
provider,
model,
cacheHit: true,
latencyMs: Date.now() - startTime,
usage: cachedUsage,
costUsd: cachedCost,
});
return {
success: true,
response: new Response(JSON.stringify(cached), {
response: new Response(cachedSse || JSON.stringify(cached), {

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

To avoid serializing the cached object twice with JSON.stringify(), we can store the serialized JSON string in a variable and reuse it.

      const cachedJson = JSON.stringify(cached);
      const cachedSse = stream ? synthesizeOpenAiSseFromJson(cachedJson) : "";
      const cacheHitMetaHeaders = buildOmniRouteResponseMetaHeaders({
        provider,
        model,
        cacheHit: true,
        latencyMs: Date.now() - startTime,
        usage: cachedUsage,
        costUsd: cachedCost,
      });
      return {
        success: true,
        response: new Response(cachedSse || cachedJson, {

Comment on lines +5502 to +5506
const isNonSseJsonBody =
!!providerResponse.body &&
upstreamContentType.includes("application/json") &&
!upstreamContentType.includes("text/event-stream") &&
!upstreamContentType.includes("application/x-ndjson");

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

To prevent consuming the response body of error responses (e.g., 4xx or 5xx status codes) unnecessarily, we should check that the response is successful (providerResponse.ok) before attempting to parse and convert it to SSE.

    const isNonSseJsonBody =
      providerResponse.ok &&
      !!providerResponse.body &&
      upstreamContentType.includes("application/json") &&
      !upstreamContentType.includes("text/event-stream") &&
      !upstreamContentType.includes("application/x-ndjson");

Comment on lines +47 to +75
choices.forEach((choice, fallbackIndex) => {
if (!isRecord(choice)) return;
const index = typeof choice.index === "number" ? choice.index : fallbackIndex;
const message = isRecord(choice.message) ? choice.message : {};

// First chunk carries role + whatever the message produced (content,
// reasoning_content, tool_calls). Putting them in one delta is valid and
// keeps downstream translation simple.
const delta: JsonRecord = { role: typeof message.role === "string" ? message.role : "assistant" };
if (typeof message.content === "string" && message.content.length > 0) {
delta.content = message.content;
}
if (typeof message.reasoning_content === "string" && message.reasoning_content.length > 0) {
delta.reasoning_content = message.reasoning_content;
}
if (Array.isArray(message.tool_calls) && message.tool_calls.length > 0) {
delta.tool_calls = message.tool_calls;
}

out += sseEvent({ ...base, choices: [{ index, delta, finish_reason: null }] });

const finishReason =
typeof choice.finish_reason === "string" && choice.finish_reason ? choice.finish_reason : "stop";
const finalChoice: JsonRecord = { index, delta: {}, finish_reason: finishReason };
const finalChunk: JsonRecord = { ...base, choices: [finalChoice] };
if (isRecord(parsed.usage)) finalChunk.usage = parsed.usage;
out += sseEvent(finalChunk);
emittedAny = true;
});

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

medium

In multi-choice responses, appending the usage object to the final chunk of every choice results in duplicate usage fields being sent in the stream. According to the OpenAI specification, usage should only be returned once at the end of the stream. We can filter the choices first to identify the last valid choice and only append usage to its final chunk.

  const validChoices = choices.filter(isRecord);
  validChoices.forEach((choice, idx) => {
    const index = typeof choice.index === "number" ? choice.index : idx;
    const message = isRecord(choice.message) ? choice.message : {};

    // First chunk carries role + whatever the message produced (content,
    // reasoning_content, tool_calls). Putting them in one delta is valid and
    // keeps downstream translation simple.
    const delta: JsonRecord = { role: typeof message.role === "string" ? message.role : "assistant" };
    if (typeof message.content === "string" && message.content.length > 0) {
      delta.content = message.content;
    }
    if (typeof message.reasoning_content === "string" && message.reasoning_content.length > 0) {
      delta.reasoning_content = message.reasoning_content;
    }
    if (Array.isArray(message.tool_calls) && message.tool_calls.length > 0) {
      delta.tool_calls = message.tool_calls;
    }

    out += sseEvent({ ...base, choices: [{ index, delta, finish_reason: null }] });

    const finishReason =
      typeof choice.finish_reason === "string" && choice.finish_reason ? choice.finish_reason : "stop";
    const finalChoice: JsonRecord = { index, delta: {}, finish_reason: finishReason };
    const finalChunk: JsonRecord = { ...base, choices: [finalChoice] };

    const isLast = idx === validChoices.length - 1;
    if (isLast && isRecord(parsed.usage)) {
      finalChunk.usage = parsed.usage;
    }
    out += sseEvent(finalChunk);
    emittedAny = true;
  });

@kilo-code-bot

kilo-code-bot Bot commented Jun 3, 2026 •

Copy link
Copy Markdown

Code Review Summary

Status: 1 Issue Found | Recommendation: Address before merge

Overview

Severity Count
CRITICAL 0
WARNING 1
SUGGESTION 0
Issue Details (click to expand)

WARNING

File Line Issue
open-sse/utils/jsonToSse.ts 59-64 Missing reasoning alias support in synthesizeOpenAiSseFromJson — the codebase elsewhere (streamPayloadCollector.ts:160-167, sseParser.ts:148-155, responseSanitizer.ts:316-323) normalizes the reasoning field alias to reasoning_content for providers like NVIDIA kimi-k2.5, but this new function only checks message.reasoning_content. A provider using reasoning in a non-streaming JSON response would lose that data during synthesis.
Other Observations (not in diff)

No issues found in unchanged code.

Files Reviewed (4 files)
  • open-sse/utils/jsonToSse.ts - 1 issue
  • open-sse/handlers/chatCore.ts - no issues
  • tests/unit/json-to-sse-3089.test.ts - no issues
  • CHANGELOG.md - no issues

Reviewed by laguna-m.1-20260312:free · 3,778,929 tokens

diegosouzapw added a commit that referenced this pull request Jun 4, 2026
#2952/#3108 made streaming cache hits SSE-wrapped (so streaming clients keep
content + reasoning_content), but two chatcore tests still asserted the pre-fix
'cache HIT returns JSON regardless of stream flag'. Update them to assert SSE
(text/event-stream) + verify the cached content appears in the SSE frames.
HouMinXi pushed a commit to HouMinXi/OmniRoute that referenced this pull request Aug 2, 2026
…rap cache hits (diegosouzapw#3089, diegosouzapw#2952) (diegosouzapw#3108)

diegosouzapw#3089: reasoning openai-compatible upstreams that ignore stream:true and return application/json produced STREAM_EARLY_EOF because readiness only scans SSE data: frames. chatCore now detects a non-SSE JSON upstream body on the streaming path and synthesizes an equivalent OpenAI SSE stream (new synthesizeOpenAiSseFromJson util), preserving content + reasoning_content. diegosouzapw#2952: semantic-cache hits returned application/json regardless of stream flag, so streaming clients lost reasoning_content; stream requests now SSE-wrap the cached completion via the same helper. Unit tests for the converter (4).
HouMinXi pushed a commit to HouMinXi/OmniRoute that referenced this pull request Aug 2, 2026
…rap behavior

diegosouzapw#2952/diegosouzapw#3108 made streaming cache hits SSE-wrapped (so streaming clients keep
content + reasoning_content), but two chatcore tests still asserted the pre-fix
'cache HIT returns JSON regardless of stream flag'. Update them to assert SSE
(text/event-stream) + verify the cached content appears in the SSE frames.
Poid-ZA pushed a commit to Poid-ZA/OmniRoute that referenced this pull request Aug 5, 2026
…rap cache hits (diegosouzapw#3089, diegosouzapw#2952) (diegosouzapw#3108)

diegosouzapw#3089: reasoning openai-compatible upstreams that ignore stream:true and return application/json produced STREAM_EARLY_EOF because readiness only scans SSE data: frames. chatCore now detects a non-SSE JSON upstream body on the streaming path and synthesizes an equivalent OpenAI SSE stream (new synthesizeOpenAiSseFromJson util), preserving content + reasoning_content. diegosouzapw#2952: semantic-cache hits returned application/json regardless of stream flag, so streaming clients lost reasoning_content; stream requests now SSE-wrap the cached completion via the same helper. Unit tests for the converter (4).
Poid-ZA pushed a commit to Poid-ZA/OmniRoute that referenced this pull request Aug 5, 2026
…rap behavior

diegosouzapw#2952/diegosouzapw#3108 made streaming cache hits SSE-wrapped (so streaming clients keep
content + reasoning_content), but two chatcore tests still asserted the pre-fix
'cache HIT returns JSON regardless of stream flag'. Update them to assert SSE
(text/event-stream) + verify the cached content appears in the SSE frames.
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…rap cache hits (diegosouzapw#3089, diegosouzapw#2952) (diegosouzapw#3108)

diegosouzapw#3089: reasoning openai-compatible upstreams that ignore stream:true and return application/json produced STREAM_EARLY_EOF because readiness only scans SSE data: frames. chatCore now detects a non-SSE JSON upstream body on the streaming path and synthesizes an equivalent OpenAI SSE stream (new synthesizeOpenAiSseFromJson util), preserving content + reasoning_content. diegosouzapw#2952: semantic-cache hits returned application/json regardless of stream flag, so streaming clients lost reasoning_content; stream requests now SSE-wrap the cached completion via the same helper. Unit tests for the converter (4).
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…rap behavior

diegosouzapw#2952/diegosouzapw#3108 made streaming cache hits SSE-wrapped (so streaming clients keep
content + reasoning_content), but two chatcore tests still asserted the pre-fix
'cache HIT returns JSON regardless of stream flag'. Update them to assert SSE
(text/event-stream) + verify the cached content appears in the SSE frames.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant