Skip to content

fix: strip internal reasoning placeholder from user-visible content (#8081) - #8162

Merged
diegosouzapw merged 5 commits into
diegosouzapw:release/v3.8.49from
Dingding-leo:fix/strip-reasoning-placeholder-from-content
Jul 22, 2026
Merged

diegosouzapw merged 5 commits into
diegosouzapw:release/v3.8.49from
Dingding-leo:fix/strip-reasoning-placeholder-from-content

Conversation

@Dingding-leo

Copy link
Copy Markdown
Contributor

Summary

Fixes #8081. The internal reasoning replay sentinel (prior reasoning summary unavailable) can leak into user-visible assistant content when a model echoes it through ordinary message.content / delta.content. Existing suppression only checked reasoning_content fields and reasoning-specific events.

Changes

  • open-sse/utils/reasoningPlaceholder.ts: Add stripInternalReasoningPlaceholder() — removes all occurrences of the sentinel and trims; returns "" when nothing meaningful remains
  • Streaming paths (3 files): Strip at the delta.content entry point in responsesTransformer.ts, openai-responses.ts, and openai-to-claude.ts; skip emission entirely when only the placeholder was present
  • Non-streaming paths (responseSanitizer.ts): Strip in sanitizeMessageContent() and sanitizeResponsesMessageContent() (all three text paths)

translateText is unaffected (uses mode='translate' via the plain newsClient). The per-provider reasoning_content check remains as defense-in-depth.

Verification

  • npx tsc --pretty false -p tsconfig.typecheck-core.json — zero errors
  • reasoning-cache.test.ts — 52/52 pass
  • translator-resp-openai-to-claude.test.ts — 11/11 pass
  • tool-request-sanitization.test.ts — 8/8 pass
  • response-sanitizer.test.ts + responses-transformer.test.ts + translator-resp-openai-responses.test.ts — 83/83 pass

5 files changed, +33 −11

…iegosouzapw#8081)

The internal reasoning replay sentinel '(prior reasoning summary unavailable)'
can leak into user-visible assistant content when a model echoes it through
ordinary message.content / delta.content. Existing suppression only checked
reasoning_content fields and reasoning-specific events.

Changes:
- Add stripInternalReasoningPlaceholder() to reasoningPlaceholder.ts —
  removes all occurrences of the sentinel and trims; returns '' when
  nothing meaningful remains
- Streaming: strip in responsesTransformer.ts, openai-responses.ts, and
  openai-to-claude.ts at the delta.content entry point; skip emission
  entirely when only the placeholder was present
- Non-streaming: strip in responseSanitizer.ts sanitizeMessageContent()
  and sanitizeResponsesMessageContent() (all three text paths)

translateText is unaffected (uses mode='translate' via plain newsClient).
The per-provider reasoning_content check remains as defense-in-depth.
Copilot AI review requested due to automatic review settings July 22, 2026 13:51
@gemini-code-assist

Copy link
Copy Markdown
Contributor

Caution

The consumer version of Gemini Code Assist on GitHub has been sunset. All code review activity has officially ceased.

Copilot AI left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@chatgpt-codex-connector chatgpt-codex-connector Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

💡 Codex Review

Here are some automated review suggestions for this pull request.

Reviewed commit: 9dd2bfd1a8

ℹ️ About Codex in GitHub

Your team has set up Codex to review pull requests in this repo. Reviews are triggered when you

  • Open a pull request for review
  • Mark a draft as ready
  • Comment "@codex review".

If Codex has suggestions, it will comment; otherwise it will react with 👍.

Codex can also answer questions or update the PR. Try commenting "@codex address that feedback".

* meaningful remains so callers can skip emission entirely.
*/
export function stripInternalReasoningPlaceholder(value: string): string {
return value.replaceAll(NON_ANTHROPIC_THINKING_PLACEHOLDER, "").trim();

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 Badge Preserve streamed token whitespace when stripping sentinel

Because stripInternalReasoningPlaceholder() is now called for every delta.content chunk in the Responses and Claude translators, the unconditional .trim() changes normal content even when the sentinel is absent. Streaming chunks commonly begin with a space (for example " world" after "Hello"), so this turns user-visible text into "Helloworld"; only use trimming to decide whether the placeholder-only result is empty, not for the returned content.

Useful? React with 👍 / 👎.

// through ordinary content (#8081).
if (delta.content) {
const strippedContent = stripInternalReasoningPlaceholder(delta.content);
if (!strippedContent) return;

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 Badge Keep processing after placeholder-only SSE events

When an upstream emits a placeholder-only delta.content, this exits the entire transform() call, not just the text-emission branch. Since a single network chunk can contain multiple complete SSE messages, a placeholder event at the front of the buffer causes later visible deltas, tool calls, or the finish event in the same buffer to be dropped; suppress this event’s text while continuing to process the remaining messages and handlers.

Useful? React with 👍 / 👎.

@diegosouzapw
diegosouzapw merged commit 07dada6 into diegosouzapw:release/v3.8.49 Jul 22, 2026
2 of 5 checks passed
@diegosouzapw

Copy link
Copy Markdown
Owner

Obrigado, @Dingding-leo — mergeado na release/v3.8.49. Corrigi o bug que você tinha aberto: os guards passaram a pular apenas a emissão do content block vazio, mantendo finish_reason/tool_calls do mesmo chunk. Adicionei o teste de regressão (placeholder-only + finish_reason no mesmo chunk ainda emite o stop) — provado RED sem o fix, GREEN com. 45/45 nas suites de tradução. Escolhemos este PR para o #8081 (o #8110 foi fechado em favor deste).

fenix007 pushed a commit to fenix007/OmniRoute that referenced this pull request Jul 23, 2026
…s + eslint baseline

The release branch accumulated deterministic unit-test failures (fast-path red on
every open PR). These are the ones with a clear, surgical root cause:

1. diegosouzapw#6863 combo model-lockout — the diegosouzapw#7940/diegosouzapw#7980 "cap exactCooldownMs against
   maxCooldownMs" clamp was also clamping an AUTHORITATIVE parsed upstream quota
   reset (e.g. "Resets in 92h27m28s") down to maxCooldownMs, so an exhausted model
   was retried far too early. recordModelLockoutFailure now takes
   exactCooldownIsUpstreamReset — set by the combo callers when the exact cooldown
   is a real upstream reset — which exempts it from the cap. The diegosouzapw#7980 computed
   until-midnight cap is unchanged (flag absent → still capped).

2. diegosouzapw#5786 streaming claude←codex — stripInternalReasoningPlaceholder (diegosouzapw#8081/diegosouzapw#8162)
   unconditionally .trim()'d every value. On the per-delta streaming path this ate
   the meaningful edge spaces of each delta ("Hello, " + "world." + " Bye." glued to
   "Hello,world.Bye."). It now only collapses to "" when whitespace is all that
   remains after removing the placeholder, preserving real content verbatim.

3. SPAWN_CAPABLE_PREFIXES test — diegosouzapw#7892 added /api/vnc-session (11th spawn-capable
   prefix, spawns Docker) but the client-safe guard test still expected 10 and did
   not list it. Aligned to 11 + added the entry to the checklist.

4. ESLint baseline — diegosouzapw#8008/diegosouzapw#8062 merged new test files with no-explicit-any without
   refreshing the frozen suppressions, so "No new ESLint warnings" went red for the
   whole branch. Regenerated the two affected entries
   (combo-routing-engine.test.ts 269→271, oauth-refresh-connection-dedup-8059.test.ts +1).

Validated: the three failing tests now pass; the sibling guards they interact with
stay green (diegosouzapw#7980 exact-cooldown-cap 4/4, diegosouzapw#8162 placeholder suites 17+12+41,
account-fallback 77); typecheck:core clean; lint:json --max-warnings 0 exits 0.

NOTE: the release branch has ~20 further real base-red failures (compression-engine
catalog, handleChat fallback, provider candidate transparency, i18n, misc). Those are
tracked separately, one focused PR per root-cause cluster; this PR is the first slice.
fenix007 pushed a commit to fenix007/OmniRoute that referenced this pull request Jul 24, 2026
…s + eslint baseline

The release branch accumulated deterministic unit-test failures (fast-path red on
every open PR). These are the ones with a clear, surgical root cause:

1. diegosouzapw#6863 combo model-lockout — the diegosouzapw#7940/diegosouzapw#7980 "cap exactCooldownMs against
   maxCooldownMs" clamp was also clamping an AUTHORITATIVE parsed upstream quota
   reset (e.g. "Resets in 92h27m28s") down to maxCooldownMs, so an exhausted model
   was retried far too early. recordModelLockoutFailure now takes
   exactCooldownIsUpstreamReset — set by the combo callers when the exact cooldown
   is a real upstream reset — which exempts it from the cap. The diegosouzapw#7980 computed
   until-midnight cap is unchanged (flag absent → still capped).

2. diegosouzapw#5786 streaming claude←codex — stripInternalReasoningPlaceholder (diegosouzapw#8081/diegosouzapw#8162)
   unconditionally .trim()'d every value. On the per-delta streaming path this ate
   the meaningful edge spaces of each delta ("Hello, " + "world." + " Bye." glued to
   "Hello,world.Bye."). It now only collapses to "" when whitespace is all that
   remains after removing the placeholder, preserving real content verbatim.

3. SPAWN_CAPABLE_PREFIXES test — diegosouzapw#7892 added /api/vnc-session (11th spawn-capable
   prefix, spawns Docker) but the client-safe guard test still expected 10 and did
   not list it. Aligned to 11 + added the entry to the checklist.

4. ESLint baseline — diegosouzapw#8008/diegosouzapw#8062 merged new test files with no-explicit-any without
   refreshing the frozen suppressions, so "No new ESLint warnings" went red for the
   whole branch. Regenerated the two affected entries
   (combo-routing-engine.test.ts 269→271, oauth-refresh-connection-dedup-8059.test.ts +1).

Validated: the three failing tests now pass; the sibling guards they interact with
stay green (diegosouzapw#7980 exact-cooldown-cap 4/4, diegosouzapw#8162 placeholder suites 17+12+41,
account-fallback 77); typecheck:core clean; lint:json --max-warnings 0 exits 0.

NOTE: the release branch has ~20 further real base-red failures (compression-engine
catalog, handleChat fallback, provider candidate transparency, i18n, misc). Those are
tracked separately, one focused PR per root-cause cluster; this PR is the first slice.
diegosouzapw added a commit that referenced this pull request Jul 24, 2026
… sentinel (#8382)

Regression: #8162 (port of #8081) added an unconditional `.trim()` to
stripInternalReasoningPlaceholder(), applied to every streaming
delta.content chunk across 3 call-sites (openai-to-claude.ts,
openai-responses.ts, responsesTransformer.ts). Leading/trailing
whitespace at a chunk boundary is a real word boundary between
streaming fragments; trimming it glues adjacent chunks together on
the client ("Hello, " + "world." + " Bye." -> "Hello,world.Bye.").

Fix: early-return via .includes() before the replaceAll+trim, so the
function is a true no-op when the sentinel is absent from the chunk.
Behavior when the sentinel IS present is unchanged.

Validation:
- tests/unit/streaming-reasoning-dedup-5786.test.ts: the "(A-guard)"
  test was RED on the base branch ('Hello,world.Bye.' vs
  'Hello, world. Bye.'); GREEN after the fix (4/4 passing).
- tests/unit/translator-resp-openai-to-claude.test.ts: added a new
  multi-chunk boundary-whitespace regression test, proven RED against
  the pre-fix code (12/13), GREEN after (13/13).
- No regressions in responses-transformer.test.ts (17/17),
  responses-transformer-dense-output.test.ts (3/3), or the other
  suites exercising the shared placeholder utility (160/160 total
  across all consumers).

Refs #8162
Refs #8081
diegosouzapw pushed a commit that referenced this pull request Jul 24, 2026
…rd spaces (#8341)

Live incident: streamed assistant text was losing the spaces BETWEEN words
(e.g. "Bilden är en riktig JPEG nu" -> "Bildenärenriktig JPEG nu") on the
Responses-API and Claude streaming paths.

stripInternalReasoningPlaceholder() (#8081/#8162) is called on every
individual delta.content chunk, and unconditionally called .trim() even when
its sentinel ("(prior reasoning summary unavailable)") was never present in
that chunk. Tokenizers commonly emit sub-word tokens with a leading space as
part of the token (e.g. " en", " riktig") -- each such chunk got its only
whitespace character (the inter-word space) silently trimmed away before
being appended to the accumulated message, while the words themselves stayed
intact. Punctuation-only chunks were largely unaffected, matching what was
observed live.

Only trims when the sentinel is actually present -- preserves the original
#8081 intent (collapse a placeholder-only chunk to "") without touching the
overwhelming majority of chunks that never contain it.

Co-authored-by: Markus Hartung <markus.hartream@gmail.com>
HouMinXi pushed a commit to HouMinXi/OmniRoute that referenced this pull request Aug 2, 2026
…iegosouzapw#8081) (diegosouzapw#8162)

* fix: strip internal reasoning placeholder from user-visible content (diegosouzapw#8081)

The internal reasoning replay sentinel '(prior reasoning summary unavailable)'
can leak into user-visible assistant content when a model echoes it through
ordinary message.content / delta.content. Existing suppression only checked
reasoning_content fields and reasoning-specific events.

Changes:
- Add stripInternalReasoningPlaceholder() to reasoningPlaceholder.ts —
  removes all occurrences of the sentinel and trims; returns '' when
  nothing meaningful remains
- Streaming: strip in responsesTransformer.ts, openai-responses.ts, and
  openai-to-claude.ts at the delta.content entry point; skip emission
  entirely when only the placeholder was present
- Non-streaming: strip in responseSanitizer.ts sanitizeMessageContent()
  and sanitizeResponsesMessageContent() (all three text paths)

translateText is unaffected (uses mode='translate' via plain newsClient).
The per-provider reasoning_content check remains as defense-in-depth.

* fix: skip only empty content block on reasoning-placeholder, keep finish_reason/tool_calls (diegosouzapw#8081)

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

* chore(quality): rebaseline openai-responses.ts own-growth (diegosouzapw#8081 guard)

---------

Co-authored-by: Austin Liu <austinliu@Austins-MacBook-Air-3.local>
Co-authored-by: Probe Test <probe@example.com>
Co-authored-by: Dingding-leo <Dingding-leo@users.noreply.github.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
HouMinXi pushed a commit to HouMinXi/OmniRoute that referenced this pull request Aug 2, 2026
… sentinel (diegosouzapw#8382)

Regression: diegosouzapw#8162 (port of diegosouzapw#8081) added an unconditional `.trim()` to
stripInternalReasoningPlaceholder(), applied to every streaming
delta.content chunk across 3 call-sites (openai-to-claude.ts,
openai-responses.ts, responsesTransformer.ts). Leading/trailing
whitespace at a chunk boundary is a real word boundary between
streaming fragments; trimming it glues adjacent chunks together on
the client ("Hello, " + "world." + " Bye." -> "Hello,world.Bye.").

Fix: early-return via .includes() before the replaceAll+trim, so the
function is a true no-op when the sentinel is absent from the chunk.
Behavior when the sentinel IS present is unchanged.

Validation:
- tests/unit/streaming-reasoning-dedup-5786.test.ts: the "(A-guard)"
  test was RED on the base branch ('Hello,world.Bye.' vs
  'Hello, world. Bye.'); GREEN after the fix (4/4 passing).
- tests/unit/translator-resp-openai-to-claude.test.ts: added a new
  multi-chunk boundary-whitespace regression test, proven RED against
  the pre-fix code (12/13), GREEN after (13/13).
- No regressions in responses-transformer.test.ts (17/17),
  responses-transformer-dense-output.test.ts (3/3), or the other
  suites exercising the shared placeholder utility (160/160 total
  across all consumers).

Refs diegosouzapw#8162
Refs diegosouzapw#8081
HouMinXi pushed a commit to HouMinXi/OmniRoute that referenced this pull request Aug 2, 2026
…rd spaces (diegosouzapw#8341)

Live incident: streamed assistant text was losing the spaces BETWEEN words
(e.g. "Bilden är en riktig JPEG nu" -> "Bildenärenriktig JPEG nu") on the
Responses-API and Claude streaming paths.

stripInternalReasoningPlaceholder() (diegosouzapw#8081/diegosouzapw#8162) is called on every
individual delta.content chunk, and unconditionally called .trim() even when
its sentinel ("(prior reasoning summary unavailable)") was never present in
that chunk. Tokenizers commonly emit sub-word tokens with a leading space as
part of the token (e.g. " en", " riktig") -- each such chunk got its only
whitespace character (the inter-word space) silently trimmed away before
being appended to the accumulated message, while the words themselves stayed
intact. Punctuation-only chunks were largely unaffected, matching what was
observed live.

Only trims when the sentinel is actually present -- preserves the original
diegosouzapw#8081 intent (collapse a placeholder-only chunk to "") without touching the
overwhelming majority of chunks that never contain it.

Co-authored-by: Markus Hartung <markus.hartream@gmail.com>
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…iegosouzapw#8081) (diegosouzapw#8162)

* fix: strip internal reasoning placeholder from user-visible content (diegosouzapw#8081)

The internal reasoning replay sentinel '(prior reasoning summary unavailable)'
can leak into user-visible assistant content when a model echoes it through
ordinary message.content / delta.content. Existing suppression only checked
reasoning_content fields and reasoning-specific events.

Changes:
- Add stripInternalReasoningPlaceholder() to reasoningPlaceholder.ts —
  removes all occurrences of the sentinel and trims; returns '' when
  nothing meaningful remains
- Streaming: strip in responsesTransformer.ts, openai-responses.ts, and
  openai-to-claude.ts at the delta.content entry point; skip emission
  entirely when only the placeholder was present
- Non-streaming: strip in responseSanitizer.ts sanitizeMessageContent()
  and sanitizeResponsesMessageContent() (all three text paths)

translateText is unaffected (uses mode='translate' via plain newsClient).
The per-provider reasoning_content check remains as defense-in-depth.

* fix: skip only empty content block on reasoning-placeholder, keep finish_reason/tool_calls (diegosouzapw#8081)

Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>

* chore(quality): rebaseline openai-responses.ts own-growth (diegosouzapw#8081 guard)

---------

Co-authored-by: Austin Liu <austinliu@Austins-MacBook-Air-3.local>
Co-authored-by: Probe Test <probe@example.com>
Co-authored-by: Dingding-leo <Dingding-leo@users.noreply.github.com>
Co-authored-by: diegosouzapw <8016841+diegosouzapw@users.noreply.github.com>
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
… sentinel (diegosouzapw#8382)

Regression: diegosouzapw#8162 (port of diegosouzapw#8081) added an unconditional `.trim()` to
stripInternalReasoningPlaceholder(), applied to every streaming
delta.content chunk across 3 call-sites (openai-to-claude.ts,
openai-responses.ts, responsesTransformer.ts). Leading/trailing
whitespace at a chunk boundary is a real word boundary between
streaming fragments; trimming it glues adjacent chunks together on
the client ("Hello, " + "world." + " Bye." -> "Hello,world.Bye.").

Fix: early-return via .includes() before the replaceAll+trim, so the
function is a true no-op when the sentinel is absent from the chunk.
Behavior when the sentinel IS present is unchanged.

Validation:
- tests/unit/streaming-reasoning-dedup-5786.test.ts: the "(A-guard)"
  test was RED on the base branch ('Hello,world.Bye.' vs
  'Hello, world. Bye.'); GREEN after the fix (4/4 passing).
- tests/unit/translator-resp-openai-to-claude.test.ts: added a new
  multi-chunk boundary-whitespace regression test, proven RED against
  the pre-fix code (12/13), GREEN after (13/13).
- No regressions in responses-transformer.test.ts (17/17),
  responses-transformer-dense-output.test.ts (3/3), or the other
  suites exercising the shared placeholder utility (160/160 total
  across all consumers).

Refs diegosouzapw#8162
Refs diegosouzapw#8081
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…rd spaces (diegosouzapw#8341)

Live incident: streamed assistant text was losing the spaces BETWEEN words
(e.g. "Bilden är en riktig JPEG nu" -> "Bildenärenriktig JPEG nu") on the
Responses-API and Claude streaming paths.

stripInternalReasoningPlaceholder() (diegosouzapw#8081/diegosouzapw#8162) is called on every
individual delta.content chunk, and unconditionally called .trim() even when
its sentinel ("(prior reasoning summary unavailable)") was never present in
that chunk. Tokenizers commonly emit sub-word tokens with a leading space as
part of the token (e.g. " en", " riktig") -- each such chunk got its only
whitespace character (the inter-word space) silently trimmed away before
being appended to the accumulated message, while the words themselves stayed
intact. Punctuation-only chunks were largely unaffected, matching what was
observed live.

Only trims when the sentinel is actually present -- preserves the original
diegosouzapw#8081 intent (collapse a placeholder-only chunk to "") without touching the
overwhelming majority of chunks that never contain it.

Co-authored-by: Markus Hartung <markus.hartream@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(backend): internal reasoning-replay placeholder leaks into assistant message content

3 participants