Skip to content

fix(responses): surface truncated generations as response.incomplete, not completed - #14806

Merged
diegosouzapw merged 6 commits into
diegosouzapw:release/v3.8.51from
hartmark:fix/openai-responses-truncated-status-incomplete
Sep 29, 2026
Merged

diegosouzapw merged 6 commits into
diegosouzapw:release/v3.8.51from
hartmark:fix/openai-responses-truncated-status-incomplete

Conversation

@hartmark

Copy link
Copy Markdown
Contributor

What Problem This Solves

A provider that stopped a generation because it hit the token limit (Chat Completions finish_reason: "length", or a content filter) got translated to the OpenAI Responses API's response.completed event with status: "completed" — indistinguishable from a genuine successful completion. A caller has no signal the output was cut off mid-generation instead of the model actually finishing.

Observed live in production: a model looping on repeated garbled/incoherent text for its entire completion-token budget (completion_tokens: 8192, exactly the cap) got reported as status: "completed" even though the real upstream response carried finish_reason: "length" the whole time.

Why This Change Was Made

The real OpenAI Responses API contract distinguishes this with a dedicated response.incomplete event carrying incomplete_details: { reason: "max_output_tokens" | "content_filter" } — already correctly implemented for the ChatGPT-web bridge (vendor/codex-chatgpt-web/bridge.ts), but the generic Chat-Completions-to-Responses stream translator (translator/response/openai-responses.ts, used by any provider routed through the standard chat-completions path, e.g. Mistral) discarded finish_reason entirely and always emitted status: "completed" regardless of its value.

Fix: remember finish_reason on stream state when it arrives, and in sendCompleted(), map "length"/"content_filter" to status: "incomplete" with the matching incomplete_details, emitting response.incomplete instead of response.completed for those cases. An upstream mid-stream error still takes priority (status: "failed"), and every other finish_reason ("stop", "tool_calls", etc.) keeps the existing completed behavior unchanged — verified by a new sibling test.

Evidence

  • Reproduced from a real production log: raw upstream SSE stream chunk shows "finish_reason":"length", but OmniRoute's own aggregated response and the client-facing response.completed event both showed status: "completed".
  • New regression test (tests/unit/translator-resp-openai-responses-roundtrip.test.ts): a finish_reason: "length" chunk now produces response.incomplete with incomplete_details: { reason: "max_output_tokens" }, and never response.completed. TDD-verified: fails on the original code (asserts response.completed is undefined, but the original code emits it), passes with the fix.
  • Sibling regression test confirms an ordinary finish_reason: "stop" is unaffected (response.completed, status: "completed", no incomplete_details).
  • Full tests/unit/translator* + tests/unit/openai-responses* suite (185 tests) green.
  • Pre-commit hooks (prettier, eslint, docs-sync, t11 any-budget, tracked-artifacts, ai-attribution) all pass clean.

Production LOC

Net +25 in translator/response/openai-responses.ts (capture finish_reason on state; branch sendCompleted()'s status/event-type/incomplete_details on it). Test file +73/-3.

… not completed

A provider that stops on the token limit (Chat Completions finish_reason:
length, or a content filter) was translated to the OpenAI Responses API's
response.completed event with status:completed -- indistinguishable from
a genuine successful completion. A caller has no signal the output was cut
off mid-generation, and in production this let a truncated, degenerating
generation (observed: a model looping on repeated garbled text for its
entire 8192-token budget) get treated as a normal, trustworthy result.

The real Responses API contract distinguishes this with a dedicated
response.incomplete event carrying incomplete_details: { reason:
max_output_tokens | content_filter } -- already correctly implemented
for the ChatGPT-web bridge (vendor/codex-chatgpt-web/bridge.ts), but the
Chat-Completions-to-Responses stream translator (translator/response/
openai-responses.ts) discarded finish_reason entirely and always emitted
status:completed regardless of its value.

Fix: remember finish_reason on stream state when it arrives, and in
sendCompleted(), map length/content_filter to status:incomplete with
the matching incomplete_details, emitting response.incomplete instead of
response.completed for those cases. An upstream mid-stream error still
takes priority (status:failed), and every other finish_reason keeps the
existing completed behavior unchanged.

New regression tests cover both the previously-broken truncated case and
confirm an ordinary finish_reason:stop is unaffected.
@diegosouzapw

Copy link
Copy Markdown
Owner

Thanks for the clear write-up and the repro. Mapping length/content_filter to response.incomplete with incomplete_details matches the Responses API contract, and the new test fails without the change. Before merging, could you (1) add a changelog fragment under changelog.d/fixes/, and (2) look at open-sse/transformer/responsesTransformer.ts (~l.608), which is the other Chat→Responses path and still emits response.completed for a length finish? Tests for content_filter and for an upstream error combined with length would be nice too.

…path; add content_filter + error-priority tests

Per review: this repo has two parallel Chat-Completions-to-Responses stream
translators (translator/response/openai-responses.ts, already fixed, and
transformer/responsesTransformer.ts) -- the second one had the identical
finish_reason-discarding bug and still unconditionally emitted
response.completed for a length/content_filter stop. Applied the same fix:
remember finish_reason on stream state, map length/content_filter to
status:incomplete with incomplete_details, emit response.incomplete instead
of response.completed for those cases.

Also added the requested test coverage:
- content_filter -> response.incomplete with incomplete_details.reason:
  content_filter, in both translators.
- An upstream error arriving after a deferred finish_reason:length still
  wins (status:failed, never incomplete) -- exercises the real
  awaitingTrailingUsage deferred-completion path, not just the immediate
  case, and proves the priority ordering the fix relies on.

Added a changelog fragment (changelog.d/fixes/) per the repo's convention.
@hartmark

Copy link
Copy Markdown
Contributor Author

Thanks for the review — addressed all four points:

  1. Changelog fragment added (changelog.d/fixes/responses-length-status-incomplete.md).
  2. open-sse/transformer/responsesTransformer.ts — same bug, same fix: captured finish_reason on stream state and branched sendCompleted()'s status/event-type/incomplete_details on it, mirroring translator/response/openai-responses.ts exactly (including comments cross-referencing each other now).
  3. content_filter tests added in both files (tests/unit/translator-resp-openai-responses-roundtrip.test.ts and tests/unit/responses-transformer.test.ts).
  4. Upstream-error + length test added — exercises the real deferred-completion path (finish_reason:"length" with no usage yet → awaitingTrailingUsage, then a mid-stream error chunk arrives instead of the expected trailing usage chunk) and confirms status:"failed" wins, never incomplete.

Full translator*/openai-responses*/responses-transformer* suite (225 tests) green. All new tests TDD-verified against the pre-fix code where practical.

hartmark added a commit to hartmark/OmniRoute that referenced this pull request Sep 27, 2026
…tions as response.incomplete, not completed) into dev/omniroute-dev-combined
…th Responses emitters

Declares finishReason on the transformer state (fixes the open-sse TS2339) and moves the
length/content_filter -> status:incomplete mapping into one helper so
openai-responses.ts stays within its frozen file-size ceiling. No behavior change.
@diegosouzapw
diegosouzapw merged commit 9fd4ff6 into diegosouzapw:release/v3.8.51 Sep 29, 2026
7 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants