Repository navigation
fix(responses): surface truncated generations as response.incomplete, not completed - #14806
Merged
diegosouzapw merged 6 commits intoSep 29, 2026
Conversation
… not completed
A provider that stops on the token limit (Chat Completions finish_reason:
length, or a content filter) was translated to the OpenAI Responses API's
response.completed event with status:completed -- indistinguishable from
a genuine successful completion. A caller has no signal the output was cut
off mid-generation, and in production this let a truncated, degenerating
generation (observed: a model looping on repeated garbled text for its
entire 8192-token budget) get treated as a normal, trustworthy result.
The real Responses API contract distinguishes this with a dedicated
response.incomplete event carrying incomplete_details: { reason:
max_output_tokens | content_filter } -- already correctly implemented
for the ChatGPT-web bridge (vendor/codex-chatgpt-web/bridge.ts), but the
Chat-Completions-to-Responses stream translator (translator/response/
openai-responses.ts) discarded finish_reason entirely and always emitted
status:completed regardless of its value.
Fix: remember finish_reason on stream state when it arrives, and in
sendCompleted(), map length/content_filter to status:incomplete with
the matching incomplete_details, emitting response.incomplete instead of
response.completed for those cases. An upstream mid-stream error still
takes priority (status:failed), and every other finish_reason keeps the
existing completed behavior unchanged.
New regression tests cover both the previously-broken truncated case and
confirm an ordinary finish_reason:stop is unaffected.
Owner
|
Thanks for the clear write-up and the repro. Mapping |
…path; add content_filter + error-priority tests Per review: this repo has two parallel Chat-Completions-to-Responses stream translators (translator/response/openai-responses.ts, already fixed, and transformer/responsesTransformer.ts) -- the second one had the identical finish_reason-discarding bug and still unconditionally emitted response.completed for a length/content_filter stop. Applied the same fix: remember finish_reason on stream state, map length/content_filter to status:incomplete with incomplete_details, emit response.incomplete instead of response.completed for those cases. Also added the requested test coverage: - content_filter -> response.incomplete with incomplete_details.reason: content_filter, in both translators. - An upstream error arriving after a deferred finish_reason:length still wins (status:failed, never incomplete) -- exercises the real awaitingTrailingUsage deferred-completion path, not just the immediate case, and proves the priority ordering the fix relies on. Added a changelog fragment (changelog.d/fixes/) per the repo's convention.
Contributor
Author
|
Thanks for the review — addressed all four points:
Full |
hartmark
added a commit
to hartmark/OmniRoute
that referenced
this pull request
Sep 27, 2026
…tions as response.incomplete, not completed) into dev/omniroute-dev-combined
…th Responses emitters Declares finishReason on the transformer state (fixes the open-sse TS2339) and moves the length/content_filter -> status:incomplete mapping into one helper so openai-responses.ts stays within its frozen file-size ceiling. No behavior change.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What Problem This Solves
A provider that stopped a generation because it hit the token limit (Chat Completions
finish_reason: "length", or a content filter) got translated to the OpenAI Responses API'sresponse.completedevent withstatus: "completed"— indistinguishable from a genuine successful completion. A caller has no signal the output was cut off mid-generation instead of the model actually finishing.Observed live in production: a model looping on repeated garbled/incoherent text for its entire completion-token budget (
completion_tokens: 8192, exactly the cap) got reported asstatus: "completed"even though the real upstream response carriedfinish_reason: "length"the whole time.Why This Change Was Made
The real OpenAI Responses API contract distinguishes this with a dedicated
response.incompleteevent carryingincomplete_details: { reason: "max_output_tokens" | "content_filter" }— already correctly implemented for the ChatGPT-web bridge (vendor/codex-chatgpt-web/bridge.ts), but the generic Chat-Completions-to-Responses stream translator (translator/response/openai-responses.ts, used by any provider routed through the standard chat-completions path, e.g. Mistral) discardedfinish_reasonentirely and always emittedstatus: "completed"regardless of its value.Fix: remember
finish_reasonon stream state when it arrives, and insendCompleted(), map"length"/"content_filter"tostatus: "incomplete"with the matchingincomplete_details, emittingresponse.incompleteinstead ofresponse.completedfor those cases. An upstream mid-stream error still takes priority (status: "failed"), and every otherfinish_reason("stop","tool_calls", etc.) keeps the existingcompletedbehavior unchanged — verified by a new sibling test.Evidence
"finish_reason":"length", but OmniRoute's own aggregated response and the client-facingresponse.completedevent both showedstatus: "completed".tests/unit/translator-resp-openai-responses-roundtrip.test.ts): afinish_reason: "length"chunk now producesresponse.incompletewithincomplete_details: { reason: "max_output_tokens" }, and neverresponse.completed. TDD-verified: fails on the original code (assertsresponse.completedisundefined, but the original code emits it), passes with the fix.finish_reason: "stop"is unaffected (response.completed,status: "completed", noincomplete_details).tests/unit/translator*+tests/unit/openai-responses*suite (185 tests) green.t11any-budget, tracked-artifacts, ai-attribution) all pass clean.Production LOC
Net +25 in
translator/response/openai-responses.ts(capturefinish_reasonon state; branchsendCompleted()'s status/event-type/incomplete_detailson it). Test file +73/-3.