fix(sse): accept a chat completion truncated at finish_reason length - #14899
diegosouzapw merged 2 commits into
Conversation
The Claude message shape already treats stop_reason max_tokens with no visible text as a valid truncated completion (diegosouzapw#12968): a thinking model can burn a one-token probe budget before emitting anything. Translating that body to chat completions maps the stop reason to finish_reason "length", and both the malformed-body check and the combo quality gate then rejected it as empty, so the same request 502'd on /v1/chat and /v1/responses while /v1/messages returned 200. A length finish with no text now passes; a plain stop with no text stays rejected. Signed-off-by: Minxi Hou <houminxi@gmail.com>
01cbd3b to
ff4ad1f
Compare
finishReason.ts normalizes max_tokens to length before diagnostics sees it. The length exemption must not widen to the raw spelling: a caller that bypasses normalization still gets empty_choices. Injection: widen the predicate to include max_tokens and this test goes red.
Review waiverTwo LOCAL review rounds on this branch. Product code: zero CONFIRMED findings across both rounds. Round 1 (commit ff4ad1f): verdict PENDING, 0 CONFIRMED, 3 UNCERTAIN. Two of the three were RECEIPT_UNTRUSTED (excerpt line-count off by one, an infrastructure artifact). The third suggested a test for raw Round 2 (commit 6483334): verdict FAIL, 1 CONFIRMED. The single CONFIRMED is RECEIPT_INVALID — the reviewer's generated excerpt declared a line count that did not match the excerpt body. This is a known infrastructure class (excerpt validation), not a product defect. The three UNCERTAIN findings were style suggestions about the new test (assertion method consistency, comment wording, edge-case coverage), none identifying a code defect. The RECEIPT_INVALID / RECEIPT_UNTRUSTED class recurs across rounds on different files. It is the review tool's excerpt formatter, not the code under review. Per the standing rule, infrastructure CONFIRMED findings do not reset product clean rounds and do not block submission when the product diff itself has zero confirmed defects. |
041e0ff
into
diegosouzapw:release/v3.8.51
…diegosouzapw#14894) Validated in the 2026-09-28 merge-batch board. Reconciled with the tip after diegosouzapw#14899 landed (both sides only appended tests to validate-response-quality.test.ts — kept both). Focused tests 118/118 on the reconciled head. Thank you @HouMinXi!
Summary
A thinking Claude model that spends its whole output budget before emitting text answers
/v1/messageswithstop_reason: "max_tokens"and an empty content array, which is a valid truncation and already accepted (#12968). The chat-completions translation of that same body isfinish_reason: "length"with no text, and both the malformed-body check and the combo quality gate rejected it as empty. The result is HTTP 502 on/v1/chat/completionsand/v1/responsesfor a request the native format returns 200 for.Both checks now accept a
lengthfinish with no visible text. A plainstopwith no text is still rejected.Closes #14898
Test plan
node --import tsx/esm --test tests/unit/diagnostics.test.ts tests/unit/validate-response-quality.test.ts— 62 pass, 0 failmax_tokens: 1against a thinking Claude model returns 200 on/v1/chat/completionsand/v1/responsesinstead of 502