fix(sse): trust finish_reason over reasoning-ratio heuristic in response quality validation - #12262
Merged
diegosouzapw merged 2 commits intoSep 1, 2026
Conversation
…tio heuristic in response quality validation A truncated response with empty content and reasoning_content present was only rejected by validateResponseQuality() when reasoning consumed >=90% of completion_tokens. A response truncated at a lower ratio (e.g. 63%) passed through as "valid" even though the caller received no usable content and finish_reason was explicitly "length" (or the alternate "max_tokens" naming some providers use) -- an unambiguous truncation signal the validator wasn't reading. Reproduced live against nvidia/nemotron-3-super-120b-a12b: content:null, finish_reason:length, reasoning_tokens 645/1024 (63%). Trust finish_reason directly when it's reported, falling back to the existing token-ratio heuristic only when it isn't. Does not affect the deliberate-tiny-probe case (e.g. max_tokens:1 connectivity pings) -- those never produce reasoning_content, so the branch this change is in doesn't run for them.
Owner
|
Validated in local merge-train on 192.168.0.113 — train of #12258 #12262 #12166 #12281 #11259 #11950 merged clean onto
|
diegosouzapw
merged commit Sep 1, 2026
0ff1647
into
diegosouzapw:release/v3.8.51
13 of 16 checks passed
This was referenced Sep 1, 2026
Merged
muhamadgalihsaputra
pushed a commit
to niyatna/NiyatnaRoute
that referenced
this pull request
Sep 27, 2026
…nse quality validation (diegosouzapw#12262) * fix(sse): trust finish_reason:length/max_tokens over the reasoning-ratio heuristic in response quality validation A truncated response with empty content and reasoning_content present was only rejected by validateResponseQuality() when reasoning consumed >=90% of completion_tokens. A response truncated at a lower ratio (e.g. 63%) passed through as "valid" even though the caller received no usable content and finish_reason was explicitly "length" (or the alternate "max_tokens" naming some providers use) -- an unambiguous truncation signal the validator wasn't reading. Reproduced live against nvidia/nemotron-3-super-120b-a12b: content:null, finish_reason:length, reasoning_tokens 645/1024 (63%). Trust finish_reason directly when it's reported, falling back to the existing token-ratio heuristic only when it isn't. Does not affect the deliberate-tiny-probe case (e.g. max_tokens:1 connectivity pings) -- those never produce reasoning_content, so the branch this change is in doesn't run for them. * docs(changelog): add fragment for diegosouzapw#12262 --------- Co-authored-by: brick30llc-ctrl <admin@brick30.com>
2 tasks done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
validateResponseQuality()'s reasoning-truncation guard (added for #3587) only rejects aresponse with empty
content+ presentreasoning_contentwhen reasoning consumed 90%+of
completion_tokens. A response truncated at a lower ratio slips through as "valid" eventhough the caller gets nothing usable — and the response already carries an unambiguous
signal that was going unread:
finish_reason: "length"(or the"max_tokens"naming someproviders use).
Reproduced live against
nvidia/nemotron-3-super-120b-a12bon a self-hosted deployment:content: null,finish_reason: "length",reasoning_tokens: 645/completion_tokens: 1024(63% — below the 90% threshold). The combo loop treated it as a successful completion and
never retried, so the end client received an empty answer for what looked like a normal,
non-degraded 200 response.
Fix
When
contentis empty andreasoning_contentis present (the existing precondition forthis whole branch), check
finish_reasonfirst:"length"or"max_tokens"is trusteddirectly as truncation, independent of the token ratio. The existing ratio heuristic remains
as a fallback for providers that don't report
finish_reasonreliably. Does not touch orweaken the deliberate-tiny-probe case (e.g.
max_tokens: 1connectivity pings, seeerrorClassifier.ts'sLEGIT_EMPTY_OPENAI_FINISH) — those never producereasoning_contentat all, so this branch's precondition already excludes them.
Test plan
tests/unit/combo-quality-validator-reasoning.test.ts: the exact63%-ratio repro (now invalid), the
max_tokensnaming variant, a regression guard thatno-
finish_reason+ <90% ratio stays valid (unchanged [BUG] Reasoning models in combos consume all tokens for reasoning_content, leaving content empty #3587 behavior), and a check thatfinish_reason: "stop"doesn't short-circuit the ratio heuristic.node --import tsx/esm --test tests/unit/combo-quality-validator-reasoning.test.ts—16/16 pass (12 pre-existing + 4 new).
validate-response-quality.test.ts(17),quality-validation-benign-error.test.ts,routing-quality.test.ts,routing-scoring-quality.test.ts,quality-rail-gate-membership.test.ts(23) — all pass.npm run typecheck:core— only pre-existing, unrelatedomniglypherrors (confirmedpresent on the base branch too, nothing from the changed files).
eslinton both changed files — clean.