Repository navigation
fix(sse): clamp max_tokens to the model output cap on every path - #8698
Merged
Merged
Conversation
enforceOutputTokenBudget only capped the three output-token fields against the remaining context window, so a request whose max_tokens exceeded the model's own output ceiling reached the upstream unchanged on the single-model path (the reasoning-token buffer covers only thinking models inside combo routing). Pass the model's explicit output cap into the budget check and use it as an extra upper bound when adjusting the fields. The reject decision stays tied to the context window: an output cap smaller than the default output budget must not turn a valid request into a 400.
The bare-string form of getExplicitModelOutputCap resolves to `provider: null`, which skips the registry cap and the operator's `max_token` capability override (#6524) — the documented escape hatch for a wrong synced `limit_output`. Clamping against a stale static spec while the operator had raised the ceiling would silently truncate output. Matches the { provider, model } form already used by the sibling capability lookups in this file (getResolvedModelCapabilities, supportsMaxTokens).
The unit tests drive enforceOutputTokenBudget() directly, so dropping the cap
argument at the handleChatCore callsite left every one of them green. Add a
wiring test that runs handleChatCore end to end against a stubbed fetch and
asserts the body actually dispatched upstream.
The cap comes from an operator `max_token` capability override rather than a
catalog model: the override table is keyed by provider, so the test also pins
the { provider, model } lookup — both the missing argument and the bare-string
form fail it (verified by mutating each in turn).
Also floor `maxOutputTokenCap` before the positivity test. A fractional cap
below 1 previously passed `> 0` and floored to an effective cap of 0, clamping
every field to zero; sub-token caps are meaningless and now read as absent.
Unreachable through the callsite (toPositiveInteger filters it) but the exported
contract was wrong.
The adjustment log now states the output ceiling in effect instead of claiming
the cap caused the adjustment — a field can also be adjusted by removal of an
invalid value, which the cap did not cause.
9 of 10 tasks
HouMinXi
pushed a commit
to HouMinXi/OmniRoute
that referenced
this pull request
Aug 2, 2026
…gosouzapw#8698) * fix(sse): clamp max_tokens to the model output cap on every path enforceOutputTokenBudget only capped the three output-token fields against the remaining context window, so a request whose max_tokens exceeded the model's own output ceiling reached the upstream unchanged on the single-model path (the reasoning-token buffer covers only thinking models inside combo routing). Pass the model's explicit output cap into the budget check and use it as an extra upper bound when adjusting the fields. The reject decision stays tied to the context window: an output cap smaller than the default output budget must not turn a valid request into a 400. * fix(sse): key the output-cap lookup by provider + model The bare-string form of getExplicitModelOutputCap resolves to `provider: null`, which skips the registry cap and the operator's `max_token` capability override (diegosouzapw#6524) — the documented escape hatch for a wrong synced `limit_output`. Clamping against a stale static spec while the operator had raised the ceiling would silently truncate output. Matches the { provider, model } form already used by the sibling capability lookups in this file (getResolvedModelCapabilities, supportsMaxTokens). * test(sse): cover the output-cap callsite; harden the sub-token cap guard The unit tests drive enforceOutputTokenBudget() directly, so dropping the cap argument at the handleChatCore callsite left every one of them green. Add a wiring test that runs handleChatCore end to end against a stubbed fetch and asserts the body actually dispatched upstream. The cap comes from an operator `max_token` capability override rather than a catalog model: the override table is keyed by provider, so the test also pins the { provider, model } lookup — both the missing argument and the bare-string form fail it (verified by mutating each in turn). Also floor `maxOutputTokenCap` before the positivity test. A fractional cap below 1 previously passed `> 0` and floored to an effective cap of 0, clamping every field to zero; sub-token caps are meaningless and now read as absent. Unreachable through the callsite (toPositiveInteger filters it) but the exported contract was wrong. The adjustment log now states the output ceiling in effect instead of claiming the cap caused the adjustment — a field can also be adjusted by removal of an invalid value, which the cap did not cause.
muhamadgalihsaputra
pushed a commit
to niyatna/NiyatnaRoute
that referenced
this pull request
Sep 27, 2026
…gosouzapw#8698) * fix(sse): clamp max_tokens to the model output cap on every path enforceOutputTokenBudget only capped the three output-token fields against the remaining context window, so a request whose max_tokens exceeded the model's own output ceiling reached the upstream unchanged on the single-model path (the reasoning-token buffer covers only thinking models inside combo routing). Pass the model's explicit output cap into the budget check and use it as an extra upper bound when adjusting the fields. The reject decision stays tied to the context window: an output cap smaller than the default output budget must not turn a valid request into a 400. * fix(sse): key the output-cap lookup by provider + model The bare-string form of getExplicitModelOutputCap resolves to `provider: null`, which skips the registry cap and the operator's `max_token` capability override (diegosouzapw#6524) — the documented escape hatch for a wrong synced `limit_output`. Clamping against a stale static spec while the operator had raised the ceiling would silently truncate output. Matches the { provider, model } form already used by the sibling capability lookups in this file (getResolvedModelCapabilities, supportsMaxTokens). * test(sse): cover the output-cap callsite; harden the sub-token cap guard The unit tests drive enforceOutputTokenBudget() directly, so dropping the cap argument at the handleChatCore callsite left every one of them green. Add a wiring test that runs handleChatCore end to end against a stubbed fetch and asserts the body actually dispatched upstream. The cap comes from an operator `max_token` capability override rather than a catalog model: the override table is keyed by provider, so the test also pins the { provider, model } lookup — both the missing argument and the bare-string form fail it (verified by mutating each in turn). Also floor `maxOutputTokenCap` before the positivity test. A fractional cap below 1 previously passed `> 0` and floored to an effective cap of 0, clamping every field to zero; sub-token caps are meaningless and now read as absent. Unreachable through the callsite (toPositiveInteger filters it) but the exported contract was wrong. The adjustment log now states the output ceiling in effect instead of claiming the cap caused the adjustment — a field can also be adjusted by removal of an invalid value, which the cap did not cause.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Problem
A client-supplied
max_tokensabove the target model's own output ceiling reached theupstream unchanged on the single-model path, where it can come back as
"response exceeded output token maximum".
Two clamps existed, neither covering the case:
enforceOutputTokenBudget()(open-sse/handlers/chatCore/outputTokenBudget.ts) capped thethree output-token fields, but only against the remaining context window
(
contextLimit − estimatedInputTokens) — not against the model's output cap.resolveReasoningBufferedMaxTokens()(open-sse/services/reasoningTokenBuffer.ts:42) doesclamp to
getExplicitModelOutputCap(), but only forsupportsThinkingmodels and onlyinside combo routing.
So a plain (non-combo) request to a non-thinking model was never bounded by the model's own
ceiling. This becomes reachable as soon as an operator raises the global output budget above
some routed model's cap — e.g. Claude Code with
CLAUDE_CODE_MAX_OUTPUT_TOKENS=128000infront of a 64K-output model.
Change
enforceOutputTokenBudget()takes an optional fifth argument, the model's output cap, anduses it as an additional upper bound when adjusting the fields only:
The accept/reject decision (
ok:false) and the returnedavailableOutputTokensstay tied tothe context window alone. That separation is deliberate: a model whose output ceiling is
smaller than
defaultOutputTokensmust not start returning400. The cap limits how much isrequested, never whether the request fits. A dedicated test guards this.
The callsite (
open-sse/handlers/chatCore.ts) resolves the cap withgetExplicitModelOutputCap({ provider, model: effectiveModel })— the object form, matchingthe sibling capability lookups in the same file. The bare-string form resolves to
provider: null, which skips both the registry cap and the operator'smax_tokencapabilityoverride (#6524); clamping against a stale static spec while the operator had raised the
ceiling would silently truncate output.
Properties: only ever reduces, never raises. Absent / null / non-positive cap (unknown model)
leaves behavior byte-identical — fail-open. Idempotent with the reasoning buffer, which already
lands at or below the same ceiling.
Tests
tests/unit/output-token-budget-model-cap.test.ts(8) — clamp above the cap; no elevationbelow it; cap absent ⇒ identical result; all three field names; window tighter than the cap
wins; cap smaller than
defaultOutputTokensstill accepted (the 400 regression guard);adjustedFieldsexactness; sub-token cap treated as absent.tests/unit/chatcore-model-output-cap-wiring.test.ts(2) — driveshandleChatCore()end toend against a stubbed
fetchand asserts the body actually dispatched upstream. The caparrives via an operator
max_tokenoverride, so the test pins the{ provider, model }keying too. Verified to fail when the cap argument is dropped and when the lookup is
switched to the bare-string form.
Existing suites pass unedited:
output-token-budget.test.ts,chatcore-combo-context-limit-8378.test.ts,chatcore-context-window-boundary.test.ts,reasoning-token-buffer-6274.test.ts.typecheck:coreandlintclean.Known remaining gap (not in this PR)
When the client sends no
max_tokensand the target is Claude format,open-sse/translator/helpers/maxTokensHelper.tsinjectsDEFAULT_MAX_TOKENS(64000) afterthis clamp, unbounded by the model cap.
adjustMaxTokens()is model-agnostic and called fromseveral translators, so bounding it means threading the cap through them — a wider change than
this fix. Filed as follow-up rather than smuggled in here.