Skip to content

fix(sse): clamp max_tokens to the model output cap on every path - #8698

Merged
diegosouzapw merged 4 commits into
release/v3.8.49from
feat/core-max-tokens-clamp
Jul 26, 2026
Merged

diegosouzapw merged 4 commits into
release/v3.8.49from
feat/core-max-tokens-clamp

Conversation

@diegosouzapw

Copy link
Copy Markdown
Owner

Problem

A client-supplied max_tokens above the target model's own output ceiling reached the
upstream unchanged on the single-model path, where it can come back as
"response exceeded output token maximum".

Two clamps existed, neither covering the case:

  • enforceOutputTokenBudget() (open-sse/handlers/chatCore/outputTokenBudget.ts) capped the
    three output-token fields, but only against the remaining context window
    (contextLimit − estimatedInputTokens) — not against the model's output cap.
  • resolveReasoningBufferedMaxTokens() (open-sse/services/reasoningTokenBuffer.ts:42) does
    clamp to getExplicitModelOutputCap(), but only for supportsThinking models and only
    inside combo routing.

So a plain (non-combo) request to a non-thinking model was never bounded by the model's own
ceiling. This becomes reachable as soon as an operator raises the global output budget above
some routed model's cap — e.g. Claude Code with CLAUDE_CODE_MAX_OUTPUT_TOKENS=128000 in
front of a 64K-output model.

Change

enforceOutputTokenBudget() takes an optional fifth argument, the model's output cap, and
uses it as an additional upper bound when adjusting the fields only:

effectiveCap = cap != null && cap > 0 ? Math.min(availableOutputTokens, cap) : availableOutputTokens

The accept/reject decision (ok:false) and the returned availableOutputTokens stay tied to
the context window alone. That separation is deliberate: a model whose output ceiling is
smaller than defaultOutputTokens must not start returning 400. The cap limits how much is
requested, never whether the request fits. A dedicated test guards this.

The callsite (open-sse/handlers/chatCore.ts) resolves the cap with
getExplicitModelOutputCap({ provider, model: effectiveModel }) — the object form, matching
the sibling capability lookups in the same file. The bare-string form resolves to
provider: null, which skips both the registry cap and the operator's max_token capability
override (#6524); clamping against a stale static spec while the operator had raised the
ceiling would silently truncate output.

Properties: only ever reduces, never raises. Absent / null / non-positive cap (unknown model)
leaves behavior byte-identical — fail-open. Idempotent with the reasoning buffer, which already
lands at or below the same ceiling.

Tests

  • tests/unit/output-token-budget-model-cap.test.ts (8) — clamp above the cap; no elevation
    below it; cap absent ⇒ identical result; all three field names; window tighter than the cap
    wins; cap smaller than defaultOutputTokens still accepted (the 400 regression guard);
    adjustedFields exactness; sub-token cap treated as absent.
  • tests/unit/chatcore-model-output-cap-wiring.test.ts (2) — drives handleChatCore() end to
    end against a stubbed fetch and asserts the body actually dispatched upstream. The cap
    arrives via an operator max_token override, so the test pins the { provider, model }
    keying too. Verified to fail when the cap argument is dropped and when the lookup is
    switched to the bare-string form.

Existing suites pass unedited: output-token-budget.test.ts,
chatcore-combo-context-limit-8378.test.ts, chatcore-context-window-boundary.test.ts,
reasoning-token-buffer-6274.test.ts. typecheck:core and lint clean.

Known remaining gap (not in this PR)

When the client sends no max_tokens and the target is Claude format,
open-sse/translator/helpers/maxTokensHelper.ts injects DEFAULT_MAX_TOKENS (64000) after
this clamp, unbounded by the model cap. adjustMaxTokens() is model-agnostic and called from
several translators, so bounding it means threading the cap through them — a wider change than
this fix. Filed as follow-up rather than smuggled in here.

enforceOutputTokenBudget only capped the three output-token fields against the
remaining context window, so a request whose max_tokens exceeded the model's own
output ceiling reached the upstream unchanged on the single-model path (the
reasoning-token buffer covers only thinking models inside combo routing).

Pass the model's explicit output cap into the budget check and use it as an extra
upper bound when adjusting the fields. The reject decision stays tied to the
context window: an output cap smaller than the default output budget must not
turn a valid request into a 400.
The bare-string form of getExplicitModelOutputCap resolves to `provider: null`,
which skips the registry cap and the operator's `max_token` capability override
(#6524) — the documented escape hatch for a wrong synced `limit_output`. Clamping
against a stale static spec while the operator had raised the ceiling would
silently truncate output.

Matches the { provider, model } form already used by the sibling capability
lookups in this file (getResolvedModelCapabilities, supportsMaxTokens).
The unit tests drive enforceOutputTokenBudget() directly, so dropping the cap
argument at the handleChatCore callsite left every one of them green. Add a
wiring test that runs handleChatCore end to end against a stubbed fetch and
asserts the body actually dispatched upstream.

The cap comes from an operator `max_token` capability override rather than a
catalog model: the override table is keyed by provider, so the test also pins
the { provider, model } lookup — both the missing argument and the bare-string
form fail it (verified by mutating each in turn).

Also floor `maxOutputTokenCap` before the positivity test. A fractional cap
below 1 previously passed `> 0` and floored to an effective cap of 0, clamping
every field to zero; sub-token caps are meaningless and now read as absent.
Unreachable through the callsite (toPositiveInteger filters it) but the exported
contract was wrong.

The adjustment log now states the output ceiling in effect instead of claiming
the cap caused the adjustment — a field can also be adjusted by removal of an
invalid value, which the cap did not cause.
@diegosouzapw
diegosouzapw merged commit 6389c5b into release/v3.8.49 Jul 26, 2026
20 checks passed
@diegosouzapw
diegosouzapw deleted the feat/core-max-tokens-clamp branch July 26, 2026 19:30
HouMinXi pushed a commit to HouMinXi/OmniRoute that referenced this pull request Aug 2, 2026
…gosouzapw#8698)

* fix(sse): clamp max_tokens to the model output cap on every path

enforceOutputTokenBudget only capped the three output-token fields against the
remaining context window, so a request whose max_tokens exceeded the model's own
output ceiling reached the upstream unchanged on the single-model path (the
reasoning-token buffer covers only thinking models inside combo routing).

Pass the model's explicit output cap into the budget check and use it as an extra
upper bound when adjusting the fields. The reject decision stays tied to the
context window: an output cap smaller than the default output budget must not
turn a valid request into a 400.

* fix(sse): key the output-cap lookup by provider + model

The bare-string form of getExplicitModelOutputCap resolves to `provider: null`,
which skips the registry cap and the operator's `max_token` capability override
(diegosouzapw#6524) — the documented escape hatch for a wrong synced `limit_output`. Clamping
against a stale static spec while the operator had raised the ceiling would
silently truncate output.

Matches the { provider, model } form already used by the sibling capability
lookups in this file (getResolvedModelCapabilities, supportsMaxTokens).

* test(sse): cover the output-cap callsite; harden the sub-token cap guard

The unit tests drive enforceOutputTokenBudget() directly, so dropping the cap
argument at the handleChatCore callsite left every one of them green. Add a
wiring test that runs handleChatCore end to end against a stubbed fetch and
asserts the body actually dispatched upstream.

The cap comes from an operator `max_token` capability override rather than a
catalog model: the override table is keyed by provider, so the test also pins
the { provider, model } lookup — both the missing argument and the bare-string
form fail it (verified by mutating each in turn).

Also floor `maxOutputTokenCap` before the positivity test. A fractional cap
below 1 previously passed `> 0` and floored to an effective cap of 0, clamping
every field to zero; sub-token caps are meaningless and now read as absent.
Unreachable through the callsite (toPositiveInteger filters it) but the exported
contract was wrong.

The adjustment log now states the output ceiling in effect instead of claiming
the cap caused the adjustment — a field can also be adjusted by removal of an
invalid value, which the cap did not cause.
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…gosouzapw#8698)

* fix(sse): clamp max_tokens to the model output cap on every path

enforceOutputTokenBudget only capped the three output-token fields against the
remaining context window, so a request whose max_tokens exceeded the model's own
output ceiling reached the upstream unchanged on the single-model path (the
reasoning-token buffer covers only thinking models inside combo routing).

Pass the model's explicit output cap into the budget check and use it as an extra
upper bound when adjusting the fields. The reject decision stays tied to the
context window: an output cap smaller than the default output budget must not
turn a valid request into a 400.

* fix(sse): key the output-cap lookup by provider + model

The bare-string form of getExplicitModelOutputCap resolves to `provider: null`,
which skips the registry cap and the operator's `max_token` capability override
(diegosouzapw#6524) — the documented escape hatch for a wrong synced `limit_output`. Clamping
against a stale static spec while the operator had raised the ceiling would
silently truncate output.

Matches the { provider, model } form already used by the sibling capability
lookups in this file (getResolvedModelCapabilities, supportsMaxTokens).

* test(sse): cover the output-cap callsite; harden the sub-token cap guard

The unit tests drive enforceOutputTokenBudget() directly, so dropping the cap
argument at the handleChatCore callsite left every one of them green. Add a
wiring test that runs handleChatCore end to end against a stubbed fetch and
asserts the body actually dispatched upstream.

The cap comes from an operator `max_token` capability override rather than a
catalog model: the override table is keyed by provider, so the test also pins
the { provider, model } lookup — both the missing argument and the bare-string
form fail it (verified by mutating each in turn).

Also floor `maxOutputTokenCap` before the positivity test. A fractional cap
below 1 previously passed `> 0` and floored to an effective cap of 0, clamping
every field to zero; sub-token caps are meaningless and now read as absent.
Unreachable through the callsite (toPositiveInteger filters it) but the exported
contract was wrong.

The adjustment log now states the output ceiling in effect instead of claiming
the cap caused the adjustment — a field can also be adjusted by removal of an
invalid value, which the cap did not cause.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant