Conversation
`max_tokens` on an OpenAI-compatible /v1/models record is the model's maximum output length, the same meaning as the request parameter -- not the context window. Since diegosouzapw#14159 it sat in the contextWindow candidate list, so a record carrying only `max_tokens` synced with `inputTokenLimit` set to that output cap: the catalog advertised a 4K window for a 128K model, the combo context-window filter excluded it for any prompt above the cap, and the compression / context-relay heuristics sized the prompt against an output limit. Drop it from the candidates and leave a comment saying why, since the sibling `max_context_window` is a real window field. A provider that genuinely reports its window that way needs a per-provider entry with a test naming it. Regression guard: a record with only `max_tokens` yields no `inputTokenLimit`; a record with `context_length` + `max_tokens` keeps `context_length`. Closes diegosouzapw#14318
…ery-max-tokens-not-context-window
|
Closing this as superseded — the base already contains the fix, and I would rather you not spend a review on two guards for one hole. I merged
const contextWindow = firstPositiveNumber(
record.context_length,
record.contextLength,
record.contextWindow,
record.max_model_len,
record.maxModelLen,
record.max_context_window,
record.max_input_tokens,
record.maxInputTokens,
topProvider.context_length
);and the base comment where mine was now reads "Anthropic Models API reports the window as
test("a record with only max_tokens does not treat the output cap as the window", () => {
const [model] = normalizeDiscoveredModels(
[{ id: "claude-opus-5", max_tokens: 128000 }],
"claude"
);
assert.equal(model.inputTokenLimit, undefined);The one shape only this PR covered is Say the word if you disagree and I will reopen it as a test-only PR. |
Summary
normalizeDiscoveredModels()no longer readsmax_tokensas a context window(
src/lib/providerModels/modelDiscovery.ts), and a regression test pins bothdirections of the candidate list. Closes #14318.
Motivation
max_tokenson an OpenAI-compatible/v1/modelsrecord is the model's maximumoutput length - the same meaning as the
max_tokensrequest parameter - notthe window. It was added to the
contextWindowcandidate list by #14159 (there-land of #12630) without a test pinning it, so a record that carries only
max_tokensnow syncs as:{ "id": "local-llm", "max_tokens": 4096 } -> inputTokenLimit: 4096Consequences for a local/LM-Studio-style endpoint that reports it that way:
combo-context-window-filterexcludes that model for any promptabove the output cap, so it silently stops being routable;
inputTokenLimitsizethe prompt against an output limit.
The real window fields already in the list (
context_length,contextWindow,max_model_lenper #12858,max_context_window,top_provider.context_length)cover every provider that legitimately reports a window; removing
max_tokensdoes not leave any of them without one.
What changed
src/lib/providerModels/modelDiscovery.ts: droppedrecord.max_tokensfromthe
contextWindowcandidates and left a comment saying why it must not bere-added, because the adjacent
max_context_windowIS a window field and thepair is easy to confuse. If a provider turns out to report its window as
max_tokens, it belongs in per-provider with a test naming that provider.tests/unit/discovery-max-tokens-context-window-14318.test.ts(new): theregression guard the issue asked for.
What is broken without it, and how it works after
Before:
inputTokenLimit= 4096 for amax_tokens-only record (assertion outputbelow). After: no
inputTokenLimitis invented from an output cap, and a recordthat reports both
context_length: 131072andmax_tokens: 4096still keeps131072 - the ordering already handled that case, and the test locks it so a
future candidate reshuffle cannot break it either way.
Commands run
Every figure above was re-verified on the branch as it now stands: the discovery family is still 150 tests / 148 pass / 2 skipped at the merged head
5a19c3b0a, so the base refresh changed no outcome. Base branch:release/v3.8.51(the highest activerelease/v*, and the repo default), per Golden Path step 1. Cut at7a921299c; refreshed onto34113170fby merging the base in (merge commit5a19c3b0a) after the branch fell 2 commits behind, with the two SSE commits that arrived touching nothing on this path. Re-verified on the merged tree:discovery-max-tokens-context-window-143182/2, andprovider-models-route+model-token-limit-catalog+openrouter-context-length-320269/69.Migrations, feature flags, generated artifacts
A changelog fragment was added to this branch after I re-read CONTRIBUTING.md line 409: user-facing changes ship a file under the changelog directory rather than editing CHANGELOG.md, so this PR now carries one. It is the only addition beyond the fix and its test.
None. No schema, catalog file, generated reference, or feature flag is touched;
docs/reference/PROVIDER_REFERENCE.mdis generated from the provider registry,not from discovery output, so no regenerated diff is expected.
CI-only validation still pending
npm run typecheck:core,npm run lint(whole repo),npm run test:unit,npm run test:vitestand the coverage ratchet: run by CI on this PR. Locally Iran the focused equivalents listed above; the change only removes one argument
from a variadic
firstPositiveNumber()call and adds a test file, so notype surface changes.
npm run test:unitwas not run locally (the suite is ~10k tests); the150-test discovery subset above is the focused gate for this module.
base-red note
open against this base. Any failure in CI that also reproduces on the base tip
7a921299cis inherited, not introduced here - my focused suite is green on topof that tip.