fix(command-code): omit max_tokens when client omits it; correct registry caps - #5221
Conversation
There was a problem hiding this comment.
Code Review
This pull request updates the Command Code executor and registry configuration to handle max_tokens more robustly. Specifically, it reduces the registered maxOutputTokens for several models and refactors the clampMaxTokens logic to omit the max_tokens parameter entirely when not supplied by the client, allowing the upstream provider to apply its native defaults. When a client-supplied value is present, it is safely clamped to the 200,000 token endpoint ceiling to prevent 400 errors. Corresponding unit tests have been updated and added to verify these behaviors. There are no review comments provided, and I have no additional feedback to offer as the changes are clean, well-documented, and properly tested.
Important
The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.
dcdc392 to
b434997
Compare
…stry caps The executor always sent params.max_tokens to /alpha/generate, fabricating a value from the registry maxOutputTokens when the client sent none. For DeepSeek V4 (registry maxOutputTokens: 384000) this produced a request the endpoint rejects with 400 "Too big: expected number to be <=200000 at params.max_tokens", breaking DeepSeek V4 Pro/Flash on command-code entirely. Root cause: getModelMaxTokensCap (added in diegosouzapw#4518) fed the model's advertised output capacity (384000) straight into the request as the cap. That capacity is distinct from the endpoint's hard per-request ceiling of 200000, and is also simply wrong in the registry (Command Code's gateway caps DeepSeek output at 131072 per /provider/v1/models). Fix (executor): max_tokens is optional on /alpha/generate. Only forward it when the client actually supplies one, clamped to the 200000 endpoint ceiling so an oversized client value degrades gracefully instead of 400ing. When the client omits it, omit the field so upstream applies the model's native default. This mirrors the provider-driven clamp convention in antigravity.ts and removes the registry dependency from the request path (getModelMaxTokensCap deleted). Fix (registry): correct maxOutputTokens to the real Command Code gateway values from /provider/v1/models: DeepSeek V4 384000->131072, Kimi 131072->65536, GLM-5/5.1 131072->32768, MiniMax M2.5/M2.7 131072->65536, Qwen 3.6 131072->32768. These feed the combo router's output-limit check and dashboard metadata. The separate direct-DeepSeek spec in modelSpecs.ts (384000) is left untouched, since the model truly supports 384K output on its native API. Tests: omit-when-absent (GLM + DeepSeek, the reported scenario), clamp oversized client value to 200000, and honor a smaller client value unchanged. Co-authored-by: Cursor <cursoragent@cursor.com>
d76115d to
509806d
Compare
|
Heads-up on the one red check (
This PR only touches Happy to rebase onto a newer release branch if the heap-4076 fix lands there. Thanks! |
Cherry-picked the corrective part of #5221 only: the executor stops fabricating `max_tokens` (= per-model registry cap) when the client omits it, which caused `400 "expected <=200000"` on /alpha/generate for high-cap models. An explicit oversized client value is clamped to the 200k endpoint ceiling. The PR's registry maxOutputTokens recaps (open-sse/config/providers/registry/command-code/index.ts) are intentionally NOT included pending reconciliation; #5221 stays open for that.
|
Thanks @adivekar-utexas! The corrective part of this PR is now in I deliberately left out the registry I'm keeping this PR open so we can settle the cap values (and optionally restore a per-model clamp-down for explicit oversized values) without losing your work. Really appreciate the fix! |
…#5221) Cherry-picked the corrective part of diegosouzapw#5221 only: the executor stops fabricating `max_tokens` (= per-model registry cap) when the client omits it, which caused `400 "expected <=200000"` on /alpha/generate for high-cap models. An explicit oversized client value is clamped to the 200k endpoint ceiling. The PR's registry maxOutputTokens recaps (open-sse/config/providers/registry/command-code/index.ts) are intentionally NOT included pending reconciliation; diegosouzapw#5221 stays open for that.
…stry caps (diegosouzapw#5221) Integrated into release/v3.8.40 — corrective max_tokens part already cherry-picked (7ffc6da); this brings the registry maxOutputTokens caps. Thanks @adivekar-utexas.
Summary
DeepSeek V4 Pro/Flash are currently broken on the
command-codeprovider. Any request without a client-suppliedmax_tokens(e.g. the dashboard test console) fails with:Root cause
CommandCodeExecutoralways sendsparams.max_tokensto/alpha/generate, fabricating a value when the client sends none.getModelMaxTokensCap(added in #4518) fed the model's advertised output capacity from the registry straight in as the cap. For DeepSeek V4 that registry value is384000, but/alpha/generateenforces a hard per-request ceiling of200000, so the request is rejected.Two conflated concepts:
maxOutputTokens) — and the registry value is also just wrong: Command Code's gateway caps DeepSeek output at131072per/provider/v1/models, not384000.200000on themax_tokensfield, regardless of model.Fix
Executor.
max_tokensis optional on/alpha/generate. Only forward it when the client actually supplies one, clamped to the200000ceiling so an oversized client value degrades gracefully instead of 400ing. When the client omits it, omit the field so upstream applies the model's native default. This mirrors the provider-driven clamp convention already used inantigravity.ts, and removes the registry dependency from the request path (getModelMaxTokensCapdeleted).Registry. Correct
maxOutputTokensto the real Command Code gateway values from/provider/v1/models:These feed the combo router's output-limit check (
exceedsKnownOutputLimit) and dashboard metadata. The separate direct-DeepSeek spec inmodelSpecs.ts(384000) is intentionally left untouched, since the model truly supports 384K output on its native API — only Command Code's gateway caps it lower.Test plan
tests/unit/command-code-executor.test.ts— new/updated:max_tokenswhen the client omits one (GLM-5.x)max_tokensfor DeepSeek V4 when the client omits one (the reported scenario)max_tokens(500000) down to 200000max_tokens(2048) unchangedMade with Cursor