fix(command-code): cap max_tokens per model using registry maxOutputTokens - #4518
Merged
diegosouzapw merged 1 commit intoJun 21, 2026
Conversation
Contributor
|
Warning You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again! |
adivekar-utexas
force-pushed
the
fix/command-code-per-model-max-tokens
branch
from
June 21, 2026 14:57
b872914 to
1065158
Compare
diegosouzapw
force-pushed
the
fix/command-code-per-model-max-tokens
branch
2 times, most recently
from
June 21, 2026 15:06
dade587 to
9d53023
Compare
…okens CommandCodeExecutor built the upstream params with a hardcoded MAX_COMMAND_CODE_TOKENS = 200_000 cap. This is too large for models like GLM-5.x that reject requests above 131_072 with "The max_tokens parameter is illegal.:限制数值范围[1,131072]". Look up the per-model maxOutputTokens from REGISTRY["command-code"] and use it as the upper bound, falling back to MAX_COMMAND_CODE_TOKENS when the model is unregistered or omits the field (preserves previous behavior for unknown models). Rebuilt onto release/v3.8.33: the passthrough/resolvedModel block from the PR base is already in release (diegosouzapw#2986), so only the per-model cap is net-new. Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>
diegosouzapw
force-pushed
the
fix/command-code-per-model-max-tokens
branch
from
June 21, 2026 15:29
9d53023 to
ae73073
Compare
Owner
|
Obrigado, @adivekar-utexas! 🙌 Reconstruí sobre a |
Merged
4 tasks done
adivekar-utexas
added a commit
to adivekar-utexas/OmniRoute
that referenced
this pull request
Jun 28, 2026
…stry caps The executor always sent params.max_tokens to /alpha/generate, fabricating a value from the registry maxOutputTokens when the client sent none. For DeepSeek V4 (registry maxOutputTokens: 384000) this produced a request the endpoint rejects with 400 "Too big: expected number to be <=200000 at params.max_tokens", breaking DeepSeek V4 Pro/Flash on command-code entirely. Root cause: getModelMaxTokensCap (added in diegosouzapw#4518) fed the model's advertised output capacity (384000) straight into the request as the cap. That capacity is distinct from the endpoint's hard per-request ceiling of 200000, and is also simply wrong in the registry (Command Code's gateway caps DeepSeek output at 131072 per /provider/v1/models). Fix (executor): max_tokens is optional on /alpha/generate. Only forward it when the client actually supplies one, clamped to the 200000 endpoint ceiling so an oversized client value degrades gracefully instead of 400ing. When the client omits it, omit the field so upstream applies the model's native default. This mirrors the provider-driven clamp convention in antigravity.ts and removes the registry dependency from the request path (getModelMaxTokensCap deleted). Fix (registry): correct maxOutputTokens to the real Command Code gateway values from /provider/v1/models: DeepSeek V4 384000->131072, Kimi 131072->65536, GLM-5/5.1 131072->32768, MiniMax M2.5/M2.7 131072->65536, Qwen 3.6 131072->32768. These feed the combo router's output-limit check and dashboard metadata. The separate direct-DeepSeek spec in modelSpecs.ts (384000) is left untouched, since the model truly supports 384K output on its native API. Tests: omit-when-absent (GLM + DeepSeek, the reported scenario), clamp oversized client value to 200000, and honor a smaller client value unchanged. Co-authored-by: Cursor <cursoragent@cursor.com>
adivekar-utexas
added a commit
to adivekar-utexas/OmniRoute
that referenced
this pull request
Jun 28, 2026
…stry caps The executor always sent params.max_tokens to /alpha/generate, fabricating a value from the registry maxOutputTokens when the client sent none. For DeepSeek V4 (registry maxOutputTokens: 384000) this produced a request the endpoint rejects with 400 "Too big: expected number to be <=200000 at params.max_tokens", breaking DeepSeek V4 Pro/Flash on command-code entirely. Root cause: getModelMaxTokensCap (added in diegosouzapw#4518) fed the model's advertised output capacity (384000) straight into the request as the cap. That capacity is distinct from the endpoint's hard per-request ceiling of 200000, and is also simply wrong in the registry (Command Code's gateway caps DeepSeek output at 131072 per /provider/v1/models). Fix (executor): max_tokens is optional on /alpha/generate. Only forward it when the client actually supplies one, clamped to the 200000 endpoint ceiling so an oversized client value degrades gracefully instead of 400ing. When the client omits it, omit the field so upstream applies the model's native default. This mirrors the provider-driven clamp convention in antigravity.ts and removes the registry dependency from the request path (getModelMaxTokensCap deleted). Fix (registry): correct maxOutputTokens to the real Command Code gateway values from /provider/v1/models: DeepSeek V4 384000->131072, Kimi 131072->65536, GLM-5/5.1 131072->32768, MiniMax M2.5/M2.7 131072->65536, Qwen 3.6 131072->32768. These feed the combo router's output-limit check and dashboard metadata. The separate direct-DeepSeek spec in modelSpecs.ts (384000) is left untouched, since the model truly supports 384K output on its native API. Tests: omit-when-absent (GLM + DeepSeek, the reported scenario), clamp oversized client value to 200000, and honor a smaller client value unchanged. Co-authored-by: Cursor <cursoragent@cursor.com>
tkgo11
pushed a commit
to tkgo11/OmniRoute
that referenced
this pull request
Sep 23, 2026
…okens (diegosouzapw#4518) clampMaxTokens now uses the per-model maxOutputTokens from REGISTRY['command-code'] as the upper bound (falls back to MAX_COMMAND_CODE_TOKENS), so GLM-5.x stops being rejected for max_tokens > 131072. Rebuilt onto release/v3.8.33 (passthrough block from the PR base is already present via diegosouzapw#2986); 3 tests added. Integrated into release/v3.8.33.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
CommandCodeExecutor.buildCommandCodeBodyclampsmax_tokensto a hardcodedMAX_COMMAND_CODE_TOKENS = 200_000. This is too high for models like GLM-5.x which reject requests above 131,072 with"The max_tokens parameter is illegal.:限制数值范围[1,131072]"→ upstream HTTP 502.Lookup the per-model
maxOutputTokensfromREGISTRY["command-code"]and use it as the cap. Falls back toMAX_COMMAND_CODE_TOKENSfor unknown models.Diff vs release/v3.8.32
open-sse/executors/commandCode.ts: +21 / -4 (clampMaxTokens accepts acaparg, newgetModelMaxTokensCaphelper, thread the cap into the call)tests/unit/command-code-executor.test.ts: +52 / -0 (3 new tests)Test plan
Manual verification
After this patch, a request to
command-code/zai-org/GLM-5.1(orcommand-code/zai-org/GLM-5) sends:instead of the previous
max_tokens: 200000. The upstream no longer rejects the request.The
release/v3.8.32base already includes PR #4473 (the reasoning/thinking field passthrough), so this PR is a clean additive change on top of that work.