Skip to content

fix(command-code): cap max_tokens per model using registry maxOutputTokens - #4518

Merged
diegosouzapw merged 1 commit into
diegosouzapw:release/v3.8.33from
adivekar-utexas:fix/command-code-per-model-max-tokens
Jun 21, 2026
Merged

diegosouzapw merged 1 commit into
diegosouzapw:release/v3.8.33from
adivekar-utexas:fix/command-code-per-model-max-tokens

Conversation

@adivekar-utexas

@adivekar-utexas adivekar-utexas commented Jun 21, 2026 •

Copy link
Copy Markdown
Contributor

Summary

CommandCodeExecutor.buildCommandCodeBody clamps max_tokens to a hardcoded MAX_COMMAND_CODE_TOKENS = 200_000. This is too high for models like GLM-5.x which reject requests above 131,072 with "The max_tokens parameter is illegal.:限制数值范围[1,131072]" → upstream HTTP 502.

Lookup the per-model maxOutputTokens from REGISTRY["command-code"] and use it as the cap. Falls back to MAX_COMMAND_CODE_TOKENS for unknown models.

Diff vs release/v3.8.32

  • open-sse/executors/commandCode.ts: +21 / -4 (clampMaxTokens accepts a cap arg, new getModelMaxTokensCap helper, thread the cap into the call)
  • tests/unit/command-code-executor.test.ts: +52 / -0 (3 new tests)

Test plan

Manual verification

After this patch, a request to command-code/zai-org/GLM-5.1 (or command-code/zai-org/GLM-5) sends:

"params": { "model": "zai-org/GLM-5.1", "max_tokens": 131072, ... }

instead of the previous max_tokens: 200000. The upstream no longer rejects the request.

The release/v3.8.32 base already includes PR #4473 (the reasoning/thinking field passthrough), so this PR is a clean additive change on top of that work.

@gemini-code-assist

Copy link
Copy Markdown
Contributor

Warning

You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again!

@diegosouzapw
diegosouzapw changed the base branch from main to release/v3.8.33 June 21, 2026 14:54
@adivekar-utexas
adivekar-utexas force-pushed the fix/command-code-per-model-max-tokens branch from b872914 to 1065158 Compare June 21, 2026 14:57
@adivekar-utexas
adivekar-utexas changed the base branch from release/v3.8.33 to release/v3.8.32 June 21, 2026 14:57
@diegosouzapw
diegosouzapw force-pushed the fix/command-code-per-model-max-tokens branch 2 times, most recently from dade587 to 9d53023 Compare June 21, 2026 15:06
…okens

CommandCodeExecutor built the upstream params with a hardcoded
MAX_COMMAND_CODE_TOKENS = 200_000 cap. This is too large for models like
GLM-5.x that reject requests above 131_072 with
"The max_tokens parameter is illegal.:限制数值范围[1,131072]".

Look up the per-model maxOutputTokens from REGISTRY["command-code"] and use
it as the upper bound, falling back to MAX_COMMAND_CODE_TOKENS when the model
is unregistered or omits the field (preserves previous behavior for unknown
models). Rebuilt onto release/v3.8.33: the passthrough/resolvedModel block from
the PR base is already in release (diegosouzapw#2986), so only the per-model cap is net-new.

Co-authored-by: diegosouzapw <diegosouza.pw@gmail.com>
@diegosouzapw
diegosouzapw force-pushed the fix/command-code-per-model-max-tokens branch from 9d53023 to ae73073 Compare June 21, 2026 15:29
@diegosouzapw
diegosouzapw changed the base branch from release/v3.8.32 to release/v3.8.33 June 21, 2026 15:33
@diegosouzapw
diegosouzapw merged commit cda5831 into diegosouzapw:release/v3.8.33 Jun 21, 2026
1 check passed
@diegosouzapw

Copy link
Copy Markdown
Owner

Obrigado, @adivekar-utexas! 🙌 Reconstruí sobre a release/v3.8.33 aplicando só o net-novo (o bloco de passthrough do PR já estava na release via #2986) e validei os 3 testes (GLM-5.x→131072, DeepSeek v4→384000, e honra max_tokens menor do cliente). Fix de alto valor — GLM-5.x estava quebrando hoje. Entra na próxima release.

@diegosouzapw diegosouzapw mentioned this pull request Jun 22, 2026
adivekar-utexas added a commit to adivekar-utexas/OmniRoute that referenced this pull request Jun 28, 2026
…stry caps

The executor always sent params.max_tokens to /alpha/generate, fabricating a
value from the registry maxOutputTokens when the client sent none. For DeepSeek
V4 (registry maxOutputTokens: 384000) this produced a request the endpoint
rejects with 400 "Too big: expected number to be <=200000 at params.max_tokens",
breaking DeepSeek V4 Pro/Flash on command-code entirely.

Root cause: getModelMaxTokensCap (added in diegosouzapw#4518) fed the model's advertised
output capacity (384000) straight into the request as the cap. That capacity is
distinct from the endpoint's hard per-request ceiling of 200000, and is also
simply wrong in the registry (Command Code's gateway caps DeepSeek output at
131072 per /provider/v1/models).

Fix (executor): max_tokens is optional on /alpha/generate. Only forward it when
the client actually supplies one, clamped to the 200000 endpoint ceiling so an
oversized client value degrades gracefully instead of 400ing. When the client
omits it, omit the field so upstream applies the model's native default. This
mirrors the provider-driven clamp convention in antigravity.ts and removes the
registry dependency from the request path (getModelMaxTokensCap deleted).

Fix (registry): correct maxOutputTokens to the real Command Code gateway values
from /provider/v1/models: DeepSeek V4 384000->131072, Kimi 131072->65536,
GLM-5/5.1 131072->32768, MiniMax M2.5/M2.7 131072->65536, Qwen 3.6 131072->32768.
These feed the combo router's output-limit check and dashboard metadata. The
separate direct-DeepSeek spec in modelSpecs.ts (384000) is left untouched, since
the model truly supports 384K output on its native API.

Tests: omit-when-absent (GLM + DeepSeek, the reported scenario), clamp oversized
client value to 200000, and honor a smaller client value unchanged.

Co-authored-by: Cursor <cursoragent@cursor.com>
adivekar-utexas added a commit to adivekar-utexas/OmniRoute that referenced this pull request Jun 28, 2026
…stry caps

The executor always sent params.max_tokens to /alpha/generate, fabricating a
value from the registry maxOutputTokens when the client sent none. For DeepSeek
V4 (registry maxOutputTokens: 384000) this produced a request the endpoint
rejects with 400 "Too big: expected number to be <=200000 at params.max_tokens",
breaking DeepSeek V4 Pro/Flash on command-code entirely.

Root cause: getModelMaxTokensCap (added in diegosouzapw#4518) fed the model's advertised
output capacity (384000) straight into the request as the cap. That capacity is
distinct from the endpoint's hard per-request ceiling of 200000, and is also
simply wrong in the registry (Command Code's gateway caps DeepSeek output at
131072 per /provider/v1/models).

Fix (executor): max_tokens is optional on /alpha/generate. Only forward it when
the client actually supplies one, clamped to the 200000 endpoint ceiling so an
oversized client value degrades gracefully instead of 400ing. When the client
omits it, omit the field so upstream applies the model's native default. This
mirrors the provider-driven clamp convention in antigravity.ts and removes the
registry dependency from the request path (getModelMaxTokensCap deleted).

Fix (registry): correct maxOutputTokens to the real Command Code gateway values
from /provider/v1/models: DeepSeek V4 384000->131072, Kimi 131072->65536,
GLM-5/5.1 131072->32768, MiniMax M2.5/M2.7 131072->65536, Qwen 3.6 131072->32768.
These feed the combo router's output-limit check and dashboard metadata. The
separate direct-DeepSeek spec in modelSpecs.ts (384000) is left untouched, since
the model truly supports 384K output on its native API.

Tests: omit-when-absent (GLM + DeepSeek, the reported scenario), clamp oversized
client value to 200000, and honor a smaller client value unchanged.

Co-authored-by: Cursor <cursoragent@cursor.com>
tkgo11 pushed a commit to tkgo11/OmniRoute that referenced this pull request Sep 23, 2026
…okens (diegosouzapw#4518)

clampMaxTokens now uses the per-model maxOutputTokens from REGISTRY['command-code'] as the upper bound (falls back to MAX_COMMAND_CODE_TOKENS), so GLM-5.x stops being rejected for max_tokens > 131072. Rebuilt onto release/v3.8.33 (passthrough block from the PR base is already present via diegosouzapw#2986); 3 tests added.

Integrated into release/v3.8.33.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants