Skip to content

fix(command-code): omit max_tokens when client omits it; correct registry caps - #5221

Merged
diegosouzapw merged 1 commit into
diegosouzapw:release/v3.8.40from
adivekar-utexas:fix/command-code-max-tokens-omit
Jun 28, 2026
Merged

diegosouzapw merged 1 commit into
diegosouzapw:release/v3.8.40from
adivekar-utexas:fix/command-code-max-tokens-omit

Conversation

@adivekar-utexas

Copy link
Copy Markdown
Contributor

Summary

DeepSeek V4 Pro/Flash are currently broken on the command-code provider. Any request without a client-supplied max_tokens (e.g. the dashboard test console) fails with:

[400] Invalid request error ... Validation error: Too big: expected number to be <=200000 at "params.max_tokens"

Root cause

CommandCodeExecutor always sends params.max_tokens to /alpha/generate, fabricating a value when the client sends none. getModelMaxTokensCap (added in #4518) fed the model's advertised output capacity from the registry straight in as the cap. For DeepSeek V4 that registry value is 384000, but /alpha/generate enforces a hard per-request ceiling of 200000, so the request is rejected.

Two conflated concepts:

  • Model output capability (registry maxOutputTokens) — and the registry value is also just wrong: Command Code's gateway caps DeepSeek output at 131072 per /provider/v1/models, not 384000.
  • Endpoint per-request hard limit — a flat 200000 on the max_tokens field, regardless of model.

Fix

Executor. max_tokens is optional on /alpha/generate. Only forward it when the client actually supplies one, clamped to the 200000 ceiling so an oversized client value degrades gracefully instead of 400ing. When the client omits it, omit the field so upstream applies the model's native default. This mirrors the provider-driven clamp convention already used in antigravity.ts, and removes the registry dependency from the request path (getModelMaxTokensCap deleted).

Registry. Correct maxOutputTokens to the real Command Code gateway values from /provider/v1/models:

Model Was Now
DeepSeek V4 Pro / Flash 384000 131072
Kimi K2.6 / K2.5 131072 65536
GLM-5.1 / GLM-5 131072 32768
MiniMax M2.7 / M2.5 131072 65536
Qwen 3.6 Max / Plus 131072 32768

These feed the combo router's output-limit check (exceedsKnownOutputLimit) and dashboard metadata. The separate direct-DeepSeek spec in modelSpecs.ts (384000) is intentionally left untouched, since the model truly supports 384K output on its native API — only Command Code's gateway caps it lower.

Test plan

  • tests/unit/command-code-executor.test.ts — new/updated:
    • omits max_tokens when the client omits one (GLM-5.x)
    • omits max_tokens for DeepSeek V4 when the client omits one (the reported scenario)
    • clamps an oversized client max_tokens (500000) down to 200000
    • honors a smaller client max_tokens (2048) unchanged
  • 22/22 command-code executor tests pass
  • 132/132 related combo/provider-validation/minimax model-spec tests pass (no regression)
  • pre-commit hooks (prettier/eslint/any-budget/tracked-artifacts) pass

Made with Cursor

@gemini-code-assist gemini-code-assist Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Code Review

This pull request updates the Command Code executor and registry configuration to handle max_tokens more robustly. Specifically, it reduces the registered maxOutputTokens for several models and refactors the clampMaxTokens logic to omit the max_tokens parameter entirely when not supplied by the client, allowing the upstream provider to apply its native defaults. When a client-supplied value is present, it is safely clamped to the 200,000 token endpoint ceiling to prevent 400 errors. Corresponding unit tests have been updated and added to verify these behaviors. There are no review comments provided, and I have no additional feedback to offer as the changes are clean, well-documented, and properly tested.

Important

The consumer version of Gemini Code Assist on GitHub is being sunset. Starting June 18, 2026, new organization installations will be blocked, and all code review activity will officially cease on July 17, 2026.
For more details on the timeline and next steps, please review the Help Documentation.

@diegosouzapw
diegosouzapw changed the base branch from release/v3.8.39 to release/v3.8.40 June 28, 2026 11:12
@adivekar-utexas
adivekar-utexas force-pushed the fix/command-code-max-tokens-omit branch from dcdc392 to b434997 Compare June 28, 2026 11:14
…stry caps

The executor always sent params.max_tokens to /alpha/generate, fabricating a
value from the registry maxOutputTokens when the client sent none. For DeepSeek
V4 (registry maxOutputTokens: 384000) this produced a request the endpoint
rejects with 400 "Too big: expected number to be <=200000 at params.max_tokens",
breaking DeepSeek V4 Pro/Flash on command-code entirely.

Root cause: getModelMaxTokensCap (added in diegosouzapw#4518) fed the model's advertised
output capacity (384000) straight into the request as the cap. That capacity is
distinct from the endpoint's hard per-request ceiling of 200000, and is also
simply wrong in the registry (Command Code's gateway caps DeepSeek output at
131072 per /provider/v1/models).

Fix (executor): max_tokens is optional on /alpha/generate. Only forward it when
the client actually supplies one, clamped to the 200000 endpoint ceiling so an
oversized client value degrades gracefully instead of 400ing. When the client
omits it, omit the field so upstream applies the model's native default. This
mirrors the provider-driven clamp convention in antigravity.ts and removes the
registry dependency from the request path (getModelMaxTokensCap deleted).

Fix (registry): correct maxOutputTokens to the real Command Code gateway values
from /provider/v1/models: DeepSeek V4 384000->131072, Kimi 131072->65536,
GLM-5/5.1 131072->32768, MiniMax M2.5/M2.7 131072->65536, Qwen 3.6 131072->32768.
These feed the combo router's output-limit check and dashboard metadata. The
separate direct-DeepSeek spec in modelSpecs.ts (384000) is left untouched, since
the model truly supports 384K output on its native API.

Tests: omit-when-absent (GLM + DeepSeek, the reported scenario), clamp oversized
client value to 200000, and honor a smaller client value unchanged.

Co-authored-by: Cursor <cursoragent@cursor.com>
@adivekar-utexas
adivekar-utexas force-pushed the fix/command-code-max-tokens-omit branch from d76115d to 509806d Compare June 28, 2026 11:50
@adivekar-utexas

Copy link
Copy Markdown
Contributor Author

Heads-up on the one red check (Unit Tests fast-path (2/2)): the two failing tests are pre-existing on release/v3.8.40 and unrelated to this PR.

  1. tests/unit/dockerfile-build-heap-4076.test.ts → "[BUG] Bug Report: Docker Build Failure due to OOM (JavaScript Heap Out of Memory) during Next.js Production Build #4076 the heap ceiling is set BEFORE npm run build". I verified this fails on the pristine release/v3.8.40 tip (c4280433c) with zero changes from this branch applied, so it is not a regression from this change.
  2. tests/unit/build/opencode-plugin-esm-only.test.ts → "@omniroute/opencode-plugin dist/ not built". This is a CI build-step dependency (the plugin dist/ wasn't built in the runner), also unrelated to command-code.

This PR only touches open-sse/executors/commandCode.ts, the command-code registry, and tests/unit/command-code-executor.test.ts. All command-code executor tests pass (13/13), and the other checks (Vitest, Unit Tests 1/2, dast-smoke, semgrep) are green.

Happy to rebase onto a newer release branch if the heap-4076 fix lands there. Thanks!

diegosouzapw pushed a commit that referenced this pull request Jun 28, 2026
Cherry-picked the corrective part of #5221 only: the executor stops fabricating
`max_tokens` (= per-model registry cap) when the client omits it, which caused
`400 "expected <=200000"` on /alpha/generate for high-cap models. An explicit
oversized client value is clamped to the 200k endpoint ceiling. The PR's registry
maxOutputTokens recaps (open-sse/config/providers/registry/command-code/index.ts)
are intentionally NOT included pending reconciliation; #5221 stays open for that.
@diegosouzapw

Copy link
Copy Markdown
Owner

Thanks @adivekar-utexas! The corrective part of this PR is now in release/v3.8.40 (cherry-picked with your authorship preserved, credited in CHANGELOG): the executor no longer fabricates max_tokens when the client omits it, so the upstream applies the model's native default — fixing the 400 "expected <=200000" on /alpha/generate for high-cap models. The oversized-client-value clamp to the 200k ceiling and the new tests came along too.

I deliberately left out the registry maxOutputTokens recaps (open-sse/config/providers/registry/command-code/index.ts) for now: those values (e.g. GLM 32768) feed routing decisions (exceedsKnownOutputLimit, task-aware routing) and one disagrees with the upstream limit your removed comment cited (GLM 131072). Since the executor no longer reads per-model caps for clamping, dropping a present large client value to the per-model cap on a direct single-model call is also no longer guarded — worth reconciling separately.

I'm keeping this PR open so we can settle the cap values (and optionally restore a per-model clamp-down for explicit oversized values) without losing your work. Really appreciate the fix!

@diegosouzapw
diegosouzapw merged commit d8a392a into diegosouzapw:release/v3.8.40 Jun 28, 2026
3 checks passed
tkgo11 pushed a commit to tkgo11/OmniRoute that referenced this pull request Sep 23, 2026
…#5221)

Cherry-picked the corrective part of diegosouzapw#5221 only: the executor stops fabricating
`max_tokens` (= per-model registry cap) when the client omits it, which caused
`400 "expected <=200000"` on /alpha/generate for high-cap models. An explicit
oversized client value is clamped to the 200k endpoint ceiling. The PR's registry
maxOutputTokens recaps (open-sse/config/providers/registry/command-code/index.ts)
are intentionally NOT included pending reconciliation; diegosouzapw#5221 stays open for that.
tkgo11 pushed a commit to tkgo11/OmniRoute that referenced this pull request Sep 23, 2026
…stry caps (diegosouzapw#5221)

Integrated into release/v3.8.40 — corrective max_tokens part already cherry-picked (7ffc6da); this brings the registry maxOutputTokens caps. Thanks @adivekar-utexas.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants