Repository navigation
fix: reasoning_effort for alibaba/minimax/xiaomi - #3075
Conversation
Fixes #3050 — reasoning_effort was silently ignored for every model routed through the alibaba provider because DashScope's chat completions API doesn't recognize the parameter; thinking is controlled via enable_thinking (boolean) and thinking_budget (max thinking tokens). - prepare-request-body.ts: dedicated alibaba case that translates the unified reasoning parameters into DashScope's native fields for mappings that declare reasoningMaxTokens: none disables thinking explicitly (DashScope thinking models think by default), other tiers enable it with a native budget mirroring the Google tier mapping, and reasoning.max_tokens forwards as thinking_budget verbatim. The budget is kept below the caller's max_tokens because DashScope rejects thinking_budget >= max_tokens for some models (verified on glm-5.2). - models: declared reasoningMaxTokens (the actual parameter the provider supports, rather than effort tiers it doesn't) and published reasoning_effort in supportedParameters on the live-verified alibaba mappings: qwen3-max, qwen3.7-max/plus, qwen3.5-397b-a17b, the four qwen3.6 models, glm-5.2, deepseek-v4-pro/flash. The cn-beijing-only mappings (glm-5, kimi-k2.5) stay untouched until they can be verified. - docs: note Alibaba among the budget-translated providers. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Xace31kcoujqY3ndZrtR7i
Fixes #3051 — reasoning_effort was silently ignored for every model routed through the minimax provider. MiniMax's API doesn't recognize the parameter; its thinking models take a binary thinking parameter ({ type: "adaptive" | "disabled" }) and think by default. - prepare-request-body.ts: dedicated minimax case translating reasoning_effort into the thinking parameter: none/minimal map to an explicit disable, low..max to an explicit adaptive enable, and unset sends nothing to keep the provider default (thinking on). Only MiniMax-M3 can actually turn thinking off — the M2.x family silently ignores "disabled" and keeps thinking (verified live) — so disable requests are gated on mappings declaring none in reasoningEfforts and collapse onto the minimum elsewhere. The existing reasoning_split extra_body behavior moved from the default case unchanged. - minimax.ts catalog: declared reasoningEfforts on the eight active minimax chat mappings — the full tier list on MiniMax-M3, low..max on the always-thinking M2.x family. MiniMax-Text-01 stays undeclared (no observable thinking to control). Verified live against api.minimax.io and with TEST_MODELS="minimax/minimax-m3,minimax/minimax-m2.7,minimax/minimax-m2.5,minimax/minimax-m2.1" FULL_MODE=true e2e (112 passed, 0 failed). Co-Authored-By: Claude Fable 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01Xace31kcoujqY3ndZrtR7i
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
WalkthroughAlibaba and MiniMax reasoning metadata, request-body translations, tests, and documentation are updated. Alibaba uses DashScope thinking fields and budgets; MiniMax uses provider-specific thinking modes and effort availability. ChangesReasoning provider mappings
Estimated code review effort: 4 (Complex) | ~45 minutes Possibly related PRs
Suggested reviewers: 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
🧹 Nitpick comments (4)
packages/actions/src/prepare-request-body.spec.ts (1)
1090-1098: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick winTest doesn't exercise the intended "no reasoningMaxTokens" scenario.
"kimi-k2.5"has noalibabaprovider mapping at all, so this test passes becausegetProviderMappingreturnsundefined(mapping not found), not because a real Alibaba mapping lacksreasoningMaxTokens. Use an actual Alibaba-provider model without the flag (e.g.qwen-max) to test the intended branch.✅ Suggested fix
test("sends nothing for mappings without budget-controlled thinking", async () => { const requestBody = await prepare({ - model: "kimi-k2.5", + model: "qwen-max", reasoningEffort: "high", });🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@packages/actions/src/prepare-request-body.spec.ts` around lines 1090 - 1098, Update the test case around prepare to use the Alibaba-provider model "qwen-max" instead of "kimi-k2.5", ensuring it exercises a real mapping without reasoningMaxTokens while preserving the existing assertions that all reasoning fields are undefined.packages/actions/src/prepare-request-body.ts (2)
2099-2116: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick winDuplicate tier→budget mapping — identical to the Google branch's
getThinkingBudget.This switch (minimal→512, low→2048, high→24576, xhigh/max→65536, medium/default→8192) is byte-for-byte the same as the existing Google
getThinkingBudgetat lines 3247-3264. Consider extracting a single shared helper (e.g.getTierThinkingBudget(effort)) used by both branches, so the two providers' tier tables can't silently drift apart in a future edit.♻️ Suggested extraction
+function getTierThinkingBudget(effort?: string): number { + switch (effort) { + case "minimal": + return 512; + case "low": + return 2048; + case "high": + return 24576; + case "xhigh": + case "max": + return 65536; + case "medium": + default: + return 8192; + } +}Then reuse
getTierThinkingBudgetin both thealibabaandgoogle-*branches instead of redefining the switch locally in each.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@packages/actions/src/prepare-request-body.ts` around lines 2099 - 2116, Extract the duplicated effort-to-thinking-budget switch into a shared helper such as getTierThinkingBudget, preserving the existing minimal, low, medium/default, high, xhigh, and max mappings. Replace the local getThinkingBudget definitions in both the Alibaba branch and the Google branch with calls to this shared helper.
2163-2185: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick winDuplicate binary-thinking mapping — identical pattern to the Moonshot branch.
The
wantsThinking/canDisableThinkinglogic here is structurally identical to the Moonshot branch above (lines 2040-2051), differing only in the enabled-state literal ("adaptive" vs "enabled"). Consider factoring this into a small shared helper (e.g.resolveBinaryThinking(reasoning_effort, canDisable, enabledValue)) to avoid maintaining the same disable/enable decision tree twice.♻️ Suggested extraction
+function resolveBinaryThinking( + reasoning_effort: string | undefined, + canDisableThinking: boolean, + enabledType: string, +): { type: string } | undefined { + const wantsThinking = + reasoning_effort !== "none" && reasoning_effort !== "minimal"; + if (wantsThinking) { + return { type: enabledType }; + } + if (canDisableThinking) { + return { type: "disabled" }; + } + return undefined; +}Then in both
moonshotandminimax:const thinking = resolveBinaryThinking(reasoning_effort, canDisableThinking, "enabled" | "adaptive"); if (thinking) requestBody.thinking = thinking;🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@packages/actions/src/prepare-request-body.ts` around lines 2163 - 2185, Extract the duplicated binary-thinking decision tree from the Moonshot and MiniMax branches into a shared helper, such as resolveBinaryThinking, accepting reasoning_effort, the can-disable flag, and the provider-specific enabled-state value. Update both branches to use the helper and assign requestBody.thinking only when it returns a value, preserving Moonshot’s "enabled" and MiniMax’s "adaptive" literals and existing disable behavior.apps/docs/content/features/reasoning.mdx (1)
169-181: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win"Supported Models" and "Provider-Specific Constraints" weren't updated to include Alibaba.
The callout above (lines 142-147) was updated in this PR to say
reasoning.max_tokensis now "Supported by Anthropic Claude and Google Gemini thinking models, plus Alibaba-hosted thinking models," but the detailed "### Supported Models" list (171-174) and "### Provider-Specific Constraints" section (180-181) right below it still only mention Anthropic and Google, with no Alibaba bullet and no mention of DashScope'sthinking_budget < max_tokensconstraint implemented inprepare-request-body.ts. This leaves the doc self-contradictory for a reader who continues past the callout.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@apps/docs/content/features/reasoning.mdx` around lines 169 - 181, Update the “Supported Models” section to add Alibaba-hosted thinking models and update “Provider-Specific Constraints” with Alibaba/DashScope’s requirement that thinking_budget be less than max_tokens, while preserving the existing Anthropic and Google details.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Nitpick comments:
In `@apps/docs/content/features/reasoning.mdx`:
- Around line 169-181: Update the “Supported Models” section to add
Alibaba-hosted thinking models and update “Provider-Specific Constraints” with
Alibaba/DashScope’s requirement that thinking_budget be less than max_tokens,
while preserving the existing Anthropic and Google details.
In `@packages/actions/src/prepare-request-body.spec.ts`:
- Around line 1090-1098: Update the test case around prepare to use the
Alibaba-provider model "qwen-max" instead of "kimi-k2.5", ensuring it exercises
a real mapping without reasoningMaxTokens while preserving the existing
assertions that all reasoning fields are undefined.
In `@packages/actions/src/prepare-request-body.ts`:
- Around line 2099-2116: Extract the duplicated effort-to-thinking-budget switch
into a shared helper such as getTierThinkingBudget, preserving the existing
minimal, low, medium/default, high, xhigh, and max mappings. Replace the local
getThinkingBudget definitions in both the Alibaba branch and the Google branch
with calls to this shared helper.
- Around line 2163-2185: Extract the duplicated binary-thinking decision tree
from the Moonshot and MiniMax branches into a shared helper, such as
resolveBinaryThinking, accepting reasoning_effort, the can-disable flag, and the
provider-specific enabled-state value. Update both branches to use the helper
and assign requestBody.thinking only when it returns a value, preserving
Moonshot’s "enabled" and MiniMax’s "adaptive" literals and existing disable
behavior.
ℹ️ Review info
⚙️ Run configuration
Configuration used: Repository UI
Review profile: CHILL
Plan: Pro
Run ID: f4f09554-fe33-421f-9f8b-872196c1a7d6
📒 Files selected for processing (7)
apps/docs/content/features/reasoning.mdxpackages/actions/src/prepare-request-body.spec.tspackages/actions/src/prepare-request-body.tspackages/models/src/models/alibaba.tspackages/models/src/models/deepseek.tspackages/models/src/models/minimax.tspackages/models/src/models/zai.ts
Fixes #3084 — Xiaomi natively accepts reasoning_effort low/medium/high (verified live: high consistently thinks longer than low) but 400s every other tier. Forward the native tiers verbatim, translate none to the documented binary disable (thinking: { type: "disabled" }, verified to zero out reasoning tokens on mimo-v2.5-pro and mimo-v2.5), and declare reasoningEfforts on the two active mappings. Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
## Summary The `novita/deepseek-v3.2` reasoning e2e tests (`basic reasoning`, `reasoning + streaming`) have been failing the aggregate e2e check on every PR (seen on theopenco#3075, theopenco#3087, and others). Root cause, verified live against Novita's API on 2026-07-16: Novita changed behavior so that sending `reasoning_effort` **alongside** the `chat_template_kwargs: { thinking: true }` flag suppresses reasoning entirely (0 reasoning tokens, no `reasoning_content`), while `reasoning_effort` on its own has no effect on this hybrid model. The gateway sends both: `requiresEnableThinking` adds the template flag, and the default OpenAI-compatible case forwards `reasoning_effort` raw because the mapping declared no `supportedParameters`. ## Fix Declare `supportedParameters` (without `reasoning_effort`) on the novita mapping so the raw effort value is dropped and only the `chat_template_kwargs.thinking` flag — the one control the deployment actually honors — is forwarded. Metadata-only change, mirroring the sibling deepseek mappings. ## Testing - `TEST_MODELS="novita/deepseek-v3.2" FULL_MODE=true pnpm test:e2e` — 88 passed, 0 failed (previously the two reasoning tests failed) - Verified live: template flag alone → 804 reasoning chars; flag + `reasoning_effort` → 0 - models/actions unit suites pass 🤖 Generated with [Claude Code](https://claude.com/claude-code) <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit * **Bug Fixes** * Improved compatibility with the Novita provider by restricting requests to supported generation and control settings. * Prevented unsupported reasoning controls from being forwarded, ensuring reasoning behavior is handled consistently. <!-- end of auto-generated comment: release notes by coderabbit.ai --> Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Summary
Fixes #3050, fixes #3051, and fixes #3084 —
reasoning_effortwas silently ignored (or wrongly forwarded) for every model routed through thealibaba,minimax, andxiaomiproviders. Alibaba and MiniMax routed through the default OpenAI-compatible case, which forwardsreasoning_effortraw (or drops it whensupportedParametersis declared); neither API recognizes the field, so the value vanished with no error. Xiaomi had a dedicated case that forwarded the raw value, which the API only partially accepts.Alibaba / DashScope (#3050)
DashScope controls thinking via
enable_thinking(boolean) andthinking_budget(max thinking tokens); thinking models think by default. Verified live:thinking_budgetcapsreasoning_tokensexactly,enable_thinking: falsecleanly disables thinking on every tested model.prepare-request-body.ts: dedicatedalibabacase. Rather than declaring effort tiers the provider doesn't support, mappings whose thinking is budget-controlled declarereasoningMaxTokens(the actual parameter the provider supports), and the case translates only for them:none→enable_thinking: falseminimal..max→enable_thinking: true+ a nativethinking_budgetmirroring the Google tier→budget mapping (512 … 65536)reasoning.max_tokens→ forwarded verbatim asthinking_budget(now also passesreasoningMaxTokensvalidation)max_tokensbecause DashScope rejectsthinking_budget >= max_tokensfor some models (verified live on glm-5.2)reasoningMaxTokens: true+reasoning_effortinsupportedParameterson the live-verified alibaba mappings: qwen3-max, qwen3.7-max/plus, qwen3.5-397b-a17b, the four qwen3.6 models, glm-5.2, deepseek-v4-pro/flash. The cn-beijing-only mappings (glm-5, kimi-k2.5) are left untouched until they can be verified.reasoning.mdx.MiniMax (#3051)
MiniMax thinking is a binary
thinkingparameter ({ type: "adaptive" | "disabled" }) and models think by default. Verified live: only MiniMax-M3 actually honors"disabled"; the whole M2.x family silently ignores it and keeps thinking, matching MiniMax's docs.prepare-request-body.ts: dedicatedminimaxcase mirroring the Moonshot binary-thinking contract:none/minimal→ explicit disable (gated on mappings declaringnoneinreasoningEfforts, so always-on M2.x collapses onto its minimum),low..max→thinking: { type: "adaptive" }, unset → nothing sent. The existingreasoning_splitextra_body behavior moved from the default case unchanged.reasoningEffortson the eight active minimax chat mappings — full tier list on MiniMax-M3,low..maxon the always-thinking M2.x family. MiniMax-Text-01 stays undeclared (no observable thinking to control).Xiaomi (#3084)
The issue assumed Xiaomi ignores
reasoning_effortentirely, but live probing shows it natively acceptslow/medium/highwith a real graduated effect (highconsistently thinks ~3-5x longer thanlowon matched prompts) and rejects every other tier with a 400 (Input should be 'low', 'medium' or 'high'). Thinking is on by default and the documented binary control isthinking: { type: "enabled" | "disabled" };"disabled"verifiably zeroes out reasoning tokens on both active models.prepare-request-body.ts: thexiaomicase forwards the native tiers verbatim (unsupported ones surface the provider's 4xx per the no-downgrade rule) and translatesnonetothinking: { type: "disabled" }, gated on mappings declaringnoneinreasoningEfforts; unset sends nothing and keeps the provider default.reasoningEfforts: ["none", "low", "medium", "high"]on the two active mappings (mimo-v2.5-pro, mimo-v2.5).Testing
prepare-request-bodysuite (170) plus models/gateway specs passTEST_MODELSpinned to all 17 touched mappings withFULL_MODE=true— every reasoning test (basic reasoning, reasoning + streaming, reasoning + tool calls) passed for all 11 alibaba mappings (singapore region), all 4 minimax mappings, and both xiaomi mappings; remaining failures wereinvalid_api_keyon cn-beijing/us-virginia region expansions (intl-only test key; also hits un-pinned models whose cheapest region is cn-beijing — pre-existing routing behavior on main, reproduced with plain requests that carry no reasoning params) plus a few content/harness flakes that passed on pinned re-runsnone→ 0 reasoning tokens,high→ reasoning present,xhigh→ provider 400 surfaced unchanged; minimax-m3none→ 0 reasoning tokenspnpm build(17/17) andpnpm formatpass🤖 Generated with Claude Code
https://claude.ai/code/session_01Xace31kcoujqY3ndZrtR7i