Skip to content

fix: reasoning_effort for alibaba/minimax/xiaomi - #3075

Merged
steebchen merged 3 commits into
mainfrom
claude/alibaba-reasoning-effort-e2e-35tnvw
Jul 16, 2026
Merged

steebchen merged 3 commits into
mainfrom
claude/alibaba-reasoning-effort-e2e-35tnvw

Conversation

@steebchen

@steebchen steebchen commented Jul 15, 2026 •

Copy link
Copy Markdown
Member

Summary

Fixes #3050, fixes #3051, and fixes #3084 — reasoning_effort was silently ignored (or wrongly forwarded) for every model routed through the alibaba, minimax, and xiaomi providers. Alibaba and MiniMax routed through the default OpenAI-compatible case, which forwards reasoning_effort raw (or drops it when supportedParameters is declared); neither API recognizes the field, so the value vanished with no error. Xiaomi had a dedicated case that forwarded the raw value, which the API only partially accepts.

Alibaba / DashScope (#3050)

DashScope controls thinking via enable_thinking (boolean) and thinking_budget (max thinking tokens); thinking models think by default. Verified live: thinking_budget caps reasoning_tokens exactly, enable_thinking: false cleanly disables thinking on every tested model.

  • prepare-request-body.ts: dedicated alibaba case. Rather than declaring effort tiers the provider doesn't support, mappings whose thinking is budget-controlled declare reasoningMaxTokens (the actual parameter the provider supports), and the case translates only for them:
    • none → enable_thinking: false
    • minimal..max → enable_thinking: true + a native thinking_budget mirroring the Google tier→budget mapping (512 … 65536)
    • reasoning.max_tokens → forwarded verbatim as thinking_budget (now also passes reasoningMaxTokens validation)
    • unset → nothing sent, provider default preserved
    • the budget is kept below the caller's max_tokens because DashScope rejects thinking_budget >= max_tokens for some models (verified live on glm-5.2)
  • Catalog: reasoningMaxTokens: true + reasoning_effort in supportedParameters on the live-verified alibaba mappings: qwen3-max, qwen3.7-max/plus, qwen3.5-397b-a17b, the four qwen3.6 models, glm-5.2, deepseek-v4-pro/flash. The cn-beijing-only mappings (glm-5, kimi-k2.5) are left untouched until they can be verified.
  • Docs: note Alibaba among the budget-translated providers in reasoning.mdx.

MiniMax (#3051)

MiniMax thinking is a binary thinking parameter ({ type: "adaptive" | "disabled" }) and models think by default. Verified live: only MiniMax-M3 actually honors "disabled"; the whole M2.x family silently ignores it and keeps thinking, matching MiniMax's docs.

  • prepare-request-body.ts: dedicated minimax case mirroring the Moonshot binary-thinking contract: none/minimal → explicit disable (gated on mappings declaring none in reasoningEfforts, so always-on M2.x collapses onto its minimum), low..max → thinking: { type: "adaptive" }, unset → nothing sent. The existing reasoning_split extra_body behavior moved from the default case unchanged.
  • Catalog: declared reasoningEfforts on the eight active minimax chat mappings — full tier list on MiniMax-M3, low..max on the always-thinking M2.x family. MiniMax-Text-01 stays undeclared (no observable thinking to control).

Xiaomi (#3084)

The issue assumed Xiaomi ignores reasoning_effort entirely, but live probing shows it natively accepts low/medium/high with a real graduated effect (high consistently thinks ~3-5x longer than low on matched prompts) and rejects every other tier with a 400 (Input should be 'low', 'medium' or 'high'). Thinking is on by default and the documented binary control is thinking: { type: "enabled" | "disabled" }; "disabled" verifiably zeroes out reasoning tokens on both active models.

  • prepare-request-body.ts: the xiaomi case forwards the native tiers verbatim (unsupported ones surface the provider's 4xx per the no-downgrade rule) and translates none to thinking: { type: "disabled" }, gated on mappings declaring none in reasoningEfforts; unset sends nothing and keeps the provider default.
  • Catalog: reasoningEfforts: ["none", "low", "medium", "high"] on the two active mappings (mimo-v2.5-pro, mimo-v2.5).

Testing

  • 16 new unit specs covering effort→budget/thinking translation, explicit disable, provider-default passthrough, budget clamping, native-tier forwarding, and ungated mappings; full prepare-request-body suite (170) plus models/gateway specs pass
  • Combined local e2e: TEST_MODELS pinned to all 17 touched mappings with FULL_MODE=true — every reasoning test (basic reasoning, reasoning + streaming, reasoning + tool calls) passed for all 11 alibaba mappings (singapore region), all 4 minimax mappings, and both xiaomi mappings; remaining failures were invalid_api_key on cn-beijing/us-virginia region expansions (intl-only test key; also hits un-pinned models whose cheapest region is cn-beijing — pre-existing routing behavior on main, reproduced with plain requests that carry no reasoning params) plus a few content/harness flakes that passed on pinned re-runs
  • Live gateway spot checks: xiaomi none → 0 reasoning tokens, high → reasoning present, xhigh → provider 400 surfaced unchanged; minimax-m3 none → 0 reasoning tokens
  • pnpm build (17/17) and pnpm format pass

🤖 Generated with Claude Code

https://claude.ai/code/session_01Xace31kcoujqY3ndZrtR7i

claude added 2 commits July 15, 2026 13:22
Fixes #3050 — reasoning_effort was silently ignored for every model
routed through the alibaba provider because DashScope's chat completions
API doesn't recognize the parameter; thinking is controlled via
enable_thinking (boolean) and thinking_budget (max thinking tokens).

- prepare-request-body.ts: dedicated alibaba case that translates the
  unified reasoning parameters into DashScope's native fields for
  mappings that declare reasoningMaxTokens: none disables thinking
  explicitly (DashScope thinking models think by default), other tiers
  enable it with a native budget mirroring the Google tier mapping, and
  reasoning.max_tokens forwards as thinking_budget verbatim. The budget
  is kept below the caller's max_tokens because DashScope rejects
  thinking_budget >= max_tokens for some models (verified on glm-5.2).
- models: declared reasoningMaxTokens (the actual parameter the
  provider supports, rather than effort tiers it doesn't) and published
  reasoning_effort in supportedParameters on the live-verified alibaba
  mappings: qwen3-max, qwen3.7-max/plus, qwen3.5-397b-a17b, the four
  qwen3.6 models, glm-5.2, deepseek-v4-pro/flash. The cn-beijing-only
  mappings (glm-5, kimi-k2.5) stay untouched until they can be verified.
- docs: note Alibaba among the budget-translated providers.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Xace31kcoujqY3ndZrtR7i
Fixes #3051 — reasoning_effort was silently ignored for every model
routed through the minimax provider. MiniMax's API doesn't recognize
the parameter; its thinking models take a binary thinking parameter
({ type: "adaptive" | "disabled" }) and think by default.

- prepare-request-body.ts: dedicated minimax case translating
  reasoning_effort into the thinking parameter: none/minimal map to an
  explicit disable, low..max to an explicit adaptive enable, and unset
  sends nothing to keep the provider default (thinking on). Only
  MiniMax-M3 can actually turn thinking off — the M2.x family silently
  ignores "disabled" and keeps thinking (verified live) — so disable
  requests are gated on mappings declaring none in reasoningEfforts and
  collapse onto the minimum elsewhere. The existing reasoning_split
  extra_body behavior moved from the default case unchanged.
- minimax.ts catalog: declared reasoningEfforts on the eight active
  minimax chat mappings — the full tier list on MiniMax-M3, low..max on
  the always-thinking M2.x family. MiniMax-Text-01 stays undeclared (no
  observable thinking to control).

Verified live against api.minimax.io and with
TEST_MODELS="minimax/minimax-m3,minimax/minimax-m2.7,minimax/minimax-m2.5,minimax/minimax-m2.1"
FULL_MODE=true e2e (112 passed, 0 failed).

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01Xace31kcoujqY3ndZrtR7i
Copilot AI review requested due to automatic review settings July 15, 2026 15:23
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

@coderabbitai

coderabbitai Bot commented Jul 15, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

Walkthrough

Alibaba and MiniMax reasoning metadata, request-body translations, tests, and documentation are updated. Alibaba uses DashScope thinking fields and budgets; MiniMax uses provider-specific thinking modes and effort availability.

Changes

Reasoning provider mappings

Layer / File(s) Summary
Model reasoning capability declarations
packages/models/src/models/alibaba.ts, packages/models/src/models/deepseek.ts, packages/models/src/models/zai.ts, packages/models/src/models/minimax.ts
Alibaba-backed models declare reasoning token support and effort parameters. MiniMax models declare supported reasoning tiers, including always-enabled M2.x behavior.
Provider-specific request translation
packages/actions/src/prepare-request-body.ts, packages/actions/src/prepare-request-body.spec.ts
Alibaba reasoning inputs map to enable_thinking and thinking_budget; MiniMax inputs map to thinking and extra_body.reasoning_split. Tests cover tier mappings, disabling, defaults, budget clamping, and always-on models.
Reasoning parameter documentation
apps/docs/content/features/reasoning.mdx
Documentation adds Alibaba to native thinking-budget mappings and describes the DashScope thinking_budget override.

Estimated code review effort: 4 (Complex) | ~45 minutes

Possibly related PRs

Suggested reviewers: copilot, rcogal, ratchaw

🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Linked Issues check ✅ Passed The request-body changes and model catalog updates address Alibaba DashScope and MiniMax reasoning_effort mismatches.
Out of Scope Changes check ✅ Passed The added docs, tests, and catalog updates are directly related to the reasoning_effort fix.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title captures the main reasoning_effort fixes for Alibaba and MiniMax, though it mentions Xiaomi, which is not part of this changeset.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch claude/alibaba-reasoning-effort-e2e-35tnvw

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (4)
packages/actions/src/prepare-request-body.spec.ts (1)

1090-1098: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Test doesn't exercise the intended "no reasoningMaxTokens" scenario.

"kimi-k2.5" has no alibaba provider mapping at all, so this test passes because getProviderMapping returns undefined (mapping not found), not because a real Alibaba mapping lacks reasoningMaxTokens. Use an actual Alibaba-provider model without the flag (e.g. qwen-max) to test the intended branch.

✅ Suggested fix
 	test("sends nothing for mappings without budget-controlled thinking", async () => {
 		const requestBody = await prepare({
-			model: "kimi-k2.5",
+			model: "qwen-max",
 			reasoningEffort: "high",
 		});
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@packages/actions/src/prepare-request-body.spec.ts` around lines 1090 - 1098,
Update the test case around prepare to use the Alibaba-provider model "qwen-max"
instead of "kimi-k2.5", ensuring it exercises a real mapping without
reasoningMaxTokens while preserving the existing assertions that all reasoning
fields are undefined.
packages/actions/src/prepare-request-body.ts (2)

2099-2116: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Duplicate tier→budget mapping — identical to the Google branch's getThinkingBudget.

This switch (minimal→512, low→2048, high→24576, xhigh/max→65536, medium/default→8192) is byte-for-byte the same as the existing Google getThinkingBudget at lines 3247-3264. Consider extracting a single shared helper (e.g. getTierThinkingBudget(effort)) used by both branches, so the two providers' tier tables can't silently drift apart in a future edit.

♻️ Suggested extraction
+function getTierThinkingBudget(effort?: string): number {
+	switch (effort) {
+		case "minimal":
+			return 512;
+		case "low":
+			return 2048;
+		case "high":
+			return 24576;
+		case "xhigh":
+		case "max":
+			return 65536;
+		case "medium":
+		default:
+			return 8192;
+	}
+}

Then reuse getTierThinkingBudget in both the alibaba and google-* branches instead of redefining the switch locally in each.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@packages/actions/src/prepare-request-body.ts` around lines 2099 - 2116,
Extract the duplicated effort-to-thinking-budget switch into a shared helper
such as getTierThinkingBudget, preserving the existing minimal, low,
medium/default, high, xhigh, and max mappings. Replace the local
getThinkingBudget definitions in both the Alibaba branch and the Google branch
with calls to this shared helper.

2163-2185: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Duplicate binary-thinking mapping — identical pattern to the Moonshot branch.

The wantsThinking/canDisableThinking logic here is structurally identical to the Moonshot branch above (lines 2040-2051), differing only in the enabled-state literal ("adaptive" vs "enabled"). Consider factoring this into a small shared helper (e.g. resolveBinaryThinking(reasoning_effort, canDisable, enabledValue)) to avoid maintaining the same disable/enable decision tree twice.

♻️ Suggested extraction
+function resolveBinaryThinking(
+	reasoning_effort: string | undefined,
+	canDisableThinking: boolean,
+	enabledType: string,
+): { type: string } | undefined {
+	const wantsThinking =
+		reasoning_effort !== "none" && reasoning_effort !== "minimal";
+	if (wantsThinking) {
+		return { type: enabledType };
+	}
+	if (canDisableThinking) {
+		return { type: "disabled" };
+	}
+	return undefined;
+}

Then in both moonshot and minimax: const thinking = resolveBinaryThinking(reasoning_effort, canDisableThinking, "enabled" | "adaptive"); if (thinking) requestBody.thinking = thinking;

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@packages/actions/src/prepare-request-body.ts` around lines 2163 - 2185,
Extract the duplicated binary-thinking decision tree from the Moonshot and
MiniMax branches into a shared helper, such as resolveBinaryThinking, accepting
reasoning_effort, the can-disable flag, and the provider-specific enabled-state
value. Update both branches to use the helper and assign requestBody.thinking
only when it returns a value, preserving Moonshot’s "enabled" and MiniMax’s
"adaptive" literals and existing disable behavior.
apps/docs/content/features/reasoning.mdx (1)

169-181: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

"Supported Models" and "Provider-Specific Constraints" weren't updated to include Alibaba.

The callout above (lines 142-147) was updated in this PR to say reasoning.max_tokens is now "Supported by Anthropic Claude and Google Gemini thinking models, plus Alibaba-hosted thinking models," but the detailed "### Supported Models" list (171-174) and "### Provider-Specific Constraints" section (180-181) right below it still only mention Anthropic and Google, with no Alibaba bullet and no mention of DashScope's thinking_budget < max_tokens constraint implemented in prepare-request-body.ts. This leaves the doc self-contradictory for a reader who continues past the callout.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@apps/docs/content/features/reasoning.mdx` around lines 169 - 181, Update the
“Supported Models” section to add Alibaba-hosted thinking models and update
“Provider-Specific Constraints” with Alibaba/DashScope’s requirement that
thinking_budget be less than max_tokens, while preserving the existing Anthropic
and Google details.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Nitpick comments:
In `@apps/docs/content/features/reasoning.mdx`:
- Around line 169-181: Update the “Supported Models” section to add
Alibaba-hosted thinking models and update “Provider-Specific Constraints” with
Alibaba/DashScope’s requirement that thinking_budget be less than max_tokens,
while preserving the existing Anthropic and Google details.

In `@packages/actions/src/prepare-request-body.spec.ts`:
- Around line 1090-1098: Update the test case around prepare to use the
Alibaba-provider model "qwen-max" instead of "kimi-k2.5", ensuring it exercises
a real mapping without reasoningMaxTokens while preserving the existing
assertions that all reasoning fields are undefined.

In `@packages/actions/src/prepare-request-body.ts`:
- Around line 2099-2116: Extract the duplicated effort-to-thinking-budget switch
into a shared helper such as getTierThinkingBudget, preserving the existing
minimal, low, medium/default, high, xhigh, and max mappings. Replace the local
getThinkingBudget definitions in both the Alibaba branch and the Google branch
with calls to this shared helper.
- Around line 2163-2185: Extract the duplicated binary-thinking decision tree
from the Moonshot and MiniMax branches into a shared helper, such as
resolveBinaryThinking, accepting reasoning_effort, the can-disable flag, and the
provider-specific enabled-state value. Update both branches to use the helper
and assign requestBody.thinking only when it returns a value, preserving
Moonshot’s "enabled" and MiniMax’s "adaptive" literals and existing disable
behavior.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro

Run ID: f4f09554-fe33-421f-9f8b-872196c1a7d6

📥 Commits

Reviewing files that changed from the base of the PR and between fc2f5f1 and 67e74af.

📒 Files selected for processing (7)
  • apps/docs/content/features/reasoning.mdx
  • packages/actions/src/prepare-request-body.spec.ts
  • packages/actions/src/prepare-request-body.ts
  • packages/models/src/models/alibaba.ts
  • packages/models/src/models/deepseek.ts
  • packages/models/src/models/minimax.ts
  • packages/models/src/models/zai.ts

Fixes #3084 — Xiaomi natively accepts reasoning_effort low/medium/high
(verified live: high consistently thinks longer than low) but 400s every
other tier. Forward the native tiers verbatim, translate none to the
documented binary disable (thinking: { type: "disabled" }, verified to
zero out reasoning tokens on mimo-v2.5-pro and mimo-v2.5), and declare
reasoningEfforts on the two active mappings.

Co-Authored-By: Claude Fable 5 <noreply@anthropic.com>
@steebchen steebchen changed the title fix(gateway): map reasoning_effort for Alibaba, MiniMax fix(gateway): reasoning_effort for Alibaba, MiniMax, Xiaomi Jul 16, 2026
@steebchen steebchen changed the title fix(gateway): reasoning_effort for Alibaba, MiniMax, Xiaomi fix: reasoning_effort for alibaba/minimax/xiaomi Jul 16, 2026
@steebchen
steebchen enabled auto-merge July 16, 2026 12:29
@steebchen
steebchen added this pull request to the merge queue Jul 16, 2026
Merged via the queue into main with commit 1b3ff98 Jul 16, 2026
23 of 24 checks passed
@steebchen
steebchen deleted the claude/alibaba-reasoning-effort-e2e-35tnvw branch July 16, 2026 12:49
pull Bot pushed a commit to soitun/llmgateway that referenced this pull request Jul 16, 2026
## Summary

The `novita/deepseek-v3.2` reasoning e2e tests (`basic reasoning`,
`reasoning + streaming`) have been failing the aggregate e2e check on
every PR (seen on theopenco#3075, theopenco#3087, and others).

Root cause, verified live against Novita's API on 2026-07-16: Novita
changed behavior so that sending `reasoning_effort` **alongside** the
`chat_template_kwargs: { thinking: true }` flag suppresses reasoning
entirely (0 reasoning tokens, no `reasoning_content`), while
`reasoning_effort` on its own has no effect on this hybrid model. The
gateway sends both: `requiresEnableThinking` adds the template flag, and
the default OpenAI-compatible case forwards `reasoning_effort` raw
because the mapping declared no `supportedParameters`.

## Fix

Declare `supportedParameters` (without `reasoning_effort`) on the novita
mapping so the raw effort value is dropped and only the
`chat_template_kwargs.thinking` flag — the one control the deployment
actually honors — is forwarded. Metadata-only change, mirroring the
sibling deepseek mappings.

## Testing

- `TEST_MODELS="novita/deepseek-v3.2" FULL_MODE=true pnpm test:e2e` — 88
passed, 0 failed (previously the two reasoning tests failed)
- Verified live: template flag alone → 804 reasoning chars; flag +
`reasoning_effort` → 0
- models/actions unit suites pass

🤖 Generated with [Claude Code](https://claude.com/claude-code)

<!-- This is an auto-generated comment: release notes by coderabbit.ai
-->

## Summary by CodeRabbit

* **Bug Fixes**
* Improved compatibility with the Novita provider by restricting
requests to supported generation and control settings.
* Prevented unsupported reasoning controls from being forwarded,
ensuring reasoning behavior is handled consistently.

<!-- end of auto-generated comment: release notes by coderabbit.ai -->

Co-authored-by: Claude Fable 5 <noreply@anthropic.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

3 participants