Skip to content

fix(gateway): forward reasoning_effort to Together AI - #3367

Merged
steebchen merged 1 commit into
mainfrom
fix/together-reasoning-effort
Aug 2, 2026
Merged

steebchen merged 1 commit into
mainfrom
fix/together-reasoning-effort

Conversation

@steebchen

@steebchen steebchen commented Aug 1, 2026 •

Copy link
Copy Markdown
Member

Fixes #3361.

The bug

case "inference.net": case "together-ai": in packages/actions/src/prepare-request-body.ts forwarded response_format, temperature, max_tokens, top_p and the penalties, then breaked before the default case's generic reasoning_effort forwarder. reasoning_effort was therefore never sent to Together AI for any model — including the gateway's own auto-routing default and deepseek-v4-pro's declared reasoningEfforts: ["high", "max"], which was dead on arrival.

What the upstream API actually does

Probed live against api.together.ai (not from docs). Together is OpenAI-compatible on reasoning_effort, but disabling thinking is not uniform — reasoning models run on two serving stacks with different switches:

Mapping reasoning_effort Disable switch
openai/gpt-oss-120b, openai/gpt-oss-20b validated, 400s outside low/medium/high n/a (none is rejected)
google/gemma-4-31b-it validated; the 400 names the literals ('none', 'low', 'medium', 'high') reasoning_effort: "none" → 0 reasoning tokens
deepseek-ai/DeepSeek-V4-Pro accepted unvalidated; xhigh/max roughly double reasoning tokens, low/medium/high land on the default thinking: { type: "disabled" }
MiniMaxAI/MiniMax-M3, moonshotai/Kimi-K2.6, moonshotai/Kimi-K3 accepted unvalidated; no tier measurably changes reasoning length thinking: { type: "disabled" }

The two switches are not interchangeable: Gemma ignores thinking (5 runs, reasoning continued at ~1.5k tokens), and DeepSeek V4 Pro / Kimi K2.6 ignore reasoning_effort: "none" (still reasoned). thinking.type is validated by the second stack — a bad value 400s with unknown variant ... expected one of enabled, disabled, adaptive.

The fix

  • Forward reasoning_effort verbatim from the together-ai case, per the repo's no-downgrade rule.
  • Add together-ai to handlesNoneNatively so none survives to the switch.
  • Translate none per serving stack, driven by a new per-mapping requiresDisableThinkingParam flag (sibling to the existing requiresEnableThinking), so the DeepSeek/MiniMax/Kimi mappings emit thinking: { type: "disabled" } and Gemma emits reasoning_effort: "none".
  • Set each mapping's reasoningEfforts from measured behaviour rather than assumption, replacing deepseek-v4-pro's unverified ["high", "max"].

Testing

  • 6 new unit tests in prepare-request-body.spec.ts covering both stacks, the none translation, and the drop when a mapping does not declare none. pnpm test:unit for packages/actions + packages/models: 546 passed.
  • pnpm build and pnpm format clean.
  • E2E against the real Together API (TEST_MODELS=... FULL_MODE=true pnpm test:e2e), which expands one case per declared effort tier. All pass: gpt-oss-120b/20b low/medium/high, gemma none/low/medium/high, deepseek-v4-pro none/xhigh/max, and none for minimax-m3, kimi-k2.6, kimi-k3.

Pre-existing failures (not from this change)

together-ai/gemma-4-31b-it times out at 60s intermittently across unrelated suites (JSON output, streaming, tool calls) and on the plain no-parameter baseline when probed directly — its effort tiers each pass in one run and time out in another, in a different combination each time. together-ai/gpt-oss-120b fails tool-calling and Responses tool-calling. Neither of those suites sends reasoning_effort at all.

Out of scope

moonshotai/Kimi-K2.5 and zai-org/GLM-4.7 on Together return Unable to access non-serverless model for every request — those mappings need a dedicated endpoint and appear unusable as configured. Left untouched here; worth a separate look.

🤖 Generated with Claude Code

Summary by CodeRabbit

  • New Features
    • Added model-specific reasoning effort support for Together AI models, including supported none, low, medium, high, xhigh, and max tiers where applicable.
    • Added support for disabling reasoning on compatible models.
    • Added Together AI configurations for additional Kimi models.
  • Bug Fixes
    • Improved request handling so reasoning settings are correctly forwarded, translated, or omitted based on model capabilities.

The together-ai case in prepare-request-body never wrote reasoning_effort
and broke before the default case's generic forwarder, so the parameter
was silently dropped for every Together mapping — including the gateway's
own auto-routing default. deepseek-v4-pro declared reasoningEfforts for
together-ai that could never take effect.

Verified live against api.together.ai: Together is OpenAI-compatible on
reasoning_effort, but disabling thinking is not uniform. The gpt-oss and
Gemma deployments validate reasoning_effort (naming their accepted
literals in the 400) and take "none" there, while thinking is ignored.
The DeepSeek V4 Pro, MiniMax M3 and Kimi deployments accept any
reasoning_effort string without complaint and keep thinking on; only
thinking: { type: "disabled" } turns it off. Mappings on the second stack
now set requiresDisableThinkingParam.

Each mapping's reasoningEfforts is set from measured behaviour:
upstream-validated sets for gpt-oss (low/medium/high) and Gemma
(none/low/medium/high), and for the rest the tiers that measurably change
the reasoning length (deepseek-v4-pro: none/xhigh/max; minimax-m3 and
kimi-k2.6/k3: none only).

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Copilot AI review requested due to automatic review settings August 1, 2026 19:57
@chatgpt-codex-connector

Copy link
Copy Markdown

You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard.

@coderabbitai

coderabbitai Bot commented Aug 1, 2026 •

Copy link
Copy Markdown
Contributor

Review Change Stack

Walkthrough

Together AI request construction now forwards supported reasoning_effort values. Model mappings define supported tiers and whether disabling requires thinking: { type: "disabled" }. Tests cover supported, unsupported, disabled, and absent effort values.

Changes

Together AI reasoning

Layer / File(s) Summary
Provider reasoning mappings
packages/models/src/models.ts, packages/models/src/models/deepseek.ts, packages/models/src/models/google.ts, packages/models/src/models/minimax.ts, packages/models/src/models/moonshot.ts, packages/models/src/models/openai.ts
Mappings now declare Together AI reasoning tiers and the requiresDisableThinkingParam flag.
Together AI request translation
packages/actions/src/prepare-request-body.ts
Together AI preserves supported none values, forwards graded tiers, and uses either reasoning_effort: "none" or thinking: { type: "disabled" }.
Reasoning request coverage
packages/actions/src/prepare-request-body.spec.ts
Tests cover tier forwarding, disabled reasoning, unsupported efforts, route-specific behavior, and missing effort values.

Estimated code review effort: 3 (Moderate) | ~20 minutes

Possibly related PRs

Suggested reviewers: copilot

Sequence Diagram(s)

sequenceDiagram
  participant Client
  participant prepareRequestBody
  participant TogetherAI
  Client->>prepareRequestBody: Send reasoning_effort
  prepareRequestBody->>TogetherAI: Forward effort or thinking.disabled
  TogetherAI-->>Client: Return model response
Loading
🚥 Pre-merge checks | ✅ 5
✅ Passed checks (5 passed)
Check name Status Explanation
Description Check ✅ Passed Check skipped - CodeRabbit’s high-level summary is enabled.
Title check ✅ Passed The title clearly summarizes the main change: forwarding reasoning_effort to Together AI.
Linked Issues check ✅ Passed The changes forward reasoning_effort for Together AI and update model support and disabling mappings as required by issue #3361.
Out of Scope Changes check ✅ Passed The tests, request handling, provider mapping flag, and model declarations are directly related to the Together AI reasoning_effort fix.
Docstring Coverage ✅ Passed No functions found in the changed files to evaluate docstring coverage. Skipping docstring coverage check.
✨ Finishing Touches
📝 Generate docstrings
  • Create stacked PR
  • Commit on current branch
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/together-reasoning-effort

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@coderabbitai coderabbitai Bot left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

🧹 Nitpick comments (1)
packages/actions/src/prepare-request-body.spec.ts (1)

4811-4842: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick win

Replace the broad any assertion.

Use a narrow request-body type for reasoning_effort and thinking. The tests only require these fields.

Proposed fix
-			) as Promise<any>;
+			) as Promise<{
+				reasoning_effort?: string;
+				thinking?: { type: "disabled" };
+			}>;

As per coding guidelines, **/*.{ts,tsx} must not use any or as any unless absolutely necessary.

🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

In `@packages/actions/src/prepare-request-body.spec.ts` around lines 4811 - 4842,
Replace the broad `as Promise<any>` assertion in togetherReasoning with a narrow
request-body type containing only reasoning_effort and thinking, and use that
type for the promise assertion. Preserve the existing test inputs and behavior
while removing any usage.

Source: Coding guidelines

🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.

Nitpick comments:
In `@packages/actions/src/prepare-request-body.spec.ts`:
- Around line 4811-4842: Replace the broad `as Promise<any>` assertion in
togetherReasoning with a narrow request-body type containing only
reasoning_effort and thinking, and use that type for the promise assertion.
Preserve the existing test inputs and behavior while removing any usage.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Repository UI

Review profile: CHILL

Plan: Pro Plus

Run ID: 77218cf2-765a-4938-a7dc-0f7262433b85

📥 Commits

Reviewing files that changed from the base of the PR and between 1b5a083 and 16e82e7.

📒 Files selected for processing (8)
  • packages/actions/src/prepare-request-body.spec.ts
  • packages/actions/src/prepare-request-body.ts
  • packages/models/src/models.ts
  • packages/models/src/models/deepseek.ts
  • packages/models/src/models/google.ts
  • packages/models/src/models/minimax.ts
  • packages/models/src/models/moonshot.ts
  • packages/models/src/models/openai.ts

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

This PR fixes Together AI support for the OpenAI-compatible reasoning_effort parameter in the gateway request builder, and updates Together-specific model catalog metadata so reasoning_effort: "none" can correctly disable thinking per Together’s serving stack behavior.

Changes:

  • Forward reasoning_effort for together-ai (and translate "none" to either reasoning_effort: "none" or thinking: { type: "disabled" } depending on mapping).
  • Add requiresDisableThinkingParam to provider mappings and set Together mappings’ reasoningEfforts based on observed Together behavior.
  • Add unit tests covering Together’s “none” translation and forwarding behavior.

Reviewed changes

Copilot reviewed 8 out of 8 changed files in this pull request and generated 4 comments.

Show a summary per file
File Description
packages/actions/src/prepare-request-body.ts Adds Together-specific forwarding/translation for reasoning_effort, and treats Together as a provider that must preserve "none" for downstream handling.
packages/actions/src/prepare-request-body.spec.ts Adds unit tests asserting Together reasoning_effort forwarding and "none" translation behavior.
packages/models/src/models.ts Adds requiresDisableThinkingParam to the mapping interface to drive Together “disable thinking” translation behavior.
packages/models/src/models/openai.ts Declares Together reasoningEfforts for gpt-oss mappings (validated low/medium/high).
packages/models/src/models/google.ts Declares Together reasoningEfforts for Gemma mapping (validated none/low/medium/high).
packages/models/src/models/deepseek.ts Updates Together DeepSeek-V4-Pro reasoningEfforts and sets requiresDisableThinkingParam.
packages/models/src/models/minimax.ts Declares Together MiniMax-M3 reasoningEfforts and sets requiresDisableThinkingParam.
packages/models/src/models/moonshot.ts Declares Together Kimi K2.6/K3 reasoningEfforts and sets requiresDisableThinkingParam.
Suppressed comments (1)

packages/models/src/models/moonshot.ts:708

  • The comment says Together's deployment "accepts any reasoning_effort string without validating it", but reasoningEfforts is exposed via /v1/models as the exact accepted reasoning_effort values for a mapping (apps/gateway/src/models/models.ts:82-90). Listing only ["none"] here is likely to under-report what the provider will accept, which can mislead clients relying on the model catalog. Consider either omitting reasoningEfforts (unknown/unbounded), or enumerating the accepted values and clarifying in the comment which tiers actually change behavior.
				// Together's deployment accepts any reasoning_effort string without
				// validating it, and no tier measurably changes the reasoning length,
				// so `none` — honoured through the `thinking` switch rather than
				// reasoning_effort — is the only effort this mapping really applies.
				reasoningEfforts: ["none"],
				requiresDisableThinkingParam: true,

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

// (which they validate) actually turns it off. Mappings on the second
// stack set `requiresDisableThinkingParam`; either way only mappings
// that declare `none` in `reasoningEfforts` can disable at all.
if (supportsReasoning && reasoning_effort !== undefined) {
Comment on lines +51 to +56
// Together's deployment accepts any reasoning_effort string without
// validating it, and no tier measurably changes the reasoning length,
// so `none` — honoured through the `thinking` switch rather than
// reasoning_effort — is the only effort this mapping really applies.
reasoningEfforts: ["none"],
requiresDisableThinkingParam: true,
Comment on lines +468 to +473
// Together's deployment accepts any reasoning_effort string without
// validating it, and no tier measurably changes the reasoning length,
// so `none` — honoured through the `thinking` switch rather than
// reasoning_effort — is the only effort this mapping really applies.
reasoningEfforts: ["none"],
requiresDisableThinkingParam: true,
Comment on lines +319 to +325
// Together's deployment accepts any reasoning_effort string without
// validating it, and only the top tiers measurably change behaviour:
// xhigh and max roughly double the reasoning tokens, while
// low/medium/high land on the provider default. `none` is honoured
// through the `thinking` switch, not through reasoning_effort.
reasoningEfforts: ["none", "xhigh", "max"],
requiresDisableThinkingParam: true,
@steebchen
steebchen added this pull request to the merge queue Aug 2, 2026
Merged via the queue into main with commit 119b3c7 Aug 2, 2026
13 checks passed
@steebchen
steebchen deleted the fix/together-reasoning-effort branch August 2, 2026 02:51
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: reasoning_effort is silently discarded for Together AI. Parameter mismatch.

2 participants