Skip to content

fix(nvidia): disable thinking for minimaxai/minimax-m2.7 on NVIDIA NIM - #2323

Open
anki1kr wants to merge 10 commits into
decolua:masterfrom
anki1kr:fix/nvidia-minimax-thinking-2268
Open

anki1kr wants to merge 10 commits into
decolua:masterfrom
anki1kr:fix/nvidia-minimax-thinking-2268

Conversation

@anki1kr

@anki1kr anki1kr commented Jul 3, 2026

Copy link
Copy Markdown
Contributor

Problem

nvidia/minimaxai/minimax-m2.7 returns 400 "Unsupported parameter(s): thinking" because the request translator injects the MiniMax-native thinking body field, which NVIDIA's OpenAI-compatible proxy rejects.

Root Cause

getCapabilitiesForModel("nvidia", "minimaxai/minimax-m2.7"):

  1. No PROVIDER_CAPABILITIES["nvidia"] entry → step 1 misses
  2. baseModel = "minimax-m2.7" → no MODEL_CAPABILITIES["minimax-m2.7"] → step 2 misses
  3. Pattern *minimax-m2.7* matches → returns { reasoning: true, thinkingFormat: "minimax" }

applyThinking then executes applyFormat("minimax", ...) which sets:

body.thinking = { type: "adaptive" }

NVIDIA's NIM proxy does not forward the MiniMax thinking field and rejects it.

Fix

Add PROVIDER_CAPABILITIES["nvidia"] with minimaxai/minimax-m2.7: { reasoning: false }.

The provider-specific override runs first in getCapabilitiesForModel, so NVIDIA-hosted MiniMax models get reasoning: false → applyThinking calls stripAll() and removes any thinking field before forwarding to NVIDIA.

Direct MiniMax API access (without the nvidia provider) continues to use thinkingFormat: "minimax" from the pattern match.

Tests

tests/unit/nvidia-minimax-thinking.test.js — 4 tests:

  • NVIDIA + minimaxai/minimax-m2.7 → reasoning: false
  • NVIDIA + minimaxai/minimax-m2.7 → thinkingFormat ≠ "minimax"
  • Direct (no provider) minimaxai/minimax-m2.7 → reasoning: true, thinkingFormat: "minimax" (unaffected)
  • Provider override takes priority over pattern for NVIDIA

Fixes #2268

anki1kr added 10 commits July 2, 2026 17:32
…olua#2224)

On a fresh install or after deleting %APPDATA%\9router, rootCA.key and
rootCA.crt are absent. server.js tried to readFileSync them and called
process.exit(1) on failure, making the MITM server permanently broken
without an obvious fix path for the user.

Root cause: generateRootCA() existed in cert/rootCA.js but was never
called at startup — only accessible via the UI 'Generate CA' button,
which itself requires the server to already be running.

Fix:
- cert/rootCA.js: adds ensureRootCASync() — same logic as generateRootCA()
  but truly synchronous (node-forge RSA is blocking), callable at module
  load time in CommonJS scripts without an async wrapper
- server.js: calls ensureRootCASync() before the readFileSync block;
  if generation fails the error is reported and the process exits cleanly;
  also regenerates automatically when an existing cert is expired/corrupt

5 unit tests cover: generate-when-absent, idempotent-when-valid,
regenerate-when-corrupt, generate-when-partial (key only), mkdir-when-dir-missing.
…patch

Kiro / CodeWhisperer rejects Claude model IDs that use dot notation for
the version number (e.g. claude-sonnet-4.5) and requires dash notation
(claude-sonnet-4-5). The upstream dispatch was forwarding the 9router
dot-format ID directly, producing INVALID_MODEL_ID (HTTP 400) for every
Claude model while non-Claude models (deepseek-3.2, glm-5) succeeded.

Add toKiroModelId() to kiroConstants.js that replaces dots with dashes
in Claude model IDs only, leaving all other model IDs unchanged. Apply
the conversion in both OpenAI→Kiro and Claude→Kiro translators right
after resolveKiroModel() strips the synthetic suffixes.

Fixes decolua#2308. Related: decolua#2257.
…alls

When the MITM dashboard button is clicked multiple times rapidly, or when
auto-restart races with a manual start, two concurrent startServer() calls
could both pass the serverProcess check (process not yet spawned) and race
to spawn two MITM servers. The second spawn fails silently and leaves the
state in an inconsistent position.

Add a module-level mitmStarting boolean that is set to true at the start
of the spawn path and cleared in a finally block on all exit paths. A
second caller that arrives while mitmStarting is true receives an explicit
"MITM server is already starting" error rather than a silent race.

The flag is reset via finally{} so crashes or thrown errors during startup
never leave it permanently set, avoiding the permanent-lock scenario
described in issue decolua#2310.

Fixes decolua#2310. Related: decolua#2286.
Gemini's function_declarations schema validator rejects any property that
contains multipleOf, returning:

  400 Invalid JSON payload received. Unknown name "multipleOf" at
  'request.tools[0].function_declarations[N].parameters.properties[K].value'

multipleOf is a valid JSON Schema keyword (used to constrain numeric
fields to multiples of a value) but is not part of Gemini's subset of
supported schema properties. Add it to UNSUPPORTED_SCHEMA_CONSTRAINTS
alongside minLength/maxLength/exclusiveMinimum/exclusiveMaximum so it
is removed recursively from all function declaration schemas before
the request reaches the Gemini backend.

Fixes decolua#2309.
…ields before forwarding

openaiResponsesToOpenAIRequest already deleted input, instructions, store,
include, prompt_cache_key and reasoning when converting the Responses API
format to Chat Completions. But it left client_metadata, background and
truncation in the outgoing body.

When Codex sends a request containing client_metadata to a non-OpenAI
upstream (e.g. NVIDIA NIM), the upstream rejects with:
  "Validation: Unsupported parameter(s): `client_metadata`" (400)

Fix: add the three missing deletes to the cleanup section.

Closes decolua#2311
… conversion

When Claude Code routes through Kiro/CodeWhisperer, system messages were
being silently converted to plain user messages with no structural marker.
This caused the full Claude Code system prompt (env info, tool definitions,
memory instructions, billing headers, etc.) to appear as raw user text in
the Kiro conversation payload — leaking context, wasting tokens, and making
the model behave unpredictably.

Root cause: in convertMessages() the line
  if (role === ROLE.SYSTEM || role === ROLE.TOOL) role = ROLE.USER;
flattened both cases identically. No wrapping was applied to system content.

Fix: split the two cases. System messages get their text extracted and
wrapped in <system-reminder>…</system-reminder> before the role is changed
to user. Tool messages continue to be promoted to user role without wrapping
(tool output is already structured and doesn't need a provenance marker).

Closes decolua#2306
…ex parser

Two bugs in the google-tts provider caused 502 errors for input >~200 chars:

1. Google Translate TTS batchexecute silently returns null as the audio payload
   when the input text exceeds its ~200-char limit. The old code did:
     JSON.parse(split[0][2])[0]
   which became JSON.parse(null) → null → null[0] throwing
   "Cannot read properties of null (reading '0')".

2. The response parser used data.split("\n")[3] to find the JSON line, which
   is fragile and breaks when Google changes whitespace/line layout in the
   response body.

Fix:
- Split text into ≤190-char chunks at sentence ("." "?" "!") then word
  boundaries; synthesize each chunk separately; concatenate the MP3 buffers
  (MP3 frames are self-delimiting so buffer concatenation produces valid audio).
- Replace the line-index parser with a scan that looks for the line containing
  the "jQ1olc" rpcId, making it robust to response format changes.

Closes decolua#2287
…models

Vertex AI (Cloud Code Assist) rejects Claude requests that end with an
assistant message, returning:

  400 "This model does not support assistant message prefill.
       The conversation must end with a user message."

Clients can send a trailing assistant message as a prefill hint (OpenAI
supports this). When 9Router converts such a request to Claude format for
Antigravity (antigravity/claude-opus-4-6, claude-sonnet-5, etc.), the
trailing assistant message is preserved verbatim by openaiToClaudeRequest
and then forwarded to Vertex — which rejects it.

Fix: after all other Antigravity-specific cleanup in
openaiToClaudeRequestForAntigravity, strip any trailing assistant messages
so the conversation always ends with a user turn.

Closes decolua#2302
…stries

Registers Claude Sonnet 5 (claude-sonnet-5) across:
- providers/registry/claude.js — Claude Code OAuth provider (first in list)
- providers/registry/antigravity.js — Antigravity Vertex provider
- providers/capabilities.js — vision + reasoning + search + 1M context + 64K output
- providers/pricing.js — same tier as Sonnet 4.6 ($3/$15 per M tokens)

Closes decolua#2267
NVIDIA's OpenAI-compatible NIM endpoint for minimaxai/minimax-m2.7 rejects
requests with 400 "Unsupported parameter(s): thinking".

Root cause: the pattern *minimax-m2.7* in PATTERN_CAPABILITIES assigned
thinkingFormat:"minimax" to the model. applyThinking() then set
body.thinking = { type: "adaptive" } — the native MiniMax wire format —
before forwarding to NVIDIA. NVIDIA's proxy does not pass through the
MiniMax thinking field and rejects the parameter.

Fix: add a PROVIDER_CAPABILITIES["nvidia"] entry that marks
minimaxai/minimax-m2.7 as reasoning:false. getCapabilitiesForModel()
checks provider-specific overrides before pattern matching, so the
NVIDIA host gets thinking stripped rather than injected.

Direct MiniMax API access (no provider prefix or provider=minimax)
continues to use thinkingFormat:"minimax" from the pattern as before.

Fixes decolua#2268
Copilot AI review requested due to automatic review settings July 3, 2026 10:43

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Copilot was unable to review this pull request because the user who requested the review has reached their quota limit.

bloodf pushed a commit to bloodf/9router that referenced this pull request Jul 3, 2026
# Conflicts:
#	open-sse/providers/capabilities.js
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

API Error: 400

2 participants