Skip to content

fix(translator): pass output_config.effort=max through verbatim - #9053

Merged
diegosouzapw merged 1 commit into
diegosouzapw:release/v3.8.50from
ikelvingo:fix/claude-to-openai-max-passthrough
Aug 4, 2026
Merged

diegosouzapw merged 1 commit into
diegosouzapw:release/v3.8.50from
ikelvingo:fix/claude-to-openai-max-passthrough

Conversation

@ikelvingo

Copy link
Copy Markdown
Contributor

Summary

Fixes a contract mismatch where output_config.effort="max" from an Anthropic-format request was incorrectly rewritten to xhigh when translating to an OpenAI-compatible upstream. This caused upstreams that accept max literally but reject xhigh — such as Ollama Cloud — to return HTTP 400 invalid reasoning value: 'xhigh'.

Root Cause

open-sse/translator/request/claude-to-openai.ts::normalizeOpenAIReasoningEffort unconditionally rewrote max → xhigh during step 1 (translation). However, provider-aware allowlist decisions belong in sanitizeReasoningEffortForProvider (the executor layer), which holds the full provider capability table — including providers that explicitly opt in to supportsMaxEffortForProvider(): ollama-cloud, opencode-go deepseek, moonshot k3, and native claude.

Because the translator runs before the sanitizer, the carrier was already xhigh by the time the sanitizer saw it, so the effortStr === "max" branch could never fire — the max information was lost before reaching the sanitizer.

Fix

Removed the if (normalized === "max") return "xhigh" line in claude-to-openai.ts. The translator now only performs form conversion (lowercase + non-empty filtering) and no longer makes allowlist decisions; max is written verbatim to result.reasoning_effort. Also updated the top comment in reasoningEffort.ts (English) to document this division of responsibility.

Impact

  • Fix: Anthropic clients → OpenAI-shape upstreams that accept max literally (e.g. ollama-cloud) now pass max through unchanged instead of returning 400.
  • No regression: the sanitizer's effortStr === "max" branch already covered the "client sends max directly in OpenAI-format" case; it now also covers the new "translator passes max through" path.

Verification

  • tests/unit/translator-claude-to-openai.test.ts — assertion updated to expect max pass-through.
  • tests/unit/base-executor-sanitize-effort.test.ts — added end-to-end test: translateRequest(CLAUDE→OPENAI) + sanitizeReasoningEffortForProvider("ollama-cloud", …) chain keeps output_config.effort="max" as max end-to-end.

Repro (before fix)

curl -X POST http://localhost:20128/v1/messages \
  -H "Content-Type: application/json" \
  -d '{
    "model": "ollama-cloud/gemma4:31b",
    "max_tokens": 4096,
    "messages": [{"role":"user","content":"hi"}],
    "output_config": {"effort": "max"}
  }'
# HTTP 400 {"error":{"message":"invalid reasoning value: 'xhigh'
#   (must be \"high\", \"medium\", \"low\", \"max\", or \"none\")"}}

After fix: HTTP 200, reasoning_effort: "max" forwarded to ollama-cloud.

@ikelvingo
ikelvingo requested a review from diegosouzapw as a code owner July 31, 2026 10:22
The claude->openai translator was unconditionally rewriting max to xhigh, which broke any OpenAI-shape upstream that accepts max literally (e.g. ollama-cloud, opencode-go deepseek, moonshot k3, native Claude). Provider-aware effort policy is owned by sanitizeReasoningEffortForProvider in the executor; the translator should only do form conversion.

Regression guard: tests/unit/base-executor-sanitize-effort.test.ts end-to-end case (claude -> ollama-cloud preserves max).
@ikelvingo
ikelvingo force-pushed the fix/claude-to-openai-max-passthrough branch from b2d59f4 to e30b950 Compare July 31, 2026 10:22
@diegosouzapw

Copy link
Copy Markdown
Owner

Merged via local merge-train on 192.168.0.113 (32 cores) @ train tip e3a7696bb969536578f6a2defeada7d04fa9d1f9 — log /srv/omniroute-train/.claude/worktrees/merge-train-20260804-120425-suite.log.

Green: typecheck:core, check-complexity, check-cognitive-complexity, check-changelog-integrity, test:vitest (34 files / 291 tests).

Reds, each discriminated against the clean release tip 2e4268003 with zero PRs boarded (merge-gates §3):

  • check-file-size — src/sse/handlers/chat.ts 1846>1845 and open-sse/executors/base.ts 1623>1578 reproduce with identical numbers on the bare tip. Baseline drift for the release captain; not introduced here.
  • test:unit — 15 failures, all reproduced identically on the bare tip. Thirteen are the bare-model routing assertions (OpenAI remains the historical default, getModelInfoCore keeps unprefixed gpt-5.5 on the OpenAI fallback, …) left stale by fix(routing): bare model ids route to codex first; validate synced candidates #9275, which intentionally moved bare gpt-5.5/gpt-5.6-sol to codex without updating them; two are the trailing-period message change already covered by test(sse): expect the trailing period in the no-credentials message #9392.
  • provider-limits-local-apikey-sync-spacing — the only failure not on the bare tip, and a load flake: the assertion requires a ~0ms gap and measured 25/49ms under a 32-core full-suite run. Passes 3/3 in isolation on this exact train tip.

No red attributable to this PR.

@diegosouzapw
diegosouzapw merged commit c790b57 into diegosouzapw:release/v3.8.50 Aug 4, 2026
15 of 16 checks passed
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…osouzapw#9053)

The claude->openai translator was unconditionally rewriting max to xhigh, which broke any OpenAI-shape upstream that accepts max literally (e.g. ollama-cloud, opencode-go deepseek, moonshot k3, native Claude). Provider-aware effort policy is owned by sanitizeReasoningEffortForProvider in the executor; the translator should only do form conversion.

Regression guard: tests/unit/base-executor-sanitize-effort.test.ts end-to-end case (claude -> ollama-cloud preserves max).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants