fix(openai): strip reasoning_effort when GPT-5.x models carry function tools - #7101
Conversation
…from 9router#2540)
Raw api.openai.com Chat Completions rejects GPT-5.x reasoning models that carry both function tools and an active reasoning_effort with HTTP 400 ("Function tools with reasoning_effort are not supported ... Please use /v1/responses instead"). The existing forceResponsesUpstream guard only reroutes openai-compatible-* connections carrying MCP/tool_search tool shapes; the plain openai provider had no equivalent guard, so gpt-5.x models used with a coding client (function tools + any explicit reasoning effort) still hit the upstream 400. Add stripGpt5ReasoningWhenTools() (gpt5SamplingGuard.ts), wired into chatCore.ts alongside the existing sampling guard, to drop reasoning_effort/reasoning when function tools are present and reasoning is active, letting the request succeed on /v1/chat/completions.
Reported-by: Tech Solution (@techsolutionmta) (decolua/9router#2540)
|
Warning You have reached your daily quota limit. Please wait up to 24 hours and I will start processing your requests again! |
|
You have reached your Codex usage limits for code reviews. You can see your limits in the Codex usage dashboard. |
|
Validado no probe: revertendo gpt5SamplingGuard.ts o import quebra (função não existe) e os 9 testes caem; com o fix, todos passam. Wiring em chatCore.ts consistente com o guard irmão (stripGpt5SamplingWhenReasoning), escopo corretamente restrito a provider==='openai'+gpt-5*. Único ponto de atenção (não bloqueante): ao stripar o campo aninhado 'reasoning' (vs. o 'reasoning_effort' string), o código apaga o objeto inteiro — se ele algum dia carregar sub-campos além de 'effort' eles somem junto. É o mesmo padrão já usado no guard irmão, então não é regressão introduzida por este PR — só deixo registrado para avaliação futura. Merge-ready. |
stripGpt5ReasoningWhenTools gated on provider+model-name alone, so once #7242 routes the public GPT-5.6 family to /v1/responses (targetFormat "openai-responses", which natively supports tools + reasoning), the two PRs would compose into the worst of both worlds: routed to the endpoint that supports reasoning, but reasoning stripped anyway. Pass the request's already-resolved targetFormat into the guard and skip the strip whenever it is not going out over /chat/completions, so the guard tracks the actual upstream surface instead of a model-name list that would need updating for every future GPT-5.x family. Reported-by: Tech Solution (@techsolutionmta) (decolua/9router#2540)
…n tools (diegosouzapw#7101) * fix(openai): strip reasoning_effort when GPT-5.x tools present (port from 9router#2540) Raw api.openai.com Chat Completions rejects GPT-5.x reasoning models that carry both function tools and an active reasoning_effort with HTTP 400 ("Function tools with reasoning_effort are not supported ... Please use /v1/responses instead"). The existing forceResponsesUpstream guard only reroutes openai-compatible-* connections carrying MCP/tool_search tool shapes; the plain openai provider had no equivalent guard, so gpt-5.x models used with a coding client (function tools + any explicit reasoning effort) still hit the upstream 400. Add stripGpt5ReasoningWhenTools() (gpt5SamplingGuard.ts), wired into chatCore.ts alongside the existing sampling guard, to drop reasoning_effort/reasoning when function tools are present and reasoning is active, letting the request succeed on /v1/chat/completions. Reported-by: Tech Solution (@techsolutionmta) (decolua/9router#2540) * fix(openai): scope reasoning-strip guard to /chat/completions only stripGpt5ReasoningWhenTools gated on provider+model-name alone, so once diegosouzapw#7242 routes the public GPT-5.6 family to /v1/responses (targetFormat "openai-responses", which natively supports tools + reasoning), the two PRs would compose into the worst of both worlds: routed to the endpoint that supports reasoning, but reasoning stripped anyway. Pass the request's already-resolved targetFormat into the guard and skip the strip whenever it is not going out over /chat/completions, so the guard tracks the actual upstream surface instead of a model-name list that would need updating for every future GPT-5.x family. Reported-by: Tech Solution (@techsolutionmta) (decolua/9router#2540)
…n tools (diegosouzapw#7101) * fix(openai): strip reasoning_effort when GPT-5.x tools present (port from 9router#2540) Raw api.openai.com Chat Completions rejects GPT-5.x reasoning models that carry both function tools and an active reasoning_effort with HTTP 400 ("Function tools with reasoning_effort are not supported ... Please use /v1/responses instead"). The existing forceResponsesUpstream guard only reroutes openai-compatible-* connections carrying MCP/tool_search tool shapes; the plain openai provider had no equivalent guard, so gpt-5.x models used with a coding client (function tools + any explicit reasoning effort) still hit the upstream 400. Add stripGpt5ReasoningWhenTools() (gpt5SamplingGuard.ts), wired into chatCore.ts alongside the existing sampling guard, to drop reasoning_effort/reasoning when function tools are present and reasoning is active, letting the request succeed on /v1/chat/completions. Reported-by: Tech Solution (@techsolutionmta) (decolua/9router#2540) * fix(openai): scope reasoning-strip guard to /chat/completions only stripGpt5ReasoningWhenTools gated on provider+model-name alone, so once diegosouzapw#7242 routes the public GPT-5.6 family to /v1/responses (targetFormat "openai-responses", which natively supports tools + reasoning), the two PRs would compose into the worst of both worlds: routed to the endpoint that supports reasoning, but reasoning stripped anyway. Pass the request's already-resolved targetFormat into the guard and skip the strip whenever it is not going out over /chat/completions, so the guard tracks the actual upstream surface instead of a model-name list that would need updating for every future GPT-5.x family. Reported-by: Tech Solution (@techsolutionmta) (decolua/9router#2540)
Summary
Raw
api.openai.comChat Completions rejects GPT-5.x reasoning models that carry both functiontoolsand an activereasoning_effortwith HTTP 400:This is hit by any coding client (e.g. Claude Code via an OpenAI-compatible bridge) that sends tool calls alongside an explicit "Thinking" level. The dashboard offers no
reasoning_effort:"none"override, so it couldn't be worked around client-side either.Root cause
OmniRoute already has a guard for a related case:
shouldForceResponsesUpstream(open-sse/executors/forceResponsesUpstream.ts) reroutesopenai-compatible-*connections carrying MCP/tool_search*tool shapes onto/v1/responses. But that guard only fires foropenai-compatible-*providers — the plain, built-inopenaiprovider (rawapi.openai.comChat Completions) has no equivalent guard, so GPT-5.x + function tools + active reasoning still reached the upstream 400 on/v1/chat/completions.Fix
New
stripGpt5ReasoningWhenTools()inopen-sse/services/gpt5SamplingGuard.ts(alongside the existingstripGpt5SamplingWhenReasoning), wired intochatCore.tsright after the sampling guard: for theopenaiprovider +gpt-5*models, when the request carries a non-empty functiontoolsarray AND an activereasoning_effort/reasoning.effort(anything other than"none"), it strips the reasoning field(s) so the request succeeds on/v1/chat/completionsinstead of 400ing.Scoped narrowly:
provider === "openai"only (raw Chat Completions surface) — other providers/executors are untouched.gpt-5*models only.reasoning_effort:"none"and tool-less requests pass through unchanged.Test plan
tests/unit/gpt5-tools-reasoning-guard.test.tswritten first, confirmed failing (stripGpt5ReasoningWhenToolsdid not exist) before the fix, now passing (9/9).tests/unit/gpt5-sampling-guard.test.tsstill green (10/10) — no regression to the sibling sampling guard.npm run typecheck:core— clean.npx eslint open-sse/services/gpt5SamplingGuard.ts open-sse/handlers/chatCore.ts tests/unit/gpt5-tools-reasoning-guard.test.ts— clean.Thanks @techsolutionmta for the detailed report.