fix(providers): learn reasoning_effort capability from upstream 4xx instead of a hardcoded/opt-out default - #11116
Merged
diegosouzapw merged 5 commits intoAug 22, 2026
Conversation
added 4 commits
August 22, 2026 12:02
diegosouzapw
merged commit Aug 22, 2026
5631e91
into
diegosouzapw:release/v3.8.50
9 of 16 checks passed
4 of 5 tasks
muhamadgalihsaputra
pushed a commit
to niyatna/NiyatnaRoute
that referenced
this pull request
Sep 27, 2026
…nstead of a hardcoded/opt-out default (diegosouzapw#11116) Validated on the combined batch board over release/v3.8.50 tip 0f43f0f: static gates clean, typecheck:core clean, focused tests green. Learned reasoning_effort caps mirror the merged learnedThinkingCaps mechanism: parse the upstream 4xx enum, clamp, retry once, consult proactively — covers custom openai-compatible connections the static registry can't. 22 new test cases + full regression list green. Fixes diegosouzapw#11111. Thank you @maxmad64bis!
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Fixes #11111
Summary
sanitizeReasoningEffortForProvider()(open-sse/executors/base/reasoningEffort.ts) only downgradesxhigh/maxwhen the static registry (supportsXHighEffort) explicitly says a provider/model doesn't support it, and when it does downgrade, the target was hardcoded to"high"— never verified. Two gaps, one root cause: nothing ever checks what the upstream actually accepts.openai-compatible-chat-<uuid>) structurally can't have a registry entry, soxhighreaches them unfiltered.ovhcloud's 3 declared models) has the same problem."high"isn't always valid either (prior incident:Unexpected reasoning effort high. Supported types are xhigh, medium, low).This mirrors the mechanism already merged for
thinking_budget(open-sse/services/learnedThinkingCaps.ts, wired intobase.ts:1492-1530): parse the upstream-advertised accepted values out of a 4xx body, store the highest one in a process-wide in-memory cache keyedprovider:model, clamp the live request, retry once, and consult the cache proactively on every future request for that pair. No new persistence layer, no DB migration.Changes
open-sse/config/constants.ts— addHTTP_STATUS.UNPROCESSABLE_ENTITY = 422(OVH's@ai-sdk/openai-compatibledeserializer returns 422, not 400, for this rejection; the literal422was already used ad hoc in 4 other files but no named constant existed).open-sse/services/learnedReasoningEffortCaps.ts(new) — ordinal scale (none < minimal < low < medium < high < xhigh < max),parseReasoningEffortEnum()(generic, vendor-agnostic extraction from a 4xx body),recordLearnedReasoningEffort()/getLearnedReasoningEffort()(monotonically non-increasing, same anti-ratchet rule aslearnedThinkingCaps.ts).open-sse/executors/base/reasoningEffort.ts— consult the learned cap before the hardcoded"high"fallback (both thexhighandmaxbranches), and let it override the registry entirely when present (covers the custom-connection case, since the registry's "supports xhigh" default istrue).open-sse/executors/base.ts— reactive 400/422 clamp-and-retry, added right after the existing thinking-budget block it mirrors: parse the accepted-values enum, record it, re-run the sanitizer (now picking up the freshly learned cap), retry the same URL once.What's deliberately untouched: the static registry stays as the free fast path for well-known providers (no wasted round-trip); the
deepseek/mistral/githubspecial cases (non-ordinal translation / removal rules) return early and never reach the new code.Tests
tests/unit/http-status-unprocessable-entity.test.ts(new)tests/unit/learned-reasoning-effort-caps.test.ts(new, 14 cases — ordinal scale, enum parsing on both prose shapes observed, record/get, monotonic decrease, case-insensitive keying)tests/unit/reasoning-effort-learned-capability.test.ts(new, 6 cases — proactive clamp behavior for custom/unregistered and registry-covered providers, deepseek untouched)tests/unit/reasoning-effort-clamp-and-retry.test.ts(new, 2 cases — reactive 422 clamp-and-retry against the real OVH error text, and the learned value being sent on the first try afterward)Regression: re-ran every pre-existing reasoning-effort and thinking-budget test file (
base-reasoning-effort-split,github-claude-reasoning-effort-granular,opencode-zen-reasoning-effort,moonshot-k3,ollama-cloud-reasoning-effort-tiers-10788,duckduckgo-reasoning-effort-required,deepseek-thinking-efforts,sensenova-reasoning-effort,gemini-thinking-budget-fallback,cap-thinking-budget-gemini-fallback,learned-thinking-caps) — all pass unmodified.Commands run