Skip to content

fix(providers): learn reasoning_effort capability from upstream 4xx instead of a hardcoded/opt-out default - #11116

Merged
diegosouzapw merged 5 commits into
diegosouzapw:release/v3.8.50from
maxmad64bis:fix/reasoning-effort-capability-discovery
Aug 22, 2026
Merged

diegosouzapw merged 5 commits into
diegosouzapw:release/v3.8.50from
maxmad64bis:fix/reasoning-effort-capability-discovery

Conversation

@maxmad64bis

Copy link
Copy Markdown
Contributor

⚠️ base-red inherited: #9985

Fixes #11111

Summary

sanitizeReasoningEffortForProvider() (open-sse/executors/base/reasoningEffort.ts) only downgrades xhigh/max when the static registry (supportsXHighEffort) explicitly says a provider/model doesn't support it, and when it does downgrade, the target was hardcoded to "high" — never verified. Two gaps, one root cause: nothing ever checks what the upstream actually accepts.

  • Custom OpenAI-compatible connections (openai-compatible-chat-<uuid>) structurally can't have a registry entry, so xhigh reaches them unfiltered.
  • A registered provider/model with no reasoning metadata (e.g. ovhcloud's 3 declared models) has the same problem.
  • Even when a downgrade does trigger, "high" isn't always valid either (prior incident: Unexpected reasoning effort high. Supported types are xhigh, medium, low).

This mirrors the mechanism already merged for thinking_budget (open-sse/services/learnedThinkingCaps.ts, wired into base.ts:1492-1530): parse the upstream-advertised accepted values out of a 4xx body, store the highest one in a process-wide in-memory cache keyed provider:model, clamp the live request, retry once, and consult the cache proactively on every future request for that pair. No new persistence layer, no DB migration.

Changes

  1. open-sse/config/constants.ts — add HTTP_STATUS.UNPROCESSABLE_ENTITY = 422 (OVH's @ai-sdk/openai-compatible deserializer returns 422, not 400, for this rejection; the literal 422 was already used ad hoc in 4 other files but no named constant existed).
  2. open-sse/services/learnedReasoningEffortCaps.ts (new) — ordinal scale (none < minimal < low < medium < high < xhigh < max), parseReasoningEffortEnum() (generic, vendor-agnostic extraction from a 4xx body), recordLearnedReasoningEffort() / getLearnedReasoningEffort() (monotonically non-increasing, same anti-ratchet rule as learnedThinkingCaps.ts).
  3. open-sse/executors/base/reasoningEffort.ts — consult the learned cap before the hardcoded "high" fallback (both the xhigh and max branches), and let it override the registry entirely when present (covers the custom-connection case, since the registry's "supports xhigh" default is true).
  4. open-sse/executors/base.ts — reactive 400/422 clamp-and-retry, added right after the existing thinking-budget block it mirrors: parse the accepted-values enum, record it, re-run the sanitizer (now picking up the freshly learned cap), retry the same URL once.

What's deliberately untouched: the static registry stays as the free fast path for well-known providers (no wasted round-trip); the deepseek/mistral/github special cases (non-ordinal translation / removal rules) return early and never reach the new code.

Tests

  • tests/unit/http-status-unprocessable-entity.test.ts (new)
  • tests/unit/learned-reasoning-effort-caps.test.ts (new, 14 cases — ordinal scale, enum parsing on both prose shapes observed, record/get, monotonic decrease, case-insensitive keying)
  • tests/unit/reasoning-effort-learned-capability.test.ts (new, 6 cases — proactive clamp behavior for custom/unregistered and registry-covered providers, deepseek untouched)
  • tests/unit/reasoning-effort-clamp-and-retry.test.ts (new, 2 cases — reactive 422 clamp-and-retry against the real OVH error text, and the learned value being sent on the first try afterward)

Regression: re-ran every pre-existing reasoning-effort and thinking-budget test file (base-reasoning-effort-split, github-claude-reasoning-effort-granular, opencode-zen-reasoning-effort, moonshot-k3, ollama-cloud-reasoning-effort-tiers-10788, duckduckgo-reasoning-effort-required, deepseek-thinking-efforts, sensenova-reasoning-effort, gemini-thinking-budget-fallback, cap-thinking-budget-gemini-fallback, learned-thinking-caps) — all pass unmodified.

Commands run

node --import tsx/esm --test tests/unit/http-status-unprocessable-entity.test.ts tests/unit/learned-reasoning-effort-caps.test.ts tests/unit/reasoning-effort-learned-capability.test.ts tests/unit/reasoning-effort-clamp-and-retry.test.ts tests/unit/base-reasoning-effort-split.test.ts tests/unit/github-claude-reasoning-effort-granular.test.ts tests/unit/opencode-zen-reasoning-effort.test.ts tests/unit/moonshot-k3.test.ts tests/unit/ollama-cloud-reasoning-effort-tiers-10788.test.ts tests/unit/duckduckgo-reasoning-effort-required.test.ts tests/unit/deepseek-thinking-efforts.test.ts tests/unit/sensenova-reasoning-effort.test.ts tests/unit/gemini-thinking-budget-fallback.test.ts tests/unit/cap-thinking-budget-gemini-fallback.test.ts tests/unit/learned-thinking-caps.test.ts
# 108/108 pass

npx eslint open-sse/config/constants.ts open-sse/services/learnedReasoningEffortCaps.ts open-sse/executors/base/reasoningEffort.ts open-sse/executors/base.ts tests/unit/http-status-unprocessable-entity.test.ts tests/unit/learned-reasoning-effort-caps.test.ts tests/unit/reasoning-effort-learned-capability.test.ts tests/unit/reasoning-effort-clamp-and-retry.test.ts
# 0 errors

@diegosouzapw
diegosouzapw merged commit 5631e91 into diegosouzapw:release/v3.8.50 Aug 22, 2026
9 of 16 checks passed
@maxmad64bis
maxmad64bis deleted the fix/reasoning-effort-capability-discovery branch September 24, 2026 21:15
muhamadgalihsaputra pushed a commit to niyatna/NiyatnaRoute that referenced this pull request Sep 27, 2026
…nstead of a hardcoded/opt-out default (diegosouzapw#11116)

Validated on the combined batch board over release/v3.8.50 tip 0f43f0f: static gates clean, typecheck:core clean, focused tests green.

Learned reasoning_effort caps mirror the merged learnedThinkingCaps mechanism: parse the upstream 4xx enum, clamp, retry once, consult proactively — covers custom openai-compatible connections the static registry can't. 22 new test cases + full regression list green. Fixes diegosouzapw#11111. Thank you @maxmad64bis!
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

2 participants