fix(aux): title generation survives custom endpoints that reject reasoning_effort (#112781, supersedes #112789) - #113114
Merged
Merged
Conversation
…rejects them
Auxiliary title generation disables reasoning (reasoning_config
{"enabled": False}, #91927); on provider=custom the profile encodes that as
top-level reasoning_effort="none" - the deliberate thinking-off wire for
Ollama /v1 (#25758), vLLM and GLM. A chat-only model behind an
OpenAI-compatible relay (gpt-4.1-mini on a one-api style relay) answers
"400 Unrecognized request argument supplied: reasoning_effort" and the
title was lost with no retry (#112781).
Add a rung to the shared aux recovery ladder (sync and async drive the
same generator): when the 400 names a reasoning field, strip top-level
reasoning_effort, the adapter's private _reasoning_config and every
extra_body reasoning key, and retry once - the same reactive shape as the
temperature and response_format rungs. The custom profile's encoding is
untouched, so Ollama/vLLM/GLM users keep thinking-off on the first
request; only routes that reject the field pay one extra round-trip.
Supersedes #112789 (@KoNit-K), which dropped the encoding for every
non-Ollama custom endpoint in the shared profile and would have silently
re-enabled thinking for vLLM/GLM users who set reasoning_effort: none.
૮ >ﻌ< ა ci reviewran on 2f2e059 — fix(aux): only treat a standalone reasoning field name as a
|
… rejection; let parameter rungs chain Review follow-up on #113114. `_is_reasoning_field_rejection` fired on any 400 whose text contained "reasoning"/"think" next to a generic unsupported marker, so a route-gating 400 naming a thinking model ("The model kimi-k2-thinking is not supported when using this account") spent a strip-retry and then re-raised from `_param_rung_accepts` — the configured fallback chain, which main consulted for that error, was never reached. Require the token to be a standalone wire-field name (not a model-id segment, not "... with reasoning models"), and let `_param_rung_accepts` accept model-incompatible, reasoning-field and structured-output rejections so a temperature-strip retry that 400s on `reasoning_effort` reaches the reasoning rung and a gating 400 after any strip still reaches provider fallback. Also mirrors the new retry paragraph into the zh-Hans configuration page.
9 tasks done
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Auxiliary title generation no longer dies with
400 Unrecognized request argument supplied: reasoning_efforton custom/OpenAI-compatible endpoints whose model does not accept the field — the call is retried once with every reasoning field omitted and the title lands.Changes
agent/auxiliary_client.py— new rung in the shared aux recovery ladder (_ladder_parameter_rungs; synccall_llmandasync_call_llmdrive the same generator, so one rung covers both): when a 400/422 names a reasoning field (reasoning_effort,reasoning,thinking/think) with an unsupported/unrecognized marker, strip top-levelreasoning_effort, the adapter-private_reasoning_configand everyextra_bodyreasoning key, then retry once. Same reactive shape as the existingtemperatureandresponse_formatrungs; an unrelated 400 or a request that carried no reasoning field never retries.plugins/model-providers/custom/__init__.py— untouched:reasoning_effort: "none"remains the deliberate thinking-off encoding for Ollama/v1([Bug]: agent.reasoning_effort: none silently ignored on Ollama — main agent stuck in medium mode, bg-review fork can spiral (up to 65k tokens / 28 min) #25758), vLLM, SGLang and GLM/ARK routes, on the first request, for aux and main-agent calls alike.tests/agent/test_auxiliary_reasoning_field_rejection_retry.py— 2 invariants (sync/async retry with the reporter's verbatim 400; unrelated 400 does not strip). Red onorigin/main(2 failed / 1 passed), green on this head.website/docs/user-guide/configuration.md— one paragraph under the auxiliaryreasoning_effortknob describing the retry.Validation
Live probe (monkeypatched
openaiCompletions.create/AsyncCompletions.createacting as a relay that 400s onreasoning_effort, realcall_llm/async_call_llmwithprovider=custom,model=gpt-4.1-mini,reasoning_config={"enabled": False}, tempHERMES_HOME):origin/main9796235), sync + async{"model": "gpt-4.1-mini", "reasoning_effort": "none", "temperature": 0.3, "extra_body": {"response_format": …}}BadRequestError 400 … reasoning_effort— title lost{"model": "gpt-4.1-mini", "temperature": 0.3, "extra_body": {"response_format": …}}{"title": "Relay Title Probe"}Tests:
scripts/run_tests.shon the new file +test_structured_output_rejection_retry.py,test_unsupported_parameter_retry.py,test_unsupported_temperature_retry.py,test_auxiliary_client.py,test_injected_param_strip_retry_registry.py,test_auxiliary_auth_rung_fallthrough.py→ 287 passed, 0 failed.Root cause
title_generator.generatepassesreasoning_config={"enabled": False}(#91927);CustomProfile.build_api_kwargs_extrasprojects that as top-levelreasoning_effort: "none"for every custom endpoint, and because the profile handles reasoning there was no fallback or recovery when a chat-only model behind a relay rejects the field.Fixes #112781
Supersedes #112789 (@KoNit-K) — that PR restricts the
reasoning_effort: "none"encoding to Ollama URLs inside the shared custom profile, which would silently re-enable thinking for vLLM/SGLang/GLM users who setreasoning_effort: none(main agent included, not only the title lane). The recovery here is provider-agnostic and leaves the deliberate encoding in place; thanks to @KoNit-K for the report-to-test mapping.Infographic
Review follow-up
_is_reasoning_field_rejectionfired on route-gating 400s that merely name a thinking model (kimi-k2-thinking is not supported when using this account), spending a strip-retry and then re-raising before the fallback chain. Reproduced on head f4c7b4c (1 primary call → raise,_try_configured_fallback_chainnever called; also'max_tokens is not supported with reasoning models'matched). Fixed @ 2f2e059: the reasoning token must now be a standalone wire-field name (not a model-id segment, not "… with reasoning models"), and_param_rung_acceptsaccepts_is_model_incompatible_errorso a gating 400 after any strip still reaches provider fallback. New testtest_model_gating_400_naming_a_thinking_model_still_reaches_the_fallback_chain(red on f4c7b4c, green now).reasoning_effortre-raised before the reasoning rung. Fixed @ 2f2e059:_param_rung_acceptsalso accepts_is_reasoning_field_rejection/_is_structured_output_rejection; probe now makes 3 calls, the third without both fields.Known residuals