Skip to content

fix(aux): title generation survives custom endpoints that reject reasoning_effort (#112781, supersedes #112789) - #113114

Merged
teknium1 merged 2 commits into
mainfrom
fix/b113-aux-moa-title-effort
Sep 17, 2026
Merged

teknium1 merged 2 commits into
mainfrom
fix/b113-aux-moa-title-effort

Conversation

@teknium1

@teknium1 teknium1 commented Sep 16, 2026 •

Copy link
Copy Markdown
Collaborator

Auxiliary title generation no longer dies with 400 Unrecognized request argument supplied: reasoning_effort on custom/OpenAI-compatible endpoints whose model does not accept the field — the call is retried once with every reasoning field omitted and the title lands.

Changes

  • agent/auxiliary_client.py — new rung in the shared aux recovery ladder (_ladder_parameter_rungs; sync call_llm and async_call_llm drive the same generator, so one rung covers both): when a 400/422 names a reasoning field (reasoning_effort, reasoning, thinking/think) with an unsupported/unrecognized marker, strip top-level reasoning_effort, the adapter-private _reasoning_config and every extra_body reasoning key, then retry once. Same reactive shape as the existing temperature and response_format rungs; an unrelated 400 or a request that carried no reasoning field never retries.
  • plugins/model-providers/custom/__init__.py — untouched: reasoning_effort: "none" remains the deliberate thinking-off encoding for Ollama /v1 ([Bug]: agent.reasoning_effort: none silently ignored on Ollama — main agent stuck in medium mode, bg-review fork can spiral (up to 65k tokens / 28 min) #25758), vLLM, SGLang and GLM/ARK routes, on the first request, for aux and main-agent calls alike.
  • tests/agent/test_auxiliary_reasoning_field_rejection_retry.py — 2 invariants (sync/async retry with the reporter's verbatim 400; unrelated 400 does not strip). Red on origin/main (2 failed / 1 passed), green on this head.
  • website/docs/user-guide/configuration.md — one paragraph under the auxiliary reasoning_effort knob describing the retry.

Validation

Live probe (monkeypatched openai Completions.create/AsyncCompletions.create acting as a relay that 400s on reasoning_effort, real call_llm/async_call_llm with provider=custom, model=gpt-4.1-mini, reasoning_config={"enabled": False}, temp HERMES_HOME):

wire calls outcome
before (origin/main 9796235), sync + async 1: {"model": "gpt-4.1-mini", "reasoning_effort": "none", "temperature": 0.3, "extra_body": {"response_format": …}} BadRequestError 400 … reasoning_effort — title lost
after (this head), sync + async 1: same as above → 400; 2: {"model": "gpt-4.1-mini", "temperature": 0.3, "extra_body": {"response_format": …}} {"title": "Relay Title Probe"}

Tests: scripts/run_tests.sh on the new file + test_structured_output_rejection_retry.py, test_unsupported_parameter_retry.py, test_unsupported_temperature_retry.py, test_auxiliary_client.py, test_injected_param_strip_retry_registry.py, test_auxiliary_auth_rung_fallthrough.py → 287 passed, 0 failed.

Root cause

title_generator.generate passes reasoning_config={"enabled": False} (#91927); CustomProfile.build_api_kwargs_extras projects that as top-level reasoning_effort: "none" for every custom endpoint, and because the profile handles reasoning there was no fallback or recovery when a chat-only model behind a relay rejects the field.

Fixes #112781

Supersedes #112789 (@KoNit-K) — that PR restricts the reasoning_effort: "none" encoding to Ollama URLs inside the shared custom profile, which would silently re-enable thinking for vLLM/SGLang/GLM users who set reasoning_effort: none (main agent included, not only the title lane). The recovery here is provider-agnostic and leaves the deliberate encoding in place; thanks to @KoNit-K for the report-to-test mapping.

Infographic

title-effort

Review follow-up

  • [MAJOR] _is_reasoning_field_rejection fired on route-gating 400s that merely name a thinking model (kimi-k2-thinking is not supported when using this account), spending a strip-retry and then re-raising before the fallback chain. Reproduced on head f4c7b4c (1 primary call → raise, _try_configured_fallback_chain never called; also 'max_tokens is not supported with reasoning models' matched). Fixed @ 2f2e059: the reasoning token must now be a standalone wire-field name (not a model-id segment, not "… with reasoning models"), and _param_rung_accepts accepts _is_model_incompatible_error so a gating 400 after any strip still reaches provider fallback. New test test_model_gating_400_naming_a_thinking_model_still_reaches_the_fallback_chain (red on f4c7b4c, green now).
  • [MINOR] temperature-strip retry that 400s on reasoning_effort re-raised before the reasoning rung. Fixed @ 2f2e059: _param_rung_accepts also accepts _is_reasoning_field_rejection / _is_structured_output_rejection; probe now makes 3 calls, the third without both fields.
  • [MINOR] zh-Hans configuration.md mirror lacked the retry paragraph. Fixed @ 2f2e059 (one-line mirror addition before the 后台审查 paragraph).

Known residuals

  • None.

…rejects them

Auxiliary title generation disables reasoning (reasoning_config
{"enabled": False}, #91927); on provider=custom the profile encodes that as
top-level reasoning_effort="none" - the deliberate thinking-off wire for
Ollama /v1 (#25758), vLLM and GLM. A chat-only model behind an
OpenAI-compatible relay (gpt-4.1-mini on a one-api style relay) answers
"400 Unrecognized request argument supplied: reasoning_effort" and the
title was lost with no retry (#112781).

Add a rung to the shared aux recovery ladder (sync and async drive the
same generator): when the 400 names a reasoning field, strip top-level
reasoning_effort, the adapter's private _reasoning_config and every
extra_body reasoning key, and retry once - the same reactive shape as the
temperature and response_format rungs. The custom profile's encoding is
untouched, so Ollama/vLLM/GLM users keep thinking-off on the first
request; only routes that reject the field pay one extra round-trip.

Supersedes #112789 (@KoNit-K), which dropped the encoding for every
non-Ollama custom endpoint in the shared profile and would have silently
re-enabled thinking for vLLM/GLM users who set reasoning_effort: none.
@github-actions

github-actions Bot commented Sep 16, 2026 •

Copy link
Copy Markdown

૮ >ﻌ< ა ci review

ran on 2f2e059 — fix(aux): only treat a standalone reasoning field name as a

⚠️ Warnings

CI timings · View report · View job

Wall time 9m14s vs 6m8s (+50.5%). 6 job(s) slower, 6 faster, 1 unchanged.

  • Python tests / e2e: +13.0s
  • Python lints / Windows footguns (blocking): -7.0s
  • Check no case-colliding filenames / check-case-collisions: +6.0s
  • Python tests / Run tests: -5.0s
  • OS-specific tests / Windows-only tests: -5.0s

… rejection; let parameter rungs chain

Review follow-up on #113114. `_is_reasoning_field_rejection` fired on any 400 whose text
contained "reasoning"/"think" next to a generic unsupported marker, so a route-gating 400
naming a thinking model ("The model kimi-k2-thinking is not supported when using this
account") spent a strip-retry and then re-raised from `_param_rung_accepts` — the configured
fallback chain, which main consulted for that error, was never reached. Require the token
to be a standalone wire-field name (not a model-id segment, not "... with reasoning
models"), and let `_param_rung_accepts` accept model-incompatible, reasoning-field and
structured-output rejections so a temperature-strip retry that 400s on `reasoning_effort`
reaches the reasoning rung and a gating 400 after any strip still reaches provider fallback.

Also mirrors the new retry paragraph into the zh-Hans configuration page.
@teknium1
teknium1 merged commit 65aaf16 into main Sep 17, 2026
34 checks passed
@teknium1
teknium1 deleted the fix/b113-aux-moa-title-effort branch September 17, 2026 00:01
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: Auxiliary title_generation sends top-level reasoning_effort on custom/OpenAI-compatible endpoints - HTTP 400 on models that don't support it

1 participant