Skip to content

fix(streaming): provider-gate title-gen reasoning extra_body (#4161) - #4162

Closed
rodboev wants to merge 1 commit into
nesquena:masterfrom
rodboev:pr/4161-reasoning-provider-gate
Closed

rodboev wants to merge 1 commit into
nesquena:masterfrom
rodboev:pr/4161-reasoning-provider-gate

Conversation

@rodboev

@rodboev rodboev commented Jun 14, 2026 •

Copy link
Copy Markdown
Contributor

Fixes #4161

Thinking Path

  • Title generation via the auxiliary route unconditionally sends extra_body={"reasoning": {"enabled": False}} to suppress thinking on reasoning models, but OpenAI Chat Completions rejects the reasoning parameter as unknown, causing a 400 and fallback to low-quality local titles.
  • The reasoning disable was introduced for local reasoning models (GPU stays always active after a prompt in the webui #2083) but applied to every provider route without checking whether the server accepts it.
  • The fix adds a provider-gate helper that skips the reasoning extra_body for routes known to reject it (OpenAI, Azure), while preserving it for local endpoints, OpenRouter proxies, and other reasoning-aware providers.

What Changed

  • api/streaming.py: added _route_rejects_reasoning_extra() helper near _is_minimax_route(), and gated the reasoning_extra assignment in generate_title_raw_via_aux() behind it
  • tests/test_issue4161_reasoning_extra_provider_gate.py: focused regression tests for the helper and the patched call path

Why It Matters

Users with OpenAI as their main or auxiliary title-generation provider get silent 400 errors on every new session, falling back to heuristic titles like partial word bags. This fix restores LLM-quality titles for OpenAI routes while preserving the reasoning-disable behavior for local models.

Verification

pytest tests/test_issue4161_reasoning_extra_provider_gate.py -v --timeout=60
pytest tests/test_title_aux_routing.py -v --timeout=60

Full-suite CI context, not a required local check unless requested: pytest tests/ -v --timeout=60.

Risks / Follow-ups

  • The gate uses provider-string and base_url heuristics. A custom OpenAI-compatible endpoint that also rejects reasoning but doesn't match the detection patterns would still get the parameter; the agent's auxiliary_client.py handles temperature proactively via model-pattern gating (_fixed_temperature_for_model and _forbids_sampling_params), not a reactive retry-on-400 loop. A defense-in-depth follow-up could add a similar proactive gate for reasoning there, or introduce a new reactive unknown_parameter retry mechanism — neither exists today.
  • generate_title_raw_via_agent uses a different mechanism (agent.reasoning_config) that is provider-aware by design, so it is not affected.

@greptile-apps

greptile-apps Bot commented Jun 14, 2026

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR fixes a silent 400 error that occurred when generate_title_raw_via_aux unconditionally sent extra_body={"reasoning": {"enabled": False}} to OpenAI and Azure endpoints, which reject unknown top-level parameters, causing every new-session title to fall back to heuristic generation.

  • Adds _route_rejects_reasoning_extra() helper that returns True for api.openai.com/openai.azure.com base URLs and known OpenAI/Azure provider strings, mirroring the structure of the existing _is_minimax_route() helper.
  • Gates the reasoning_extra["reasoning"] assignment behind the new helper and passes reasoning_extra or None to extra_body so OpenAI routes receive None instead of an empty dict.
  • Adds a focused regression test file covering both the helper and the patched call path for OpenAI, Azure, local, OpenRouter, and MiniMax configurations.

Confidence Score: 4/5

Safe to merge; the change is narrowly scoped to the title-generation path and the fallback behavior on any missed route is the same silent 400 that already existed before this fix.

The core fix is correct and well-tested. The main open question is whether _split_webui_provider_model_value can transform a provider string into a form that bypasses the gate before _route_rejects_reasoning_extra is called — this path has no test coverage. Additionally, CHANGELOG.md was not updated despite the project's stated contribution guidelines.

api/streaming.py — specifically the ordering of provider normalization (lines 2596–2604) relative to the gate call (line 2611); confirm normalized provider values still match the rejection list.

Important Files Changed

Filename Overview
api/streaming.py Adds _route_rejects_reasoning_extra() helper and gates the reasoning extra_body assignment; the logic is sound for the covered provider strings and base URLs, though CHANGELOG.md is missing and provider normalization ordering is a minor untested edge case.
tests/test_issue4161_reasoning_extra_provider_gate.py Comprehensive regression suite covering True/False paths of the new helper and the full generate_title_raw_via_aux call path for OpenAI, Azure, local, OpenRouter, and MiniMax routes; patching strategy correctly targets agent.auxiliary_client._get_auxiliary_task_config via _get_aux_title_config.

Flowchart

%%{init: {'theme': 'neutral'}}%%
flowchart TD
    A[generate_title_raw_via_aux] --> B[Resolve provider / model / base_url]
    B --> C{_route_rejects_reasoning_extra}
    C -->|True - OpenAI / Azure| D["reasoning_extra = {}"]
    C -->|False - local / openrouter / other| E["reasoning_extra = {reasoning: {enabled: false}}"]
    D --> F{_is_minimax_route?}
    E --> F
    F -->|Yes| G["reasoning_extra += {reasoning_split: true}"]
    F -->|No| H["extra_body = reasoning_extra or None"]
    G --> H
    H --> I[call_llm with extra_body]
    I --> J{Response OK?}
    J -->|Yes| K["Return title, llm_aux"]
    J -->|No| L[Retry / fallback]
Loading

Reviews (1): Last reviewed commit: "fix(streaming): provider-gate title-gen ..." | Re-trigger Greptile

Comment thread api/streaming.py
Comment thread api/streaming.py
@nesquena-hermes

Copy link
Copy Markdown
Collaborator

Review — the gate is correct and the tests pin the right behavior

Pulled the branch into a read-only worktree and read the full generate_title_raw_via_aux() context (api/streaming.py:2607-2640), the new _route_rejects_reasoning_extra() helper (L2416-2425), the regression suite, and cross-checked the agent-side agent/auxiliary_client.py. The fix is right and the diagnosis in #4161 is accurate.

Code reference

The gate replaces the unconditional disable with a route check, and crucially flips extra_body=reasoning_extra to extra_body=reasoning_extra or None so an empty dict doesn't get forwarded as extra_body={}:

reasoning_extra = {}
if not _route_rejects_reasoning_extra(provider, model, base_url):
    reasoning_extra["reasoning"] = {"enabled": False}
if _is_minimax_route(provider, model, base_url):
    reasoning_extra["reasoning_split"] = True
...
extra_body=reasoning_extra or None,

The or None matters — passing a bare {} would still reach auxiliary_client.call_llm, and while that's harmless today, None is the cleaner contract. Good catch baking it in.

The test matrix is thorough: the allow/deny split covers openai/openai-api/openai-codex/azure/azure/deployment plus api.openai.com and openai.azure.com base_urls as True, and openrouter/anthropic/ollama/lmstudio/minimax/empty as False. The false-positive guard test_openrouter_in_model_does_not_trigger_true (provider openrouter, model openai/gpt-4o) is exactly the trap to pin — the helper only inspects provider and base_url, never model, so a model slug containing openai can't poison the gate.

One inaccuracy in the PR's own risk note

The "Risks / Follow-ups" section says the agent's call_llm "already strips temperature on rejection" and proposes extending that retry to strip reasoning. That mischaracterizes the agent mechanism. Reading agent/auxiliary_client.py:4945-4961, the temperature handling is proactive model-pattern gating, not a reactive retry on a 400:

fixed_temperature = _fixed_temperature_for_model(model, base_url)
if fixed_temperature is OMIT_TEMPERATURE:
    temperature = None  # strip — let server choose
...
if temperature is not None:
    from agent.anthropic_adapter import _forbids_sampling_params
    if _forbids_sampling_params(model):
        temperature = None

There is no except BadRequest → strip param → retry loop in call_llm for unknown_parameter/unsupported_parameter. So the defense-in-depth follow-up you describe would be a new reactive mechanism, not an extension of an existing one. Worth correcting in the PR body so a future reader doesn't go looking for a retry hook that isn't there. (The proactive-gate approach this PR takes is the better match for the existing agent style anyway.)

Minor: openai-codex gating is safe but redundant

openai-codex routes through the Responses API on chatgpt.com/backend-api/codex, and that path already no-ops on {"reasoning": {"enabled": False}} — auxiliary_client.py:705-707 reads extra_body.reasoning and does pass when enabled is False rather than forwarding it. So gating codex off the reasoning extra_body changes nothing observable for that route; it's harmless and consistent to include, just not load-bearing the way the openai-api Chat Completions case is. No change needed — flagging only so the test name test_openai_codex_provider_returns_true isn't read as "codex was also 400ing," which it wasn't.

Verification

The unit + integration split looks complete for the helper and the patched call path. No execution from my side (review is read-only) — the green suite is on your CI. LGTM once the risk-note wording is corrected.

nesquena-hermes added a commit that referenced this pull request Jun 14, 2026
Release NV (v0.51.409): provider-gate title-gen reasoning extra_body (#4161/#2083, consolidates #4162+#3944)
Hinotoi-agent pushed a commit to Hinotoi-agent/hermes-webui that referenced this pull request Jun 14, 2026
…a#4161, nesquena#2083)

Consolidates nesquena#4162 + nesquena#3944. Aux path: skip the reasoning-disable inject for
routes that 400 on it via _route_rejects_reasoning_extra — hostname-matched
OpenAI / Azure OpenAI / Azure AI Foundry (+ provider aliases openai*/azure*/
azure-foundry/azure-ai*) and OpenRouter Anthropic mandatory-reasoning; keep it
for local/OpenRouter-non-anthropic/other reasoning routes (nesquena#4161). Agent path:
rely on the existing agent.reasoning_config={enabled:False} + _build_api_kwargs()
(route-correct per provider profile) rather than re-injecting a generic disable
on top of it; only MiniMax reasoning_split is added (unchanged from master).

Co-authored-by: Rod Boev <rod.boev@gmail.com>
SysAdminDoc pushed a commit to SysAdminDoc/hermes-webui that referenced this pull request Jun 26, 2026
…a#4161, nesquena#2083)

Consolidates nesquena#4162 + nesquena#3944. Aux path: skip the reasoning-disable inject for
routes that 400 on it via _route_rejects_reasoning_extra — hostname-matched
OpenAI / Azure OpenAI / Azure AI Foundry (+ provider aliases openai*/azure*/
azure-foundry/azure-ai*) and OpenRouter Anthropic mandatory-reasoning; keep it
for local/OpenRouter-non-anthropic/other reasoning routes (nesquena#4161). Agent path:
rely on the existing agent.reasoning_config={enabled:False} + _build_api_kwargs()
(route-correct per provider profile) rather than re-injecting a generic disable
on top of it; only MiniMax reasoning_split is added (unchanged from master).

Co-authored-by: Rod Boev <rod.boev@gmail.com>
SysAdminDoc pushed a commit to SysAdminDoc/hermes-webui that referenced this pull request Jun 26, 2026
Release NV (v0.51.409): provider-gate title-gen reasoning extra_body (nesquena#4161/nesquena#2083, consolidates nesquena#4162+nesquena#3944)
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

WebUI auto-title generation sends unsupported reasoning parameter to OpenAI Chat Completions, causing llm_error_aux and poor fallback titles

2 participants