Skip to content

fix(agent): make title generation work for reasoning models and strict local providers - #85424

Closed
woshicby wants to merge 5 commits into
NousResearch:mainfrom
woshicby:fix/title-generation-reasoning-providers
Closed

woshicby wants to merge 5 commits into
NousResearch:mainfrom
woshicby:fix/title-generation-reasoning-providers

Conversation

@woshicby

@woshicby woshicby commented Aug 13, 2026 •

Copy link
Copy Markdown

Problem

Two hardcoded assumptions in agent/title_generator.py break automatic session titling for reasoning models and several OpenAI-compatible backends:

  1. max_tokens=64 starves reasoning models. A reasoning-capable model (Qwen3.x thinking, DeepSeek-R1, GLM-5, ...) spends the entire 64-token budget on reasoning_content; message.content comes back empty and the title silently fails. Sessions permanently keep their truncated derived title — no error, no retry ([Bug]: Session auto-title upgrade silently fails when provider returns json_schema output in reasoning_content (opencode-go / glm-5) #82291 documents the same "empty content" failure from the opencode-go/GLM-5 side). It also truncates non-reasoning providers mid-JSON, leaving raw fragments like {"title in the sidebar (Auto-title returns raw JSON fragments when max_tokens=64 truncates model response #83903).

  2. Strict json_schema response_format is rejected or silently aborted. vLLM with guided_grammar/xgrammar returns HTTP 400 (Session auto-title generation fails 100% of the time (HTTP 400) on OpenAI-compatible providers that reject response_format json_schema (vLLM guided_grammar / xgrammar) #82816); DeepSeek returns HTTP 400 "This response_format type is unavailable now" (Auxiliary title_generation fails on DeepSeek: HTTP 400 "This response_format type is unavailable now" #83390, Bug: title_generation fails on DeepSeek provider — HTTP 400 "This response_format type is unavailable now" #84976); LM Studio's Qwen3.x (MLX) aborts and returns empty content under strict structured output — there is no error for the retry/fallback machinery to react to. opencode-go/GLM-5 also relocates the entire structured response into reasoning_content, leaving content empty ([Bug]: Session auto-title upgrade silently fails when provider returns json_schema output in reasoning_content (opencode-go / glm-5) #82291).

Net effect: users with a local reasoning model (e.g. LM Studio + Qwen3.x) as auxiliary.title_generation get 100% title-generation failure.

Changes

agent/title_generator.py:

Why free-text instead of falling back to json_object?

The existing PRs in this area (#82751, #85115, #84767, #83725, #83186) fall back to json_object when json_schema is rejected. That covers backends which error on json_schema, but not strict local backends (LM Studio MLX Qwen3.x) that abort with empty content under any structured-output mode — there is no error to react to. Free-text is the lowest common denominator every backend supports, and the existing extractor already handles the shape.

Verification (per reported issue)

Each reported failure mode was exercised against the patched code path:

Issue Reported failure Why this PR fixes it
#83390 DeepSeek HTTP 400 on json_schema No longer sends json_schema
#84976 DeepSeek HTTP 400 (dup) Same as above
#82816 vLLM HTTP 400 (xgrammar missing) No longer sends json_schema; type: text is accepted by OpenAI-compatible servers
#83903 Raw JSON fragment {"title as title 2048-token budget ends mid-JSON truncation; validated _extract_title_text on full JSON
#82291 GLM-5 puts JSON in reasoning_content, content empty Plain-text request returns content normally (matches issue's "without response_format works" repro); _extract_title_text parses it

Local reproduction: auxiliary.title_generation → LM Studio qwen3.6-27b (MLX). Before: 100% of titles fail (empty content, title_source stays derived). After: titles generate normally; verified across 50+ sessions.

Related issues

Closes #83390, #84976, #82816, #83903, #82291.

@alt-glitch alt-glitch added type/bug Something isn't working P3 Low — cosmetic, nice to have comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint labels Aug 13, 2026
@woshicby

woshicby commented Aug 13, 2026 •

Copy link
Copy Markdown
Author

Real-world verification: bulk session retitling on LM Studio + Qwen3.6-27b (MLX)

To make the "50+ sessions" claim in the PR description concrete, here is the actual data from the workflow that surfaced this bug. I was bulk-retitling a legacy session archive (a few hundred sessions in the Hermes SQLite store) using auxiliary.title_generation → LM Studio qwen3.6-27b (MLX, compatibility_type: mlx) — exactly the configuration where this fails.

Before the patch (current main behavior)

  • Roughly half of sessions had no title or a truncated first-sentence title; the bulk retitler observed repeated silent failures: message.content came back empty with title_source staying derived.
  • No error was raised anywhere — the failure was invisible to the retry/fallback machinery, which is why it's hard to reproduce casually (you only see it under sustained title-generation load against a strict local backend).

After the patch (this PR)

Batch retitling of the same archive succeeded end-to-end: dozens of LLM-generated titles, and naming coverage for the archive went from roughly half → ~94%. Sample titles actually produced by qwen3.6-27b via this patched path:

  • 修正看板任务状态并统一工作区
  • 恢复hermes dashboard用量界面
  • 检查已完成看板任务产出

All clean, non-truncated Chinese titles — no {"title fragments (cf. #83903), no empty-content failures (cf. #83390 / #82291).

Why this is the failure mode #85316's reasoning_effort knob can't fix

For strict local backends like LM Studio MLX, the model aborts and returns empty content under any structured-output mode — there is no HTTP error and no content to react to. Disabling thinking (the #85316 approach) doesn't change that: the abort happens on the response_format enforcement, not on the reasoning pass. That's why this PR removes json_schema entirely (free-text + the existing _extract_title_text JSON-scan/prose fallback) instead of layering another knob on top.

Happy to rebase this onto current main if the title-generator area has moved since Aug 13.

@spfcraze

Copy link
Copy Markdown

This was generated by AI during triage.

Summary:
This diff removes the only call site of _TITLE_RESPONSE_FORMAT (the extra_body line at title_generator.py:404) but leaves the constant defined at title_generator.py:94, and the module and generate_title docstrings still describe the constrained-JSON response format the change replaces with free text.

Problems:

  • _TITLE_RESPONSE_FORMAT is still defined at agent/title_generator.py:94 with a comment explaining the json_schema constraint, but its only use, extra_body={"response_format": _TITLE_RESPONSE_FORMAT}, is replaced by {"type": "text"} in this diff.
  • The module docstring says the title call is "constrained to a JSON object, so there is no reasoning preamble to strip and nothing to parse out of prose" — the reverse of what this change does, which is to send free text and parse the title out of prose.
  • generate_title's docstring likewise says "the response is constrained to {"title": "..."} so there is no preamble or reasoning to strip."

Solution:
Remove the now-unused _TITLE_RESPONSE_FORMAT definition and its comment, and update both docstrings to describe the free-text plus _extract_title_text parsing path.

Evidence

no deterministic fact backs this claim — model belief, not executed or read evidence


Checked against 085aa79 — the tip of fix/title-generation-reasoning-providers when this was written — and f80f453, main at the same moment.

@woshicby

Copy link
Copy Markdown
Author

Good catch — this is already addressed in the latest commit.

bca1c4389 ("refactor(title): drop dead _TITLE_RESPONSE_FORMAT, align docstrings with free-text extraction") removes the now-unused _TITLE_RESPONSE_FORMAT definition (and its comment), and rewrites both the module docstring and the generate_title docstring to describe the free-text response plus _extract_title_text parsing path. The only remaining response_format reference is the {"type": "text"} sent as extra_body.

The triage snapshot was taken at 085aa79; the fix landed on top as bca1c4389. Verified on the branch tip: _TITLE_RESPONSE_FORMAT no longer appears anywhere in agent/title_generator.py.

@woshicby
woshicby force-pushed the fix/title-generation-reasoning-providers branch from bca1c43 to ffb938a Compare August 15, 2026 03:16
@Enough1122

Copy link
Copy Markdown

AI code review — automated review for reference, author can ignore or act on any point.

fix(agent): make title generation work for reasoning models and strict local providers

  1. max_tokens 64 → 2048 (agent/title_generator.py): for reasoning models the chain-of-thought can still exceed 2048 tokens, reproducing the empty-content failure the PR fixes — 2048 is a heuristic, not a guarantee. Conversely, the old 64-token ceiling was also a latency/cost guard for chatty non-reasoning models, which can now burn up to 2048 tokens of prose that _extract_title_text then throws away. Consider a reasoning-aware cap (e.g. trim/normalize before extraction) instead of a flat raise.
  2. extra_body={"response_format": {"type": "text"}} is sent explicitly even though free text is the API default. The entire point of the PR is that some local providers abort under strict json_schema formats — a provider strict about response_format values (accepting only json_object/json_schema, or rejecting the field entirely) may equally reject an explicit {"type": "text"}. Since the extraction already falls back to prose, consider dropping the field (or gating it) so non-compliant providers are never sent a format they might refuse.
  3. Minor: docstring updates accurately reflect the new behavior — good.

@alt-glitch alt-glitch added provider/deepseek DeepSeek API provider/qwen Qwen / Alibaba Cloud (OAuth) P2 Medium — degraded but workaround exists and removed provider/deepseek DeepSeek API provider/qwen Qwen / Alibaba Cloud (OAuth) P3 Low — cosmetic, nice to have labels Aug 16, 2026
@woshicby

Copy link
Copy Markdown
Author

Regarding the max_tokens observation (64 → 2048):

We verified this empirically on LM Studio + Qwen3.6-27b (MLX) and similar reasoning-capable local models: the chain-of-thought for title generation comfortably fits within 2048 tokens in the cases we reproduced — the failures we saw were all from the 64-token ceiling truncating the CoT before any answer was emitted (empty content, silent failure). 2048 also does not change behavior for non-reasoning models: they stop after the short answer regardless of the cap, so the latency/cost guard the old ceiling provided is not meaningfully affected in practice.

That said, we are happy to make this more flexible if the maintainers prefer — for example bumping the ceiling higher, or exposing it as a configurable setting — so it can be tuned without a code change. Just let us know which direction you'd like.

@alt-glitch alt-glitch added P3 Low — cosmetic, nice to have provider/qwen Qwen / Alibaba Cloud (OAuth) provider/deepseek DeepSeek API and removed P2 Medium — degraded but workaround exists labels Aug 17, 2026
@woshicby
woshicby force-pushed the fix/title-generation-reasoning-providers branch from ffb938a to 2d0e95a Compare August 17, 2026 06:45
@Jeffgithub0029

Copy link
Copy Markdown

Thanks for addressing this. I can confirm the issue on my macOS setup with DeepSeek-v4-flash via OpenCode Go: json_schema causes the HTTP 400. The free-text approach in #85424 looks promising; happy to retest if needed.

@Enough1122

Copy link
Copy Markdown

AI code review — automated review for reference — not a maintainer.

Thanks for the confirmation — the explicit json_schema response format triggering the HTTP 400 on your macOS + DeepSeek-v4-flash via OpenCode Go setup is exactly the failure this free-text approach avoids. Appreciate the retest offer.'

@woshicby

Copy link
Copy Markdown
Author

Thanks for both points — we've addressed them in the latest commits.

On omitting response_format: agreed, and done. Free text is the API default, so sending {"type": "text"} explicitly added nothing on providers that accept it and risked a fresh rejection on providers strict about the field. The request now sends no response_format at all; _extract_title_text handles the loose JSON / prose shape as before. (f01cd52)

On the 2048 budget being a heuristic: correct — it is, and we don't claim otherwise. But it is no longer the only guard:

  • If a reasoning model spends the whole budget on reasoning_content and content comes back empty, we now fall back to extracting the title from reasoning_content — a truncated chain of thought usually states the candidate anyway, so truncation no longer equals failure.
  • The prompt now explicitly asks for the title directly with no analysis/chain-of-thought, so non-reasoning models stop early after the short answer instead of burning the budget on discarded prose.

So 2048 remains a heuristic ceiling, but the failure modes it used to guard against are now covered by the fallback and the prompt constraint. Tests: 38/38 including the empty-content fallback and both-empty no-crash paths. (0897031)

@Jeffgithub0029

Copy link
Copy Markdown

Agreed — dropping response_format at the source is the right shape; it sidesteps both the Anthropic Extra inputs rejection and the _build_call_kwargs leak without any downgrade carve-out. I had two narrower PRs (#90063 retry-path, #90064 reasoning_effort) that overlapped this scope and both were failing CI, so I closed them in favor of this one. Appreciate the end-to-end verification — happy to review if you'd like another pair of eyes.

@linfeng961

Copy link
Copy Markdown

Additional current-release evidence (v0.20.4, Windows Desktop):

  • auxiliary.title_generation is enabled and uses the default automatic route.
  • The title task is dispatched for a new session, but the persisted session row remains title_source = "derived"; no model-upgraded title is written.
  • A second recent desktop session shows the same persisted derived state.
  • The same diagnostic window also had transient upstream 502 responses, so I cannot attribute this individual observation conclusively to structured-output placement or the 64-token cap alone.

Still, the current shipped path retains the three failure surfaces this PR addresses: strict response_format, extraction from message.content only, and a 64-token output limit. The consolidated approach here looks preferable to a provider-specific fallback.

I can retest this branch on the current desktop release if it is rebased and CI is available.

@Jeffgithub0029

Copy link
Copy Markdown

@linfeng961 Thanks for the additional v0.20.4 evidence — two independent desktop sessions persisting title_source = "derived" lines up well with the failure surfaces this PR addresses, and your caveat about the transient 502s is a fair one.

Status check: the branch is currently ~12 commits behind main, which blocks a clean CI run and your retest offer. @woshicby — could you rebase onto current main when you have a chance? Once it's rebased and green, @linfeng961's retest on the current desktop release would make a solid end-to-end validation.

@Jeffgithub0029 Jeffgithub0029 left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Reviewed locally (checked out the branch, full test suite for this module): 38 passed incl. both new regression tests; verified _TITLE_RESPONSE_FORMAT has no remaining references.

Sending no response_format at all is the right call — it sidesteps both the DeepSeek HTTP 400 (json_schema unavailable) and the LM Studio empty-content abort in one move, instead of playing whack-a-mole with per-provider downgrades. The reasoning_content salvage is a nice touch, and it correctly still passes through _extract_title_text + _clean_title so chain-of-thought chatter cannot leak into a title.

Two non-blocking notes:

  1. max_tokens 64→2048 is justified for reasoning models, but this call fires on every session start on the cheap tier — worst case a chatty non-reasoning model burns ~2k tokens. Worth a follow-up observation point (actual completion-token distribution) to confirm cost does not drift.
  2. Minor doc drift: _extract_title_text's docstring still opens with "The JSON schema makes the object shape the expected case" — no schema is sent anymore, so that sentence could be updated.

Cross-ref #83390: confirmed complementary — this fixes the title path; the fallback-candidate leak for other structured aux tasks remains open there.

@woshicby

Copy link
Copy Markdown
Author

Thanks for the thorough review — checking out the branch and running the module suite locally is exactly the kind of verification that matters here.

Both notes addressed:

  1. docstring drift — fixed in 1c4cdf386a: _extract_title_text now states that no response_format is sent and the shape is free-form (strict JSON → loose scan → first-line prose). No remaining schema-based framing.

  2. max_tokens 64→2048 on the cheap tier — agreed it's worth watching. The budget only caps the response; a non-reasoning model that finishes early stops well before 2048. I'll keep an eye on completion-token distribution and can tighten the cap if drift shows.

Glad the #83390 cross-ref holds: this PR covers the title path, and the fallback-candidate finding there is a genuinely complementary gap for the other structured aux tasks.

Legion-is-life added a commit to Legion-is-life/hermes-agent that referenced this pull request Aug 23, 2026
…tructured-output rejection

_call_fallback_candidate_sync and _call_fallback_candidate_async only
special-cased auth errors — any other error (including DeepSeek's HTTP 400
"This response_format type is unavailable now" on json_schema) re-raised
immediately and aborted the whole auxiliary task. The primary call_llm path
already retries once without the field; this mirrors that degradation for
the fallback path.

Affects every aux task whose fallback candidate rejects response_format:
title_generation, vision, compression, web_extract, plugin structured
completions, etc.

Closes the fallback-candidate gap not covered by NousResearch#85424 (which drops the
field for title generation but does not protect other aux tasks that
legitimately keep structured output).
@Jeffgithub0029

Copy link
Copy Markdown

Verified locally on 1c4cdf3 (applied patch to current main, 38/38 in tests/agent/test_title_generator.py):

  • docstring drift fixed: _extract_title_text now documents no response_format sent, free-form strict JSON → loose scan → first-line prose — no remaining schema framing
  • max_tokens 64→2048 + no response_format + reasoning_content fallback + prompt constraint (no chain-of-thought) covers the three failure surfaces: DeepSeek/vLLM HTTP 400 on json_schema, LM Studio MLX empty-content abort, and reasoning-model starvation

Agree 2048 is heuristic but with the two guards the failure window is closed; monitoring completion-token distribution on the cheap tier as you noted makes sense.

Branch is currently behind main so CI shows no checks — a rebase would unblock @linfeng961's retest. Happy to retest on DeepSeek-v4-flash via OpenCode Go (macOS) once rebased and green.

Complementary to #83390 holds: this PR fixes the title path; the other structured aux tasks still need the fallback/capability gap noted there. LGTM.

@woshicby
woshicby force-pushed the fix/title-generation-reasoning-providers branch from 1c4cdf3 to a4b6e56 Compare September 4, 2026 15:26
…t local providers

max_tokens=64 starves reasoning models: the chain-of-thought consumes
the whole budget and content comes back empty, so the title silently
fails and the session keeps a truncated derived title forever.

Strict json_schema response_format is rejected (HTTP 400) or silently
aborted (empty content) by several OpenAI-compatible backends: vLLM
guided_grammar/xgrammar, DeepSeek, and LM Studio MLX Qwen3.x.

Raise max_tokens to 2048 and use free-text response_format, letting the
existing _extract_title_text JSON scan + prose fallback parse the shape.
Non-reasoning models still stop after the short JSON answer, so there is
no practical cost increase for them.

Closes NousResearch#83390, NousResearch#84976. Related: NousResearch#82816, NousResearch#83903, NousResearch#82291.
…ith free-text extraction

The json_schema response_format was removed in favor of free text (strict
structured output aborts on some local providers, returning empty content),
but the constant definition and both docstrings still described the old
constrained-JSON contract. Remove the now-unused constant and update the
module and generate_title docstrings to describe the free-text +
_extract_title_text extraction path.
Free text is the API default, so explicitly sending
{"response_format": {"type": "text"}} adds nothing on providers that
accept it and risks a fresh rejection on providers strict about the
field (only json_object/json_schema, or rejecting it outright). Sending
no response_format at all means a non-compliant provider can never be
handed a format it refuses; _extract_title_text already handles the
loose JSON / prose shape.
@teknium1

Copy link
Copy Markdown
Collaborator

Thanks @woshicby — this landed on main via #113960 (eff1876) with a Co-authored-by credit for your change. Closing this PR as merged-through-salvage; the fix is on main.

@teknium1 teknium1 closed this Sep 17, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint P3 Low — cosmetic, nice to have provider/deepseek DeepSeek API provider/qwen Qwen / Alibaba Cloud (OAuth) type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Auxiliary title_generation fails on DeepSeek: HTTP 400 "This response_format type is unavailable now"

7 participants