Conversation
Real-world verification: bulk session retitling on LM Studio + Qwen3.6-27b (MLX)To make the "50+ sessions" claim in the PR description concrete, here is the actual data from the workflow that surfaced this bug. I was bulk-retitling a legacy session archive (a few hundred sessions in the Hermes SQLite store) using Before the patch (current
|
|
This was generated by AI during triage. Summary: Problems:
Solution: Evidenceno deterministic fact backs this claim — model belief, not executed or read evidence Checked against |
|
Good catch — this is already addressed in the latest commit.
The triage snapshot was taken at |
bca1c43 to
ffb938a
Compare
fix(agent): make title generation work for reasoning models and strict local providers
|
|
Regarding the We verified this empirically on LM Studio + Qwen3.6-27b (MLX) and similar reasoning-capable local models: the chain-of-thought for title generation comfortably fits within 2048 tokens in the cases we reproduced — the failures we saw were all from the 64-token ceiling truncating the CoT before any answer was emitted (empty That said, we are happy to make this more flexible if the maintainers prefer — for example bumping the ceiling higher, or exposing it as a configurable setting — so it can be tuned without a code change. Just let us know which direction you'd like. |
ffb938a to
2d0e95a
Compare
|
Thanks for addressing this. I can confirm the issue on my macOS setup with DeepSeek-v4-flash via OpenCode Go: |
Thanks for the confirmation — the explicit |
|
Thanks for both points — we've addressed them in the latest commits. On omitting On the 2048 budget being a heuristic: correct — it is, and we don't claim otherwise. But it is no longer the only guard:
So 2048 remains a heuristic ceiling, but the failure modes it used to guard against are now covered by the fallback and the prompt constraint. Tests: 38/38 including the empty-content fallback and both-empty no-crash paths. (0897031) |
|
Agreed — dropping |
|
Additional current-release evidence (v0.20.4, Windows Desktop):
Still, the current shipped path retains the three failure surfaces this PR addresses: strict I can retest this branch on the current desktop release if it is rebased and CI is available. |
|
@linfeng961 Thanks for the additional v0.20.4 evidence — two independent desktop sessions persisting Status check: the branch is currently ~12 commits behind |
Jeffgithub0029
left a comment
There was a problem hiding this comment.
Reviewed locally (checked out the branch, full test suite for this module): 38 passed incl. both new regression tests; verified _TITLE_RESPONSE_FORMAT has no remaining references.
Sending no response_format at all is the right call — it sidesteps both the DeepSeek HTTP 400 (json_schema unavailable) and the LM Studio empty-content abort in one move, instead of playing whack-a-mole with per-provider downgrades. The reasoning_content salvage is a nice touch, and it correctly still passes through _extract_title_text + _clean_title so chain-of-thought chatter cannot leak into a title.
Two non-blocking notes:
max_tokens64→2048 is justified for reasoning models, but this call fires on every session start on the cheap tier — worst case a chatty non-reasoning model burns ~2k tokens. Worth a follow-up observation point (actual completion-token distribution) to confirm cost does not drift.- Minor doc drift:
_extract_title_text's docstring still opens with "The JSON schema makes the object shape the expected case" — no schema is sent anymore, so that sentence could be updated.
Cross-ref #83390: confirmed complementary — this fixes the title path; the fallback-candidate leak for other structured aux tasks remains open there.
|
Thanks for the thorough review — checking out the branch and running the module suite locally is exactly the kind of verification that matters here. Both notes addressed:
Glad the #83390 cross-ref holds: this PR covers the title path, and the fallback-candidate finding there is a genuinely complementary gap for the other structured aux tasks. |
…tructured-output rejection _call_fallback_candidate_sync and _call_fallback_candidate_async only special-cased auth errors — any other error (including DeepSeek's HTTP 400 "This response_format type is unavailable now" on json_schema) re-raised immediately and aborted the whole auxiliary task. The primary call_llm path already retries once without the field; this mirrors that degradation for the fallback path. Affects every aux task whose fallback candidate rejects response_format: title_generation, vision, compression, web_extract, plugin structured completions, etc. Closes the fallback-candidate gap not covered by NousResearch#85424 (which drops the field for title generation but does not protect other aux tasks that legitimately keep structured output).
|
Verified locally on 1c4cdf3 (applied patch to current main, 38/38 in tests/agent/test_title_generator.py):
Agree 2048 is heuristic but with the two guards the failure window is closed; monitoring completion-token distribution on the cheap tier as you noted makes sense. Branch is currently behind Complementary to #83390 holds: this PR fixes the title path; the other structured aux tasks still need the fallback/capability gap noted there. LGTM. |
1c4cdf3 to
a4b6e56
Compare
…t local providers max_tokens=64 starves reasoning models: the chain-of-thought consumes the whole budget and content comes back empty, so the title silently fails and the session keeps a truncated derived title forever. Strict json_schema response_format is rejected (HTTP 400) or silently aborted (empty content) by several OpenAI-compatible backends: vLLM guided_grammar/xgrammar, DeepSeek, and LM Studio MLX Qwen3.x. Raise max_tokens to 2048 and use free-text response_format, letting the existing _extract_title_text JSON scan + prose fallback parse the shape. Non-reasoning models still stop after the short JSON answer, so there is no practical cost increase for them. Closes NousResearch#83390, NousResearch#84976. Related: NousResearch#82816, NousResearch#83903, NousResearch#82291.
…ith free-text extraction The json_schema response_format was removed in favor of free text (strict structured output aborts on some local providers, returning empty content), but the constant definition and both docstrings still described the old constrained-JSON contract. Remove the now-unused constant and update the module and generate_title docstrings to describe the free-text + _extract_title_text extraction path.
Free text is the API default, so explicitly sending
{"response_format": {"type": "text"}} adds nothing on providers that
accept it and risks a fresh rejection on providers strict about the
field (only json_object/json_schema, or rejecting it outright). Sending
no response_format at all means a non-compliant provider can never be
handed a format it refuses; _extract_title_text already handles the
loose JSON / prose shape.
a4b6e56 to
3774d93
Compare
Problem
Two hardcoded assumptions in
agent/title_generator.pybreak automatic session titling for reasoning models and several OpenAI-compatible backends:max_tokens=64starves reasoning models. A reasoning-capable model (Qwen3.x thinking, DeepSeek-R1, GLM-5, ...) spends the entire 64-token budget onreasoning_content;message.contentcomes back empty and the title silently fails. Sessions permanently keep their truncatedderivedtitle — no error, no retry ([Bug]: Session auto-title upgrade silently fails when provider returns json_schema output in reasoning_content (opencode-go / glm-5) #82291 documents the same "empty content" failure from the opencode-go/GLM-5 side). It also truncates non-reasoning providers mid-JSON, leaving raw fragments like{"titlein the sidebar (Auto-title returns raw JSON fragments when max_tokens=64 truncates model response #83903).Strict
json_schemaresponse_formatis rejected or silently aborted. vLLM with guided_grammar/xgrammar returns HTTP 400 (Session auto-title generation fails 100% of the time (HTTP 400) on OpenAI-compatible providers that reject response_format json_schema (vLLM guided_grammar / xgrammar) #82816); DeepSeek returns HTTP 400 "This response_format type is unavailable now" (Auxiliary title_generation fails on DeepSeek: HTTP 400 "This response_format type is unavailable now" #83390, Bug: title_generation fails on DeepSeek provider — HTTP 400 "This response_format type is unavailable now" #84976); LM Studio's Qwen3.x (MLX) aborts and returns emptycontentunder strict structured output — there is no error for the retry/fallback machinery to react to. opencode-go/GLM-5 also relocates the entire structured response intoreasoning_content, leavingcontentempty ([Bug]: Session auto-title upgrade silently fails when provider returns json_schema output in reasoning_content (opencode-go / glm-5) #82291).Net effect: users with a local reasoning model (e.g. LM Studio + Qwen3.x) as
auxiliary.title_generationget 100% title-generation failure.Changes
agent/title_generator.py:max_tokens=64→2048— gives reasoning models headroom to finish thinking and emit the tiny JSON title. Non-reasoning models still stop right after the short answer, so there is no practical cost increase. This also eliminates the mid-JSON truncation that produced raw fragments as titles (Auto-title returns raw JSON fragments when max_tokens=64 truncates model response #83903).response_formatjson_schema→{"type": "text"}— the existing_extract_title_text()already does a JSON scan with a prose fallback (it was written for non-compliant providers), so the structured title is still parsed out of free text. Sending plain text means:json_schema, so the HTTP 400 (Session auto-title generation fails 100% of the time (HTTP 400) on OpenAI-compatible providers that reject response_format json_schema (vLLM guided_grammar / xgrammar) #82816) cannot occur.reasoning_content; it returnscontentnormally, matching the issue's own "Without response_format: works" reproduction ([Bug]: Session auto-title upgrade silently fails when provider returns json_schema output in reasoning_content (opencode-go / glm-5) #82291).Why free-text instead of falling back to
json_object?The existing PRs in this area (#82751, #85115, #84767, #83725, #83186) fall back to
json_objectwhenjson_schemais rejected. That covers backends which error onjson_schema, but not strict local backends (LM Studio MLX Qwen3.x) that abort with emptycontentunder any structured-output mode — there is no error to react to. Free-text is the lowest common denominator every backend supports, and the existing extractor already handles the shape.Verification (per reported issue)
Each reported failure mode was exercised against the patched code path:
type: textis accepted by OpenAI-compatible servers{"titleas title_extract_title_texton full JSONreasoning_content,contentemptycontentnormally (matches issue's "without response_format works" repro);_extract_title_textparses itLocal reproduction:
auxiliary.title_generation→ LM Studioqwen3.6-27b(MLX). Before: 100% of titles fail (empty content,title_sourcestaysderived). After: titles generate normally; verified across 50+ sessions.Related issues
Closes #83390, #84976, #82816, #83903, #82291.