fix(agent): retry title generation with json_object on HTTP 400 - #82868
Closed
chelsealong wants to merge 1 commit into
Closed
chelsealong wants to merge 1 commit into
chelsealong wants to merge 1 commit into
Conversation
generate_title() hardcodes a strict json_schema response_format with no escape hatch. Providers that reject it (vLLM translating json_schema into an uncompilable guided_grammar when xgrammar isn't installed, DeepSeek/Kimi rejecting the type outright) 400 on every call, and the 400 lands in the generic except bucket as non-retryable — the derived title sticks forever with no user-visible error. Retry once with response_format: json_object (which these providers accept and the existing loose-JSON extraction already tolerates) when the json_schema attempt fails with an HTTP 400. Detection is via status_code (set by the openai SDK) with a text fallback for the "Error code: 400 - ..." message shape, so it also catches rejections that never mention "response_format" by name, like vLLM's guided_grammar/xgrammar error. Fixes NousResearch#82816
3 tasks done
viz-A-viz
added a commit
to viz-A-viz/hermes-agent
that referenced
this pull request
Aug 10, 2026
Title generation hardcodes a strict json_schema response_format with no fallback. Providers without structured-output support (DeepSeek returns HTTP 400 "This response_format type is unavailable now") fail the whole call and the session keeps its truncated derived name. Walk a constraint ladder instead: json_schema -> json_object -> no response_format, pinning thinking off on retries so default-on reasoning models (DeepSeek V4) don't burn the 64-token budget on reasoning and return an empty content field. Failures unrelated to response_format (auth, quota, network) break out immediately - a different format cannot fix those. Unlike the other open PRs for this bug (NousResearch#82073, NousResearch#82372, NousResearch#82751, NousResearch#82868, NousResearch#82890), the retried calls also send thinking: {"type": "disabled"} - without it DeepSeek answers with an empty content and the title still never appears, even though the 400 is gone.
Collaborator
|
Thanks @chelsealong for working on the title-generation 400 on json_schema. This landed on main through #89589 and #113966 (5cc8177), which covers the same symptom on the primary and fallback auxiliary paths and omits the field up front for providers known to reject it. Closing as superseded by the landed fix — the tracking issue (#83390 cluster) is closed with the same references. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
generate_title()hardcodes a strictresponse_format: json_schemaon everytitle-generation call, with no fallback. Providers that reject
json_schemaget a 400 that lands in the generic
exceptbucket as non-retryable — thesession keeps its truncated derived title forever, with no user-visible
error.
This adds a single retry: if the
json_schemaattempt fails with an HTTP400, retry once with
response_format: json_object, which these providersaccept and which the existing loose-JSON extraction (
_extract_title_text)already tolerates. Detection is via
status_code(set by the openai SDK)with a text fallback for the
"Error code: 400 - ..."message shape, so italso catches rejections that never mention
response_formatby name — e.g.vLLM translating
json_schemainto an uncompilableguided_grammarand400ing with a
compile_grammar_error: No module named 'xgrammar'message.Related Issue
Fixes #82816
Type of Change
Changes Made
agent/title_generator.py: added_response_format_rejected();generate_title()now retries the title call withjson_objectwhen thejson_schemaattempt 400s, instead of giving up.tests/agent/test_title_generator.py: added tests for the SDK-status_code-400 case, the vLLMguided_grammar/xgrammarmessage shape (nostatus_codeattribute, no mention ofresponse_format), and a regression guard that an unrelated failure (e.g. a 402) does not retry.How to Test
auxiliary.title_generationto point at an OpenAI-compatible endpoint that rejectsresponse_format: json_schema(any vLLM gateway withoutxgrammar, or DeepSeek/Kimi).agent.loglogsTitle generation failed: Error code: 400 - ...on every new session.json_objectand the session gets an LLM-generated title.Test output (in a local venv with the project's pinned core deps + pytest):
Confirmed the two new tests fail without the fix (
git checkout HEAD~1 -- agent/title_generator.py— the commit before this one — then rerun):Also ran the neighboring suite to check for regressions:
ruff check agent/title_generator.py tests/agent/test_title_generator.py— all checks passed.Checklist
Code
fix(scope):)response_formatfallback problem, but each is scoped to a specific provider (DeepSeek/Kimi hardcoded list, or Anthropic'sextra_bodypassthrough) or gates the retry on the error message containing the literal string"response_format". None of those cover this issue's vLLMguided_grammar/xgrammarfailure, whose error text never mentionsresponse_format. This PR'sstatus_code-based detection is provider-agnostic and covers that case.Documentation & Housekeeping
AI assistance disclosure
This PR was prepared with AI assistance (Claude, Anthropic) under human supervision: the agent read the issue and the four related open PRs, identified the gap (none of them catch the vLLM
guided_grammar/xgrammar400, since its message never saysresponse_format), wrote the minimal fix and tests, and verified the tests fail on the pre-fix code and pass after.