Skip to content

fix(agent): retry title generation with json_object on HTTP 400 - #82868

Closed
chelsealong wants to merge 1 commit into
NousResearch:mainfrom
chelsealong:fix/title-generation-json-schema-400-fallback
Closed

chelsealong wants to merge 1 commit into
NousResearch:mainfrom
chelsealong:fix/title-generation-json-schema-400-fallback

Conversation

@chelsealong

Copy link
Copy Markdown

What does this PR do?

generate_title() hardcodes a strict response_format: json_schema on every
title-generation call, with no fallback. Providers that reject json_schema
get a 400 that lands in the generic except bucket as non-retryable — the
session keeps its truncated derived title forever, with no user-visible
error.

This adds a single retry: if the json_schema attempt fails with an HTTP
400, retry once with response_format: json_object, which these providers
accept and which the existing loose-JSON extraction (_extract_title_text)
already tolerates. Detection is via status_code (set by the openai SDK)
with a text fallback for the "Error code: 400 - ..." message shape, so it
also catches rejections that never mention response_format by name — e.g.
vLLM translating json_schema into an uncompilable guided_grammar and
400ing with a compile_grammar_error: No module named 'xgrammar' message.

Related Issue

Fixes #82816

Type of Change

  • 🐛 Bug fix (non-breaking change that fixes an issue)

Changes Made

  • agent/title_generator.py: added _response_format_rejected(); generate_title() now retries the title call with json_object when the json_schema attempt 400s, instead of giving up.
  • tests/agent/test_title_generator.py: added tests for the SDK-status_code-400 case, the vLLM guided_grammar/xgrammar message shape (no status_code attribute, no mention of response_format), and a regression guard that an unrelated failure (e.g. a 402) does not retry.

How to Test

  1. Configure auxiliary.title_generation to point at an OpenAI-compatible endpoint that rejects response_format: json_schema (any vLLM gateway without xgrammar, or DeepSeek/Kimi).
  2. Before the fix: every session's title stays the truncated first-message text; agent.log logs Title generation failed: Error code: 400 - ... on every new session.
  3. After the fix: the title call retries with json_object and the session gets an LLM-generated title.

Test output (in a local venv with the project's pinned core deps + pytest):

$ python -m pytest tests/agent/test_title_generator.py -q
....................................                                     [100%]
36 passed in 1.24s

Confirmed the two new tests fail without the fix (git checkout HEAD~1 -- agent/title_generator.py — the commit before this one — then rerun):

FAILED tests/agent/test_title_generator.py::TestGenerateTitle::test_retries_with_json_object_on_http_400_status_code - AssertionError: assert None == 'Fix login button'
FAILED tests/agent/test_title_generator.py::TestGenerateTitle::test_retries_on_vllm_guided_grammar_rejection - AssertionError: assert None == 'Debug xgrammar error'
2 failed, 1 passed, 33 deselected in 0.39s

Also ran the neighboring suite to check for regressions:

$ python -m pytest tests/agent/test_auxiliary_client.py -q
180 passed in 7.41s

ruff check agent/title_generator.py tests/agent/test_title_generator.py — all checks passed.

Checklist

Code

Documentation & Housekeeping

  • N/A — no config keys, docs, or architecture changed

AI assistance disclosure

This PR was prepared with AI assistance (Claude, Anthropic) under human supervision: the agent read the issue and the four related open PRs, identified the gap (none of them catch the vLLM guided_grammar/xgrammar 400, since its message never says response_format), wrote the minimal fix and tests, and verified the tests fail on the pre-fix code and pass after.

generate_title() hardcodes a strict json_schema response_format with no
escape hatch. Providers that reject it (vLLM translating json_schema into
an uncompilable guided_grammar when xgrammar isn't installed, DeepSeek/Kimi
rejecting the type outright) 400 on every call, and the 400 lands in the
generic except bucket as non-retryable — the derived title sticks forever
with no user-visible error.

Retry once with response_format: json_object (which these providers accept
and the existing loose-JSON extraction already tolerates) when the
json_schema attempt fails with an HTTP 400. Detection is via status_code
(set by the openai SDK) with a text fallback for the "Error code: 400 - ..."
message shape, so it also catches rejections that never mention
"response_format" by name, like vLLM's guided_grammar/xgrammar error.

Fixes NousResearch#82816
@alt-glitch alt-glitch added type/bug Something isn't working P3 Low — cosmetic, nice to have comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint provider/openai OpenAI / Codex Responses API labels Aug 10, 2026
@alt-glitch

Copy link
Copy Markdown

This was generated by AI during triage.

Related: #82073 uses a narrower response-format rejection predicate, while #82372 uses a static provider-aware downgrade. This PR's HTTP-400 fallback also covers the vLLM guided_grammar case from #82816; maintainers should choose the retry scope.

viz-A-viz added a commit to viz-A-viz/hermes-agent that referenced this pull request Aug 10, 2026
Title generation hardcodes a strict json_schema response_format with no
fallback. Providers without structured-output support (DeepSeek returns
HTTP 400 "This response_format type is unavailable now") fail the whole
call and the session keeps its truncated derived name.

Walk a constraint ladder instead: json_schema -> json_object -> no
response_format, pinning thinking off on retries so default-on reasoning
models (DeepSeek V4) don't burn the 64-token budget on reasoning and
return an empty content field. Failures unrelated to response_format
(auth, quota, network) break out immediately - a different format cannot
fix those.

Unlike the other open PRs for this bug (NousResearch#82073, NousResearch#82372, NousResearch#82751, NousResearch#82868,
NousResearch#82890), the retried calls also send thinking: {"type": "disabled"} -
without it DeepSeek answers with an empty content and the title still
never appears, even though the 400 is gone.
@teknium1

Copy link
Copy Markdown
Collaborator

Thanks @chelsealong for working on the title-generation 400 on json_schema. This landed on main through #89589 and #113966 (5cc8177), which covers the same symptom on the primary and fallback auxiliary paths and omits the field up front for providers known to reject it. Closing as superseded by the landed fix — the tracking issue (#83390 cluster) is closed with the same references.

@teknium1 teknium1 closed this Sep 17, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint P3 Low — cosmetic, nice to have provider/openai OpenAI / Codex Responses API type/bug Something isn't working

Projects

None yet

3 participants