Skip to content

fix(title): omit unsupported structured output - #84767

Closed
EAbaracus wants to merge 1 commit into
NousResearch:mainfrom
EAbaracus:title-response-format-fix
Closed

EAbaracus wants to merge 1 commit into
NousResearch:mainfrom
EAbaracus:title-response-format-fix

Conversation

@EAbaracus

@EAbaracus EAbaracus commented Aug 12, 2026 •

Copy link
Copy Markdown

What does this PR do?

Fixes auxiliary session-title generation failures on providers/models that do not support OpenAI structured-output schemas.

When Hermes generates a session title, it previously always sent:

{
  "response_format": {
    "type": "json_schema",
    "json_schema": {
      "name": "session_title"
    }
  }
}

Nous + deepseek/deepseek-v4-flash-0731 rejects that request with HTTP 400, while the normal chat request succeeds.

This PR makes response_format capability-aware:

  • True → send the existing JSON Schema request format
  • False → omit response_format
  • None / unknown → omit response_format

The existing JSON/prose title parser remains unchanged.

Related Issue

Fixes #

Type of Change

  • 🐛 Bug fix (non-breaking change that fixes an issue)
  • ✨ New feature (non-breaking change that adds functionality)
  • 🔒 Security fix
  • 📝 Documentation update
  • ✅ Tests (adding or improving test coverage)
  • ♻️ Refactor (no behavior change)
  • 🎯 New skill (bundled or hub)

Changes Made

  • Added _title_request_extra_body() in agent/title_generator.py.
  • Added tri-state response_format_supported handling to generate_title().
  • Omitted empty extra_body entirely when structured output is unsupported or unknown.
  • Preserved the existing _TITLE_RESPONSE_FORMAT for explicitly supported routes.
  • Added tests covering:
    • unknown capability;
    • unsupported capability;
    • supported capability;
    • plain/prose title parsing without response_format.

Provider/model resolution, retry classification, delegation, and 503 handling are unchanged.

How to Test

Focused title-generation tests:

venv/Scripts/python.exe -m pytest \
  tests/agent/test_title_generator.py \
  -v --tb=short

Result:

37 passed

Gateway regression tests:

venv/Scripts/python.exe -m pytest \
  tests/gateway/test_compression_progress_notices.py \
  -v --tb=short

Result:

38 passed

Main delegation smoke test:

hermes chat -q "Reply with exactly: P0_DELEGATION_OK" \
  --provider nous \
  -m deepseek/deepseek-v4-flash-0731

Result:

P0_DELEGATION_OK

The full auxiliary-client suite was also run:

177 passed, 4 unavailable

The four unavailable tests are async tests requiring pytest-asyncio, which is not installed in the test environment. They were not failures caused by this change.

Checklist

Code

  • I've read the Contributing Guide
  • My commit messages follow Conventional Commits (fix(scope):, feat(scope):, etc.)
  • I searched for existing PRs to make sure this isn't a duplicate
  • My PR contains only changes related to this fix/feature (no unrelated commits)
  • I've run pytest tests/ -q and all tests pass
  • I've added tests for my changes (required for bug fixes, strongly encouraged for features)
  • I've tested on my platform: Windows 10

Documentation & Housekeeping

  • I've updated relevant documentation (README, docs/, docstrings) — N/A; behavior is covered by the implementation docstring and tests
  • I've updated cli-config.yaml.example if I added/changed config keys — N/A; no config keys changed
  • I've updated CONTRIBUTING.md or AGENTS.md if I changed architecture or workflows — N/A
  • I've considered cross-platform impact (Windows, macOS) per the compatibility guide — request-shape logic is platform-independent
  • I've updated tool descriptions/schemas if I changed tool behavior — N/A; no tool schema changed

For New Skills

  • This skill is broadly useful to most users (bundled) — N/A
  • SKILL.md follows the standard format — N/A
  • No external dependencies that aren't already available — N/A
  • I've tested the skill end-to-end — N/A

Screenshots / Logs

Before this change:

Auxiliary title generation failed:
HTTP 400: This request is not valid.
Additional info: Provider returned error

The failing route was:

provider=nous
model=deepseek/deepseek-v4-flash-0731
base_url=https://inference-api.nousresearch.com/v1

The failure occurred only on the title-generation request containing response_format=json_schema.

After this change, the request-shape tests pass for both unknown and unsupported capabilities, while explicitly supported capability still sends the existing schema.

@alt-glitch alt-glitch added type/bug Something isn't working comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint P3 Low — cosmetic, nice to have labels Aug 12, 2026
@spfcraze

Copy link
Copy Markdown

This was generated by AI during triage.

Summary:
This PR's new response_format_supported gate is never passed by a live caller, so the JSON-Schema title request is now omitted on every route and structured title output is dropped everywhere, not just on the failing provider.

Problems:

  • hermes_cli/sessions_cmd.py:961 calls generate_title(typed) with a single argument, and the auto-titler at agent/title_generator.py:604 passes failure_callback/main_runtime/runtime_validator but not the new response_format_supported.
  • With the None default, _title_request_extra_body returns {} and extra_body is omitted, so the branch that would send the JSON Schema is unreachable from production code.

Solution:
Resolve the capability at the call sites from the model's advertised supported_parameters / structured_output (parsed in hermes_cli/models.py and agent/models_dev.py) and pass response_format_supported, so routes that support structured output keep sending the JSON Schema.

Evidence

no deterministic fact backs this claim — model belief, not executed or read evidence


Checked against 9d3207b — the PR head when this was written — and fa83af3, main at the same moment.

@Enough1122

Copy link
Copy Markdown

AI code review — automated review for reference, author can ignore or act on any point.

Main concern: the new response_format_supported parameter defaults to None, and neither production call site passes it — agent/title_generator.py:604 and hermes_cli/sessions_cmd.py:953 both call generate_title(...) without it. _title_request_extra_body(None) returns {}, so response_format is now never sent: the change silently disables structured output everywhere rather than just "omitting it when unsupported." If the flag is meant to be derived from the runtime/validator, the call sites (or a resolver inside generate_title) need to supply it; otherwise this is a behavior regression beyond the stated intent.

  • No test exercises the real caller wiring — only the helper and generate_title with an explicit flag. An integration test asserting the capability flag is actually derived from the runtime would catch the silent-disable regression above.
  • Minor: _title_request_extra_body returning {} and the caller's truthiness check works, but returning None when unsupported would make the "omit entirely" intent slightly more explicit.

@EAbaracus EAbaracus closed this by deleting the head repository Sep 16, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint P3 Low — cosmetic, nice to have type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants