You signed in with another tab or window. Reload to refresh your session.You signed out in another tab or window. Reload to refresh your session.You switched accounts on another tab or window. Reload to refresh your session.Dismiss alert
hermes chat -q "Reply with exactly: P0_DELEGATION_OK" \
--provider nous \
-m deepseek/deepseek-v4-flash-0731
Result:
P0_DELEGATION_OK
The full auxiliary-client suite was also run:
177 passed, 4 unavailable
The four unavailable tests are async tests requiring pytest-asyncio, which is not installed in the test environment. They were not failures caused by this change.
Checklist
Code
I've read the Contributing Guide
My commit messages follow Conventional Commits (fix(scope):, feat(scope):, etc.)
I searched for existing PRs to make sure this isn't a duplicate
My PR contains only changes related to this fix/feature (no unrelated commits)
I've run pytest tests/ -q and all tests pass
I've added tests for my changes (required for bug fixes, strongly encouraged for features)
I've tested on my platform: Windows 10
Documentation & Housekeeping
I've updated relevant documentation (README, docs/, docstrings) — N/A; behavior is covered by the implementation docstring and tests
I've updated cli-config.yaml.example if I added/changed config keys — N/A; no config keys changed
I've updated CONTRIBUTING.md or AGENTS.md if I changed architecture or workflows — N/A
I've considered cross-platform impact (Windows, macOS) per the compatibility guide — request-shape logic is platform-independent
I've updated tool descriptions/schemas if I changed tool behavior — N/A; no tool schema changed
For New Skills
This skill is broadly useful to most users (bundled) — N/A
SKILL.md follows the standard format — N/A
No external dependencies that aren't already available — N/A
I've tested the skill end-to-end — N/A
Screenshots / Logs
Before this change:
Auxiliary title generation failed:
HTTP 400: This request is not valid.
Additional info: Provider returned error
The failure occurred only on the title-generation request containing response_format=json_schema.
After this change, the request-shape tests pass for both unknown and unsupported capabilities, while explicitly supported capability still sends the existing schema.
Summary:
This PR's new response_format_supported gate is never passed by a live caller, so the JSON-Schema title request is now omitted on every route and structured title output is dropped everywhere, not just on the failing provider.
Problems:
hermes_cli/sessions_cmd.py:961 calls generate_title(typed) with a single argument, and the auto-titler at agent/title_generator.py:604 passes failure_callback/main_runtime/runtime_validator but not the new response_format_supported.
With the None default, _title_request_extra_body returns {} and extra_body is omitted, so the branch that would send the JSON Schema is unreachable from production code.
Solution:
Resolve the capability at the call sites from the model's advertised supported_parameters / structured_output (parsed in hermes_cli/models.py and agent/models_dev.py) and pass response_format_supported, so routes that support structured output keep sending the JSON Schema.
Evidence
no deterministic fact backs this claim — model belief, not executed or read evidence
Checked against 9d3207b — the PR head when this was written — and fa83af3, main at the same moment.
AI code review — automated review for reference, author can ignore or act on any point.
Main concern: the new response_format_supported parameter defaults to None, and neither production call site passes it — agent/title_generator.py:604 and hermes_cli/sessions_cmd.py:953 both call generate_title(...) without it. _title_request_extra_body(None) returns {}, so response_format is now never sent: the change silently disables structured output everywhere rather than just "omitting it when unsupported." If the flag is meant to be derived from the runtime/validator, the call sites (or a resolver inside generate_title) need to supply it; otherwise this is a behavior regression beyond the stated intent.
No test exercises the real caller wiring — only the helper and generate_title with an explicit flag. An integration test asserting the capability flag is actually derived from the runtime would catch the silent-disable regression above.
Minor: _title_request_extra_body returning {} and the caller's truthiness check works, but returning None when unsupported would make the "omit entirely" intent slightly more explicit.
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
comp/agentCore agent runtime: loop, agent_init, prompt builder, context-compression, responses endpointP3Low — cosmetic, nice to havetype/bugSomething isn't working
4 participants
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What does this PR do?
Fixes auxiliary session-title generation failures on providers/models that do not support OpenAI structured-output schemas.
When Hermes generates a session title, it previously always sent:
{ "response_format": { "type": "json_schema", "json_schema": { "name": "session_title" } } }Nous +
deepseek/deepseek-v4-flash-0731rejects that request with HTTP 400, while the normal chat request succeeds.This PR makes
response_formatcapability-aware:True→ send the existing JSON Schema request formatFalse→ omitresponse_formatNone/ unknown → omitresponse_formatThe existing JSON/prose title parser remains unchanged.
Related Issue
Fixes #
Type of Change
Changes Made
_title_request_extra_body()inagent/title_generator.py.response_format_supportedhandling togenerate_title().extra_bodyentirely when structured output is unsupported or unknown._TITLE_RESPONSE_FORMATfor explicitly supported routes.response_format.Provider/model resolution, retry classification, delegation, and 503 handling are unchanged.
How to Test
Focused title-generation tests:
Result:
Gateway regression tests:
Result:
Main delegation smoke test:
hermes chat -q "Reply with exactly: P0_DELEGATION_OK" \ --provider nous \ -m deepseek/deepseek-v4-flash-0731Result:
The full auxiliary-client suite was also run:
The four unavailable tests are async tests requiring
pytest-asyncio, which is not installed in the test environment. They were not failures caused by this change.Checklist
Code
fix(scope):,feat(scope):, etc.)pytest tests/ -qand all tests passDocumentation & Housekeeping
cli-config.yaml.exampleif I added/changed config keys — N/A; no config keys changedCONTRIBUTING.mdorAGENTS.mdif I changed architecture or workflows — N/AFor New Skills
SKILL.mdfollows the standard format — N/AScreenshots / Logs
Before this change:
The failing route was:
The failure occurred only on the title-generation request containing
response_format=json_schema.After this change, the request-shape tests pass for both unknown and unsupported capabilities, while explicitly supported capability still sends the existing schema.