fix(hermes): hallucination guard — strip fabricated tool output - #1
Merged
Conversation
Add a three-layer defense against LLM hallucinated tool output:
1. Streaming token buffer: detects and strips fake {"output": ..., "exit_code": ...}
JSON blocks from the SSE token stream in real-time using pattern matching
2. Post-run scrub: scans all assistant messages after agent.run_conversation()
and replaces any remaining fabricated output blocks
3. Anti-hallucination system prompt: ephemeral instruction injected into every
agent run that explicitly forbids writing fake tool output
Includes 25 unit tests covering detection, stripping, and prompt validation.
Ported from upstream hermes-webui PR nesquena#133.
Co-Authored-By: Claude Opus 4.6 <noreply@anthropic.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Summary
api/streaming.pythat detects and strips fabricated tool output ({"output": ..., "exit_code": ...}) from LLM responsesrun_conversation()completes, replacing any fabricated output blocksPorted from upstream hermes-webui PR nesquena#133.
Test plan
pytest tests/test_hallucination_guard.py -v)cd ~/zen-console && git pull origin main && sudo systemctl restart zen-console🤖 Generated with Claude Code