Repository navigation
feat: add self-evolution observability for normal users - #49009
doubleheiker wants to merge 4 commits into
Conversation
There was a problem hiding this comment.
Pull request overview
Adds an opt-in, local “self-evolution” event log so normal users can inspect what durable changes Hermes made over time (memory + skill mutations), including a CLI to list/show/stats/clear events and documentation/tests to support the feature.
Changes:
- Introduces
agent/evolution_log.py(JSONL event log in$HERMES_HOME/evolution/events.jsonl) plus filtering/ID resolution and retention helpers. - Adds
hermes evolution ...CLI commands and config defaults underevolution.*. - Wires evolution event recording into
memoryandskill_managetool flows, and adds docs + test coverage.
Reviewed changes
Copilot reviewed 14 out of 14 changed files in this pull request and generated 4 comments.
Show a summary per file
| File | Description |
|---|---|
agent/evolution_log.py |
New JSONL event log implementation, event building (diff/redaction/truncation), list/filter/clear helpers. |
hermes_cli/evolution.py |
New hermes evolution subcommand implementation (enable/disable/list/show/stats/clear). |
hermes_cli/main.py |
Registers the evolution CLI subcommand. |
hermes_cli/config.py |
Adds evolution defaults to DEFAULT_CONFIG. |
tools/memory_tool.py |
Adds summary/reason args and attempts to record memory mutation events. |
tools/skill_manager_tool.py |
Adds summary/reason args and attempts to record skill mutation events (including per-file targets). |
website/docs/user-guide/features/evolution.md |
User-facing documentation for enabling and using the feature. |
tests/agent/test_evolution_log.py |
Unit tests for event ID/timestamps/diff/truncate/redaction/filter/resolve. |
tests/hermes_cli/test_evolution_cli.py |
CLI behavior tests for enable/disable/list/show/stats/clear. |
tests/hermes_cli/test_evolution_config.py |
Tests default evolution config is disabled. |
tests/tools/test_memory_evolution.py |
Tests memory mutations produce evolution events (when enabled). |
tests/tools/test_memory_tool_schema.py |
Ensures schema includes summary/reason guidance. |
tests/tools/test_skill_manager_evolution.py |
Tests skill mutations produce evolution events (when enabled). |
tests/tools/test_skill_manager_tool.py |
Ensures schema includes summary/reason guidance. |
💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.
| new_string: str = None, | ||
| replace_all: bool = False, | ||
| absorbed_into: str = None, | ||
| summary: str = None, | ||
| reason: str = None, |
There was a problem hiding this comment.
Agreed. summary and reason should be preserved through staging and replay so approved skill writes emit the same metadata as direct writes. I’ll include those fields in the staged payload and pass them through apply_skill_pending(), with a test for the write-gate path.
| if apply: | ||
| path = get_events_path() | ||
| ensure_evolution_dir() | ||
| tmp = path.with_suffix(".jsonl.tmp") | ||
| with tmp.open("w", encoding="utf-8") as f: | ||
| for event in retained: | ||
| f.write(json.dumps(event, ensure_ascii=False, sort_keys=True) + "\n") | ||
| f.flush() | ||
| os.fsync(f.fileno()) | ||
| tmp.replace(path) | ||
| return deleted, len(retained) |
There was a problem hiding this comment.
Agreed. Replacing the log file while appenders lock the opened events file can lose events in that race. I’ll switch append/clear coordination to a shared lock file acquired before opening or rewriting events.jsonl, and add a regression test around clear preserving append-safe behavior where practical.
update this to count categories per event and add/adjust a CLI stats test that covers repeated events of the same type. Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>
teknium1
left a comment
There was a problem hiding this comment.
Thanks for the thorough local, opt-in observability implementation. The memory batch and approved-pending paths are covered in tests/tools/test_memory_evolution.py:95-286, and the log uses profile-safe get_hermes_home() in agent/evolution_log.py:69-76.
Problems
- The PR's
tools/memory_tool.py:964validatestargetbefore normalizingNone. Current main deliberately normalizestarget is Noneattools/memory_tool.py:980-984(commit07d93413e, #46356), with regression coverage intests/tools/test_memory_tool.py:538. Porting the PR's function body as-is would restore rejection of strict-provider calls that emittarget: null.
Suggested changes
- Preserve the current null-target normalization while integrating the evolution hooks, and add an evolution-enabled regression case for it.
Automated hermes-sweeper review.
| old_text: str = None, | ||
| operations: Optional[List[Dict[str, Any]]] = None, | ||
| store: Optional[MemoryStore] = None, | ||
| summary: str = None, |
There was a problem hiding this comment.
When porting this signature change, retain current main's target is None normalization before validation (tools/memory_tool.py:980-984, commit 07d93413e). This PR head validates target directly at line 964, which would regress the strict-provider target: null fix.
What does this PR do?
Problem:
Hermes can improve its future behavior through durable self-evolution mechanisms like memory and skills, but users currently have no direct way to see what Hermes evolved.
Today, users can inspect separate low-level artifacts:
But there is no user-facing timeline that answers:
This makes Hermes's self-evolution powerful but opaque.
Solution:
Adds PR1 of Hermes Self-Evolution Observability.
This introduces a local, opt-in evolution event log for durable agent self-modifications made through memory and
skill_manage. When enabled, Hermes records successful memory and skill mutations to$HERMES_HOME/evolution/events.jsonlwith event metadata, summaries, optional reasons, redacted/truncated unified diffs, and CLI inspection commands.The implementation is intentionally local and fail-open: if evolution logging fails, memory and skill operations still succeed. Evolution is disabled by default.
Related Issue
Fixes #
Type of Change
Changes Made
agent/evolution_log.py$HERMES_HOME/evolution/events.jsonl--- before/+++ afterhermes_cli/evolution.pyhermes evolution enablehermes evolution disablehermes evolution listhermes evolution timelinehermes evolution show <event-id-or-short-id>hermes evolution statshermes evolution clear --older-than DAYS [--yes]hermes_cli/main.pyhermes_cli/config.pyevolution.enabled: falseevolution.record_diff: trueevolution.redact: trueevolution.max_diff_chars: 20000tools/memory_tool.pymemory.addmemory.replacememory.removesummaryandreasonschema fieldstools/skill_manager_tool.pyskill.createskill.patchskill.editskill.deleteskill.write_fileskill.remove_filesummaryandreasonschema fields[skill deleted: content omitted]for deleted skill contenttests/agent/test_evolution_log.pytests/hermes_cli/test_evolution_cli.pytests/hermes_cli/test_evolution_config.pytests/tools/test_memory_evolution.pytests/tools/test_skill_manager_evolution.pywebsite/docs/user-guide/features/evolution.mdPROJECT_DELIVERY_VERIFICATION_REPORT.mdPROJECT_DELIVERY_VERIFICATION_REPORT.zh-CN.mdHow to Test
.venv/bin/python3 -m pytest \ tests/agent/test_evolution_log.py \ tests/hermes_cli/test_evolution_cli.py \ tests/tools/test_memory_evolution.py \ tests/tools/test_skill_manager_evolution.py \ tests/tools/test_memory_tool_schema.py \ tests/tools/test_skill_manager_tool.py \ -q -o 'addopts=' Expected result: 131 passed.venv/bin/python3 -m pytest \ tests/tools/test_memory_tool.py \ tests/tools/test_skill_manager_tool.py \ -q -o 'addopts=' Expected result: 160 passedThen verify:
Checklist
Code
I've read the Contributing Guide (https://github.com/NousResearch/hermes-agent/blob/main/CONTRIBUTING.md)
My commit messages follow Conventional Commits (https://www.conventionalcommits.org/) (fix(scope):, feat(scope):,
etc.)
I searched for existing PRs (https://github.com/NousResearch/hermes-agent/pulls) to make sure this isn't a
duplicate
My PR contains only changes related to this fix/feature (no unrelated commits)
I've run pytest tests/ -q and all tests pass
I've added tests for my changes (required for bug fixes, strongly encouraged for features)
I've tested on my platform: Ubuntu/Linux
Documentation & Housekeeping
I've updated relevant documentation (README, docs/, docstrings) — or N/A
I've updated cli-config.yaml.example if I added/changed config keys — or N/A
I've updated CONTRIBUTING.md or AGENTS.md if I changed architecture or workflows — or N/A
I've considered cross-platform impact (Windows, macOS) per the compatibility guide
(https://github.com/NousResearch/hermes-agent/blob/main/CONTRIBUTING.md#cross-platform-compatibility) — or N/A
I've updated tool descriptions/schemas if I changed tool behavior — or N/A
For New Skills
N/A
Screenshots / Logs
Focused PR1 suite:
Follow-up
A follow-up PR can extend this same event log to curator-driven self-evolution:
curator.mark_stalecurator.archivecurator.run_summaryNo-op curator runs should not emit events.
How It Works in Practice
User enables self-evolution logging:
hermes evolution enableLater, Hermes saves a new durable memory:
Hermes appends a local event:
{ "schema_version": 1, "id": "evt_20260608_073000_a1b2c3", "timestamp": "2026-06-08T07:30:00Z", "type": "memory.add", "target": "memories/USER.md", "target_kind": "memory", "target_name": "user", "summary": "Recorded user's communication preference", "reason": "User asked Hermes to keep future answers concise.", "diff_format": "unified", "redaction_enabled": true, "redaction_applied": false, "diff_truncated": false }The user can then inspect it: