Skip to content

feat: add self-evolution observability for normal users - #49009

Open
doubleheiker wants to merge 4 commits into
NousResearch:mainfrom
doubleheiker:feat/evolution-observability
Open

doubleheiker wants to merge 4 commits into
NousResearch:mainfrom
doubleheiker:feat/evolution-observability

Conversation

@doubleheiker

Copy link
Copy Markdown

What does this PR do?

Problem:

Hermes can improve its future behavior through durable self-evolution mechanisms like memory and skills, but users currently have no direct way to see what Hermes evolved.

Today, users can inspect separate low-level artifacts:

  • ~/.hermes/memories/MEMORY.md
  • ~/.hermes/memories/USER.md
  • ~/.hermes/skills/
  • session history
  • insights
  • curator state

But there is no user-facing timeline that answers:

  • What did Hermes learn or change?
  • Which memory or skill changed?
  • What exactly changed?
  • Why did Hermes make that durable update?
  • How has Hermes's self-evolution developed over time?

This makes Hermes's self-evolution powerful but opaque.

Solution:

Adds PR1 of Hermes Self-Evolution Observability.

This introduces a local, opt-in evolution event log for durable agent self-modifications made through memory and skill_manage. When enabled, Hermes records successful memory and skill mutations to $HERMES_HOME/evolution/events.jsonl with event metadata, summaries, optional reasons, redacted/truncated unified diffs, and CLI inspection commands.

The implementation is intentionally local and fail-open: if evolution logging fails, memory and skill operations still succeed. Evolution is disabled by default.

Related Issue

Fixes #

Type of Change

  • 🐛 Bug fix (non-breaking change that fixes an issue)
  • ✨ New feature (non-breaking change that adds functionality)
  • 🔒 Security fix
  • 📝 Documentation update
  • ✅ Tests (adding or improving test coverage)
  • ♻️ Refactor (no behavior change)
  • 🎯 New skill (bundled or hub)

Changes Made

  • Added core evolution event logging in agent/evolution_log.py
    • JSONL storage at $HERMES_HOME/evolution/events.jsonl
    • event IDs, timestamps, event schema versioning
    • unified diffs with --- before / +++ after
    • redaction and truncation metadata
    • event filtering, ID resolution, stats support, and explicit cleanup helpers
  • Added top-level CLI in hermes_cli/evolution.py
    • hermes evolution enable
    • hermes evolution disable
    • hermes evolution list
    • hermes evolution timeline
    • hermes evolution show <event-id-or-short-id>
    • hermes evolution stats
    • hermes evolution clear --older-than DAYS [--yes]
  • Wired the CLI through hermes_cli/main.py
  • Added config defaults in hermes_cli/config.py
    • evolution.enabled: false
    • evolution.record_diff: true
    • evolution.redact: true
    • evolution.max_diff_chars: 20000
  • Integrated memory observability in tools/memory_tool.py
    • records memory.add
    • records memory.replace
    • records memory.remove
    • adds optional summary and reason schema fields
  • Integrated skill observability in tools/skill_manager_tool.py
    • records skill.create
    • records skill.patch
    • records skill.edit
    • records skill.delete
    • records skill.write_file
    • records skill.remove_file
    • adds optional summary and reason schema fields
    • uses [skill deleted: content omitted] for deleted skill content
  • Added tests:
    • tests/agent/test_evolution_log.py
    • tests/hermes_cli/test_evolution_cli.py
    • tests/hermes_cli/test_evolution_config.py
    • tests/tools/test_memory_evolution.py
    • tests/tools/test_skill_manager_evolution.py
    • updated memory and skill schema tests
  • Added documentation:
    • website/docs/user-guide/features/evolution.md
  • Added delivery verification reports:
    • PROJECT_DELIVERY_VERIFICATION_REPORT.md
    • PROJECT_DELIVERY_VERIFICATION_REPORT.zh-CN.md

How to Test

  1. Run the focused PR1 suite:
.venv/bin/python3 -m pytest \
  tests/agent/test_evolution_log.py \
  tests/hermes_cli/test_evolution_cli.py \
  tests/tools/test_memory_evolution.py \
  tests/tools/test_skill_manager_evolution.py \
  tests/tools/test_memory_tool_schema.py \
  tests/tools/test_skill_manager_tool.py \
  -q -o 'addopts='

Expected result:

131 passed
  1. Run changed-area existing tests:
  .venv/bin/python3 -m pytest \
    tests/tools/test_memory_tool.py \
    tests/tools/test_skill_manager_tool.py \
    -q -o 'addopts='

  Expected result:

  160 passed
  1. Manually verify with a temporary HERMES_HOME:
  HERMES_HOME="$(mktemp -d /tmp/hermes-evolution.XXXXXX)"
  .venv/bin/python3 -m hermes_cli.main evolution list
  .venv/bin/python3 -m hermes_cli.main evolution enable

Then verify:

  - evolution list shows disabled guidance before enable
  - evolution enable creates $HERMES_HOME/evolution
  - enable does not create an empty events.jsonl
  - a successful memory.add creates an event
  - evolution list, evolution stats, and evolution show <short-id> display the event correctly

Checklist

Code

Documentation & Housekeeping

For New Skills

N/A

Screenshots / Logs

Focused PR1 suite:

  131 passed in 1.68s

  Conflict-area verification after upstream merge resolution:

  109 passed in 1.31s

  Git sanity check:

  git diff --check
  # passed

Follow-up

A follow-up PR can extend this same event log to curator-driven self-evolution:

  • curator.mark_stale
  • curator.archive
  • curator.run_summary

No-op curator runs should not emit events.

How It Works in Practice

User enables self-evolution logging:

hermes evolution enable

Later, Hermes saves a new durable memory:

User prefers concise technical explanations.

Hermes appends a local event:

{
  "schema_version": 1,
  "id": "evt_20260608_073000_a1b2c3",
  "timestamp": "2026-06-08T07:30:00Z",
  "type": "memory.add",
  "target": "memories/USER.md",
  "target_kind": "memory",
  "target_name": "user",
  "summary": "Recorded user's communication preference",
  "reason": "User asked Hermes to keep future answers concise.",
  "diff_format": "unified",
  "redaction_enabled": true,
  "redaction_applied": false,
  "diff_truncated": false
}

The user can then inspect it:

hermes evolution list
hermes evolution show a1b2c3

Copilot AI review requested due to automatic review settings June 19, 2026 12:29
@alt-glitch alt-glitch added type/feature New feature or request comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint comp/cli CLI entry point, hermes_cli/, setup wizard tool/memory Memory tool and memory providers tool/skills Skills system (list, view, manage) P3 Low — cosmetic, nice to have labels Jun 19, 2026

Copilot AI left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Pull request overview

Adds an opt-in, local “self-evolution” event log so normal users can inspect what durable changes Hermes made over time (memory + skill mutations), including a CLI to list/show/stats/clear events and documentation/tests to support the feature.

Changes:

  • Introduces agent/evolution_log.py (JSONL event log in $HERMES_HOME/evolution/events.jsonl) plus filtering/ID resolution and retention helpers.
  • Adds hermes evolution ... CLI commands and config defaults under evolution.*.
  • Wires evolution event recording into memory and skill_manage tool flows, and adds docs + test coverage.

Reviewed changes

Copilot reviewed 14 out of 14 changed files in this pull request and generated 4 comments.

Show a summary per file
File Description
agent/evolution_log.py New JSONL event log implementation, event building (diff/redaction/truncation), list/filter/clear helpers.
hermes_cli/evolution.py New hermes evolution subcommand implementation (enable/disable/list/show/stats/clear).
hermes_cli/main.py Registers the evolution CLI subcommand.
hermes_cli/config.py Adds evolution defaults to DEFAULT_CONFIG.
tools/memory_tool.py Adds summary/reason args and attempts to record memory mutation events.
tools/skill_manager_tool.py Adds summary/reason args and attempts to record skill mutation events (including per-file targets).
website/docs/user-guide/features/evolution.md User-facing documentation for enabling and using the feature.
tests/agent/test_evolution_log.py Unit tests for event ID/timestamps/diff/truncate/redaction/filter/resolve.
tests/hermes_cli/test_evolution_cli.py CLI behavior tests for enable/disable/list/show/stats/clear.
tests/hermes_cli/test_evolution_config.py Tests default evolution config is disabled.
tests/tools/test_memory_evolution.py Tests memory mutations produce evolution events (when enabled).
tests/tools/test_memory_tool_schema.py Ensures schema includes summary/reason guidance.
tests/tools/test_skill_manager_evolution.py Tests skill mutations produce evolution events (when enabled).
tests/tools/test_skill_manager_tool.py Ensures schema includes summary/reason guidance.

💡 Add Copilot custom instructions for smarter, more guided reviews. Learn how to get started.

Comment thread hermes_cli/evolution.py
Comment thread tools/memory_tool.py
Comment on lines 1059 to +1063
new_string: str = None,
replace_all: bool = False,
absorbed_into: str = None,
summary: str = None,
reason: str = None,

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed. summary and reason should be preserved through staging and replay so approved skill writes emit the same metadata as direct writes. I’ll include those fields in the staged payload and pass them through apply_skill_pending(), with a test for the write-gate path.

Comment thread agent/evolution_log.py Outdated
Comment on lines +326 to +336
if apply:
path = get_events_path()
ensure_evolution_dir()
tmp = path.with_suffix(".jsonl.tmp")
with tmp.open("w", encoding="utf-8") as f:
for event in retained:
f.write(json.dumps(event, ensure_ascii=False, sort_keys=True) + "\n")
f.flush()
os.fsync(f.fileno())
tmp.replace(path)
return deleted, len(retained)

Copy link
Copy Markdown
Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Agreed. Replacing the log file while appenders lock the opened events file can lose events in that race. I’ll switch append/clear coordination to a shared lock file acquired before opening or rewriting events.jsonl, and add a regression test around clear preserving append-safe behavior where practical.

doubleheiker and others added 3 commits June 21, 2026 11:15
update this to count categories per event and add/adjust a CLI stats test that covers repeated events of the same type.

Co-authored-by: Copilot Autofix powered by AI <175728472+Copilot@users.noreply.github.com>

@teknium1 teknium1 left a comment

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for the thorough local, opt-in observability implementation. The memory batch and approved-pending paths are covered in tests/tools/test_memory_evolution.py:95-286, and the log uses profile-safe get_hermes_home() in agent/evolution_log.py:69-76.

Problems

  • The PR's tools/memory_tool.py:964 validates target before normalizing None. Current main deliberately normalizes target is None at tools/memory_tool.py:980-984 (commit 07d93413e, #46356), with regression coverage in tests/tools/test_memory_tool.py:538. Porting the PR's function body as-is would restore rejection of strict-provider calls that emit target: null.

Suggested changes

  • Preserve the current null-target normalization while integrating the evolution hooks, and add an evolution-enabled regression case for it.

Automated hermes-sweeper review.

Comment thread tools/memory_tool.py
old_text: str = None,
operations: Optional[List[Dict[str, Any]]] = None,
store: Optional[MemoryStore] = None,
summary: str = None,

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

When porting this signature change, retain current main's target is None normalization before validation (tools/memory_tool.py:980-984, commit 07d93413e). This PR head validates target directly at line 964, which would regress the strict-provider target: null fix.

@teknium1 teknium1 added sweeper:risk-security-boundary Sweeper risk: may affect sandboxing, auth, credentials, or sensitive data sweeper:risk-compatibility Sweeper risk: may break existing users, config, migrations, defaults, or upgrades sweeper:blast-massive Sweeper blast radius: massive — everyone, every turn (invariant surface) labels Jul 14, 2026

This branch has not been deployed

No deployments
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/agent Core agent runtime: loop, agent_init, prompt builder, context-compression, responses endpoint comp/cli CLI entry point, hermes_cli/, setup wizard P3 Low — cosmetic, nice to have sweeper:blast-massive Sweeper blast radius: massive — everyone, every turn (invariant surface) sweeper:risk-compatibility Sweeper risk: may break existing users, config, migrations, defaults, or upgrades sweeper:risk-security-boundary Sweeper risk: may affect sandboxing, auth, credentials, or sensitive data tool/memory Memory tool and memory providers tool/skills Skills system (list, view, manage) type/feature New feature or request

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants