Skip to content

feat(research): AutoResearch loop — Karpathy + Autogenesis AOOR, TaskSpec, researcher profile - #1

Merged
nicoechaniz merged 150 commits into
nicoechaniz:mainfrom
Fede654:feat/hermes-autoresearch-upstream
May 17, 2026
Merged

nicoechaniz merged 150 commits into
nicoechaniz:mainfrom
Fede654:feat/hermes-autoresearch-upstream

Conversation

@Fede654

@Fede654 Fede654 commented Apr 22, 2026

Copy link
Copy Markdown

Summary

Adds a self-improving research loop to Hermes Agent, wired as a first-class LLM-callable tool and a bootstrappable researcher profile.

The loop is domain-agnostic: any task with a measurable deliverable can be run iteratively. It combines two frameworks:

  • Karpathy inner loop: baseline → iterate → keep/discard (think before coding, surgical changes, verifiable metric movements)
  • Autogenesis AOOR (Act → Observe → Optimize → Remember): each round produces a structured learnings.jsonl entry; on early stop a reflection pass diagnoses why the metric stalled

New files

File Purpose
agent/research_supervisor.py Core: TaskSpec + ResearchSupervisor.run(), _observe(), _reflect(), _improve_attempt(), _score_with_llm_judge()
agent/research_runner.py ExperimentRunner: Karpathy keep/discard loop, ExperimentHistory
agent/research_metrics.py UniversalMetricParser: JSON → CSV → stdout fallback
agent/research_evolution.py Ported from AutoResearchClaw (MIT)
tools/research_tool.py run_research tool — LLM calls it like delegate_task; registers in "research" toolset
hermes_cli/researcher_scaffold.py hermes profile setup researcher — writes config, SOUL.md, MEMORY.md
prompts/autoresearch.yaml Worker prompt templates (compute budget, Karpathy guidelines, metric reporting)
skills/autoresearch/ 13 skills (a-evolve, karpathy-guidelines, hypothesis-formulation, literature-search, scientific-writing, statistical-reporting, 7 domain skills)
HERMES_RESEARCH.md Architecture doc for the research system
RESEARCH_AGENTS.md Worker agent protocol
tests/agent/test_research_supervisor.py 18 tests (9 unit, 9 integration)

Supported task types

TaskSpec(
    topic="Find attention mechanism papers post-2022",
    deliverable="Ranked list with abstracts",
    metric_key="relevance_score",
    task_type="search",           # code | search | research | generic
    evaluation_mode="llm_judge",  # self_report | llm_judge
    evaluation_prompt="Score 0-1: coverage of attention mechanisms after 2022",
)

Bootstrap flow

hermes profile create researcher
hermes profile setup researcher   # writes config.yaml, SOUL.md, MEMORY.md
researcher chat                   # run_research tool available

Autogenesis integration

  • _observe(): HeartbeatMemorySystem schema {type, key, insight, confidence, source} → learnings.jsonl
  • _reflect(): SEPL reflection optimizer on early stop — LLM diagnoses bottleneck, posts to Lattice
  • SEPL rollback: on regression, restores best on-disk artifact (not seed string)

Credits

Test plan

  • pytest tests/agent/test_research_supervisor.py -m "not integration" — 9 unit tests pass
  • pytest tests/agent/test_research_supervisor.py::TestResearchSupervisorBaseline — integration tests pass
  • python -c "from tools.research_tool import run_research; from tools.registry import registry; print(registry._tools['run_research'].toolset)" → research
  • hermes profile create researcher && hermes profile setup researcher — creates profile with correct SOUL.md
  • hermes -p researcher chat with a research prompt triggers run_research

🤖 Generated with Claude Code

nicoechaniz pushed a commit that referenced this pull request Apr 22, 2026
- entry.tsx no longer writes bootBanner() to the main screen before the
  alt-screen enters. The <Banner> renders inside the alt screen via the
  seeded intro row, so nothing is lost — just the flash that preceded it.
  Fixes the torn first frame reported on Alacritty (blitz row 5 #17) and
  shaves the 'starting agent' hang perception (row 5 #1) since the UI
  paints straight into the steady-state view
- AlternateScreen prefixes ERASE_SCROLLBACK (\x1b[3J) to its entry so
  strict emulators start from a pristine grid; named constants replace
  the inline sequences for clarity
- bootBanner.ts deleted — dead code
@Fede654

Fede654 commented Apr 23, 2026

Copy link
Copy Markdown
Author

🔄 Revisión compactada — cambios empujados

Commits nuevos en esta rama (18 total)

7ea25c46 docs(autoresearch): update research docs with latest architecture and procedures
61c6e994 fix(autoresearch): restore LLM judge on every iteration
8e68e23d perf(autoresearch): implement 5 optimization fixes from log analysis
cc929c0a feat(autoresearch): add research_job orchestration for long-running loops
a72abd4c feat(autoresearch): add research_job tool for detached long-running loops
4d64d568 fix(autoresearch): include terminal in research/search default toolsets
0da707a3 fix(autoresearch): add Tools Available + anti-XML guard to task briefs
f08dc63c fix(autoresearch): correct sandbox messaging and add partial recovery
8c417d45 fix(autoresearch): apply analysis fixes from full cycle validation
7b158eed fix(autoresearch): apply Codex ia-bridge review findings
694ed47a feat(research): wire ResearchSupervisor as tool + researcher profile scaffold
b42b1a60 fix(research): address audit findings — rollback, observe, reflect guards
165f9bf4 feat(research): incorporate Autogenesis AOOR loop and HeartbeatMemorySystem
c5d9da4d feat(research): generalize supervisor loop to any measurable task type
e1b0c0bc feat(autoresearch): apply Karpathy guidelines to experiment loop
25a61a56 test(autoresearch): add integration tests for ResearchSupervisor
3a5e6977 feat(autoresearch): implement ResearchSupervisor and program.md template
c51a2d2c feat(autoresearch): vendor AutoResearchClaw skills and prompts

Validación end-to-end (4 jobs completados)

Job Iteraciones Best Metric Wall Time
smoke-test-001 2 pass_rate=1.0 ~2 min
meta-benchmark-v2 4 completeness_score=0.95 ~20 min
perf-analysis 4 completeness_score=1.0 ~46 min
perf-validation 4 completeness_score=0.96 ~27 min

Fixes críticos aplicados post-validación

  1. Terminal en default toolsets — workers ahora pueden ejecutar código
  2. Anti-XML guard — previne function_calls de kimi-coding
  3. Partial recovery — última iteración fallida no mata todo el job
  4. Lock file (O_EXCL) — previene restarts duplicados (~16 min ahorrados)
  5. Provider cache — ~14 resoluciones → 1 por config
  6. Subdirectory hints cache — elimina lecturas repetidas de disco
  7. Early stop agresivo — baseline >=0.9 → stop en 1 iteración sin mejora
  8. Judge en todas las iteraciones — revirió optimización que saltaba iteraciones impares

Documentación actualizada

  • HERMES_RESEARCH.md — arquitectura desacoplada, TaskSpec, checkpoints
  • RESEARCH_AGENTS.md — contrato worker actualizado (task_brief.md, JSON tools)
  • RESEARCH_OPERATIONS.md — guía de operación nueva (launch, monitor, recovery)

Estado

Rama empujada a Fede654/hermes-agent:feat/hermes-autoresearch-upstream. Lista para review.

@nicoechaniz

Copy link
Copy Markdown
Owner

╭─ ⚕ Hermes review ──────────────────────────────
Now I have a comprehensive view of the PR. Let me provide a detailed evaluation:

Pull Request Evaluation: #1 — AutoResearch Loop

Author: Fede654 (Federico Bonino)  
Branch: feat/hermes-autoresearch-upstream  
Status: Review in progress



Summary

This PR adds a substantial AutoResearch system to Hermes Agent — an iterative self-improving research loop combining the "Karpathy inner loop" (baseline → iterate → keep/discard) with Autogenesis AOOR (Act → Observe → Optimize → Remember). The system is exposed as a new run_research tool and a researcher profile scaffold.

Size: +5,579 lines / -12 lines across 29 files  
New files: 22 (core modules, skills, docs, tests)  
Modified files: 7 (existing codebase)



What Works Well

1. Architecture is sound — Clean separation between ResearchSupervisor (orchestration), ExperimentRunner (Karpathy loop), and worker subagents via delegate_task. The checkpoint system enables durable, resumable research jobs.

2. Test coverage exists — 18 tests (9 unit + 9 integration), all passing:
   bash
   9 passed (unit) + 9 passed (integration) = 18/18 tests passing
   

3. Existing tests unaffected — 133 toolset-related tests pass; auxiliary_client tests (80) pass.

4. Good documentation — Three comprehensive docs:
   - HERMES_RESEARCH.md — Architecture and usage guide
   - RESEARCH_AGENTS.md — Worker agent contract
   - RESEARCH_OPERATIONS.md — Operational procedures and anti-patterns

5. Key fixes in existing code:
   - auxiliary_client.py: Added provider client caching (reduces auth resolution calls)
   - delegate_tool.py: Safer getattr() for parent_agent attributes
   - subdirectory_hints.py: Added caching for hint loading
   - code_execution_tool.py: Increased timeout from 5min → 15min + hot-reload config loading



⚠️ Issues Requiring Changes

1. Hardcoded paths (blocking)

Three hardcoded paths to the author's environment need to be made configurable:

| File | Line | Issue | Suggested Fix |
|------|------|-------|---------------|
| agent/research_supervisor.py | 403 | lattice_root: str = "/home/fede/.hermes/org" | Use environment variable or config with sensible default |
| tools/research_job_tool.py | 168 | hermes_root = Path("/home/fede/.hermes/hermes-agent") | Use get_hermes_home() or HERMES_HOME env var |
| tools/research_job_tool.py | 291 | Same issue as above | Same fix |

2. Non-profile-aware config loading

tools/research_job_tool.py:38 uses Path.home() / ".hermes" instead of get_hermes_home():
python
Current (incorrect):
config_path = Path.home() / ".hermes" / "config.yaml"

Should be:
from hermes_constants import get_hermes_home
config_path = get_hermes_home() / "config.yaml"


3. Code execution timeout change is global

The 5min → 15min timeout increase in code_execution_tool.py affects all execute_code calls, not just research jobs. This could mask performance issues in unrelated code. Consider making this configurable per-task or via config.yaml override.

4. Missing .lattice/ handling

The research_job_tool.py spawns background processes that expect a Hermes checkout at a specific path. If the user's install is different (e.g., pip install vs git clone), the job runner will fail. The tool should either:
- Accept hermes_root as a parameter
- Detect the running Hermes installation automatically
- Document this requirement clearly



💡 Suggestions (Non-blocking)

1. Karpathy guidelines skill — The PR includes skills/autoresearch/karpathy-guidelines/ which is excellent. Consider making this auto-injected into researcher profiles.

2. Lattice integration — The Lattice comment function in research_supervisor.py shells out to lattice CLI. This is fine but could be more robust with a Python API fallback.

3. Lock file cleanup — The research job runner uses .runner.lock files. Add documentation on how users should clean up stale locks (e.g., after crashes).



Test Results


Unit tests (9):           PASSED
Integration tests (9):    PASSED
Toolset tests (133):      PASSED
Auxiliary client (80):    PASSED




Verdict

Request Changes — The PR is architecturally sound and well-tested, but the hardcoded paths (/home/fede/.hermes/...) make it non-portable. These need to be configurable before merge.

Recommended fixes:
1. Replace hardcoded paths with environment variables or config options
2. Use get_hermes_home() instead of Path.home() / ".hermes"
3. Document the HERMES_HOME requirement for research jobs
4. Consider making the code execution timeout configurable rather than global

After these changes, this PR should be ready for merge. The AutoResearch system is a valuable addition to Hermes.

@Fede654
Fede654 force-pushed the feat/hermes-autoresearch-upstream branch from 7ea25c4 to 0fe9364 Compare April 24, 2026 03:25
@Fede654

Fede654 commented Apr 24, 2026

Copy link
Copy Markdown
Author

🔧 Review fixes empujados

Aplicados los 4 issues bloqueantes + 2 sugerencias de valor.

Bloqueantes resueltos

# Issue Commit
1 Hardcoded paths en research_supervisor.py y research_job_tool.py d29cd35b + a86d9328 (workspace fallback residual)
2 Path.home() / ".hermes" → get_hermes_home() en research_job_tool.py:38 d29cd35b
3 Timeout global 5→15min en code_execution_tool.py 51502d22 (revertido a 300s, RPC timeout configurable por task)
4 .lattice/ requirement no documentado 70f4e4ea (RESEARCH_OPERATIONS.md actualizado)

Sugerencias no-bloqueantes aplicadas (0fe93646)

  • Karpathy guidelines surfaced: el task brief ahora apunta explícitamente a skills/autoresearch/karpathy-guidelines/SKILL.md desde Step 0, así los workers ven el ruleset completo (no solo "Principle 1").
  • Stale-lock recovery docs: añadida sección con check de PID antes de borrar .runner.lock, para evitar que el usuario mate un runner vivo y corrompa el checkpoint.

Skipped

  • Lattice Python API fallback: el shell-out al CLI funciona y un binding Python sería refactor mayor por valor marginal. Si lo querés como hard requirement, abro PR aparte.

Validación

  • 18/18 tests pasan en tests/agent/test_research_supervisor.py
  • Sin regresiones en auxiliary_client ni toolset tests

Listo para re-review.

@Fede654
Fede654 force-pushed the feat/hermes-autoresearch-upstream branch from fb115af to 24154da Compare April 25, 2026 01:08
@Fede654

Fede654 commented Apr 25, 2026

Copy link
Copy Markdown
Author

🔄 Update — segundo round de cambios desde tu review

Tu review (Apr 23) cubrió ~25 commits. Pusheamos 12 más desde entonces, agrupados en dos batches:

Batch 1 — fixes review (los 4 bloqueantes + 2 sugerencias) ✅

Batch 2 — refactors arquitecturales auto-iniciados

Producto de un análisis fresco del PR (audit en vault/raw/hermes-repo/autoresearch-pr-analysis.md) y 5 spawns de researcher autónomos sobre Lattice tasks HRM-57..61.

Commit Cambio Por qué
b89883cb researcher profile drops mcp-obsidian Vault es plain Markdown + git; MCP layer = opacidad sin valor
9891835c researcher profile drops mcp-lattice Lattice tiene CLI nativo
37e0292e default kimi-k2.6 / kimi-coding El default claude-sonnet-4-6 rompe en planes Anthropic sin extra-usage
ee764dc4 SOUL.md/MEMORY.md operational patterns Lecciones del swarm de 5 researchers
24832976 delete prompts/autoresearch.yaml (HRM-60) Vendoreado, nunca cargado, redundante con inline templates
75437153 inherit_profile opt-in en delegate_task (HRM-58) Research workers heredan SOUL/MEMORY del researcher profile en vez de blank-slate
c106c6ce EvolutionStore wiring v1 (HRM-59) _evolve() post-loop persiste lessons a ~/.hermes/evolution/lessons.jsonl
3e331a27 agent/factory.py (HRM-57 partial) Centraliza construcción detached + post-init patching para delegate_task
343e8d08 consolidar agent/research_*.py → agent/research/ Pattern parity con cron/, gateway/, plugins/memory/

Trabajo upstream que necesita coordinación

HRM-57 está parcial. El cierre full implica extender AIAgent.__init__ para absorber 5 atributos internos (_delegate_depth, terminal_cwd, cwd, _subdirectory_hints, _delegate_spinner) que hoy se patchean post-construcción en el factory. Es upstream-territory (run_agent.py) — preferí dejarlo flagged y conversable acá antes de tocarlo.

Validación

  • 132 tests pasan (supervisor + factory + delegate suites)
  • 1 research job end-to-end ejecutado contra el código nuevo (haiku sobre CPU caches, kimi-k2.6, judge parser parseó quality_score=0.85 sin issues)
  • 5 researcher spawns autónomos validaron tanto el spawn mechanism como las propuestas que devolvieron (visibles en Lattice HRM-57..61)

Listo para re-review cuando puedas.

…ture pinning

- Reads Kimi CLI tokens from ~/.kimi/credentials/kimi-code.json
- resolve_kimi_coding_runtime_credentials(): OAuth first, refresh token support, fallback to KIMI_API_KEY
- kimi_coding_default_headers(): proper User-Agent and X-Msh-* headers for coding endpoint
- kimi_coding_required_temperature(): pins temperature to 0.6 for kimi-k2.6 on coding endpoint
- 401 retry with token refresh before aborting
- Integrates into run_agent.py and auxiliary_client.py
The upstream _fixed_temperature_for_model() already omits temperature for
Kimi models, letting the server choose the correct value. Our manual
0.6 pinning was unnecessary and could conflict with server-side mode
selection (thinking vs non-thinking). Verified working without it.
nicoechaniz pushed a commit that referenced this pull request Apr 25, 2026
…matrix, troubleshooting (NousResearch#15135)

The initial Spotify docs page shipped in NousResearch#15130 was a setup guide. This
expands it into a full feature reference:

- Per-tool parameter table for all 9 tools, extracted from the real
  schemas in tools/spotify_tool.py (actions, required/optional args,
  premium gating).
- Free vs Premium feature matrix — which actions work on which tier,
  so Free users don't assume Spotify tools are useless to them.
- Active-device prerequisite called out at the top; this is the #1
  cause of '403 no active device' reports for every Spotify
  integration.
- SSH / headless section explaining that browser auto-open is skipped
  when SSH_CLIENT/SSH_TTY is set, and how to tunnel the callback port.
- Token lifecycle: refresh on 401, persistence across restarts, how
  to revoke server-side via spotify.com/account/apps.
- Example prompt list so users know what to ask the agent.
- Troubleshooting expanded: no-active-device, Premium-required, 204
  now_playing, INVALID_CLIENT, 429, 401 refresh-revoked, wizard not
  opening browser.
- 'Where things live' table mapping auth.json / .env / Spotify app.

Verified with 'node scripts/prebuild.mjs && npx docusaurus build'
— page compiles, no new warnings.
The auxiliary client's _refresh_provider_credentials() handled auth
refresh for Codex, Nous, and Anthropic, but not for Kimi. When the
Kimi OAuth token expired, auxiliary calls (memory flush, compression,
session search) failed with HTTP 401 while the main client recovered
automatically.

Add a kimi-coding / kimi-coding-cn case that calls
resolve_kimi_coding_runtime_credentials(force_refresh=True) and evicts
the cached auxiliary client, mirroring the existing provider refresh
paths.

Fixes auxiliary memory flush failures when using Kimi OAuth.
@Fede654
Fede654 force-pushed the feat/hermes-autoresearch-upstream branch from 8eaec07 to 0523576 Compare April 28, 2026 06:01
Fede654 pushed a commit to Fede654/hermes-agent that referenced this pull request Apr 28, 2026
The scheme-validation commit (e77a3f2c) was too strict: a user with
legacy ''baseUrl: localhost:8000'' (no ''http://'' prefix) in their
''~/.honcho/config.json'' would get ''No API key configured'' from the
CLI after that change, even though their setup worked before.

urlparse on a schemeless host:port treats the host segment as the
scheme and leaves netloc empty, so the http/https check rejected it.

Falls back to a lenient check for schemeless strings that look like
hosts: contain '.' or ':', aren't a boolean/null literal, aren't pure
digits. The SDK still rejects truly malformed URLs at connect time
with a clearer error than ours.

Three new tests: legacy schemeless hosts accepted; obvious garbage
literals (''true'', ''null'', ''12345'') still rejected.  Reviewer
noted concern nicoechaniz#1: schemeless regression for self-hosters with old
configs.
- Add protect_first_n to DEFAULT_CONFIG['compression'] (default 3, allows 0)
- Add compression.prompt {preamble, template} for custom summary prompts
- Bump config version 22 -> 23
- Extract hardcoded prompt constants in ContextCompressor to module-level defaults
- Pass custom preamble/template through ContextCompressor constructor
- Read protect_first_n and prompt config in run_agent.py, pass to compressor
- Update status display to show protect_first_n
- Add tests for protect_first_n=0 and custom prompts
- Always preserve system prompt as literal even when protect_first_n=0
- Fix missing import resolve_kimi_coding_runtime_credentials in runtime_provider.py
- Update wiki and website docs
# Conflicts:
#	hermes_cli/config.py
#	tests/agent/test_auxiliary_client.py
#	tests/hermes_cli/test_runtime_provider_resolution.py
#	ui-tui/src/app/uiStore.ts
#	ui-tui/src/app/useConfigSync.ts
# Conflicts:
#	hermes_cli/config.py
nicoechaniz pushed a commit that referenced this pull request Apr 29, 2026
* ci(nix): auto-fix stale npm hashes on push to main

When a PR merges to main with updated package-lock.json or package.json
in ui-tui/ or web/, the new auto-fix-main job detects stale npmDepsHash
values and pushes a fix commit directly to main.

This eliminates the recurring manual hash-bump PRs (NousResearch#15420, NousResearch#15314,
NousResearch#15272, NousResearch#15244) by reusing the existing fix-lockfiles --apply pipeline.

The fix commit only touches nix/*.nix files, which are outside the push
path filter (package-lock.json / package.json), so it cannot re-trigger
itself.

Closes NousResearch#15314

* fix(ci): use GitHub App token for auto-fix-main push

GITHUB_TOKEN commits are invisible to workflow triggers (GitHub's
infinite-loop prevention). The auto-fix-main job pushes directly to
main, so the fix commit never triggered downstream nix.yml verification.

Mint a short-lived token via the repo's GitHub App (daimon-nous, APP_ID
+ APP_PRIVATE_KEY secrets) so the push is treated as a real event and
nix.yml fires to verify the corrected hashes.

Tested via workflow_dispatch dry-run: app token minted successfully,
checkout with app token succeeded, fix job correctly gated.

Resolves review feedback from Bugbot (r3144569551).

* ci(nix): rename lockfile check job for required status check

Rename 'check' → 'nix-lockfile-check' so the status check name is
unambiguous when added as a required check on main.

* fix(ci): harden auto-fix-main against races, loops, and silent failures

Address adversarial review findings:

1. Race condition (#1): Job-level concurrency with cancel-in-progress
   collapses back-to-back pushes; ref: main checkout always gets latest
   branch state; explicit push target (origin HEAD:main).

2. Loop prevention (#2): File-whitelist check before commit aborts if
   any file outside nix/{tui,web}.nix was modified, preventing
   accidental self-triggering.

3. Silent infra failures (#8): nix-lockfile-check now fails explicitly
   when fix-lockfiles exits without reporting stale status (catches nix
   setup failures, network errors, script bugs that bypass continue-on-error).

4. Commit traceability (#11): Auto-fix commits include source SHA and
   workflow run URL in the commit body.

5. Explicit push target (#12): git push origin HEAD:main instead of
   bare git push.

---------

Co-authored-by: alt-glitch <alt-glitch@users.noreply.github.com>
nicoechaniz pushed a commit that referenced this pull request Apr 29, 2026
…ch#16706)

* fix(tui): drop stale stream events after ctrl-c interrupt

Once interruptTurn() flips this.interrupted, only recordMessageDelta
short-circuited.  recordReasoningDelta/Available, recordToolStart/
Progress/Complete, and recordInlineDiffToolComplete kept populating
turnState until the python loop reached its next _interrupt_requested
check (~1s on busy turns), making it look like ctrl-c was ignored
while late "thinking" + tool calls kept landing in the UI.

Add the same interrupted guard to every stream-side recorder, and
clear the flag at startMessage() so the next turn isn't suppressed
if the previous turn never delivered message.complete.

* fix(tui): guard recordTodos against post-interrupt mutation; fake-timers in test

Copilot review on PR NousResearch#16706:

1. `recordToolStart` is interruption-guarded, but `tool.start`
   handler also calls `recordTodos(payload.todos)` first — so a
   late tool.start carrying todos could still mutate `turnState.todos`
   after Ctrl-C, leaving ghost rows in the panel.  Adds the same
   `if (this.interrupted) return` early-exit to `recordTodos` so
   *all* tool.start side-effects are dropped post-interrupt.

2. The interrupt test was leaking a real `setTimeout` (interrupt
   cooldown) across test files, which could fire later and mutate
   uiStore from the wrong test context.  Wraps the test in
   `vi.useFakeTimers()` + `vi.runAllTimers()` and restores real
   timers in finally.

3. Extends the same test with a todos payload on the post-interrupt
   tool.start so we have explicit regression coverage for #1.

* fix(tui): guard pushTrail post-interrupt; harden interrupt-test cleanup

Round 2 Copilot review on PR NousResearch#16706:

1. `tool.generating` events route through `pushTrail`, which was not
   interruption-guarded — late events could still write 'drafting …'
   into `turnTrail` after Ctrl-C, leaving a stale shimmer in the UI.
   Adds the same `if (this.interrupted) return` early-exit.

2. Test cleanup moved `vi.runAllTimers()` into `finally` (before
   `vi.useRealTimers()`) so a mid-test assertion failure can't leak
   the interrupt-cooldown setTimeout across other test files.

3. Replaced the misleading 'pre-interrupt todos … expected to be
   cleared by the interrupt cycle' comment with an accurate one
   reflecting current behaviour (interrupt does NOT clear todos).

4. Added an explicit assertion that a post-interrupt `tool.generating`
   event does not extend `turnTrail` — regression coverage for #1.
…_messages transport

PR NousResearch#12846 enabled Anthropic prompt caching for third-party gateways,
but gated it on is_claude, which excluded providers like MiniMax
that serve their own model families (MiniMax-M2.7, etc.) through the
native Anthropic protocol.

MiniMax documents full cache_control support on its /anthropic
endpoints (global and China). This patch adds MiniMax detection to
_anthropic_prompt_cache_policy() using:

- Built-in provider id (minimax, minimax-cn), or
- Known Anthropic-compatible hostname (api.minimax.io,
  api.minimaxi.com)

Both paths receive the native cache_control layout.

Refs: NousResearch#8294 (related, but only covered Claude-named models on
third-party gateways).
Closes NousResearch#17332
Fede654 and others added 6 commits May 13, 2026 00:12
research_job is the durable detached variant of run_research. It was
registered as a tool but missing from TOOLSETS['research'], so the
researcher profile (and any agent that loaded the 'research' toolset)
could not call it.

Also fix the schema description which referred to a non-existent
research_job_status tool — the actual polling interface is
research_job(action='status'), and result collection is
research_job(action='collect').

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
The May 9 ProgressSink/KanbanSink refactor removed lattice_task_id from
the public API (run_research, research_job, ResearchABTester) but the
researcher profile scaffold and the ops/agents docs still walked users
through the old Lattice CLI workflow. This commit aligns the docs and
the scaffold:

  * RESEARCH_OPERATIONS.md: "Lattice Integration (Optional)" section
    rewritten as "Kanban Integration (Optional)" — kanban_task_id,
    KanbanSink/StubSink fallback, kanban DB resolution. MCP init line
    no longer claims lattice/obsidian MCP servers (researcher profile
    dropped those in commits 9891835c / b89883cb).

  * RESEARCH_AGENTS.md: "Lattice State Transitions" → "Kanban State
    Transitions". Same contract semantics (in_progress → done /
    archived) but driven by KanbanSink, not lattice CLI.

  * hermes_cli/researcher_scaffold.py: config.yaml comments, SOUL.md
    operational context, tracking workflow, integration patterns —
    all rewritten to use `hermes kanban` CLI and kanban_task_id.
    "Long lattice comments" tooling note is now "Long kanban comments".

  * tests/agent/research/test_kanban_rename.py (new): pins the public
    surface — run_research, research_job schema, and the
    _action_start handler must all use kanban_task_id, never
    lattice_task_id. Regression-proofs the rename.

Verified no `lattice` substring remains in any of the four target
files. Full research test suite passes (129 tests).

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Two coupled fixes around the detached research-job spawn path, plus
tests pinning both.

  * Kimi OAuth path (PR nicoechaniz#1 review): the spec generator was unconditionally
    pinning api_key="" into the job spec (KIMI_API_KEY default empty). On
    this fork the default Kimi auth is OAuth via
    ~/.kimi/credentials/kimi-code.json — pinning an empty api_key in the
    spec made the auxiliary_client treat the empty string as an explicit
    override on some paths instead of falling through to OAuth resolution.
    Fix: only pin api_key in the spec when the env var is actually set
    (and non-whitespace). When unset, the field is omitted entirely and
    AIAgent → resolve_kimi_coding_runtime_credentials handles auth.

  * terminal_tool import (latent bug): _action_start and _action_resume
    both did `from tools.terminal_tool import terminal`, but the module
    exports `terminal_tool`, not `terminal`. The import would have raised
    ImportError every time research_job(action='start') was called — the
    detached job path was effectively unusable as shipped. Rename the
    import; signature is otherwise unchanged. Surfaced by the new unit
    test that stubbed the spawn and exercised _action_start end-to-end.

Tests:
  * test_spec_omits_api_key_when_env_unset
  * test_spec_pins_api_key_when_env_set
  * test_spec_strips_whitespace_api_key
  * test_action_start_does_not_write_null_api_key
  * test_detached_job_resolves_kimi_oauth_without_api_key (integration,
    opt-in via -m integration; requires real Kimi OAuth credentials)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…fresh

nicoechaniz's PR nicoechaniz#1 review (2026-05-10) asked whether the cache might
serve stale clients across OAuth credential refresh. Investigation
turned up a real bug, not just a documentation gap:

  * _resolve_provider_cache and _client_cache are two separate caches.
  * Both cache OpenAI client objects whose api_key is snapshotted at
    construction time (the OpenAI SDK does not auto-refresh the
    api_key field — it sends the bound value forever).
  * _evict_cached_clients() was only clearing _client_cache, not
    _resolve_provider_cache. So after _refresh_provider_credentials()
    rotated a Kimi OAuth access_token, the next call to
    resolve_provider_client() served a client bound to the now-stale
    token and requests would 401 indefinitely.

Fix: extend _evict_cached_clients() to clear matching entries from
_resolve_provider_cache as well, under its own lock. Both layers now
invalidate in lockstep on credential refresh.

Also add a safety docstring above _resolve_provider_cache that spells
out the contract:
  * Process lifetime, no TTL — eviction is triggered by credential
    refresh, not time.
  * Credential-distinguishing fields (explicit_api_key, base_url,
    oauth_credential_path inside main_runtime) are part of the key.
  * The main_runtime flattening is one-level deep; nested dicts/lists
    will raise TypeError at key construction. Pinned by tests so a
    regression surfaces loudly.

Tests in tests/agent/test_auxiliary_client_cache.py cover:
  * Hashability for typical runtime shapes
  * Key distinguishes explicit_api_key
  * Key distinguishes oauth_credential_path in main_runtime
  * Nested runtime values raise TypeError (limitation pin)
  * _evict_cached_clients clears _resolve_provider_cache
  * Eviction does not collateral-clear other providers

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
Two CI follow-ups on the cleanup round:

1. tools/research_job_tool.py:356 used time.time() in _action_resume
   without importing time. Surfaced by upstream's ty type-checker
   ("Name 'time' used when not defined"). Add the import.

2. scripts/release.py AUTHOR_MAP missing entries for the two git
   identities Fede654 commits through (buzondefede@gmail.com,
   fede@localhost.localdomain). Surfaced by the check-attribution
   workflow. Map both to the Fede654 GitHub username.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
…lock-when-yielding

Captures two operational rules discovered during multi-run, autonomous
research sessions where the worker exhausts its per-run iteration budget
mid-experiment and a fresh worker reclaims the task on the next
dispatcher tick.

**Resume protocol (Steps 0-3):** check workspace/STATE.json before any
action; verify each detached_jobs[].pid is alive; launch heavy jobs
under setsid so they outlive worker exits; update STATE.json on every
iteration before the kanban write. Documents the workspace state file
schema so successor workers can pick up exactly where the previous run
left off.

**Closing your run — block, don't text-only-exit:** the kanban-worker
harness treats any run that exits without calling kanban_complete or
kanban_block as `crashed`. Three consecutive `crashed` outcomes auto-
block the task with `gave_up` and the dispatcher stops reclaiming it.
The fix is for the worker to explicitly emit a cooperative
`kanban block` when yielding to a detached job — the detached job
then calls `kanban unblock` on success to wake the next worker.

Both rules were proven under fire on a long-running daemoncraft
research task that ran across ~20 worker reclaims over a day. Without
the resume protocol the worker re-ran experiments that had already
produced artifacts; without the explicit kanban_block the dispatcher
hit gave_up and the task stalled silently for hours.

No code changes; this is purely a scaffold-template doc update so any
new researcher profile bootstrapped via `hermes profile setup
researcher` inherits the rules.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@Fede654
Fede654 force-pushed the feat/hermes-autoresearch-upstream branch from 76b2ec0 to 382366d Compare May 13, 2026 03:13
@nicoechaniz
nicoechaniz force-pushed the main branch 2 times, most recently from facc91a to ab9c249 Compare May 13, 2026 09:58
@nicoechaniz
nicoechaniz merged commit a55cd3c into nicoechaniz:main May 17, 2026
9 of 13 checks passed
AnnieScigliano pushed a commit to AnnieScigliano/hermes-agent that referenced this pull request May 17, 2026
…registries

Both web_search_registry._resolve() and image_gen_registry.get_active_provider()
walked their registered providers and returned the first one matching the
capability flag — without checking whether that provider was actually
usable. On a fresh install with no credentials at all, this meant
get_active_search_provider() returned `brave-free` (legacy preference
order) even though BRAVE_SEARCH_API_KEY was unset, leading the
dispatcher to surface a "BRAVE_SEARCH_API_KEY is not set" error for a
provider the user never chose. Same bug shape in image_gen for FAL.

Resolution semantics now match tools.web_tools._get_backend():

  1. Explicit config name wins, ignoring is_available() — the dispatcher
     surfaces a precise "X_API_KEY is not set" error rather than silently
     switching backends. Matches user expectation: "I configured X, tell
     me what's wrong with X."
  2. Fallback (no explicit config) walks the legacy preference order
     filtered by is_available() — pick the highest-priority backend the
     user actually has credentials for.

is_available() is wrapped in a try/except so a buggy provider doesn't
brick resolution.

E2E verified:
  - No creds + no config: get_active_search_provider() -> None
  - Explicit brave-free + no key: get_active_search_provider() -> brave-free
    (and .is_available() correctly reports False)

This fix was identified during the spike (NousResearch#25182 finding nicoechaniz#1) and is
fold-in to the same PR rather than a follow-up.
AnnieScigliano pushed a commit to AnnieScigliano/hermes-agent that referenced this pull request May 30, 2026
Three issues flagged by the Copilot review on this PR:

1. Double JSON emit on stage failure (Copilot nicoechaniz#1, nicoechaniz#2). When -Stage <name>
   ran a worker that threw, Invoke-Stage's finally emitted a JSON result
   frame AND the entry-point catch emitted a second error frame --
   producing two concatenated JSON objects on stdout and breaking the
   one-line-per-invocation contract that drivers parse against. Same
   issue applied to -Json mode on a full install (every stage's finally
   plus a final error frame missing duration_ms/skipped).

   Fix: Invoke-Stage's finally now sets $script:_StageEmittedErrorFrame
   when it emits a failure frame; the entry-point catch checks the flag
   and skips its own emit, still exit 1.

2. $prevEAP uninitialized on early try-block throw (Copilot nicoechaniz#3). In
   Install-Uv, Test-Python, Test-Node's winget fallback,
   _Run-NpmInstall, and the playwright block, '$prevEAP =
   $ErrorActionPreference' lived as the first statement INSIDE the
   try. If anything between 'try {' and that line threw (Write-Info on
   an unusual host, the npx-finding loop, etc.), the catch's
   'if ($prevEAP) { ... }' restore was a no-op and EAP could remain
   relaxed.

   Fix: hoist '$prevEAP = $ErrorActionPreference' to the line
   immediately before 'try {' in all five sites. Catch's restore is
   now always meaningful regardless of where in the try the throw
   originated.

No change to Invoke-Stage's success path or to the four lint-clean EAP
sites (Test-Node was the only winget-related catch). All 19 metadata
smoke tests still pass.
AnnieScigliano pushed a commit to AnnieScigliano/hermes-agent that referenced this pull request May 30, 2026
Four findings from Copilot's review on PR NousResearch#22891, all in the AX
elements-array cap added by 22fa1ed:

1. The truncation note ("response truncated to N of M elements") was
   appended unconditionally — including in the som/vision multimodal
   path, whose response carries a screenshot rather than an `elements`
   array. The note described a payload field that wasn't present.
   Moved the note into the AX-text branch where the array actually
   appears.

2. `_format_elements(cap.elements)` ran on the full untrimmed list with
   its own `max_lines=40` cap, so a caller passing `max_elements=10`
   would see summary lines referencing `nicoechaniz#11..NousResearch#40` even though the JSON
   `elements` array only held nicoechaniz#1..nicoechaniz#10. Format on `visible_elements`
   instead so the summary indices always exist in the response.

3. `_coerce_max_elements` enforced a lower bound but no upper bound,
   so `max_elements=10_000_000` silently disabled the safeguard and
   reintroduced the original context-blow-up. Added a hard cap
   (`_MAX_ALLOWED_MAX_ELEMENTS = 1000`) that clamps oversized values.

4. The schema string said "Default 100" but the property carried no
   `default` field, and claimed `max_elements` had no effect on som/
   vision while the image-missing fallback path can still return an
   elements array. Added `"default": 100`, `"maximum": 1000`, and
   clarified the fallback-path wording.

Each finding gets a regression test:

- test_capture_ax_clamps_oversized_max_elements_to_hard_cap
- test_capture_ax_summary_indices_match_returned_elements
- test_capture_multimodal_summary_omits_truncation_note
- test_schema_max_elements_documents_default_and_upper_bound

Verified with `pytest tests/tools/test_computer_use.py` (53 passed,
including the 5 new cases). Confirmed each new test fails on the
pre-fix code path before applying the production change.

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
nicoechaniz pushed a commit that referenced this pull request Jun 1, 2026
…NousResearch#31416)

PR NousResearch#31416 (avoid persisting borrowed credential secrets) added
sanitize_borrowed_credential_payload, which strips access_token from
any auth.json pool entry whose (provider, source) isn't in the
_PERSISTABLE_PROVIDER_SOURCES allowlist.

(copilot, gh_cli) is borrowed (not in the allowlist), so the test
fixture's pre-seeded access_token now gets stripped at load_pool()
time, leaving the pool empty. resolve_target('1') then fails with
'No credential #1. Provider: copilot.'

Fix: align the test with the new contract. At runtime, copilot tokens
are hydrated by resolve_copilot_token() — mock that path so the pool
gets an entry the test can remove. The behavior under test
(suppression of gh_cli + env variants on remove) is unchanged.

CI repro on origin/main HEAD; reproduced locally with stock checkout.
nicoechaniz pushed a commit that referenced this pull request Jun 1, 2026
…s reached

After key #1 is marked exhausted the retry still called the API with key #1
due to env-var bias in _get_cached_client / resolve_api_key_provider_credentials.
Fix: peek the pool and pass the active entry's key as explicit_api_key.
Secondary: api_key_hint in mark_exhausted_and_rotate pins the correct entry
under concurrent CLI+gateway calls; _is_payment_error matches GoUsageLimitError;
extract_api_error_context parses "Resets in Xhr Ymin".
nicoechaniz pushed a commit that referenced this pull request Jun 1, 2026
…ookies

Mission-control style deploys reverse-proxy the dashboard at a path
prefix (e.g. mission-control.tilos.com/hermes/* -> :9119) and inject
X-Forwarded-Prefix: /hermes on every request. The SPA mount already
honoured this for asset URLs and the bootstrap __HERMES_BASE_PATH__,
but the OAuth gate didn't:

  1. The gate's Location: header to /login and the 401 envelope's
     login_url were built bare ("/login?next=..."). Under a /hermes
     prefix the browser follows that to mission-control.tilos.com/login
     which the proxy doesn't route to the dashboard.
  2. _redirect_uri (the OAuth callback URL handed to the IDP) used
     request.url_for() which doesn't honour X-Forwarded-Prefix
     (Starlette/uvicorn only proxy_headers Host + Proto + For). The
     IDP redirects back to /auth/callback instead of /hermes/auth/
     callback → 404 in the user's browser.
  3. Cookies were set with Path=/ which leaks them to other apps on
     the same origin and won't be sent back on requests under the
     prefix in the first place.

Fix threads the normalised prefix through every boundary:

  * New hermes_cli/dashboard_auth/prefix.py — single source of truth
    for X-Forwarded-Prefix parsing. web_server._normalise_prefix
    becomes a re-export so the SPA mount, the gate, and the cookies
    helper all agree.
  * middleware._unauth_response builds login_url = f"{prefix}/login".
  * routes._redirect_uri splices the prefix into the path component
    of the IDP-bound URL (with full validation of the header).
  * cookies.{set,clear}_{session,pkce}_cookie now take prefix="".
    Path attribute switches to /hermes when set; cookie name switches
    name variant (see below). Every caller passes the request's
    normalised prefix.

Cookie hardening (Teknium's lesser-note #1 in the PR review): adopt
the __Host- / __Secure- cookie name prefixes per draft-west-cookie-
prefixes. The variant is selected from (use_https, prefix):

  * Loopback HTTP → bare "hermes_session_at" (both prefixes require
    Secure, incompatible with HTTP).
  * HTTPS, direct deploy (Path=/) → "__Host-hermes_session_at".
    Strongest spec: bound to exact origin, no Domain attribute, Secure
    required.
  * HTTPS, behind a proxy prefix (Path=/hermes) →
    "__Secure-hermes_session_at". __Host- forbids Path != "/"; the
    explicit Path=/hermes covers same-origin app isolation.

Setter and reader BOTH consult the prefix because the cookie *name*
changes — a reader that looked up the bare name when the setter wrote
__Secure- would never find the value. The reader falls back across
all three variants so a request whose shape changed mid-session (e.g.
post-deploy from no-prefix to /hermes) still picks up the existing
cookie until it expires.

Test coverage:

  - tests/hermes_cli/test_dashboard_auth_prefix.py — new file. 11 tests
    pinning:
      • Location: /hermes/login on the gate's HTML redirect
      • 401 envelope login_url carries the prefix
      • Malformed X-Forwarded-Prefix is ignored (header-injection
        defence; the script-tag value is normalised to empty string)
      • _redirect_uri splices /hermes into the path (the property
        that prevents the IDP-returns-to-404 failure)
      • PKCE cookie uses Path=/hermes + __Secure- when proxied
      • Session cookies use __Host- when direct, __Secure- when
        proxied, bare on loopback HTTP
      • End-to-end round trip with hand-managed PKCE cookie carriage
        (TestClient can't simulate a Path=/hermes cookie automatically)
  - tests/hermes_cli/test_dashboard_auth_cookies.py — rewritten to pin
    each (use_https, prefix) shape produces its expected cookie name,
    plus reader-side coverage that __Host- and __Secure- variants are
    both recognised.
  - Existing tests across middleware / 401-reauth / etc. updated to
    match the new cookie names (substring contains instead of
    startswith).

Mutation-tested: reverting _unauth_response to build the bare
"/login" URL trips exactly the two tests that pin the prefix
carriage, confirming the suite discriminates the regression.
nicoechaniz pushed a commit that referenced this pull request Jun 1, 2026
Two CI flakes surfaced on PR NousResearch#34572 (both in files this PR doesn't touch;
pre-existing host-dependent flakes):

1. test_process_registry::TestPopenLeakOnSetupFailure — the failure-cleanup
   tests use a fake proc.pid (8888/9999) and assert proc.kill() runs. But
   spawn_local's primary cleanup is os.killpg(os.getpgid(pid), SIGKILL),
   falling back to proc.kill() only on ProcessLookupError/PermissionError/
   OSError. When the fake PID happens to exist on a busy host, os.getpgid
   succeeds, os.killpg fires against an UNRELATED real process group, and
   proc.kill() is never reached -> flaky AssertionError (and a real risk of
   SIGKILLing an innocent process group from a unit test). Patch os.getpgid
   to raise ProcessLookupError so the fallback path runs deterministically
   and no real killpg is ever issued.

2. test_web_server::test_resize_escape_is_forwarded — the receive loop calls
   the blocking conn.receive_bytes() with no exception guard. Once the child
   prints its winsize and exits, the PTY closes; on a missed-marker run the
   next recv blocks until the 30s pytest-timeout instead of failing fast.
   Add a try/except break (matching the working sibling tests) and bump the
   child's pre-read sleep 0.15s -> 0.5s so the resize reliably lands first.

Verified: 4/4 pass across 3 consecutive runs; root cause for #1 reproduced
(os.getpgid(1) succeeds -> old code skips proc.kill).
nicoechaniz pushed a commit that referenced this pull request Jun 7, 2026
Seven Copilot inline review comments on NousResearch#37679, four worth landing
in a polish pass before merge:

1. _dispose_unused_adapter signature: 'BasePlatformAdapter' ->
   'BasePlatformAdapter | None'. The function explicitly handles
   None and the reconnect watcher calls it with None in the
   except arm, so the annotation now matches the actual contract.

2. (duplicate of #1 on a different line) — same fix.

3. except Exception in _dispose_unused_adapter — the reviewer
   asked about asyncio.CancelledError swallowing. On Python 3.8+
   (Hermes requires 3.13, see pyproject.toml), CancelledError
   inherits from BaseException, NOT Exception, so the existing
   'except Exception' does NOT swallow task cancellation. Added
   an explicit comment explaining the contract so future readers
   don't repeat the analysis. We don't re-raise because the
   watcher loop intentionally treats dispose failures as
   best-effort: a failed dispose on an unowned adapter should not
   take down the watcher that's keeping the gateway alive.

4. _response_store = None after close in api_server.py — the
   reviewer flagged this for idempotency. Decided to keep the
   non-None state intentionally: setting it to None cascades
   to ~9 callers that access self._response_store without a
   None check, and 'close() is idempotent on a closed sqlite3
   Connection' means the current code is already safe. The
   type stays stable; LSP doesn't flag a cascade of
   reportOptionalMemberAccess errors. (This matches the
   pre-existing pattern in the codebase — e.g.
   _mark_disconnected doesn't reset state to None either.)

5. _build_adapter_with_store: reviewer worried about
   disconnect() failing on the self.name property if
   __init__ wasn't called. Already handled: we set
   'adapter.platform = Platform.API_SERVER' so the
   'self.platform.value.title()' property returns
   'Api_Server' without raising. The exception-swallowing
   branch in disconnect() does call self.name via the
   logger.debug format, so this is a real path that needs
   the platform attribute, and we have it.

6. test_disconnect_closes_response_store: bare 'pytest.raises(Exception)'
   -> 'pytest.raises(sqlite3.ProgrammingError)'. The bare
   Exception matcher would silently accept AttributeError,
   OperationalError, env-related issues, etc. The specific
   exception type ('Cannot operate on a closed database') is
   the actual signal we want — proves the SQLite conn is
   closed, not just that *something* raised.

7. test_nonretryable_failure_disposes_unowned_adapter:
   assertion tightened from '>= 1' to '== 1' on
   adapter._disconnect_calls. The docstring said 'exactly once',
   the assertion now matches. Catches the hypothetical
   'watcher disposes the same adapter twice' regression that
   '>=' would have missed.
nicoechaniz pushed a commit that referenced this pull request Jun 7, 2026
…ch#37677)

Anthropic enforces two independent ceilings per image:
1. 5 MB encoded byte size
2. 8000 px longest side

Hermes only guarded #1. A tall screenshot (e.g. 1200x12000 at 0.06 MB)
passes every byte check but fails the pixel check, returning a
non-retryable HTTP 400 that permanently bricks the conversation thread.

Fixes:
- error_classifier: add 'image dimensions exceed' pattern to
  _IMAGE_TOO_LARGE_PATTERNS so the 400 is classified as image_too_large
  and triggers the shrink/retry path instead of falling through to
  non-retryable error.
- conversation_compression: check pixel dimensions (via Pillow) even
  when byte size is under the 4 MB target. If max(dims) > 8000, force
  shrink.
- vision_tools._resize_image_for_vision: add optional max_dimension param.
  When set, images exceeding the pixel cap are downscaled even if they're
  under the byte budget. The resize loop now checks both byte AND pixel
  limits before accepting a candidate.

Closes NousResearch#37677
nicoechaniz pushed a commit that referenced this pull request Jun 7, 2026
…bes + test-leak fix (NousResearch#40909)

* fix(gateway,windows): reliability — supervisor task, JOB breakaway, status --deep

Three coordinated fixes for the Windows gateway reliability story:

1. CREATE_BREAKAWAY_FROM_JOB on every detached spawn

   The 'hermes update' triggered from the Electron Desktop GUI ran inside
   Electron's job object. Without breakaway, the post-update gateway
   watcher spawned by update — already DETACHED_PROCESS — was still
   reaped when Electron's job tore down, so the gateway never came back
   after a GUI-initiated update. Adds CREATE_BREAKAWAY_FROM_JOB (0x01000000)
   to:
     - hermes_cli/_subprocess_compat.py::windows_detach_flags() — used by
       every helper that calls windows_detach_popen_kwargs(), including
       launch_detached_profile_gateway_restart()
     - The watcher subprocess's own respawn snippet in
       hermes_cli/gateway.py (inlined flags so the watcher's child
       respawn also breaks away)

   _spawn_detached() in gateway_windows.py already had the flag; this
   change brings the rest of the codebase to parity.

2. Per-minute supervisor Scheduled Task — Windows equivalent of
   systemd Restart=always

   Introduces hermes_cli/gateway_supervisor.py and registers it as a
   second Scheduled Task ('Hermes_Gateway_Supervisor', SC MINUTE /MO 1,
   LIMITED rights) alongside the existing ONLOGON task. Every minute,
   the supervisor uses the same gateway.status.get_running_pid() probe
   as 'hermes gateway status' and, if no gateway is alive, calls
   gateway_windows._spawn_detached() (which now includes BREAKAWAY) to
   bring one back.

   Covers every crash mode, not just 'machine rebooted': taskkill,
   OOM, GUI update SIGTERM, parent job teardown. Cheap — one pythonw
   startup per minute when down, one PID-existence check per minute
   when up.

   Wired into both the schtasks-success and Startup-folder-fallback
   install paths via _install_supervisor_best_effort(), and removed in
   uninstall(). Best-effort: a failing supervisor install logs a
   warning but doesn't roll back the primary install.

3. 'hermes gateway status --deep' shows per-probe PASS/FAIL

   Replaces the existing terse '--deep' output (which only printed
   paths) with an actual diagnostic table:
     [1] PID file present
     [2] Lock file held by a live process
     [3] get_running_pid() result
     [4] _pid_exists(pid) — OS-level liveness
     [5] gateway_state.json (state + age)
     [6] Last lifecycle event from gateway-exit-diag.log

   When the high-level summary disagrees with reality, the user can
   see exactly which signal is lying.

Test-leak fix
-------------

tests/hermes_cli/test_gateway_wsl.py::TestGatewayCommandWSLMessages
monkey-patched is_linux/is_wsl/supports_systemd_services to simulate
WSL but did NOT stub is_windows(). On a Windows host, the dispatcher
in _gateway_command_inner takes the is_windows() branch BEFORE the
WSL guidance branch, so the test invoked gateway_windows.install()
for real. install() writes to %APPDATA%\...\Startup\Hermes_Gateway.cmd
— the REAL user Startup folder, never sandboxed by tmp_path — pointing
at the test's pytest-of-<user>/pytest-<N>/.../gateway-service/ wrapper.
When pytest tore down the tmp_path, every subsequent Windows login
flashed a cmd.exe window that failed to find the missing target.

Stubs is_windows=False on all four affected tests:
  test_install_wsl_no_systemd
  test_start_wsl_no_systemd
  test_status_wsl_running_manual
  test_status_wsl_not_running

Defense-in-depth: _build_startup_launcher() now prefixes the launcher
with 'if not exist <target> exit /b 0', so any future stale Startup
entry silently no-ops instead of flashing a console window.

Status enhancements
-------------------

- status() now reports supervisor task presence alongside the existing
  schtasks/Startup info, and nudges the user to reinstall if the
  supervisor isn't registered.
- Deep mode dumps both the supervisor task name + script path.

* fix(gateway,windows): drop the per-minute supervisor task — keep breakaway + deep probes

Earlier in this branch we added a per-minute schtasks-based supervisor to
respawn the gateway after crashes / GUI-update SIGTERMs. The implementation
flashed a brief console window on every firing, which stole window focus.
We tried several variants:

  - cmd.exe wrapper invoking pythonw  -> flashes (cmd.exe is console-subsystem)
  - schtasks /TR pointing at pythonw  -> flashes (uv venv launcher pythonw is
    actually subsystem=Console, not GUI; it respawns the real pythonw)
  - schtasks /TR pointing at base uv  -> still flashes (Task Scheduler-side
    conhost preallocation; documented Windows quirk)
  - XML registration with <Hidden>true>  -> still flashes (<Hidden> only hides
    the task in the Task Scheduler UI, not the spawned window)

Researched what leading projects do:

  - Ollama: GUI-subsystem tray exe + Startup-folder shortcut. No supervisor.
  - Tailscale: real Windows Service via SCM. Session 0, no console possible.
  - Syncthing: --no-console flag inside the binary + Startup folder.
  - openclaw: VBS Run(..., 0, False) wrapper. Suppresses the *window* but
    Super User Q971162 confirms focus-steal still occurs in some cases.

None of these use a per-minute polling scheduled task. The 'auto-restart on
crash' responsibility belongs INSIDE the daemon (Tailscale's in-process
recovery / Ollama's monitor+worker pair) OR is delegated to the Windows
Service Control Manager — not Task Scheduler.

So this commit drops the supervisor entirely. The CREATE_BREAKAWAY_FROM_JOB
fix in _subprocess_compat.py (from commit c1e5fa4) survives — that is the
*real* fix for problem #2 (GUI-update kills gateway): the post-update
watcher in launch_detached_profile_gateway_restart() now breaks out of
Electron's job object, so the gateway respawn watcher survives the GUI
quit and successfully respawns the gateway.

Surviving from c1e5fa4:
  * CREATE_BREAKAWAY_FROM_JOB in hermes_cli/_subprocess_compat.py (fixes #2)
  * Inlined breakaway flag in the watcher respawn snippet in gateway.py
  * hermes gateway status --deep PASS/FAIL probes (fixes #1 — visibility)
  * 'if not exist <target> exit /b 0' guard in _build_startup_launcher
    (fixes #3 — silent no-op for stale Startup entries)
  * tests/hermes_cli/test_gateway_wsl.py is_windows=False stubs (root cause
    of #3 — pytest WSL tests no longer leak Startup entries on Win hosts)

Removed in this commit:
  * hermes_cli/gateway_supervisor.py (entire file)
  * Supervisor section in hermes_cli/gateway_windows.py (~180 lines):
      get_supervisor_task_name, get_supervisor_script_path,
      _build_supervisor_cmd_script, _write_supervisor_script,
      _install_supervisor_task, is_supervisor_task_registered,
      _install_supervisor_best_effort
  * _install_supervisor_best_effort() calls in install() (3 spots)
  * supervisor cleanup block in uninstall()
  * supervisor display lines in status() / status(deep=True)

Future direction (out of scope for this PR): the right place for Windows
'Restart=always' semantics is a real Windows Service installed via
pywin32's win32serviceutil.ServiceFramework — session-0 isolation, SCM
auto-restart, no console window possible. That's a meaningful next-PR
project, not a band-aid.

Tests: 51 pass / 2 pre-existing failures in
tests/hermes_cli/test_gateway_{windows,wsl}.py (the 2 failures are
TestSupportsSystemdServicesWSL cases that fail on origin/main too —
unrelated to this PR).
nicoechaniz pushed a commit that referenced this pull request Jun 14, 2026
Add an official, production-grade WhatsApp integration via Meta's
Business Cloud API as a complement to the existing Baileys bridge.
No bridge subprocess, no QR codes, no account-ban risk — at the cost
of a Meta Business account and a public HTTPS webhook URL.

Setup is fully wizard-driven: 'hermes whatsapp-cloud' walks through
every credential with paste-time validation (catches the #1 trap of
pasting a phone number into the Phone Number ID field), generates a
verify token, and ends with copy-paste instructions for the
cloudflared / Meta-dashboard / Business Manager pieces that can't be
automated. The wizard also points users at Meta's Business Manager
for setting the bot's display name and profile picture.

Feature set:

- Inbound: text, images (with native-vision routing), voice notes
  (STT), documents (small text inlined, larger cached), reply context.
- Outbound: text with WhatsApp-flavored markdown conversion, images,
  videos, documents, opus voice notes via ffmpeg with MP3 fallback.
- Native interactive buttons for clarify, dangerous-command approval,
  and slash-command confirmation flows — matches the Telegram /
  Discord UX, graceful degrades to plain text.
- Read receipts (blue double-checkmarks) and typing indicator,
  using Meta's combined endpoint so they fire in a single API call.
- Webhook security: X-Hub-Signature-256 HMAC verification (raw body,
  constant-time), wamid deduplication, group-shaped-message refusal
  (groups deferred to v2 — Baileys still covers them).
- Full integration with the gateway's session, cron, display-tier,
  prompt-hint, and auth-allowlist systems. Cloud and Baileys can run
  side-by-side against different phone numbers.

Also wires STT (speech-to-text) through Nous's managed audio gateway
for Nous subscribers — previously the default stt.provider=local
required a separate faster-whisper install. New subscribers now get
voice-note transcription out of the box.

Docs: 418-line user guide at website/docs/user-guide/messaging/
whatsapp-cloud.md, sidebar entry, environment-variables reference,
ADDING_A_PLATFORM.md updated with the optional interactive-UX
contract for future adapter authors.

Tests: 100 dedicated tests for the adapter, 32 for the setup wizard,
20 for the Nous subscription STT wiring, plus regression coverage
across display_config, prompt_builder, and the cron scheduler.

Known limitations (deferred until clear demand signal):
- Group chats — use the Baileys bridge if you need them.
- Message templates for 24-hour-window outside-conversation sends —
  reactive chat is unaffected; cron / delegate_task with gaps > 24h
  will fail with a clear error. The agent's system prompt warns the
  model about this so it knows to mention it when scheduling delayed
  messages.
Fede654 added a commit to Fede654/hermes-agent that referenced this pull request Jun 25, 2026
…0.19 relocation)

kimi-code CLI moved its config from ~/.kimi to ~/.kimi-code in v0.19. The engine
only checked ~/.kimi/credentials/kimi-code.json, so a fresh `kimi login` (which
now writes ~/.kimi-code/credentials/) left the gateway throwing
"Provider 'kimi-coding' is set ... but no API key was found" — the nicoechaniz#1 fresh-agent
setup failure.

_kimi_cli_credentials_path / _kimi_cli_device_id_path now resolve an existing file
across both layouts (prefer current ~/.kimi-code; fall back to ~/.kimi; default to
~/.kimi-code for new installs). Read + write/refresh both go through the resolver,
so existing ~/.kimi (older CLI) installs keep working and new ones just work.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
nicoechaniz pushed a commit that referenced this pull request Jun 28, 2026
Phase 1 of the pluggable cron-scheduler refactor (Axis B — the trigger).
No call-site changes; this phase only makes the abstraction exist + tested
in isolation.

Task 1.1: cron/scheduler_provider.py — the EXPERIMENTAL CronScheduler ABC.
  Required surface is name + start; is_available()/stop() carry safe defaults.
  is_available has a no-network invariant. Docstring marks it experimental
  until the Chronos provider (Phase 4) validates the shape.
Task 1.2: InProcessCronScheduler wraps the historical 60s ticker loop, calling
  cron.scheduler.tick(sync=False) exactly as the raw ticker does. Uses
  stop_event.wait(interval) for responsive stop (both raw tickers already do).

Tests: ABC-is-abstract, default-is_available, the InProcess loop drives tick
and stops, stop() no-op, and test_abc_growth_stays_additive (the forward-compat
guard: required abstractmethods must stay exactly {name, start}, so the three
Phase-4 hooks land as NON-abstract additions).

tick() internals in cron/scheduler.py are byte-unchanged (only new file added).
Phase 0 characterization tests still green. Full tests/cron/: 445 passed.
nicoechaniz pushed a commit that referenced this pull request Jun 28, 2026
* fix(windows): harden gateway scheduled task

* fix(windows): launch gateway scheduled task via console-less wscript

The Scheduled Task ran the gateway through cmd.exe, which allocates a
console. During logon Windows broadcasts CTRL_CLOSE_EVENT to console
process groups, reaping cmd.exe and the half-initialized gateway with
STATUS_CONTROL_C_EXIT (0xC000013A) - which Task Scheduler treats as a
user cancel, so RestartOnFailure never fires and the gateway vanishes on
every reboot (issue NousResearch#45599 root cause #1).

Add a console-less .vbs launcher (wscript.exe -> pythonw.exe, both
GUI-subsystem) mirroring the gateway.cmd env + argv, and point the task
action at it. The .cmd stays for the Startup-folder fallback and /Run.

Co-Authored-By: Claude Opus 4.8 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Jeff <jeffrobodie@gmail.com>
Co-authored-by: Claude Opus 4.8 (1M context) <noreply@anthropic.com>
nicoechaniz pushed a commit that referenced this pull request Jun 28, 2026
…eation snapshot (NousResearch#44585)

An unpinned cron job follows the global default provider (config.yaml
model.default + resolve_runtime_provider). If that global state is changed
after the job is created — e.g. a temporary switch to a paid provider like
nous/claude-fable-5 — the job silently inherits it on its next tick and spends
real money. This is the reported $7.73 incident: a job created under a
free/default provider later inherited a temporary paid switch.

Fix (ask #1 only) preserves the legitimate "unpinned job should follow
model.default" use case by detecting *drift* rather than freezing the model:

- create_job (cron/jobs.py): for UNPINNED, agent-backed jobs (no explicit
  provider, not no_agent), snapshot the provider that resolution WOULD pick
  right now into a new optional `provider_snapshot` field, resolved via the
  same resolve_runtime_provider() path the ticker uses. Fail-open to None on
  any resolution error so job creation never breaks.

- run_job (cron/scheduler.py): right after runtime resolution, if the job has
  a provider_snapshot AND is unpinned AND the currently-resolved provider
  DIFFERS from the snapshot, fail closed for that run — make no paid call and
  deliver a loud, actionable alert naming both providers and telling the user
  to pin explicitly (`cronjob action=update job_id=.. provider=..`).

Back-compat: jobs with no snapshot (pre-existing jobs, no_agent jobs, or any
job whose creation-time resolution failed) behave exactly as before — the
guard only engages when a snapshot exists. Explicitly-pinned jobs (job.provider
set) are unaffected since they don't drift with global state.

Tests: tests/cron/test_cron_provider_pin.py covers snapshot-matches (runs),
snapshot-differs (fail closed, no agent constructed), no-snapshot back-compat,
None-snapshot back-compat, explicitly-pinned (runs regardless), plus create_job
snapshot capture/skip/fail-open. The fail-closed case is load-bearing (fails
without the guard).

Issue NousResearch#44585 asks #2-4 (hard-stop a running job, gateway-stop containment,
fail-closed on provider mutation) are out of scope for this change.
nicoechaniz pushed a commit that referenced this pull request Jul 16, 2026
…ture

get_copilot_api_token now returns (api_token, base_url); the auth-remove
suppression test still mocked it as a bare string, mis-unpacking into the
credential-pool seed path and failing with 'No credential #1'.
nicoechaniz pushed a commit that referenced this pull request Jul 16, 2026
…_id signature churn

Two independent bugs evicted the cached gateway AIAgent on every turn,
preventing the prompt cache from ever warming:

1. Model normalization mismatch: the post-run fallback-eviction check
   compared _agent.model (stripped in AIAgent.__init__) against the raw
   _resolve_gateway_model() config string. For vendor-prefixed config on
   native providers (e.g. 'deepseek/deepseek-v4-pro' vs 'deepseek-v4-pro')
   this was always unequal, so the agent was evicted after every
   successful run. Normalize _cfg_model the same way (skip aggregators).

2. Discord triggering message_id leaked into the cached system prompt via
   build_session_context_prompt()'s Discord IDs block. message_id changes
   every turn, so the agent-cache signature (computed from the ephemeral
   prompt) changed every Discord turn -> rebuild every message. The id is
   now injected per-turn into the user message (where per-turn content
   belongs and does not touch the cache signature); the cached IDs block
   carries a static pointer to it, preserving reply/react/pin via the
   discord tools.

Adapted from NousResearch#28846. Bug #1 fix is the contributor's; bug #2 reworked to
be non-destructive (keeps the triggering-id capability instead of deleting
it). Redundant auto-reset eviction (already on main via NousResearch#9893/NousResearch#48031) and
the wrong-premise reset_context_note plumbing from the original PR were
dropped.

Co-authored-by: Hermes Agent <hermes@nousresearch.com>
nicoechaniz pushed a commit that referenced this pull request Jul 16, 2026
… fail on '(empty)' sentinel

Two related bugs caused subagent delegation to silently return empty summaries
with 0 tokens when the user configured delegation.provider=bedrock alongside
delegation.base_url=https://bedrock-runtime.<region>.amazonaws.com.

Root cause #1 — misrouting in _resolve_delegation_credentials():
  The configured_base_url branch unconditionally forced provider='custom' and
  api_mode='chat_completions', only specializing for chatgpt.com, anthropic,
  and kimi hosts. Bedrock (and other native-SDK providers) fell through as
  'custom' + chat_completions, which then POSTed OpenAI-shaped JSON at
  Bedrock's native API. Bedrock rejected the payload and returned nothing,
  which looked like an empty LLM response to the child agent.

  Fix: when provider is one of {bedrock, vertex, google, google-genai}, skip
  the base_url short-circuit and fall through to resolve_runtime_provider(),
  which knows how to construct the proper SDK client. base_url can still be
  forwarded through that path for regional overrides.

Root cause #2 — '(empty)' sentinel accepted as success:
  After N retries of empty LLM responses, run_agent.py emits the literal
  string '(empty)' as final_response. _run_single_child then hit
  `elif summary:` — '(empty)' is truthy, so status became 'completed' and
  the parent surfaced a blank result with no error. Users saw api_calls=4,
  tokens=0, duration~0.4s, status=completed.

  Fix: treat final_response.strip() == '(empty)' as a failure so the parent
  surfaces it instead of silently accepting zero-content 'success'.

Both paths were reproduced in a live Hermes TUI session on us-west-2 Bedrock
(provider=bedrock, model=us.anthropic.claude-sonnet-4-6) and are covered by
new tests in tests/tools/test_delegate.py.
@Fede654
Fede654 deleted the feat/hermes-autoresearch-upstream branch August 6, 2026 17:43
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants