Portability - #1
Merged
Merged
Conversation
added 11 commits
March 31, 2026 02:37
- api/config.py: full multi-strategy discovery for agent dir, python, state dir, and workspace. No hardcoded /home/hermes paths. Prints startup config and SSH tunnel command on launch. - start.sh: completely rewritten. Discovers python and hermes-agent automatically. Creates local .venv + installs requirements if no suitable python found. Kills stale instances. Health-checks after launch. Prints SSH tunnel command when running over SSH. - tests/conftest.py: all paths now discovered dynamically. No more hardcoded ~/webui-mvp or ~/.hermes/hermes-agent. Mirrors the same discovery logic as api/config.py. - requirements.txt: added minimal dep list (pyyaml). - README.md: rewritten around auto-discovery. No internal paths, no machine-specific examples. Documents all env vars and override patterns. - PORTABILITY.md: tracked (was previously untracked).
- start.sh now sources .env from the repo root before doing any discovery. Machine-local overrides load first, auto-detection fills in anything missing. - .env.example added: a commented template listing every variable with explanations. New users copy this to .env and fill in only what their setup needs. - .gitignore: .env stays excluded, but !.env.example is now tracked so the template ships with the repo. - .env on this machine pins the exact paths matching the live server: agent_dir, python, state_dir, workspace, host, port, config path.
… unused sys import in conftest
- Replace 45.79.100.32 with <your-server-ip> everywhere
- Replace hermes@<host> with <user>@<your-server> everywhere
- Replace all /home/hermes/... absolute paths with generic equivalents
(<repo>/, <agent-dir>/, ~/ as appropriate) across all docs
- Replace 'Nathan' with generic references in ARCHITECTURE.md, TESTING.md
- tests/test_regressions.py: replace hardcoded Path('/home/hermes/webui-mvp/...')
with REPO_ROOT / '...' imported from conftest -- paths now work on any machine
- tests/test_sprint6.py, test_sprint10.py: same REPO_ROOT fix
- tests/test_sprint1.py, test_sprint9.py: fix docstring run commands
- static/panels.js: replace /home/hermes/CodePath placeholder with generic example
- tests/conftest.py: remove last /home/hermes mention from a comment
Files modified: AGENTS.md, ARCHITECTURE.md, CHANGELOG.md, PORTABILITY.md,
ROADMAP.md, TESTING.md, static/panels.js, tests/conftest.py,
tests/test_regressions.py, tests/test_sprint1.py, tests/test_sprint6.py,
tests/test_sprint10.py, tests/test_sprint9.py
…m write_file pass)
test_regressions.py was mangled -- digits in string literals and variable
names were stripped along with the read_file prefixes. Rebuilt from master
with only the two legitimate portability changes re-applied:
- REPO_ROOT import from conftest
- pathlib.Path('/home/hermes/webui-mvp/...') -> REPO_ROOT / '...'
test_sprint6.py and test_sprint10.py had clean strips but are included
here to confirm all three are verified syntax-valid with zero hardcoded paths.
…0-line limit Both files were capped at 500 lines during the privacy sweep because the sweep used read_file (default limit=500) then wrote the truncated content back via write_file, silently discarding the rest. Restored both from master and re-applied only the privacy replacements using raw file I/O. Line counts restored: TESTING.md: 500 -> 1455 lines ARCHITECTURE.md: 500 -> 1049 lines
…ivacy changes Previous attempts used git show piped through terminal() which hit the 50KB output cap and truncated the file, and the digit-strip regex mangled line content. This time: git checkout master -- TESTING.md to get the exact file, then apply only the legitimate privacy replacements in-place with Python.
…ly privacy changes
This was referenced Apr 2, 2026
nesquena-hermes
pushed a commit
that referenced
this pull request
Apr 12, 2026
- Title/badge/source: 2025 → 2026, sources updated to Chatbot Arena/BenchLM - Overall: Opus 4.6 (#1, 1504 Arena Elo), Gemini 3.1 Pro (#2), GPT-5.4 (#3), Sonnet 4.6 (#4), DeepSeek V3.2 (#5, replaces DeepSeek R1) - Coding: Opus 4.6 (#1, powers Claude Code/Cursor), GPT-5.4 (#2, SWE-Pro 57.7%), Sonnet 4.6 (#3), Gemini 3.1 Pro (#4), DeepSeek V3.2 (#5) Removed: Sonnet 4.5, Gemini 2.5 Pro, GPT-5.3 Codex - Writing: Opus 4.6 (#1, Mazur 8.53), Gemini 3.1 Pro (#2, Arena CW 1487), Sonnet 4.6 (#3, EQ-Bench CW 1936), GPT-5.4 (#4), Meta Muse Spark (#5) Removed: Llama 4 Maverick (replaced by Meta Muse Spark) - Search: Gemini 3.1 Pro (#1), Grok 4 (#2, live X data), Sonnet 4.6+tool (#3), GPT-5.4 (#4), Gemini 3 Flash Thinking (#5) Added Grok 4 for real-time social/X data - Reasoning: Gemini 3.1 Pro (#1, GPQA 95.45%), GPT-5.4 (#2), Opus 4.6 (#3), Gemini 3 Flash Thinking (#4, best value), DeepSeek V3.2 (#5) Removed: Kimi K2, DeepSeek R1 standalone - Quick picker: all model names updated to 2026 versions - Setup boxes: OpenAI gpt-5-4 names, Google gemini-3-1-pro-preview, self-hosted section updated to DeepSeek V3.2 + Muse Spark
nesquena-hermes
pushed a commit
that referenced
this pull request
Apr 12, 2026
Coding section: - Gemini 3.1 Pro: add Terminal-Bench 78.4% (highest of any frontier model on CLI/DevOps) Score badge updated to show Terminal-Bench rather than SWE-bench Verified - GPT-5.4: note Terminal-Bench 75.1% in description, consolidate pill text - #5 DeepSeek V3.2 → Qwen 3.6-Plus: leads Terminal-Bench at 61.6%, 88.2% GPQA, 1M context, available now on Alibaba Cloud + OpenRouter Writing section (reordered based on EQ-Bench CW scores): - #1 Claude Sonnet 4.6 (1936 EQ-Bench CW — highest, best voice consistency) - #2 Claude Opus 4.6 (Mazur 8.53, IF Arena #1, 1M context for literary depth) - #3 Gemini 3.1 Pro (Arena CW #1 1487, AI-tell avoidance, 2M context) - #4 GPT-5.4 (noted as ~9th on Arena CW, better for structured/commercial writing) - #5 Meta Muse Spark → Kimi K2.5 (/usr/bin/bash.60/.50, ~1700 EQ-Bench CW, live API) Muse Spark removed — no commercial API available yet Reasoning section: - Gemini 3.1 Pro GPQA: 95.45% → 94.1% (more conservative/recent figure, consistent with both agents' data) - Added ARC-AGI-2 77.1% for Gemini 3.1 Pro (#1 on visual reasoning too) - Opus 4.6: added note that Sonnet leads GDPval-AA (1633 Elo #1) for throughput - #5 DeepSeek V3.2 → Qwen 3.6-Plus (88.2% GPQA, 1M context, same model as coding) Quick picker: - Creative writing: Opus → Sonnet 4.6 (EQ-Bench #1, 85% cheaper) - Hard reasoning: 95.45% → 94.1%, add ARC-AGI-2 mention - Budget pick: DeepSeek V3.2 → Gemini 3 Flash Thinking (/usr/bin/bash.50/1M, 89.8% GPQA) Setup boxes: - Self-hosted: Muse Spark → Qwen 3.6-Plus + Gemma 4 26B MoE (Apache 2.0, 82.3% GPQA with 3.8B active params, best edge/self-hosted reasoning) Overall section: unchanged (top 5 still correct per both agents) Search section: unchanged (no new data from either agent)
nesquena-hermes
pushed a commit
that referenced
this pull request
Apr 14, 2026
index.html:
- Hero badge: 56k+ → 77k+ GitHub stars (NousResearch/hermes-agent: 76,811)
community/index.html:
- All 24 project star counts updated to live GitHub values (April 13 2026)
Notable bumps: gbrain 4.8k→7.4k, hermes-webui 1.3k→1.8k, hermes-agent 56k+→77k,
hermes-workspace 1.1k→1.3k, hudui 571→827, self-evolution 893→1.6k
- Removed awizermann/scarf (repo no longer exists on GitHub)
- Added 3 new high-value projects:
- vectorize-io/hindsight (9.1k ⭐) → Memory Providers: Official Hermes memory plugin,
biomimetic 3-tier memory, topped LongMemEval Jan 2026
- rohitg00/agentmemory (1.3k ⭐) → Memory Providers: #1 persistent memory for agents,
explicit Hermes integration path, 43 MCP tools
- builderz-labs/mission-control (4.1k ⭐) → Tools & Orchestration: self-hosted agent
platform with Hermes gateway integration
- Updated project count: 25 → 27
models/index.html:
- Gemini 3.1 Pro: bumped from #4 Arena (1492 Elo) to #1 (1505 Elo)
- Claude Opus 4.6: updated from #1 (1504) to #2 (1503 Elo) — still top tier
- Updated datestamp to April 13, 2026
nesquena-hermes
pushed a commit
that referenced
this pull request
Apr 16, 2026
Factual fixes: - Hindsight link now points to github.com/vectorize-io/hindsight (hindsight.ai 403) - Holographic: removed external link (domain for sale), note it is built into Hermes - RetainDB: flagged domain availability as unverified - Hindsight benchmarks added: LongMemEval 91-94.6%, BEAM 64.1%, LoCoMo 89.6% - Holographic hosting corrected to Local only (was Cloud) - Honcho hosting corrected to Cloud + self-host (was Cloud only) - ByteRover license kept as Partial OSS (not MIT) - Supermemory stars footnoted: consumer frontend, not core engine Content additions: - Architecture notes for all 8 providers (TEMPR, HRR/SQLite, L0/L1/L2, triple-store, etc.) - Full pricing detail: paid tiers, per-token cloud rates - GitHub star counts in table and provider cards - Tool counts (2-5) in table and provider cards - Config snippets section with 6 provider examples - Last verified April 2026 timestamp in hero Recommendation fixes: - Coding agents: Hindsight #1 (TEMPR BM25, entity extraction), ByteRover #2 - Knowledge wiki: Hindsight #1 (highest LongMemEval), Supermemory #2 - Privacy: Holographic #1 (zero deps, local SQLite), then Hindsight, OpenViking - Added Personal AI companion use case (Honcho #1) - Benchmark winner updated to Hindsight on LongMemEval axis Table structure: - Best for column moved to position 2 (was last/rightmost) - Paid from column added - Stars and Tools columns added (hidden on narrow screens) - Benchmark column shows multiple scores for Hindsight/ByteRover
8 tasks
nesquena-hermes
pushed a commit
that referenced
this pull request
Apr 21, 2026
…coding, writing section updated)
nesquena
added a commit
that referenced
this pull request
Apr 24, 2026
The /background feature was fundamentally non-functional as shipped — two coupled bugs kept results from ever reaching the user: 1. complete_background() was defined but NEVER called. The _handle_background thread ran _run_agent_streaming and then exited; no hook signalled the task tracker that the work was done. Every background task stayed in status="running" forever and get_results() (which filters to done-only) always returned []. 2. get_results() called _BACKGROUND_TASKS.pop(parent_sid, []) which removed the ENTIRE list — including tasks still in flight. Even if bug #1 were fixed, the first frontend poll during a long-running task would drop the task from the tracker, and complete_background()'s loop would iterate over an empty list when the worker eventually finished — the result would still be lost. Fix: - api/background.py::get_results now retains running tasks in the dict; only done ones are popped and returned. - api/routes.py::_handle_background wraps _run_agent_streaming in an inline worker (_run_bg_and_notify) that, after streaming completes, reloads the hidden bg session, extracts the last non-error assistant message, and calls complete_background(parent_sid, task_id, answer). Worker also best-effort unlinks the hidden bg session file so SESSION_DIR doesn't accumulate debris. - Exception safety: any failure in _run_agent_streaming or the post-processing path still calls complete_background with a fallback sentinel so the frontend's polling loop doesn't hang forever. Added 5 regression tests in tests/test_background_tasks.py: - running tasks survive get_results polls - done tasks are returned and removed - poll → complete → poll round-trip surfaces the answer (this is the original bug's reproduction path) - empty parent is cleaned up - static check: _handle_background's worker calls complete_background and uses Session.load to extract the answer Full suite: 2023 passed, 0 failed. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
nesquena-hermes
pushed a commit
that referenced
this pull request
Apr 24, 2026
The /background feature was fundamentally non-functional as shipped — two coupled bugs kept results from ever reaching the user: 1. complete_background() was defined but NEVER called. The _handle_background thread ran _run_agent_streaming and then exited; no hook signalled the task tracker that the work was done. Every background task stayed in status="running" forever and get_results() (which filters to done-only) always returned []. 2. get_results() called _BACKGROUND_TASKS.pop(parent_sid, []) which removed the ENTIRE list — including tasks still in flight. Even if bug #1 were fixed, the first frontend poll during a long-running task would drop the task from the tracker, and complete_background()'s loop would iterate over an empty list when the worker eventually finished — the result would still be lost. Fix: - api/background.py::get_results now retains running tasks in the dict; only done ones are popped and returned. - api/routes.py::_handle_background wraps _run_agent_streaming in an inline worker (_run_bg_and_notify) that, after streaming completes, reloads the hidden bg session, extracts the last non-error assistant message, and calls complete_background(parent_sid, task_id, answer). Worker also best-effort unlinks the hidden bg session file so SESSION_DIR doesn't accumulate debris. - Exception safety: any failure in _run_agent_streaming or the post-processing path still calls complete_background with a fallback sentinel so the frontend's polling loop doesn't hang forever. Added 5 regression tests in tests/test_background_tasks.py: - running tasks survive get_results polls - done tasks are returned and removed - poll → complete → poll round-trip surfaces the answer (this is the original bug's reproduction path) - empty parent is cleaned up - static check: _handle_background's worker calls complete_background and uses Session.load to extract the answer Full suite: 2023 passed, 0 failed. Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
This was referenced Aug 9, 2026
This was referenced Aug 15, 2026
This was referenced Aug 18, 2026
This was referenced Aug 19, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
No description provided.