chore: merge upstream Hermes-Agent — 658 commits (2026-04-19) - #3
Conversation
…NousResearch#11485) * feat(skills): add 'hermes skills reset' to un-stick bundled skills When a user edits a bundled skill, sync flags it as user_modified and skips it forever. The problem: if the user later tries to undo the edit by copying the current bundled version back into ~/.hermes/skills/, the manifest still holds the old origin hash from the last successful sync, so the fresh bundled hash still doesn't match and the skill stays stuck as user_modified. Adds an escape hatch for this case. hermes skills reset <name> Drops the skill's entry from ~/.hermes/skills/.bundled_manifest and re-baselines against the user's current copy. Future 'hermes update' runs accept upstream changes again. Non-destructive. hermes skills reset <name> --restore Also deletes the user's copy and re-copies the bundled version. Use when you want the pristine upstream skill back. Also available as /skills reset in chat. - tools/skills_sync.py: new reset_bundled_skill(name, restore=False) - hermes_cli/skills_hub.py: do_reset() + wired into skills_command and handle_skills_slash; added to the slash /skills help panel - hermes_cli/main.py: argparse entry for 'hermes skills reset' - tests/tools/test_skills_sync.py: 5 new tests covering the stuck-flag repro, --restore, unknown-skill error, upstream-removed-skill, and no-op on already-clean state - website/docs/user-guide/features/skills.md: new 'Bundled skill updates' section explaining the origin-hash mechanic + reset usage * fix(auth): codex auth remove no longer silently undone by auto-import 'hermes auth remove openai-codex' appeared to succeed but the credential reappeared on the next command. Two compounding bugs: 1. _seed_from_singletons() for openai-codex unconditionally re-imports tokens from ~/.codex/auth.json whenever the Hermes auth store is empty (by design — the Codex CLI and Hermes share that file). There was no suppression check, unlike the claude_code seed path. 2. auth_remove_command's cleanup branch only matched removed.source == 'device_code' exactly. Entries added via 'hermes auth add openai-codex' have source 'manual:device_code', so for those the Hermes auth store's providers['openai-codex'] state was never cleared on remove — the next load_pool() re-seeded straight from there. Net effect: there was no way to make a codex removal stick short of manually editing both ~/.hermes/auth.json and ~/.codex/auth.json before opening Hermes again. Fix: - Add unsuppress_credential_source() helper (mirrors suppress_credential_source()). - Gate the openai-codex branch in _seed_from_singletons() with is_source_suppressed(), matching the claude_code pattern. - Broaden auth_remove_command's codex match to handle both 'device_code' and 'manual:device_code' (via endswith check), always call suppress_credential_source(), and print guidance about the unchanged ~/.codex/auth.json file. - Clear the suppression marker in auth_add_command's openai-codex branch so re-linking via 'hermes auth add openai-codex' works. ~/.codex/auth.json is left untouched — that's the Codex CLI's own credential store, not ours to delete. Tests cover: unsuppress helper behavior, remove of both source variants, add clears suppression, seed respects suppression. E2E verified: remove → load → add → load flow now behaves correctly.
Follow-up to the reply-reference fix: ensure errors unrelated to the reply reference (e.g. 50013 Missing Permissions) do NOT trigger the no-reference retry path and still surface as a failed SendResult. Keeps the wider retry condition from silently swallowing unrelated API errors. Proposed in the original issue writeup (NousResearch#11342) as test case `test_non_reference_errors_still_propagate`.
Follow-up to the reply-reference fix: `_make_discord_adapter` used to return the raw fetched `Message` as the expected reference, but the adapter now wraps it via `ref_msg.to_reference(fail_if_not_exists=False)` so Discord treats a deleted target as 'send without reply chip'. Update the fixture to return the MessageReference sentinel so the 4 chunk-reference-identity tests assert against the right object. No production behavior change; only aligns the stale test fixture.
…ousResearch#6879) (NousResearch#11561) The Copilot API returns HTTP 400 "model_not_supported" when it receives a model ID it doesn't recognize (vendor-prefixed like `anthropic/claude-sonnet-4.6` or dash-notation like `claude-sonnet-4-6`). Two bugs combined to leave both formats unhandled: 1. `_COPILOT_MODEL_ALIASES` in hermes_cli/models.py only covered bare dot-notation and vendor-prefixed dot-notation. Hermes' default Claude IDs elsewhere use hyphens (anthropic native format), and users with an aggregator-style config who switch `model.provider` to `copilot` inherit `anthropic/claude-X-4.6` — neither case was in the table. 2. The Copilot branch of `normalize_model_for_provider()` only stripped the vendor prefix when it matched the target provider (`copilot/`) or was the special-cased `openai/` for openai-codex. Every other vendor prefix survived to the Copilot request unchanged. Fix: - Add dash-notation aliases (`claude-{opus,sonnet,haiku}-4-{5,6}` and the `anthropic/`-prefixed variants) to the alias table. - Rewire the Copilot / Copilot-ACP branch of `normalize_model_for_provider()` to delegate to the existing `normalize_copilot_model_id()`. That function already does alias lookups, catalog-aware resolution, and vendor-prefix fallback — it was being bypassed for the generic normalisation entry point. Because `switch_model()` already calls `normalize_model_for_provider()` for every `/model` switch (line 685 in model_switch.py), this single fix covers the CLI startup path (cli.py), the `/model` slash command path, and the gateway load-from-config path. Closes NousResearch#6879 Credits dsr-restyn (NousResearch#6743) who independently diagnosed the dash-notation case; their aliases are folded into this consolidated fix alongside the vendor-prefix stripping repair.
DingTalk was the only messaging platform without group-mention gating or a per-user allowlist. Slack, Telegram, Discord, WhatsApp, Matrix, and Mattermost all support these via config.yaml + matching env vars; this change closes the gap for DingTalk using the same surface: Config: platforms.dingtalk.require_mention: bool (env: DINGTALK_REQUIRE_MENTION) platforms.dingtalk.mention_patterns: list (env: DINGTALK_MENTION_PATTERNS) platforms.dingtalk.free_response_chats: list (env: DINGTALK_FREE_RESPONSE_CHATS) platforms.dingtalk.allowed_users: list (env: DINGTALK_ALLOWED_USERS) Semantics mirror Telegram's implementation: - DMs are always accepted (subject to allowed_users). - Group messages are accepted only when the chat is allowlisted, mention is not required, the bot was @mentioned (dingtalk_stream sets is_in_at_list), or the text matches a configured regex wake-word. - allowed_users matches sender_id / sender_staff_id case-insensitively; a single "*" disables the check. Rationale: without this, any DingTalk user in a group chat can trigger the bot, which makes DingTalk less safe to deploy than the other platforms. A user's config.yaml already accepts require_mention for dingtalk but the value was silently ignored.
Adds 16 regression tests for the gating logic introduced in the
salvaged commit:
* TestAllowedUsersGate — empty/wildcard/case-insensitive matching,
staff_id vs sender_id, env var CSV population
* TestMentionPatterns — compilation, case-insensitivity, invalid
regex is skipped-not-raised, JSON env var, newline fallback
* TestShouldProcessMessage — DM always accepted, group gating via
require_mention / is_in_at_list / wake-word pattern / free_response_chats
Also adds yule975 to scripts/release.py AUTHOR_MAP (release CI blocks
unmapped emails).
When a WebSocket-based platform adapter (e.g. QQ Bot) temporarily loses its connection, send() now polls is_connected for up to 15s instead of immediately returning a non-retryable failure. If the auto-reconnect completes within the window, the message is delivered normally. On timeout, the SendResult is marked retryable=True so the base class retry mechanism can attempt re-delivery. Same treatment applied to _send_media(). Adds 4 async tests covering: - Successful send after simulated reconnection - Retryable failure on timeout - Immediate success when already connected - _send_media reconnection wait Fixes NousResearch#11163
…path (NousResearch#11569) The send_message tool's direct-REST QQBot path used "QQBotAccessToken {token}" which QQ's API rejects with 401. The correct format is "QQBot {token}" — the gateway adapter at gateway/platforms/qqbot.py uses this format in all 5 header sites (lines 341, 551, 579, 1068, 1467); this was the one outlier. Credit to @Quon for surfacing this in NousResearch#10257 (that PR had unrelated issues in its media-upload logic and was closed; this salvages the genuine 1-line fix).
…ssion (NousResearch#11568) Three open issues — NousResearch#8242, NousResearch#6587, NousResearch#11345 — all trace to the same root cause: the image / audio / document download paths in `DiscordAdapter._handle_message` used plain, unauthenticated HTTP to fetch `att.url`. That broke in three independent ways: NousResearch#8242 cdn.discordapp.com attachment URLs increasingly require the bot session to download; unauthenticated httpx sees 403 Forbidden, image/voice analysis fail silently. NousResearch#6587 Some user environments (VPNs, corporate DNS, tunnels) resolve cdn.discordapp.com to private-looking IPs. Our is_safe_url() guard correctly blocks them as SSRF risks, but the user environment is legitimate — image analysis and voice STT die. NousResearch#11345 The document download path skipped is_safe_url() entirely — raw aiohttp.ClientSession.get(att.url) with no SSRF check, inconsistent with the image/audio branches. Unified fix: use `discord.Attachment.read()` as the primary download path on all three branches. `att.read()` routes through discord.py's own authenticated HTTPClient, so: - Discord CDN auth is handled (NousResearch#8242 resolved). - Our is_safe_url() gate isn't consulted for the attachment path at all — the bot session handles networking internally (NousResearch#6587 resolved). - All three branches now share the same code path, eliminating the document-path SSRF gap (NousResearch#11345 resolved). Falls back to the existing cache_*_from_url helpers (image/audio) or an SSRF-gated aiohttp fetch (documents) when `att.read()` is unavailable or fails — preserves defense-in-depth for any future payload-schema drift that could slip a non-CDN URL into att.url. New helpers on DiscordAdapter: - _read_attachment_bytes(att) — safe att.read() wrapper - _cache_discord_image(att, ext) — primary + URL fallback - _cache_discord_audio(att, ext) — primary + URL fallback - _cache_discord_document(att, ext) — primary + SSRF-gated aiohttp fallback Tests: - tests/gateway/test_discord_attachment_download.py — 12 new cases covering all three helpers: primary path, fallback on missing .read(), fallback on validator rejection, SSRF guard on document fallback, aiohttp fallback happy-path, and an E2E case via _handle_message confirming cache_image_from_url is never invoked when att.read() succeeds. - All 11 existing document-handling tests continue to pass via the aiohttp fallback path (their SimpleNamespace attachments have no .read(), which triggers the fallback — now SSRF-gated). Closes NousResearch#8242, closes NousResearch#6587, closes NousResearch#11345.
…path The Weixin adapter's send() method previously split and delivered the raw response text without first extracting MEDIA: tags or bare local file paths. This meant images, documents, and voice files referenced by the agent were silently dropped in normal (non-streaming, non-background) conversations. Changes: - In WeixinAdapter.send(), call extract_media() and extract_local_files() before formatting/splitting text. - Deliver extracted files via send_image_file(), send_document(), send_voice(), or send_video() prior to sending text chunks. - Also fix two minor typing issues in gateway/run.py where extract_media() tuples were not unpacked correctly in background and /btw task handlers. Fixes missing media delivery on Weixin personal accounts.
- stop rewriting markdown tables, headings, and links before delivery - keep markdown table blocks and headings together during chunking - update Weixin tests and docs for native markdown rendering Closes NousResearch#10308
- feat: support one-click QR scan to create DingTalk bot and establish connection - fix(gateway): wrap blocking DingTalkStreamClient.start() with asyncio.to_thread() - fix(gateway): extract message fields from CallbackMessage payload instead of ChatbotMessage - fix(gateway): add oapi.dingtalk.com to allowed webhook URL domains
Adds 15 regression tests for hermes_cli/dingtalk_auth.py covering:
* _api_post — network error mapping, errcode-nonzero mapping, success path
* begin_registration — 2-step chain, missing-nonce/device_code/uri
error cases
* wait_for_registration_success — success path, missing-creds guard,
on_waiting callback invocation
* render_qr_to_terminal — returns False when qrcode missing, prints
when available
* Configuration — BASE_URL default + override, SOURCE default
Also adds a one-line disclosure in dingtalk_qr_auth() telling users
the scan page will be OpenClaw-branded. Interim measure: DingTalk's
registration portal is hardcoded to route all sources to /openapp/
registration/openClaw, so users see OpenClaw branding regardless of
what 'source' value we send. We keep 'openClaw' as the source token
until DingTalk-Real-AI registers a Hermes-specific template.
Also adds meng93 to scripts/release.py AUTHOR_MAP.
…trivially (NousResearch#11580) Closes NousResearch#11321, closes NousResearch#10259. ## Problem The nested /skill command group (category subcommand groups + skill subcommands) serialized to ~14KB with the default 75-skill catalog, exceeding Discord's ~8000-byte per-command registration payload. The entire tree.sync() rejected with error 50035 — ALL slash commands including the 27 base commands failed to register. ## Fix Replace the nested Group layout with a single flat Command: /skill name:<autocomplete> args:<optional string> Autocomplete options are fetched dynamically by Discord when the user types — they do NOT count against the per-command registration budget. So this single command registers at ~200 bytes regardless of how many skills exist. Scales to thousands of skills with no size calculations, no splitting, no hidden skills. UX improvements: - Discord live-filters by user's typed prefix against BOTH name and description, so '/skill pdf' finds 'ocr-and-documents' via its description. More discoverable than clicking through category menus. - Unknown skill name → ephemeral error pointing user at autocomplete. - Stable alphabetical ordering across restarts. ## Why not the other proposed approaches Three prior PRs tried to fit within the 8KB limit by modifying the nested layout: - NousResearch#10214 (njiangk): truncated all descriptions to 'Run <name>' and category descriptions to 'Skills'. Works but destroys slash picker UX. - NousResearch#11385 (LeonSGP43): 40-char description clamp + iterative trim-largest-category fallback. Works but HIDES skills the user can no longer invoke via slash — functional regression. - NousResearch#10261 (zeapsu): adaptive split into /skill-<cat> top-level groups. Preserves all skills but pollutes the slash namespace with 20 top-level commands. All three work around the symptom. The flat autocomplete design dissolves the problem — there is no payload-size pressure to manage. ## Tests tests/gateway/test_discord_slash_commands.py — 5 new test cases replace the 3 old nested-structure tests: - flat-not-nested structure assertion - empty skills → no command registered - callback dispatches the right cmd_key by name - unknown name → ephemeral error, no dispatch - large-catalog regression guard (500 skills) — command payload stays under 500 bytes regardless E2E validated against real discord.py 2.7.1: - Command registers as discord.app_commands.Command (not Group). - Autocomplete filters by name AND description (verified across several queries including description-only matches like 'pdf' → OCR skill). - 500-skill catalog returns max 25 results per autocomplete query (Discord's hard cap), filtered correctly. - Choice labels formatted as 'name — description' clamped to 100 chars.
…RD_ALLOWED_USERS Fixes NousResearch#4466. Root cause: two sequential authorization gates both independently rejected bot messages, making DISCORD_ALLOW_BOTS completely ineffective. Gate 1 — `discord.py` `on_message`: _is_allowed_user ran BEFORE the bot filter, so bot senders were dropped before the DISCORD_ALLOW_BOTS policy was ever evaluated. Gate 2 — `gateway/run.py` _is_user_authorized: The gateway-level allowlist check rejected bot IDs with 'Unauthorized user: <bot_id>' even if they passed Gate 1. Fix: gateway/platforms/discord.py — reorder on_message so DISCORD_ALLOW_BOTS runs BEFORE _is_allowed_user. Bots permitted by the filter skip the user allowlist; non-bots are still checked. gateway/session.py — add is_bot: bool = False to SessionSource so the gateway layer can distinguish bot senders. gateway/platforms/base.py — expose is_bot parameter in build_source. gateway/platforms/discord.py _handle_message — set is_bot=True when building the SessionSource for bot authors. gateway/run.py _is_user_authorized — when source.is_bot is True AND DISCORD_ALLOW_BOTS is 'mentions' or 'all', return True early. Platform filter already validated the message at on_message; don't re-reject. Behavior matrix: | Config | Before | After | | DISCORD_ALLOW_BOTS=none (default) | Blocked | Blocked | | DISCORD_ALLOW_BOTS=all | Blocked | Allowed | | DISCORD_ALLOW_BOTS=mentions + @mention | Blocked | Allowed | | DISCORD_ALLOW_BOTS=mentions, no mention | Blocked | Blocked | | Human in DISCORD_ALLOWED_USERS | Allowed | Allowed | | Human NOT in DISCORD_ALLOWED_USERS | Blocked | Blocked | Co-authored-by: Hermes Maintainer <hermes@nousresearch.com>
Six test cases covering: - DISCORD_ALLOW_BOTS=mentions + bot not in DISCORD_ALLOWED_USERS → authorized - DISCORD_ALLOW_BOTS=all + bot not in DISCORD_ALLOWED_USERS → authorized - DISCORD_ALLOW_BOTS=none → bots still rejected (preserves security) - DISCORD_ALLOW_BOTS unset → same as 'none' - Humans still checked against allowlist even with allow_bots=all - Bot bypass is Discord-specific — doesn't leak to other platforms Guards against a regression where the is_bot bypass in _is_user_authorized gets moved, removed, or accidentally extended to other platforms.
…s control Adds a new DISCORD_ALLOWED_ROLES environment variable that allows filtering bot interactions by Discord role ID. Uses OR semantics with the existing DISCORD_ALLOWED_USERS - if a user matches either allowlist, they're permitted. Changes: - Parse DISCORD_ALLOWED_ROLES comma-separated role IDs on connect - Enable members intent when roles are configured (needed for role lookup) - Update _is_allowed_user() to accept optional author param for direct role check - Fallback to scanning mutual guilds when author object lacks roles (DMs, voice) - Fully backwards compatible: no behavior change when env var is unset
…th.json (NousResearch#12360) Codex OAuth refresh tokens are single-use and rotate on every refresh. Sharing them with the Codex CLI / VS Code via ~/.codex/auth.json made concurrent use of both tools a race: whoever refreshed last invalidated the other side's refresh_token. On top of that, the silent auto-import path picked up placeholder / aborted-auth data from ~/.codex/auth.json (e.g. literal {"access_token":"access-new","refresh_token":"refresh-new"}) and seeded it into the Hermes pool as an entry the selector could eventually pick. Hermes now owns its own Codex auth state end-to-end: Removed - agent/credential_pool.py: _sync_codex_entry_from_cli() method, its pre-refresh + retry + _available_entries call sites, and the post-refresh write-back to ~/.codex/auth.json. - agent/credential_pool.py: auto-import from ~/.codex/auth.json in _seed_from_singletons() — users now run `hermes auth openai-codex` explicitly. - hermes_cli/auth.py: silent runtime migration in resolve_codex_runtime_credentials() — now surfaces `codex_auth_missing` directly (message already points to `hermes auth`). - hermes_cli/auth.py: post-refresh write-back in _refresh_codex_auth_tokens(). - hermes_cli/auth.py: dead helper _write_codex_cli_tokens() and its 4 tests in test_auth_codex_provider.py. Kept - hermes_cli/auth.py: _import_codex_cli_tokens() — still used by the interactive `hermes auth openai-codex` setup flow for a user-gated one-time import (with "a separate login is recommended" messaging). User-visible impact - On existing installs with Hermes auth already present: no change. - On a fresh install where the user has only logged in via Codex CLI: `hermes chat --provider openai-codex` now fails with "No Codex credentials stored. Run `hermes auth` to authenticate." The interactive setup flow then detects ~/.codex/auth.json and offers a one-time import. - On an install where Codex CLI later refreshes its token: Hermes is unaffected (we no longer read from that file at runtime). Tests - tests/hermes_cli/test_auth_codex_provider.py: 15/15 pass. - tests/hermes_cli/test_auth_commands.py: 20/20 pass. - tests/agent/test_credential_pool.py: 31/31 pass. - Live E2E on openai-codex/gpt-5.4: 1 API call, 1.7s latency, 3 log lines, no refresh events, no auth drama. The related 14:52 refresh-loop bug (hundreds of rotations/minute on a single entry) is a separate issue — that requires a refresh-attempt cap on the auth-recovery path in run_agent.py, which remains open.
Add approvals.cron_mode config option that controls how cron jobs handle
dangerous commands. Previously, cron jobs silently auto-approved all
dangerous commands because there was no user present to approve them.
Now the behavior is configurable:
- deny (default): block dangerous commands and return a message telling
the agent to find an alternative approach. The agent loop continues —
it just can't use that specific command.
- approve: auto-approve all dangerous commands (previous behavior).
When a command is blocked, the agent receives the same response format as
a user denial in the CLI — exit_code=-1, status=blocked, with a message
explaining why and pointing to the config option. This keeps the agent
loop running and encourages it to adapt.
Implementation:
- config.py: add approvals.cron_mode to DEFAULT_CONFIG
- scheduler.py: set HERMES_CRON_SESSION=1 env var before agent runs
- approval.py: both check_command_approval() and check_all_command_guards()
now check for cron sessions and apply the configured mode
- 21 new tests covering config parsing, deny/approve behavior, and
interaction with other bypass mechanisms (yolo, containers)
…ter (NousResearch#12371) Two related race conditions in gateway/platforms/base.py that could produce duplicate agent runs or silently drop messages. Neither is specific to any one platform — all adapters inherit this logic. R5 (HIGH) — duplicate agent spawn on turn chain In _process_message_background, the pending-drain path deleted _active_sessions[session_key] before awaiting typing_task.cancel() and then recursively awaiting _process_message_background for the queued event. During the typing_task await, a fresh inbound message M3 could pass the Level-1 guard (entry now missing), set its own Event, and spawn a second _process_message_background for the same session_key — two agents running simultaneously, duplicate responses, duplicate tool calls. Fix: keep the _active_sessions entry populated and only clear() the Event. The guard stays live, so any concurrent inbound message takes the busy-handler path (queue + interrupt) as intended. R6 (MED-HIGH) — message dropped during finally cleanup The finally block has two await points (typing_task, stop_typing) before it deletes _active_sessions. A message arriving in that window passes the guard (entry still live), lands in _pending_messages via the busy-handler — and then the unconditional del removes the guard with that message still queued. Nothing drains it; the user never gets a reply. Fix: before deleting _active_sessions in finally, pop any late pending_messages entry and spawn a drain task for it. Only delete _active_sessions when no pending is waiting. Tests: tests/gateway/test_pending_drain_race.py — three regression cases. Validated: without the fix, two of the three fail exactly where the races manifest (duplicate-spawn guard loses identity, late-arrival 'LATE' message not in processed list).
…ousResearch#12416) Fixes silent data loss in the TUI when /undo, /compress, /retry, or rollback.restore runs during an in-flight agent turn. The version- guard at prompt.submit:1449 would fail the version check and silently skip writing the agent's result — UI showed the assistant reply but DB / backend history never received it, causing UI↔backend desync that persisted across session resume. Changes (tui_gateway/server.py): - session.undo, session.compress, /retry, rollback.restore (full-history only — file-scoped rollbacks still allowed): reject with 4009 when session.running is True. Users can /interrupt first. - prompt.submit: on history_version mismatch (defensive backstop), attach a 'warning' field to message.complete and log to stderr instead of silently dropping the agent's output. The UI can surface the warning to the user; the operator can spot it in logs. Tests (tests/test_tui_gateway_server.py): 6 new cases. - test_session_undo_rejects_while_running - test_session_undo_allowed_when_idle (regression guard) - test_session_compress_rejects_while_running - test_rollback_restore_rejects_full_history_while_running - test_prompt_submit_history_version_mismatch_surfaces_warning - test_prompt_submit_history_version_match_persists_normally (regression) Validated: against unpatched server.py the three 'rejects_while_running' tests fail and the version-mismatch test fails (no 'warning' field). With the fix, all 6 pass, all 33 tests in the file pass, 74 TUI tests in total pass. Live E2E against the live Python environment confirmed all 5 patches present and guards enforce 4009 exactly as designed.
Several correctness and cost-safety fixes to the Honcho dialectic path after a multi-turn investigation surfaced a chain of silent failures: - dialecticCadence default flipped 3 → 1. PR NousResearch#10619 changed this from 1 to 3 for cost, but existing installs with no explicit config silently went from per-turn dialectic to every-3-turns on upgrade. Restores pre-NousResearch#10619 behavior; 3+ remains available for cost-conscious setups. Docs + wizard + status output updated to match. - Session-start prewarm now consumed. Previously fired a .chat() on init whose result landed in HonchoSessionManager._dialectic_cache and was never read — pop_dialectic_result had zero call sites. Turn 1 paid for a duplicate synchronous dialectic. Prewarm now writes directly to the plugin's _prefetch_result via _prefetch_lock so turn 1 consumes it with no extra call. - Prewarm is now dialecticDepth-aware. A single-pass prewarm can return weak output on cold peers; the multi-pass audit/reconcile cycle is exactly the case dialecticDepth was built for. Prewarm now runs the full configured depth in the background. - Silent dialectic failure no longer burns the cadence window. _last_dialectic_turn now advances only when the result is non-empty. Empty result → next eligible turn retries immediately instead of waiting the full cadence gap. - Thread pile-up guard. queue_prefetch skips when a prior dialectic thread is still in-flight, preventing stacked races on _prefetch_result. - First-turn sync timeout is recoverable. Previously on timeout the background thread's result was stored in a dead local list. Now the thread writes into _prefetch_result under lock so the next turn picks it up. - Cadence gate applies uniformly. At cadence=1 the old "cadence > 1" guard let first-turn sync + same-turn queue_prefetch both fire. Gate now always applies. - Restored query-length reasoning-level scaling, dropped in 9a0ab34c. Scales dialecticReasoningLevel up on longer queries (+1 at ≥120 chars, +2 at ≥400), clamped at reasoningLevelCap. Two new config keys: `reasoningHeuristic` (bool, default true) and `reasoningLevelCap` (string, default "high"; previously parsed but never enforced). Respects dialecticDepthLevels and proportional lighter-early passes. - Restored short-prompt skip, dropped in ef7f315. One-word acknowledgements ("ok", "y", "thanks") and slash commands bypass both injection and dialectic fire. - Purged dead code in session.py: prefetch_dialectic, _dialectic_cache, set_dialectic_result, pop_dialectic_result — all unused after prewarm refactor. Tests: 542 passed across honcho_plugin/, agent/test_memory_provider.py, and run_agent/test_run_agent.py. New coverage: - TestTrivialPromptHeuristic (classifier + prefetch/queue skip) - TestDialecticCadenceAdvancesOnSuccess (empty-result retry, pile-up guard) - TestSessionStartDialecticPrewarm (prewarm consumed, sync fallback) - TestReasoningHeuristic (length bumps, cap clamp, interaction with depth) - TestDialecticLifecycleSmoke (end-to-end 8-turn session walk)
- Revert website/docs and SKILL.md changes; docs unification handled separately - Scrub commit/PR refs and process narration from code comments and test docstrings (no behavior change)
… multi-peer - cli: setup wizard pre-fills dialecticCadence=2 (code default stays 1 so unset → every turn) - honcho.md: fix stale dialecticCadence default in tables, add Session-Start Prewarm subsection (depth runs at init), add Query-Adaptive Reasoning Level subsection, expand Observation section with directional vs unified semantics and per-peer patterns - memory-providers.md: fix stale default, rename Multi-agent/Profiles to Multi-peer setup, add concrete walkthrough for new profiles and sync, document observation toggles + presets, link to honcho.md - SKILL.md: fix stale defaults, add Depth at session start callout
…t discard, empty-streak backoff Hardens the dialectic lifecycle against three failure modes that could leave the prefetch pipeline stuck or injecting stale content: - Stale-thread watchdog: _thread_is_live() treats any prefetch thread older than timeout × 2.0 as dead. A hung Honcho call can no longer block subsequent fires indefinitely. - Stale-result discard: pending _prefetch_result is tagged with its fire turn. prefetch() discards the result if more than cadence × 2 turns passed before a consumer read it (e.g. a run of trivial-prompt turns between fire and read). - Empty-streak backoff: consecutive empty dialectic returns widen the effective cadence (dialectic_cadence + streak, capped at cadence × 8). A healthy fire resets the streak. Prevents the plugin from hammering the backend every turn when the peer graph is cold. - liveness_snapshot() on the provider exposes current turn, last fire, pending fire-at, empty streak, effective cadence, and thread status for in-process diagnostics. - system_prompt_block: nudge the model that honcho_reasoning accepts reasoning_level minimal/low/medium/high/max per call. - hermes honcho status: surface base reasoning level, cap, and heuristic toggle so config drift is visible at a glance. Tests: 550 passed. - TestDialecticLiveness (8 tests): stale-thread recovery, stale-result discard, fresh-result retention, backoff widening, backoff ceiling, streak reset on success, streak increment on empty, snapshot shape. - Existing TestDialecticCadenceAdvancesOnSuccess::test_in_flight_thread_is_not_stacked updated to set _prefetch_thread_started_at so it tests the fresh-thread-blocks branch (stale path covered separately). - test_cli TestCmdStatus fake updated with the new config attrs surfaced in the status block.
…overage - TestDialecticDepth::test_first_turn_runs_dialectic_synchronously: covered by TestSessionStartDialecticPrewarm::test_turn1_falls_back_to_sync_when_prewarm_missing (more realistic — exercises the empty-prewarm → sync-fallback path) - TestDialecticDepth::test_first_turn_dialectic_does_not_double_fire: covered by TestDialecticLifecycleSmoke (turn 1 flow) and TestDialecticCadenceAdvancesOnSuccess::test_empty_dialectic_result_does_not_advance_cadence Both predate the prewarm refactor and test paths that are now fallback behaviors already covered elsewhere.
…wards-compat fallback
Setup wizard now always writes dialecticCadence=2 on new configs and
surfaces the reasoning level as an explicit step with all five options
(minimal / low / medium / high / max), always writing
dialecticReasoningLevel.
Code keeps a backwards-compat fallback of 1 when dialecticCadence is
unset so existing honcho.json configs that predate the setting keep
firing every turn on upgrade. New setups via the wizard get 2
explicitly; docs show 2 as the default.
Also scrubs editorial lines from code and docs ("max is reserved for
explicit tool-path selection", "Unset → every turn; wizard pre-fills 2",
and similar process-exposing phrasing) and adds an inline link to
app.honcho.dev where the server-side observation sync is mentioned in
honcho.md. Recommended cadence range updated to 1-5 across docs and
wizard copy.
The cherry-picked commit from NousResearch#11434 uses the 154585401+ prefixed noreply format. Add it alongside the existing bare entry so the contributor audit passes.
…sResearch#12369) Agents can now send arbitrary CDP commands to the browser. The tool is gated on a reachable CDP endpoint at session start — it only appears in the toolset when BROWSER_CDP_URL is set (from '/browser connect') or 'browser.cdp_url' is configured in config.yaml. Backends that don't currently expose CDP to the Python side (Camofox, default local agent-browser, cloud providers whose per-session cdp_url is not yet surfaced) do not see the tool at all. Tool schema description links to the CDP method reference at https://chromedevtools.github.io/devtools-protocol/ so the agent can web_extract specific method docs on demand. Stateless per call. Browser-level methods (Target.*, Browser.*, Storage.*) omit target_id. Page-level methods attach to the target with flatten=true and dispatch the method on the returned sessionId. Clean errors when the endpoint becomes unreachable mid-session or the URL isn't a WebSocket. Tests: 19 unit (mock CDP server + gate checks) + E2E against real headless Chrome (Target.getTargets, Browser.getVersion, Runtime.evaluate with target_id, Page.navigate + re-eval, bogus method, bogus target_id, missing endpoint) + E2E of the check_fn gate (tool hidden without CDP URL, visible with it, hidden again after unset).
…ng session (NousResearch#12441) session.interrupt on session A was blast-resolving pending clarify/sudo/secret prompts on ALL sessions sharing the same tui_gateway process. Other sessions' agent threads unblocked with empty-string answers as if the user had cancelled — silent cross-session corruption. Root cause: _pending and _answers were globals keyed by random rid with no record of the owning session. _clear_pending() iterated every entry, so the session.interrupt handler had no way to limit the release to its own sid. Fix: - tui_gateway/server.py: _pending now maps rid to (sid, Event) tuples. _clear_pending takes an optional sid argument and filters by owner_sid when provided. session.interrupt passes the calling sid so unrelated sessions are untouched. _clear_pending(None) remains the shutdown path for completeness. - _block and _respond updated to pack/unpack the new tuple format. Tests (tests/test_tui_gateway_server.py): 4 new cases. - test_interrupt_only_clears_own_session_pending: two sessions with pending prompts, interrupting one must not release the other. - test_interrupt_clears_multiple_own_pending: same-sid multi-prompt release works. - test_clear_pending_without_sid_clears_all: shutdown path preserved. - test_respond_unpacks_sid_tuple_correctly: _respond handles the tuple format. Also updated tests/tui_gateway/test_protocol.py to use the new tuple format for test_block_and_respond and test_clear_pending. Live E2E against the live Python environment confirmed cross-session isolation: interrupting sid_a released its own pending prompt without touching sid_b's. All 78 related tests pass.
…arch#12444) When Discord splits a long message at 2000 chars, _enqueue_text_event buffers each chunk and schedules a _flush_text_batch task with a short delay. If another chunk lands while the prior flush task is already inside handle_message, _enqueue_text_event calls prior_task.cancel() — and without asyncio.shield, CancelledError propagates from the flush task into handle_message → the agent's streaming request, aborting the response the user was waiting on. Reproducer: user sends a 3000-char prompt (split by Discord into 2 messages). Chunk 1 lands, flush delay starts, chunk 2 lands during the brief window when chunk 1's flush has already committed to handle_message. Agent's current streaming response is cancelled with CancelledError, user sees a truncated or missing reply. Fix (gateway/platforms/discord.py): - Wrap the handle_message call in asyncio.shield so the inner dispatch is protected from the outer task's cancel. - Add an except asyncio.CancelledError clause so the outer task still exits cleanly when cancel lands during the sleep window (before the pop) — semantics for that path are unchanged. The new flush task spawned by the follow-up chunk still handles its own batch via the normal pending-message / active-session machinery in base.py, so follow-ups are not lost. Tests: tests/gateway/test_text_batching.py — test_shield_protects_handle_message_from_cancel. Tracks a distinct first_handle_cancelled event so the assertion fails cleanly when the shield is missing (verified by stashing the fix and re-running). Live E2E on the live-loaded DiscordAdapter: first_handle_cancelled: False (shield worked) first_handle_completed: True (handle_message ran to completion)
Follow-up for the helix4u easy-fix salvage batch: - route remaining context-engine quiet-mode output through _should_emit_quiet_tool_messages() so non-CLI/library callers stay silent consistently - drop the extra senderAliases computation from WhatsApp allowlist-drop logging and remove the now-unused import This keeps the batch scoped to the intended fixes while avoiding leaked quiet-mode output and unnecessary duplicate work in the bridge.
…ay/run.py Tier 1 (config + tests + run_agent): - hermes_cli/config.py: adopt upstream auto-provider policy for aux tasks; title_generation and follow_up_generation now use provider=auto so they inherit the user's configured provider via UserLLMKeys (no silent cross-billing through OpenRouter for users who bring their own Anthropic/OpenAI key); bump _config_version from 18 to 19 to match upstream - tests/hermes_cli/test_config_v18_migration.py: rewrite to assert auto-provider defaults and _config_version=19; drop stale openrouter/gemini-2.5-flash pin - tests/hermes_cli/test_tools_config.py: concatenate upstream TestImagegenBackend- Registry/TestImagegenModelPicker with Myah TestMyahPlatformSecrets (additive) - run_agent.py: take upstream verbatim — Myah's OpenRouter fallback was dead code since gateway resolves credentials before AIAgent.__init__ via runtime_provider - tests/run_agent/test_run_agent.py: update test_tool_call_accumulation to reflect upstream's intentional = (not +=) assignment for tool function names; add test_tool_name_not_duplicated_when_provider_repeats_it to document the fix - tests/run_agent/test_streaming.py: add missing api_key/base_url to upstream's new test_tool_name_not_duplicated_when_resent_per_chunk - tests/run_agent/test_concurrent_interrupt.py: add _apply_pending_steer_to_tool_results no-op to _Stub (method added upstream; not exercised in these tests) Tier 2 (gateway/run.py — 2 hunks): - Hunk 1 (Sentry AI monitoring): keep our Sentry AI monitoring block before upstream's new bg_review queueing structure (_bg_review_release threading.Event, _bg_review_pending, _deliver_bg_review_message); tool_progress callback chain verified still works with new upstream callback assignment ordering - Hunk 2 (already_sent): take upstream's richer check (empty-sentinel + streamed + previewed) and add _native_streaming_used as additional OR clause to cover the Myah native-SSE path where GatewayStreamConsumer never sees tokens Also fix test expectations that asserted old hard-pinned defaults: - tests/gateway/test_myah_management_config.py: update schema + reset assertions to expect provider=auto, model='' for title_generation
…orms/myah.py The block marker at line 332 (media attachments ingestion) had a non-standard 'End Myah:' comment instead of the canonical closer pattern (# ──── 40+ dashes). Replace with the standard form so the automated marker parity checker now reports 11/11 for this file.
|
…Task 2A.7 + drop #3) Previously gateway/platforms/api_server.py:76-82 had an inline 'try: from logging_setup import setup_sentry; setup_sentry(); except ImportError: pass' Myah marker block. logging_setup.py was Myah-specific (agent-container Sentry init + SentryHook adapter for agent.telemetry.TelemetryHook protocol) but lived in an upstream-counterpart top-level path. This commit: - Renames logging_setup.py -> myah_hermes_plugin/sentry_init.py. - Calls sentry_init.setup_sentry() from the plugin's register() so hosted-Myah agent containers still get Sentry on gateway boot. - Drops the inline init Myah-marker block from api_server.py. Idempotent: setup_sentry() returns silently when SENTRY_DSN_AGENT is unset (the OSS-user case). Closes drop candidate #3 from spec 2026-05-06 §2.3 (deferred from Tier 2A Task 2A.1 to 2A.7 because it depended on having a plugin-side sentry_init.py to call into). Implements Tier 2A Task 2A.7 from spec 2026-05-06 §3.
…ner + Sentry/admin moves (#24) * chore(hermes): drop unused _get_adapter_for_platform() helper Verified zero callers via repo-wide grep. Function was a leftover from earlier dispatch refactoring. Drop candidate #2 from spec 2026-05-06 §2.3 (Tier 2A Task 2A.1.1). * chore(hermes): drop os.environ CHAT_ID writer (concurrency bug) Per-request mutation of process-wide environment, racing with concurrent requests. The TODO comment confirmed this was a known defect awaiting migration. set_session_vars() in session_context.py is the contextvar-safe replacement. Drop candidate #5 from spec 2026-05-06 §2.3 (Tier 2A Task 2A.1.2). * chore(hermes): catch up MAX_REQUEST_BYTES to upstream (10MB) Verified via git blame: the 1MB value at line 88 was authored by Teknium upstream at commit 80cc27e (2026-03-24) and later bumped to 10MB upstream. Fork has the older value due to drift, not a Myah security decision. Drop candidate #6 from spec 2026-05-06 §2.3 (Tier 2A Task 2A.1.3). * chore(hermes): drop /health Myah credential-validation extension The MYAH_ADAPTER_ENABLED-gated credential check was specific to the fork-bundled adapter and goes away with the plugin-based OSS architecture. /health/detailed (line 832 below) provides rich runtime status; 'hermes doctor' covers credential validation. Drop candidate #4 from spec 2026-05-06 §2.3 (Tier 2A Task 2A.1.4). * chore(hermes): drop cronjob_tools origin-unresolved WARNING Pure debug logging — the function returns None either way; the WARNING was added for diagnostic purposes during a 2025-era cron delivery bug. Function behavior unchanged. Drop candidate #7 from spec 2026-05-06 §2.3 (Tier 2A Task 2A.1.5). * feat(plugin): move myah_overrides.py from hermes_cli/ into plugin Provider catalog augmentation (230 LOC) was Myah-only data living in an upstream-counterpart directory. Moves it into the pip-installed plugin at myah_hermes_plugin/myah_admin/myah_overrides.py and updates the two known importers (plugins/myah-admin/dashboard/_providers.py and the plugin tests). After this commit, hermes_cli/ contains zero Myah-only files — one fewer cardinal-counterpart file to keep marker hygiene on. Implements Tier 2A Task 2A.6 from spec 2026-05-06 §3. * feat(plugin): move Sentry init from upstream into plugin register() (Task 2A.7 + drop #3) Previously gateway/platforms/api_server.py:76-82 had an inline 'try: from logging_setup import setup_sentry; setup_sentry(); except ImportError: pass' Myah marker block. logging_setup.py was Myah-specific (agent-container Sentry init + SentryHook adapter for agent.telemetry.TelemetryHook protocol) but lived in an upstream-counterpart top-level path. This commit: - Renames logging_setup.py -> myah_hermes_plugin/sentry_init.py. - Calls sentry_init.setup_sentry() from the plugin's register() so hosted-Myah agent containers still get Sentry on gateway boot. - Drops the inline init Myah-marker block from api_server.py. Idempotent: setup_sentry() returns silently when SENTRY_DSN_AGENT is unset (the OSS-user case). Closes drop candidate #3 from spec 2026-05-06 §2.3 (deferred from Tier 2A Task 2A.1 to 2A.7 because it depended on having a plugin-side sentry_init.py to call into). Implements Tier 2A Task 2A.7 from spec 2026-05-06 §3. * feat(plugin): add pre_gateway_dispatch hook (Tier 2A Task 2A.4) Replaces the skip_user_authorization semantics that PR #20 removed. Adds myah_hermes_plugin/myah_platform/pre_dispatch_hook.py with a single hook callback that returns {'action': 'allow'} for Myah-platform messages and None (passthrough) for everything else. Architectural note: returning 'allow' does NOT bypass the gateway-level user-allowlist check at gateway/run.py:3655. Auth bypass for Myah's single-tenant deployment is handled by allow_all_env=MYAH_ALLOW_ALL_USERS on the platform registration. The hook exists as a documented choke point for future Myah-specific routing logic (rate limiting, silent ingest, content rewrite) that must NOT live in upstream gateway/run.py. Wired via ctx.register_hook('pre_gateway_dispatch', ...) in the plugin's register() — guarded with hasattr(ctx, 'register_hook') so older PluginContext versions don't fail. Implements Tier 2A Task 2A.4 from spec 2026-05-06 §3. * feat(plugin): vendor _dispatch_approval_notify + register_gateway_notify Plugin owns its own copy of the notify-dispatch chain. Without it, the plugin's vendored request_action_confirmation has no transport to the platform adapter and the agent silently auto-approves. Variadic dispatch (1/2/3-arity callbacks) matches upstream's gateway/run.py:_dispatch_approval_notify behavior at lines 334-460. Uses threading.Lock for the registry to match upstream's thread-safe semantics in tools/approval.py. Implements docs/superpowers/plans/2026-05-06-myah-oss-completion-2a-plugin-stock-upstream.md Task 2A.2.0. * feat(plugin): vendor request_action_confirmation + _action_queues registry Plugin owns its own action confirmation primitives mirroring upstream's tools/approval.py:request_action_confirmation (sync, threading.Event-based). The sync contract is required because cronjob_tools._execute() runs in a tool-runner threadpool and blocks on event.wait(timeout) — an asyncio version would force rewriting the cron tool. Two of the six tests intentionally fail until Task 2A.2.2 lands the plugin's cron_tool.py shadow. Implements docs/superpowers/plans/2026-05-06-myah-oss-completion-2a-plugin-stock-upstream.md Task 2A.2.1. * feat(plugin): shadow upstream cronjob_tools.py with plugin-owned cron tool Verbatim copy of upstream's tools/cronjob_tools.py, with the single change of importing request_action_confirmation from the plugin-vendored myah_hermes_plugin.cron_approval. Tool registration uses the same 'cronjob' name as upstream so last-writer-wins on import order shadows upstream's handler with the plugin's. All 6 tests in test_cron_approval.py now pass (the two cron_tool sanity checks added in 2A.2.1 had been failing pending this commit). Implements docs/superpowers/plans/2026-05-06-myah-oss-completion-2a-plugin-stock-upstream.md Task 2A.2.2. * refactor(plugin): adapter.py uses plugin-vendored approval primitives Move 4 import sites in adapter.py from upstream's tools.approval to the plugin's vendored modules: - resolve_action_confirmation → myah_hermes_plugin.cron_approval - resolve_action_confirmation_by_session → myah_hermes_plugin.cron_approval - unregister_gateway_notify (×2) → myah_hermes_plugin.dispatcher Update mock targets in test_myah_confirm_dispatch.py to follow the new import paths. Comment block in send_action_confirmation updated to reference the plugin's cron_approval module. resolve_gateway_approval (legacy terminal-command approval) stays on tools.approval — it's outside Tier 2A's cron-approval scope. Implements docs/superpowers/plans/2026-05-06-myah-oss-completion-2a-plugin-stock-upstream.md Task 2A.2.3. * feat(plugin): adapter standalone aiohttp runner on MYAH_GATEWAY_PORT Tier 2A Task 2A.3 — collapse the previous hosted/standalone split. The plugin's MyahAdapter now ALWAYS owns its own aiohttp AppRunner + TCPSite via the new MyahStandaloneRunner helper, removing the dependency on upstream's gateway/platforms/api_server.py register_pre_setup_hook + get_shared_app. This is a one-way door for hosted Myah: once shipped, hosted Myah uses the standalone runner forever. See spec docs/superpowers/specs/2026-05-06-myah-oss-completion-design.md §3 Task 2A.3 for the rationale. Default port is now 8643 via MYAH_GATEWAY_PORT (was 8642 hard-coded). Port resolution still respects config.extra.port and MYAH_ADAPTER_PORT overrides for backward compat with existing yaml. Tests cover ephemeral binding (port=0 → OS-assigned), live HTTP route dispatch, and env-var resolution fallback. Zero new failures in the gateway/tools test directories vs baseline (55 pre-existing failures unchanged). * test: fix Tier 2A regressions surfaced by full CI suite The PR's local verification only ran tests/gateway/ + tests/tools/ — three tests in tests/cron/, tests/plugins/, and tests/tools/test_resolve_path.py broke under the full CI run. 1. tests/cron/test_origin_capture_log.py — DELETED. Task 2A.1.5 dropped the WARNING that this entire 'Bug E' file was designed to verify. The file's purpose is gone with the warning; keeping the happy-path companion test would just duplicate coverage that lives elsewhere. 2. tests/plugins/test_myah_admin_providers.py::test_list_providers_against_real_catalog — import switched to pytest.importorskip(). Task 2A.6 moved myah_overrides.py from hermes_cli/ into the plugin distribution, but the hermes Tests workflow only runs 'uv pip install -e ".[all,dev]"' against the core package — it never installs the plugin. Local dev runs (where the plugin is editable-installed) still exercise the real-catalog assertion; CI now skips it cleanly. The plugin's own pytest suite covers the catalog path independently. 3. tests/tools/test_resolve_path.py — added isolated_resolve_state fixture that clears tools.file_tools._file_ops_cache and tools.terminal_tool._active_environments before the two tests that rely on falling through to TERMINAL_CWD. Adding the new plugin modules to the suite shifted xdist worker assignment in a way that exposed pre-existing pollution from other tests leaving a 'default' task_id entry in those module-level dicts. The fixture is per-test (monkeypatch.setattr) so it doesn't disrupt other tests. * test: gentler test_resolve_path isolation (don't rebind module attrs) Previous fix used monkeypatch.setattr to swap _file_ops_cache and _active_environments with empty dicts. This unintentionally caused test_file_state_registry tests to fail on Linux CI under xdist — rebinding the module attribute discarded entries other tests had written under non-'default' keys, and the cross-worker timing exposed the leak. The replacement fixture mutates the existing dicts in place — only pops the 'default' key (which is what _resolve_path reads when no explicit task_id is passed) and restores the prior value after the test. No module-attribute rebinding, no impact on entries other tests created. * ci: trigger re-run to confirm test flakiness Recent runs showed 3 new tests intermittently failing on Linux CI that pass locally and pass in isolation: - test_concurrent_inserts_settle_at_cap (30s timeout) - test_modal_sandbox_fixes::TestToolResolution::test_terminal_tool_present - test_modal_sandbox_fixes::TestToolResolution::test_terminal_and_file_toolsets_resolve_all_tools This empty commit re-runs CI to confirm whether they're deterministic regressions or pre-existing test-ordering flakes exposed by the changed test count from this PR.
Merges 658 upstream commits from NousResearch/Hermes-Agent into our fork's
main. Merge-base722331a57→ upstream tip175cf7e6b(tagv2026.4.16).Upstream highlights
tui_gateway/+ui-tui/terminal UI rendering layerbrowser_cdpraw DevTools Protocol toolapprovals.cron_mode: deny/approveplugins/example-dashboard/gateway/run.pyConflicts resolved (4 files)
hermes_cli/config.py— Adopted upstreamprovider: autofor title_generation and follow_up_generation. Drops hard-coded openrouter+gemini pin that cross-billed users with other providers.run_agent.py— Took upstream verbatim. Our OpenRouter fallback was dead code (gateway always passes explicit creds viaresolve_runtime_provider()).gateway/run.py(×2) — Re-applied Sentry AI monitoring block before upstream's new_bg_review_releasequeueing; added_native_streaming_usedOR-clause to upstream's richeralready_sentlogic.tests/hermes_cli/test_tools_config.py— Concatenated both sides' test classes.Myah code integrity
Test delta
Upstream fixed 30 pre-existing failures — net improvement.
E2E verification
Passed all 10 checks on
x-ai/grok-4.20via OpenRouter from the parent Myah PR (#23):/healthreturns{"status":"ok","platform":"hermes-agent"}/myah/api/toolsetsreturns 22 toolsets (not 404)HERMES_MODEL=x-ai/grok-4.20Follow-ups (not on this branch)
approvals.cron_modein Myah Agent settings UI (new upstream feature)cron/scheduler.pyvsprocesses.py)reasoning.deltastreaming, per-request model on/v1/runs, enriched tool events, generic action confirmation