Skip to content

chore: merge upstream Hermes-Agent — 658 commits (2026-04-19) - #3

Merged
deestax merged 660 commits into
mainfrom
merge-upstream
Apr 19, 2026
Merged

deestax merged 660 commits into
mainfrom
merge-upstream

Conversation

@deestax

@deestax deestax commented Apr 19, 2026

Copy link
Copy Markdown
Member

Merges 658 upstream commits from NousResearch/Hermes-Agent into our fork's main. Merge-base 722331a57 → upstream tip 175cf7e6b (tag v2026.4.16).

Upstream highlights

  • Ink/TUI gateway — new tui_gateway/ + ui-tui/ terminal UI rendering layer
  • Honcho dialectic liveness — stale-thread watchdog, dialectic retry, prewarm, gateway session scoping
  • AWS Bedrock adapter — full Bedrock Converse API support
  • Browser CDP passthrough — new browser_cdp raw DevTools Protocol tool
  • Feishu doc/drive tools
  • Cron approval mode — approvals.cron_mode: deny/approve
  • Plugin dashboard system — plugins/example-dashboard/
  • Gateway pending-drain race fix in gateway/run.py
  • Message deduplicator, session store prune
  • title_generation now upstream — Myah adopts provider: auto, drops hard-coded openrouter+gemini pin

Conflicts resolved (4 files)

  • hermes_cli/config.py — Adopted upstream provider: auto for title_generation and follow_up_generation. Drops hard-coded openrouter+gemini pin that cross-billed users with other providers.
  • run_agent.py — Took upstream verbatim. Our OpenRouter fallback was dead code (gateway always passes explicit creds via resolve_runtime_provider()).
  • gateway/run.py (×2) — Re-applied Sentry AI monitoring block before upstream's new _bg_review_release queueing; added _native_streaming_used OR-clause to upstream's richer already_sent logic.
  • tests/hermes_cli/test_tools_config.py — Concatenated both sides' test classes.

Myah code integrity

  • 11/11 block-marker parity across all 15 Myah-marker files
  • 26 inline markers all present
  • 19/19 Myah-only files intact
  • Zero banned patterns (langfuse, render_ui, render_custom)

Test delta

pass fail skip
Baseline (pre-merge) 2969 59 39
Post-merge 3408 29 39

Upstream fixed 30 pre-existing failures — net improvement.

E2E verification

Passed all 10 checks on x-ai/grok-4.20 via OpenRouter from the parent Myah PR (#23):

  • /health returns {"status":"ok","platform":"hermes-agent"}
  • /myah/api/toolsets returns 22 toolsets (not 404)
  • Chat end-to-end: message streams with CoT reasoning visible
  • Container spawns with updated image and HERMES_MODEL=x-ai/grok-4.20

Follow-ups (not on this branch)

  1. Expose approvals.cron_mode in Myah Agent settings UI (new upstream feature)
  2. Fix pre-existing cron webhook payload mismatch (cron/scheduler.py vs processes.py)
  3. Upstream PR candidates: reasoning.delta streaming, per-request model on /v1/runs, enriched tool events, generic action confirmation

teknium1 and others added 30 commits April 17, 2026 04:10
…NousResearch#11485)

* feat(skills): add 'hermes skills reset' to un-stick bundled skills

When a user edits a bundled skill, sync flags it as user_modified and
skips it forever. The problem: if the user later tries to undo the edit
by copying the current bundled version back into ~/.hermes/skills/, the
manifest still holds the old origin hash from the last successful
sync, so the fresh bundled hash still doesn't match and the skill stays
stuck as user_modified.

Adds an escape hatch for this case.

  hermes skills reset <name>
      Drops the skill's entry from ~/.hermes/skills/.bundled_manifest and
      re-baselines against the user's current copy. Future 'hermes update'
      runs accept upstream changes again. Non-destructive.

  hermes skills reset <name> --restore
      Also deletes the user's copy and re-copies the bundled version.
      Use when you want the pristine upstream skill back.

Also available as /skills reset in chat.

- tools/skills_sync.py: new reset_bundled_skill(name, restore=False)
- hermes_cli/skills_hub.py: do_reset() + wired into skills_command and
  handle_skills_slash; added to the slash /skills help panel
- hermes_cli/main.py: argparse entry for 'hermes skills reset'
- tests/tools/test_skills_sync.py: 5 new tests covering the stuck-flag
  repro, --restore, unknown-skill error, upstream-removed-skill, and
  no-op on already-clean state
- website/docs/user-guide/features/skills.md: new 'Bundled skill updates'
  section explaining the origin-hash mechanic + reset usage

* fix(auth): codex auth remove no longer silently undone by auto-import

'hermes auth remove openai-codex' appeared to succeed but the credential
reappeared on the next command.  Two compounding bugs:

1. _seed_from_singletons() for openai-codex unconditionally re-imports
   tokens from ~/.codex/auth.json whenever the Hermes auth store is
   empty (by design — the Codex CLI and Hermes share that file).  There
   was no suppression check, unlike the claude_code seed path.

2. auth_remove_command's cleanup branch only matched
   removed.source == 'device_code' exactly.  Entries added via
   'hermes auth add openai-codex' have source 'manual:device_code', so
   for those the Hermes auth store's providers['openai-codex'] state was
   never cleared on remove — the next load_pool() re-seeded straight
   from there.

Net effect: there was no way to make a codex removal stick short of
manually editing both ~/.hermes/auth.json and ~/.codex/auth.json before
opening Hermes again.

Fix:

- Add unsuppress_credential_source() helper (mirrors
  suppress_credential_source()).
- Gate the openai-codex branch in _seed_from_singletons() with
  is_source_suppressed(), matching the claude_code pattern.
- Broaden auth_remove_command's codex match to handle both
  'device_code' and 'manual:device_code' (via endswith check), always
  call suppress_credential_source(), and print guidance about the
  unchanged ~/.codex/auth.json file.
- Clear the suppression marker in auth_add_command's openai-codex
  branch so re-linking via 'hermes auth add openai-codex' works.

~/.codex/auth.json is left untouched — that's the Codex CLI's own
credential store, not ours to delete.

Tests cover: unsuppress helper behavior, remove of both source
variants, add clears suppression, seed respects suppression.  E2E
verified: remove → load → add → load flow now behaves correctly.
Follow-up to the reply-reference fix: ensure errors unrelated to the reply
reference (e.g. 50013 Missing Permissions) do NOT trigger the no-reference
retry path and still surface as a failed SendResult. Keeps the wider retry
condition from silently swallowing unrelated API errors.

Proposed in the original issue writeup (NousResearch#11342) as test case
`test_non_reference_errors_still_propagate`.
Follow-up to the reply-reference fix: `_make_discord_adapter` used to return
the raw fetched `Message` as the expected reference, but the adapter now
wraps it via `ref_msg.to_reference(fail_if_not_exists=False)` so Discord
treats a deleted target as 'send without reply chip'. Update the fixture
to return the MessageReference sentinel so the 4 chunk-reference-identity
tests assert against the right object.

No production behavior change; only aligns the stale test fixture.
…ousResearch#6879) (NousResearch#11561)

The Copilot API returns HTTP 400 "model_not_supported" when it receives a
model ID it doesn't recognize (vendor-prefixed like
`anthropic/claude-sonnet-4.6` or dash-notation like `claude-sonnet-4-6`).
Two bugs combined to leave both formats unhandled:

1. `_COPILOT_MODEL_ALIASES` in hermes_cli/models.py only covered bare
   dot-notation and vendor-prefixed dot-notation.  Hermes' default Claude
   IDs elsewhere use hyphens (anthropic native format), and users with an
   aggregator-style config who switch `model.provider` to `copilot`
   inherit `anthropic/claude-X-4.6` — neither case was in the table.

2. The Copilot branch of `normalize_model_for_provider()` only stripped
   the vendor prefix when it matched the target provider (`copilot/`) or
   was the special-cased `openai/` for openai-codex.  Every other vendor
   prefix survived to the Copilot request unchanged.

Fix:

- Add dash-notation aliases (`claude-{opus,sonnet,haiku}-4-{5,6}` and the
  `anthropic/`-prefixed variants) to the alias table.
- Rewire the Copilot / Copilot-ACP branch of
  `normalize_model_for_provider()` to delegate to the existing
  `normalize_copilot_model_id()`.  That function already does alias
  lookups, catalog-aware resolution, and vendor-prefix fallback — it was
  being bypassed for the generic normalisation entry point.

Because `switch_model()` already calls `normalize_model_for_provider()`
for every `/model` switch (line 685 in model_switch.py), this single fix
covers the CLI startup path (cli.py), the `/model` slash command path,
and the gateway load-from-config path.

Closes NousResearch#6879

Credits dsr-restyn (NousResearch#6743) who independently diagnosed the dash-notation
case; their aliases are folded into this consolidated fix alongside the
vendor-prefix stripping repair.
DingTalk was the only messaging platform without group-mention gating or a
per-user allowlist. Slack, Telegram, Discord, WhatsApp, Matrix, and Mattermost
all support these via config.yaml + matching env vars; this change closes the
gap for DingTalk using the same surface:

Config:
  platforms.dingtalk.require_mention: bool   (env: DINGTALK_REQUIRE_MENTION)
  platforms.dingtalk.mention_patterns: list  (env: DINGTALK_MENTION_PATTERNS)
  platforms.dingtalk.free_response_chats: list  (env: DINGTALK_FREE_RESPONSE_CHATS)
  platforms.dingtalk.allowed_users: list     (env: DINGTALK_ALLOWED_USERS)

Semantics mirror Telegram's implementation:
- DMs are always accepted (subject to allowed_users).
- Group messages are accepted only when the chat is allowlisted, mention is
  not required, the bot was @mentioned (dingtalk_stream sets is_in_at_list),
  or the text matches a configured regex wake-word.
- allowed_users matches sender_id / sender_staff_id case-insensitively;
  a single "*" disables the check.

Rationale: without this, any DingTalk user in a group chat can trigger the
bot, which makes DingTalk less safe to deploy than the other platforms. A
user's config.yaml already accepts require_mention for dingtalk but the value
was silently ignored.
Adds 16 regression tests for the gating logic introduced in the
salvaged commit:

  * TestAllowedUsersGate — empty/wildcard/case-insensitive matching,
    staff_id vs sender_id, env var CSV population
  * TestMentionPatterns — compilation, case-insensitivity, invalid
    regex is skipped-not-raised, JSON env var, newline fallback
  * TestShouldProcessMessage — DM always accepted, group gating via
    require_mention / is_in_at_list / wake-word pattern / free_response_chats

Also adds yule975 to scripts/release.py AUTHOR_MAP (release CI blocks
unmapped emails).
When a WebSocket-based platform adapter (e.g. QQ Bot) temporarily
loses its connection, send() now polls is_connected for up to 15s
instead of immediately returning a non-retryable failure. If the
auto-reconnect completes within the window, the message is delivered
normally. On timeout, the SendResult is marked retryable=True so the
base class retry mechanism can attempt re-delivery.

Same treatment applied to _send_media().

Adds 4 async tests covering:
- Successful send after simulated reconnection
- Retryable failure on timeout
- Immediate success when already connected
- _send_media reconnection wait

Fixes NousResearch#11163
…path (NousResearch#11569)

The send_message tool's direct-REST QQBot path used "QQBotAccessToken {token}"
which QQ's API rejects with 401. The correct format is "QQBot {token}" — the
gateway adapter at gateway/platforms/qqbot.py uses this format in all 5 header
sites (lines 341, 551, 579, 1068, 1467); this was the one outlier.

Credit to @Quon for surfacing this in NousResearch#10257 (that PR had unrelated issues in
its media-upload logic and was closed; this salvages the genuine 1-line fix).
…ssion (NousResearch#11568)

Three open issues — NousResearch#8242, NousResearch#6587, NousResearch#11345 — all trace to the same root
cause: the image / audio / document download paths in
`DiscordAdapter._handle_message` used plain, unauthenticated HTTP to
fetch `att.url`. That broke in three independent ways:

  NousResearch#8242  cdn.discordapp.com attachment URLs increasingly require the
         bot session to download; unauthenticated httpx sees 403
         Forbidden, image/voice analysis fail silently.

  NousResearch#6587  Some user environments (VPNs, corporate DNS, tunnels) resolve
         cdn.discordapp.com to private-looking IPs. Our is_safe_url()
         guard correctly blocks them as SSRF risks, but the user
         environment is legitimate — image analysis and voice STT die.

  NousResearch#11345 The document download path skipped is_safe_url() entirely —
         raw aiohttp.ClientSession.get(att.url) with no SSRF check,
         inconsistent with the image/audio branches.

Unified fix: use `discord.Attachment.read()` as the primary download
path on all three branches. `att.read()` routes through discord.py's
own authenticated HTTPClient, so:

  - Discord CDN auth is handled (NousResearch#8242 resolved).
  - Our is_safe_url() gate isn't consulted for the attachment path at
    all — the bot session handles networking internally (NousResearch#6587 resolved).
  - All three branches now share the same code path, eliminating the
    document-path SSRF gap (NousResearch#11345 resolved).

Falls back to the existing cache_*_from_url helpers (image/audio) or an
SSRF-gated aiohttp fetch (documents) when `att.read()` is unavailable
or fails — preserves defense-in-depth for any future payload-schema
drift that could slip a non-CDN URL into att.url.

New helpers on DiscordAdapter:
  - _read_attachment_bytes(att)  — safe att.read() wrapper
  - _cache_discord_image(att, ext)     — primary + URL fallback
  - _cache_discord_audio(att, ext)     — primary + URL fallback
  - _cache_discord_document(att, ext)  — primary + SSRF-gated aiohttp fallback

Tests:
  - tests/gateway/test_discord_attachment_download.py — 12 new cases
    covering all three helpers: primary path, fallback on missing
    .read(), fallback on validator rejection, SSRF guard on document
    fallback, aiohttp fallback happy-path, and an E2E case via
    _handle_message confirming cache_image_from_url is never invoked
    when att.read() succeeds.
  - All 11 existing document-handling tests continue to pass via the
    aiohttp fallback path (their SimpleNamespace attachments have no
    .read(), which triggers the fallback — now SSRF-gated).

Closes NousResearch#8242, closes NousResearch#6587, closes NousResearch#11345.
…path

The Weixin adapter's send() method previously split and delivered the
raw response text without first extracting MEDIA: tags or bare local
file paths. This meant images, documents, and voice files referenced
by the agent were silently dropped in normal (non-streaming,
non-background) conversations.

Changes:
- In WeixinAdapter.send(), call extract_media() and
  extract_local_files() before formatting/splitting text.
- Deliver extracted files via send_image_file(), send_document(),
  send_voice(), or send_video() prior to sending text chunks.
- Also fix two minor typing issues in gateway/run.py where
  extract_media() tuples were not unpacked correctly in background
  and /btw task handlers.

Fixes missing media delivery on Weixin personal accounts.
- stop rewriting markdown tables, headings, and links before delivery
- keep markdown table blocks and headings together during chunking
- update Weixin tests and docs for native markdown rendering

Closes NousResearch#10308
- feat: support one-click QR scan to create DingTalk bot and establish connection
- fix(gateway): wrap blocking DingTalkStreamClient.start() with asyncio.to_thread()
- fix(gateway): extract message fields from CallbackMessage payload instead of ChatbotMessage
- fix(gateway): add oapi.dingtalk.com to allowed webhook URL domains
Adds 15 regression tests for hermes_cli/dingtalk_auth.py covering:
  * _api_post — network error mapping, errcode-nonzero mapping, success path
  * begin_registration — 2-step chain, missing-nonce/device_code/uri
    error cases
  * wait_for_registration_success — success path, missing-creds guard,
    on_waiting callback invocation
  * render_qr_to_terminal — returns False when qrcode missing, prints
    when available
  * Configuration — BASE_URL default + override, SOURCE default

Also adds a one-line disclosure in dingtalk_qr_auth() telling users
the scan page will be OpenClaw-branded. Interim measure: DingTalk's
registration portal is hardcoded to route all sources to /openapp/
registration/openClaw, so users see OpenClaw branding regardless of
what 'source' value we send. We keep 'openClaw' as the source token
until DingTalk-Real-AI registers a Hermes-specific template.

Also adds meng93 to scripts/release.py AUTHOR_MAP.
…trivially (NousResearch#11580)

Closes NousResearch#11321, closes NousResearch#10259.

## Problem

The nested /skill command group (category subcommand groups + skill
subcommands) serialized to ~14KB with the default 75-skill catalog,
exceeding Discord's ~8000-byte per-command registration payload. The
entire tree.sync() rejected with error 50035 — ALL slash commands
including the 27 base commands failed to register.

## Fix

Replace the nested Group layout with a single flat Command:

    /skill name:<autocomplete> args:<optional string>

Autocomplete options are fetched dynamically by Discord when the user
types — they do NOT count against the per-command registration budget.
So this single command registers at ~200 bytes regardless of how many
skills exist. Scales to thousands of skills with no size calculations,
no splitting, no hidden skills.

UX improvements:
- Discord live-filters by user's typed prefix against BOTH name and
  description, so '/skill pdf' finds 'ocr-and-documents' via its
  description. More discoverable than clicking through category menus.
- Unknown skill name → ephemeral error pointing user at autocomplete.
- Stable alphabetical ordering across restarts.

## Why not the other proposed approaches

Three prior PRs tried to fit within the 8KB limit by modifying the
nested layout:

- NousResearch#10214 (njiangk): truncated all descriptions to 'Run <name>' and
  category descriptions to 'Skills'. Works but destroys slash picker UX.
- NousResearch#11385 (LeonSGP43): 40-char description clamp + iterative
  trim-largest-category fallback. Works but HIDES skills the user can
  no longer invoke via slash — functional regression.
- NousResearch#10261 (zeapsu): adaptive split into /skill-<cat> top-level groups.
  Preserves all skills but pollutes the slash namespace with 20
  top-level commands.

All three work around the symptom. The flat autocomplete design
dissolves the problem — there is no payload-size pressure to manage.

## Tests

tests/gateway/test_discord_slash_commands.py — 5 new test cases replace
the 3 old nested-structure tests:

- flat-not-nested structure assertion
- empty skills → no command registered
- callback dispatches the right cmd_key by name
- unknown name → ephemeral error, no dispatch
- large-catalog regression guard (500 skills) — command payload stays
  under 500 bytes regardless

E2E validated against real discord.py 2.7.1:
- Command registers as discord.app_commands.Command (not Group).
- Autocomplete filters by name AND description (verified across several
  queries including description-only matches like 'pdf' → OCR skill).
- 500-skill catalog returns max 25 results per autocomplete query
  (Discord's hard cap), filtered correctly.
- Choice labels formatted as 'name — description' clamped to 100 chars.
…RD_ALLOWED_USERS

Fixes NousResearch#4466.

Root cause: two sequential authorization gates both independently rejected
bot messages, making DISCORD_ALLOW_BOTS completely ineffective.

Gate 1 — `discord.py` `on_message`:
    _is_allowed_user ran BEFORE the bot filter, so bot senders were dropped
    before the DISCORD_ALLOW_BOTS policy was ever evaluated.

Gate 2 — `gateway/run.py` _is_user_authorized:
    The gateway-level allowlist check rejected bot IDs with 'Unauthorized
    user: <bot_id>' even if they passed Gate 1.

Fix:

  gateway/platforms/discord.py — reorder on_message so DISCORD_ALLOW_BOTS
  runs BEFORE _is_allowed_user. Bots permitted by the filter skip the
  user allowlist; non-bots are still checked.

  gateway/session.py — add is_bot: bool = False to SessionSource so the
  gateway layer can distinguish bot senders.

  gateway/platforms/base.py — expose is_bot parameter in build_source.

  gateway/platforms/discord.py _handle_message — set is_bot=True when
  building the SessionSource for bot authors.

  gateway/run.py _is_user_authorized — when source.is_bot is True AND
  DISCORD_ALLOW_BOTS is 'mentions' or 'all', return True early. Platform
  filter already validated the message at on_message; don't re-reject.

Behavior matrix:

  | Config                                     | Before  | After   |
  | DISCORD_ALLOW_BOTS=none (default)          | Blocked | Blocked |
  | DISCORD_ALLOW_BOTS=all                     | Blocked | Allowed |
  | DISCORD_ALLOW_BOTS=mentions + @mention     | Blocked | Allowed |
  | DISCORD_ALLOW_BOTS=mentions, no mention    | Blocked | Blocked |
  | Human in DISCORD_ALLOWED_USERS             | Allowed | Allowed |
  | Human NOT in DISCORD_ALLOWED_USERS         | Blocked | Blocked |

Co-authored-by: Hermes Maintainer <hermes@nousresearch.com>
Six test cases covering:
- DISCORD_ALLOW_BOTS=mentions + bot not in DISCORD_ALLOWED_USERS → authorized
- DISCORD_ALLOW_BOTS=all + bot not in DISCORD_ALLOWED_USERS → authorized
- DISCORD_ALLOW_BOTS=none → bots still rejected (preserves security)
- DISCORD_ALLOW_BOTS unset → same as 'none'
- Humans still checked against allowlist even with allow_bots=all
- Bot bypass is Discord-specific — doesn't leak to other platforms

Guards against a regression where the is_bot bypass in _is_user_authorized
gets moved, removed, or accidentally extended to other platforms.
…s control

Adds a new DISCORD_ALLOWED_ROLES environment variable that allows filtering
bot interactions by Discord role ID. Uses OR semantics with the existing
DISCORD_ALLOWED_USERS - if a user matches either allowlist, they're permitted.

Changes:
- Parse DISCORD_ALLOWED_ROLES comma-separated role IDs on connect
- Enable members intent when roles are configured (needed for role lookup)
- Update _is_allowed_user() to accept optional author param for direct role check
- Fallback to scanning mutual guilds when author object lacks roles (DMs, voice)
- Fully backwards compatible: no behavior change when env var is unset
teknium1 and others added 23 commits April 18, 2026 19:19
…th.json (NousResearch#12360)

Codex OAuth refresh tokens are single-use and rotate on every refresh.
Sharing them with the Codex CLI / VS Code via ~/.codex/auth.json made
concurrent use of both tools a race: whoever refreshed last invalidated
the other side's refresh_token.  On top of that, the silent auto-import
path picked up placeholder / aborted-auth data from ~/.codex/auth.json
(e.g. literal {"access_token":"access-new","refresh_token":"refresh-new"})
and seeded it into the Hermes pool as an entry the selector could
eventually pick.

Hermes now owns its own Codex auth state end-to-end:

Removed
- agent/credential_pool.py: _sync_codex_entry_from_cli() method,
  its pre-refresh + retry + _available_entries call sites, and the
  post-refresh write-back to ~/.codex/auth.json.
- agent/credential_pool.py: auto-import from ~/.codex/auth.json in
  _seed_from_singletons() — users now run `hermes auth openai-codex`
  explicitly.
- hermes_cli/auth.py: silent runtime migration in
  resolve_codex_runtime_credentials() — now surfaces
  `codex_auth_missing` directly (message already points to `hermes auth`).
- hermes_cli/auth.py: post-refresh write-back in
  _refresh_codex_auth_tokens().
- hermes_cli/auth.py: dead helper _write_codex_cli_tokens() and its 4
  tests in test_auth_codex_provider.py.

Kept
- hermes_cli/auth.py: _import_codex_cli_tokens() — still used by the
  interactive `hermes auth openai-codex` setup flow for a user-gated
  one-time import (with "a separate login is recommended" messaging).

User-visible impact
- On existing installs with Hermes auth already present: no change.
- On a fresh install where the user has only logged in via Codex CLI:
  `hermes chat --provider openai-codex` now fails with "No Codex
  credentials stored. Run `hermes auth` to authenticate." The
  interactive setup flow then detects ~/.codex/auth.json and offers a
  one-time import.
- On an install where Codex CLI later refreshes its token: Hermes is
  unaffected (we no longer read from that file at runtime).

Tests
- tests/hermes_cli/test_auth_codex_provider.py: 15/15 pass.
- tests/hermes_cli/test_auth_commands.py: 20/20 pass.
- tests/agent/test_credential_pool.py: 31/31 pass.
- Live E2E on openai-codex/gpt-5.4: 1 API call, 1.7s latency,
  3 log lines, no refresh events, no auth drama.

The related 14:52 refresh-loop bug (hundreds of rotations/minute on a
single entry) is a separate issue — that requires a refresh-attempt
cap on the auth-recovery path in run_agent.py, which remains open.
Add approvals.cron_mode config option that controls how cron jobs handle
dangerous commands. Previously, cron jobs silently auto-approved all
dangerous commands because there was no user present to approve them.

Now the behavior is configurable:
  - deny (default): block dangerous commands and return a message telling
    the agent to find an alternative approach. The agent loop continues —
    it just can't use that specific command.
  - approve: auto-approve all dangerous commands (previous behavior).

When a command is blocked, the agent receives the same response format as
a user denial in the CLI — exit_code=-1, status=blocked, with a message
explaining why and pointing to the config option. This keeps the agent
loop running and encourages it to adapt.

Implementation:
  - config.py: add approvals.cron_mode to DEFAULT_CONFIG
  - scheduler.py: set HERMES_CRON_SESSION=1 env var before agent runs
  - approval.py: both check_command_approval() and check_all_command_guards()
    now check for cron sessions and apply the configured mode
  - 21 new tests covering config parsing, deny/approve behavior, and
    interaction with other bypass mechanisms (yolo, containers)
…ter (NousResearch#12371)

Two related race conditions in gateway/platforms/base.py that could
produce duplicate agent runs or silently drop messages. Neither is
specific to any one platform — all adapters inherit this logic.

R5 (HIGH) — duplicate agent spawn on turn chain
  In _process_message_background, the pending-drain path deleted
  _active_sessions[session_key] before awaiting typing_task.cancel()
  and then recursively awaiting _process_message_background for the
  queued event. During the typing_task await, a fresh inbound message
  M3 could pass the Level-1 guard (entry now missing), set its own
  Event, and spawn a second _process_message_background for the same
  session_key — two agents running simultaneously, duplicate responses,
  duplicate tool calls.

  Fix: keep the _active_sessions entry populated and only clear() the
  Event. The guard stays live, so any concurrent inbound message takes
  the busy-handler path (queue + interrupt) as intended.

R6 (MED-HIGH) — message dropped during finally cleanup
  The finally block has two await points (typing_task, stop_typing)
  before it deletes _active_sessions. A message arriving in that
  window passes the guard (entry still live), lands in
  _pending_messages via the busy-handler — and then the unconditional
  del removes the guard with that message still queued. Nothing
  drains it; the user never gets a reply.

  Fix: before deleting _active_sessions in finally, pop any late
  pending_messages entry and spawn a drain task for it. Only delete
  _active_sessions when no pending is waiting.

Tests: tests/gateway/test_pending_drain_race.py — three regression
cases. Validated: without the fix, two of the three fail exactly
where the races manifest (duplicate-spawn guard loses identity,
late-arrival 'LATE' message not in processed list).
…ousResearch#12416)

Fixes silent data loss in the TUI when /undo, /compress, /retry, or
rollback.restore runs during an in-flight agent turn.  The version-
guard at prompt.submit:1449 would fail the version check and silently
skip writing the agent's result — UI showed the assistant reply but
DB / backend history never received it, causing UI↔backend desync
that persisted across session resume.

Changes (tui_gateway/server.py):
- session.undo, session.compress, /retry, rollback.restore (full-history
  only — file-scoped rollbacks still allowed): reject with 4009 when
  session.running is True.  Users can /interrupt first.
- prompt.submit: on history_version mismatch (defensive backstop),
  attach a 'warning' field to message.complete and log to stderr
  instead of silently dropping the agent's output.  The UI can surface
  the warning to the user; the operator can spot it in logs.

Tests (tests/test_tui_gateway_server.py): 6 new cases.
- test_session_undo_rejects_while_running
- test_session_undo_allowed_when_idle (regression guard)
- test_session_compress_rejects_while_running
- test_rollback_restore_rejects_full_history_while_running
- test_prompt_submit_history_version_mismatch_surfaces_warning
- test_prompt_submit_history_version_match_persists_normally (regression)

Validated: against unpatched server.py the three 'rejects_while_running'
tests fail and the version-mismatch test fails (no 'warning' field).
With the fix, all 6 pass, all 33 tests in the file pass, 74 TUI tests
in total pass.  Live E2E against the live Python environment confirmed
all 5 patches present and guards enforce 4009 exactly as designed.
Several correctness and cost-safety fixes to the Honcho dialectic path
after a multi-turn investigation surfaced a chain of silent failures:

- dialecticCadence default flipped 3 → 1. PR NousResearch#10619 changed this from 1 to
  3 for cost, but existing installs with no explicit config silently went
  from per-turn dialectic to every-3-turns on upgrade. Restores pre-NousResearch#10619
  behavior; 3+ remains available for cost-conscious setups. Docs + wizard
  + status output updated to match.

- Session-start prewarm now consumed. Previously fired a .chat() on init
  whose result landed in HonchoSessionManager._dialectic_cache and was
  never read — pop_dialectic_result had zero call sites. Turn 1 paid for
  a duplicate synchronous dialectic. Prewarm now writes directly to the
  plugin's _prefetch_result via _prefetch_lock so turn 1 consumes it with
  no extra call.

- Prewarm is now dialecticDepth-aware. A single-pass prewarm can return
  weak output on cold peers; the multi-pass audit/reconcile cycle is
  exactly the case dialecticDepth was built for. Prewarm now runs the
  full configured depth in the background.

- Silent dialectic failure no longer burns the cadence window.
  _last_dialectic_turn now advances only when the result is non-empty.
  Empty result → next eligible turn retries immediately instead of
  waiting the full cadence gap.

- Thread pile-up guard. queue_prefetch skips when a prior dialectic
  thread is still in-flight, preventing stacked races on _prefetch_result.

- First-turn sync timeout is recoverable. Previously on timeout the
  background thread's result was stored in a dead local list. Now the
  thread writes into _prefetch_result under lock so the next turn
  picks it up.

- Cadence gate applies uniformly. At cadence=1 the old "cadence > 1"
  guard let first-turn sync + same-turn queue_prefetch both fire.
  Gate now always applies.

- Restored query-length reasoning-level scaling, dropped in 9a0ab34c.
  Scales dialecticReasoningLevel up on longer queries (+1 at ≥120 chars,
  +2 at ≥400), clamped at reasoningLevelCap. Two new config keys:
  `reasoningHeuristic` (bool, default true) and `reasoningLevelCap`
  (string, default "high"; previously parsed but never enforced).
  Respects dialecticDepthLevels and proportional lighter-early passes.

- Restored short-prompt skip, dropped in ef7f315. One-word
  acknowledgements ("ok", "y", "thanks") and slash commands bypass
  both injection and dialectic fire.

- Purged dead code in session.py: prefetch_dialectic, _dialectic_cache,
  set_dialectic_result, pop_dialectic_result — all unused after prewarm
  refactor.

Tests: 542 passed across honcho_plugin/, agent/test_memory_provider.py,
and run_agent/test_run_agent.py. New coverage:
- TestTrivialPromptHeuristic (classifier + prefetch/queue skip)
- TestDialecticCadenceAdvancesOnSuccess (empty-result retry, pile-up guard)
- TestSessionStartDialecticPrewarm (prewarm consumed, sync fallback)
- TestReasoningHeuristic (length bumps, cap clamp, interaction with depth)
- TestDialecticLifecycleSmoke (end-to-end 8-turn session walk)
- Revert website/docs and SKILL.md changes; docs unification handled separately
- Scrub commit/PR refs and process narration from code comments and test
  docstrings (no behavior change)
… multi-peer

- cli: setup wizard pre-fills dialecticCadence=2 (code default stays 1
  so unset → every turn)
- honcho.md: fix stale dialecticCadence default in tables, add
  Session-Start Prewarm subsection (depth runs at init), add
  Query-Adaptive Reasoning Level subsection, expand Observation
  section with directional vs unified semantics and per-peer patterns
- memory-providers.md: fix stale default, rename Multi-agent/Profiles
  to Multi-peer setup, add concrete walkthrough for new profiles and
  sync, document observation toggles + presets, link to honcho.md
- SKILL.md: fix stale defaults, add Depth at session start callout
…t discard, empty-streak backoff

Hardens the dialectic lifecycle against three failure modes that could
leave the prefetch pipeline stuck or injecting stale content:

- Stale-thread watchdog: _thread_is_live() treats any prefetch thread
  older than timeout × 2.0 as dead. A hung Honcho call can no longer
  block subsequent fires indefinitely.

- Stale-result discard: pending _prefetch_result is tagged with its
  fire turn. prefetch() discards the result if more than cadence × 2
  turns passed before a consumer read it (e.g. a run of trivial-prompt
  turns between fire and read).

- Empty-streak backoff: consecutive empty dialectic returns widen the
  effective cadence (dialectic_cadence + streak, capped at cadence × 8).
  A healthy fire resets the streak. Prevents the plugin from hammering
  the backend every turn when the peer graph is cold.

- liveness_snapshot() on the provider exposes current turn, last fire,
  pending fire-at, empty streak, effective cadence, and thread status
  for in-process diagnostics.

- system_prompt_block: nudge the model that honcho_reasoning accepts
  reasoning_level minimal/low/medium/high/max per call.

- hermes honcho status: surface base reasoning level, cap, and heuristic
  toggle so config drift is visible at a glance.

Tests: 550 passed.
- TestDialecticLiveness (8 tests): stale-thread recovery, stale-result
  discard, fresh-result retention, backoff widening, backoff ceiling,
  streak reset on success, streak increment on empty, snapshot shape.
- Existing TestDialecticCadenceAdvancesOnSuccess::test_in_flight_thread_is_not_stacked
  updated to set _prefetch_thread_started_at so it tests the
  fresh-thread-blocks branch (stale path covered separately).
- test_cli TestCmdStatus fake updated with the new config attrs surfaced
  in the status block.
…overage

- TestDialecticDepth::test_first_turn_runs_dialectic_synchronously:
  covered by TestSessionStartDialecticPrewarm::test_turn1_falls_back_to_sync_when_prewarm_missing
  (more realistic — exercises the empty-prewarm → sync-fallback path)
- TestDialecticDepth::test_first_turn_dialectic_does_not_double_fire:
  covered by TestDialecticLifecycleSmoke (turn 1 flow) and
  TestDialecticCadenceAdvancesOnSuccess::test_empty_dialectic_result_does_not_advance_cadence

Both predate the prewarm refactor and test paths that are now
fallback behaviors already covered elsewhere.
…wards-compat fallback

Setup wizard now always writes dialecticCadence=2 on new configs and
surfaces the reasoning level as an explicit step with all five options
(minimal / low / medium / high / max), always writing
dialecticReasoningLevel.

Code keeps a backwards-compat fallback of 1 when dialecticCadence is
unset so existing honcho.json configs that predate the setting keep
firing every turn on upgrade. New setups via the wizard get 2
explicitly; docs show 2 as the default.

Also scrubs editorial lines from code and docs ("max is reserved for
explicit tool-path selection", "Unset → every turn; wizard pre-fills 2",
and similar process-exposing phrasing) and adds an inline link to
app.honcho.dev where the server-side observation sync is mentioned in
honcho.md. Recommended cadence range updated to 1-5 across docs and
wizard copy.
The cherry-picked commit from NousResearch#11434 uses the 154585401+ prefixed
noreply format. Add it alongside the existing bare entry so the
contributor audit passes.
…sResearch#12369)

Agents can now send arbitrary CDP commands to the browser. The tool is
gated on a reachable CDP endpoint at session start — it only appears in
the toolset when BROWSER_CDP_URL is set (from '/browser connect') or
'browser.cdp_url' is configured in config.yaml. Backends that don't
currently expose CDP to the Python side (Camofox, default local
agent-browser, cloud providers whose per-session cdp_url is not yet
surfaced) do not see the tool at all.

Tool schema description links to the CDP method reference at
https://chromedevtools.github.io/devtools-protocol/ so the agent can
web_extract specific method docs on demand.

Stateless per call. Browser-level methods (Target.*, Browser.*,
Storage.*) omit target_id. Page-level methods attach to the target
with flatten=true and dispatch the method on the returned sessionId.
Clean errors when the endpoint becomes unreachable mid-session or
the URL isn't a WebSocket.

Tests: 19 unit (mock CDP server + gate checks) + E2E against real
headless Chrome (Target.getTargets, Browser.getVersion,
Runtime.evaluate with target_id, Page.navigate + re-eval, bogus
method, bogus target_id, missing endpoint) + E2E of the check_fn
gate (tool hidden without CDP URL, visible with it, hidden again
after unset).
…ng session (NousResearch#12441)

session.interrupt on session A was blast-resolving pending
clarify/sudo/secret prompts on ALL sessions sharing the same
tui_gateway process.  Other sessions' agent threads unblocked with
empty-string answers as if the user had cancelled — silent
cross-session corruption.

Root cause: _pending and _answers were globals keyed by random rid
with no record of the owning session.  _clear_pending() iterated
every entry, so the session.interrupt handler had no way to limit
the release to its own sid.

Fix:
- tui_gateway/server.py: _pending now maps rid to (sid, Event)
  tuples.  _clear_pending takes an optional sid argument and filters
  by owner_sid when provided.  session.interrupt passes the calling
  sid so unrelated sessions are untouched.  _clear_pending(None)
  remains the shutdown path for completeness.
- _block and _respond updated to pack/unpack the new tuple format.

Tests (tests/test_tui_gateway_server.py): 4 new cases.
- test_interrupt_only_clears_own_session_pending: two sessions with
  pending prompts, interrupting one must not release the other.
- test_interrupt_clears_multiple_own_pending: same-sid multi-prompt
  release works.
- test_clear_pending_without_sid_clears_all: shutdown path preserved.
- test_respond_unpacks_sid_tuple_correctly: _respond handles the
  tuple format.

Also updated tests/tui_gateway/test_protocol.py to use the new tuple
format for test_block_and_respond and test_clear_pending.

Live E2E against the live Python environment confirmed cross-session
isolation: interrupting sid_a released its own pending prompt without
touching sid_b's.  All 78 related tests pass.
…arch#12444)

When Discord splits a long message at 2000 chars, _enqueue_text_event
buffers each chunk and schedules a _flush_text_batch task with a
short delay.  If another chunk lands while the prior flush task is
already inside handle_message, _enqueue_text_event calls
prior_task.cancel() — and without asyncio.shield, CancelledError
propagates from the flush task into handle_message → the agent's
streaming request, aborting the response the user was waiting on.

Reproducer: user sends a 3000-char prompt (split by Discord into 2
messages).  Chunk 1 lands, flush delay starts, chunk 2 lands during
the brief window when chunk 1's flush has already committed to
handle_message.  Agent's current streaming response is cancelled
with CancelledError, user sees a truncated or missing reply.

Fix (gateway/platforms/discord.py):
- Wrap the handle_message call in asyncio.shield so the inner
  dispatch is protected from the outer task's cancel.
- Add an except asyncio.CancelledError clause so the outer task
  still exits cleanly when cancel lands during the sleep window
  (before the pop) — semantics for that path are unchanged.

The new flush task spawned by the follow-up chunk still handles its
own batch via the normal pending-message / active-session machinery
in base.py, so follow-ups are not lost.

Tests: tests/gateway/test_text_batching.py —
test_shield_protects_handle_message_from_cancel.  Tracks a distinct
first_handle_cancelled event so the assertion fails cleanly when the
shield is missing (verified by stashing the fix and re-running).

Live E2E on the live-loaded DiscordAdapter:
  first_handle_cancelled: False  (shield worked)
  first_handle_completed: True   (handle_message ran to completion)
Follow-up for the helix4u easy-fix salvage batch:
- route remaining context-engine quiet-mode output through
  _should_emit_quiet_tool_messages() so non-CLI/library callers stay
  silent consistently
- drop the extra senderAliases computation from WhatsApp allowlist-drop
  logging and remove the now-unused import

This keeps the batch scoped to the intended fixes while avoiding
leaked quiet-mode output and unnecessary duplicate work in the bridge.
…ay/run.py

Tier 1 (config + tests + run_agent):
- hermes_cli/config.py: adopt upstream auto-provider policy for aux tasks;
  title_generation and follow_up_generation now use provider=auto so they
  inherit the user's configured provider via UserLLMKeys (no silent cross-billing
  through OpenRouter for users who bring their own Anthropic/OpenAI key);
  bump _config_version from 18 to 19 to match upstream
- tests/hermes_cli/test_config_v18_migration.py: rewrite to assert auto-provider
  defaults and _config_version=19; drop stale openrouter/gemini-2.5-flash pin
- tests/hermes_cli/test_tools_config.py: concatenate upstream TestImagegenBackend-
  Registry/TestImagegenModelPicker with Myah TestMyahPlatformSecrets (additive)
- run_agent.py: take upstream verbatim — Myah's OpenRouter fallback was dead code
  since gateway resolves credentials before AIAgent.__init__ via runtime_provider
- tests/run_agent/test_run_agent.py: update test_tool_call_accumulation to reflect
  upstream's intentional = (not +=) assignment for tool function names; add
  test_tool_name_not_duplicated_when_provider_repeats_it to document the fix
- tests/run_agent/test_streaming.py: add missing api_key/base_url to upstream's
  new test_tool_name_not_duplicated_when_resent_per_chunk
- tests/run_agent/test_concurrent_interrupt.py: add _apply_pending_steer_to_tool_results
  no-op to _Stub (method added upstream; not exercised in these tests)

Tier 2 (gateway/run.py — 2 hunks):
- Hunk 1 (Sentry AI monitoring): keep our Sentry AI monitoring block before
  upstream's new bg_review queueing structure (_bg_review_release threading.Event,
  _bg_review_pending, _deliver_bg_review_message); tool_progress callback chain
  verified still works with new upstream callback assignment ordering
- Hunk 2 (already_sent): take upstream's richer check (empty-sentinel + streamed +
  previewed) and add _native_streaming_used as additional OR clause to cover the
  Myah native-SSE path where GatewayStreamConsumer never sees tokens

Also fix test expectations that asserted old hard-pinned defaults:
- tests/gateway/test_myah_management_config.py: update schema + reset assertions
  to expect provider=auto, model='' for title_generation
…orms/myah.py

The block marker at line 332 (media attachments ingestion) had a non-standard
'End Myah:' comment instead of the canonical closer pattern
(# ──── 40+ dashes). Replace with the standard form so the automated marker
parity checker now reports 11/11 for this file.
@github-actions

Copy link
Copy Markdown

⚠️ Supply Chain Risk Detected

This PR contains patterns commonly associated with supply chain attacks. This does not mean the PR is malicious — but these patterns require careful human review before merging.

⚠️ WARNING: base64 encoding/decoding detected

Base64 has legitimate uses (images, JWT, etc.) but is also commonly used to obfuscate malicious payloads. Verify the usage is appropriate.

Matches (first 20):

6965:+        payload = base64.b64encode(text.encode("utf-8")).decode("ascii")
17008:+    return base64.b64encode(os.urandom(32)).decode()
17029:+    key = base64.b64decode(key_base64)
17030:+    raw = base64.b64decode(encrypted_base64)
22940:+    image_bytes = base64.b64decode(b64_data, validate=True)
77389:+        expected_aes_key = base64.b64encode(aes_key.hex().encode("ascii")).decode("ascii")
92169:+        b64_png = base64.b64encode(FAKE_PNG).decode()
92205:+        b64_png = base64.b64encode(FAKE_PNG).decode()
96699:+                                "data": base64.b64encode(fake_pcm_bytes).decode(),
96890:+                                    "data": base64.b64encode(fake_pcm_bytes).decode()
104363:+    pcm_bytes = base64.b64decode(audio_b64)
136848:+  const b64 = Buffer.from(text, 'utf8').toString('base64')
153919:+  process.stdout.write(`\x1b]52;c;${Buffer.from(s, 'utf8').toString('base64')}\x07`)

⚠️ WARNING: exec() or eval() usage

Dynamic code execution can hide malicious behavior, especially when combined with base64 or network fetches.

Matches (first 20):

28904:+       across ``exec()``, so pip and git subprocesses also stop dying on
43795:+op('/project1/noise1').par.period.expr = "tdu.remap(op('/project1/spectrum_scale')['chan1'].eval(), 0, 1, 1, 8)"
45241:+`me.inputVal`, `me.chanIndex`, `me.sampleIndex` work ONLY in cook-context. Calling `par.expr0expr.eval()` from outside always raises an error — this is NOT a real operator error. Ignore these in error scans.
57173:+exec(open(os.path.join(os.environ.get("HERMES_HOME", os.path.expanduser("~/.hermes")), "skills/red-teaming/godmode/scripts/parseltongue.py")).read())
57182:+exec(open(os.path.join(os.environ.get("HERMES_HOME", os.path.expanduser("~/.hermes")), "skills/red-teaming/godmode/scripts/godmode_race.py")).read())
57195:+exec(open(os.path.join(os.environ.get("HERMES_HOME", os.path.expanduser("~/.hermes")), "skills/red-teaming/godmode/scripts/godmode_race.py")).read())
57208:+exec(open(os.path.join(os.environ.get("HERMES_HOME", os.path.expanduser("~/.hermes")), "skills/red-teaming/godmode/scripts/godmode_race.py")).read())
57234:+    exec(open(os.path.join(os.environ.get("HERMES_HOME", os.path.expanduser("~/.hermes")), "skills/red-teaming/godmode/scripts/godmode_race.py")).read())
57260:+    exec(open(os.path.join(os.environ.get("HERMES_HOME", os.path.expanduser("~/.hermes")), "skills/red-teaming/godmode/scripts/parseltongue.py")).read())
93122:+time and used that stale value for every subsequent _exec() call.  When
93132:+Fix: _exec() now prefers the LIVE ``env.cwd`` over the init-time
93184:+    """_exec() must use live env.cwd, not the init-time cached cwd."""
93204:+        result = ops._exec("cat target.txt")
93254:+        result = ops._exec("cat target.txt", cwd=str(dir_c))
93272:+        result = ops._exec("cat target.txt")
99543:+        self._sandbox.process.exec(
99549:+            self._sandbox.process.exec(f"rm -f {shlex.quote(remote_tar)}")
100580:+            Every _exec() call prefers the LIVE ``terminal_env.cwd`` over
100587:+            init-time cwd for every _exec() call, which caused relative
117693:+    const matches = ANSI_REGEX.exec(color)

⚠️ WARNING: Outbound network calls (POST/PUT)

Outbound POST/PUT requests in new code could be data exfiltration. Verify the destination URLs are legitimate.

Matches (first 10):

4634:+        with urllib.request.urlopen(request, timeout=timeout) as response:
5476:+        with urllib.request.urlopen(request, timeout=timeout) as response:
5556:+        with urllib.request.urlopen(request, timeout=timeout) as response:
24275:+    with urllib.request.urlopen(req, timeout=30) as resp:
24436:+        resp = requests.post(url, json=payload, timeout=15)
79366:+        with patch("hermes_cli.debug.urllib.request.urlopen",
104222:+    response = requests.post(
104331:+    response = requests.post(
107617:+                    with urllib.request.urlopen(probe, timeout=2.0) as resp:

⚠️ WARNING: Install hook files modified

These files can execute code during package installation or interpreter startup.

Files:

hermes_cli/memory_setup.py
hermes_cli/setup.py
skills/productivity/google-workspace/scripts/setup.py
tests/hermes_cli/test_setup.py
tests/skills/test_google_oauth_setup.py

⚠️ WARNING: CI/CD workflow files modified

Changes to workflow files can alter build pipelines, inject steps, or modify permissions. Verify no unauthorized actions or secrets access were added.

Files:

.github/workflows/deploy-site.yml

⚠️ WARNING: Container build files modified

Changes to Dockerfiles or compose files can alter base images, add build steps, or expose ports. Verify base image pins and build commands.

Files:

Dockerfile

⚠️ WARNING: Dependency manifest files modified

Changes to dependency files can introduce new packages or change version pins. Verify all dependency changes are intentional and from trusted sources.

Files:

pyproject.toml
ui-tui/package.json
ui-tui/packages/hermes-ink/package.json

Automated scan triggered by supply-chain-audit. If this is a false positive, a maintainer can approve after manual review.

@deestax
deestax merged commit 166decb into main Apr 19, 2026
5 of 7 checks passed
@deestax
deestax deleted the merge-upstream branch April 19, 2026 09:54
deestax added a commit that referenced this pull request May 7, 2026
…Task 2A.7 + drop #3)

Previously gateway/platforms/api_server.py:76-82 had an inline
'try: from logging_setup import setup_sentry; setup_sentry(); except ImportError: pass'
Myah marker block. logging_setup.py was Myah-specific (agent-container
Sentry init + SentryHook adapter for agent.telemetry.TelemetryHook
protocol) but lived in an upstream-counterpart top-level path.

This commit:
- Renames logging_setup.py -> myah_hermes_plugin/sentry_init.py.
- Calls sentry_init.setup_sentry() from the plugin's register() so
  hosted-Myah agent containers still get Sentry on gateway boot.
- Drops the inline init Myah-marker block from api_server.py.

Idempotent: setup_sentry() returns silently when SENTRY_DSN_AGENT is
unset (the OSS-user case).

Closes drop candidate #3 from spec 2026-05-06 §2.3 (deferred from
Tier 2A Task 2A.1 to 2A.7 because it depended on having a plugin-side
sentry_init.py to call into).

Implements Tier 2A Task 2A.7 from spec 2026-05-06 §3.
deestax added a commit that referenced this pull request May 7, 2026
…ner + Sentry/admin moves (#24)

* chore(hermes): drop unused _get_adapter_for_platform() helper

Verified zero callers via repo-wide grep. Function was a leftover
from earlier dispatch refactoring.

Drop candidate #2 from spec 2026-05-06 §2.3 (Tier 2A Task 2A.1.1).

* chore(hermes): drop os.environ CHAT_ID writer (concurrency bug)

Per-request mutation of process-wide environment, racing with
concurrent requests. The TODO comment confirmed this was a known
defect awaiting migration. set_session_vars() in session_context.py
is the contextvar-safe replacement.

Drop candidate #5 from spec 2026-05-06 §2.3 (Tier 2A Task 2A.1.2).

* chore(hermes): catch up MAX_REQUEST_BYTES to upstream (10MB)

Verified via git blame: the 1MB value at line 88 was authored by
Teknium upstream at commit 80cc27e (2026-03-24) and later bumped
to 10MB upstream. Fork has the older value due to drift, not a Myah
security decision.

Drop candidate #6 from spec 2026-05-06 §2.3 (Tier 2A Task 2A.1.3).

* chore(hermes): drop /health Myah credential-validation extension

The MYAH_ADAPTER_ENABLED-gated credential check was specific to the
fork-bundled adapter and goes away with the plugin-based OSS
architecture. /health/detailed (line 832 below) provides rich
runtime status; 'hermes doctor' covers credential validation.

Drop candidate #4 from spec 2026-05-06 §2.3 (Tier 2A Task 2A.1.4).

* chore(hermes): drop cronjob_tools origin-unresolved WARNING

Pure debug logging — the function returns None either way; the
WARNING was added for diagnostic purposes during a 2025-era cron
delivery bug. Function behavior unchanged.

Drop candidate #7 from spec 2026-05-06 §2.3 (Tier 2A Task 2A.1.5).

* feat(plugin): move myah_overrides.py from hermes_cli/ into plugin

Provider catalog augmentation (230 LOC) was Myah-only data living in an
upstream-counterpart directory. Moves it into the pip-installed plugin
at myah_hermes_plugin/myah_admin/myah_overrides.py and updates the two
known importers (plugins/myah-admin/dashboard/_providers.py and the
plugin tests).

After this commit, hermes_cli/ contains zero Myah-only files — one
fewer cardinal-counterpart file to keep marker hygiene on.

Implements Tier 2A Task 2A.6 from spec 2026-05-06 §3.

* feat(plugin): move Sentry init from upstream into plugin register() (Task 2A.7 + drop #3)

Previously gateway/platforms/api_server.py:76-82 had an inline
'try: from logging_setup import setup_sentry; setup_sentry(); except ImportError: pass'
Myah marker block. logging_setup.py was Myah-specific (agent-container
Sentry init + SentryHook adapter for agent.telemetry.TelemetryHook
protocol) but lived in an upstream-counterpart top-level path.

This commit:
- Renames logging_setup.py -> myah_hermes_plugin/sentry_init.py.
- Calls sentry_init.setup_sentry() from the plugin's register() so
  hosted-Myah agent containers still get Sentry on gateway boot.
- Drops the inline init Myah-marker block from api_server.py.

Idempotent: setup_sentry() returns silently when SENTRY_DSN_AGENT is
unset (the OSS-user case).

Closes drop candidate #3 from spec 2026-05-06 §2.3 (deferred from
Tier 2A Task 2A.1 to 2A.7 because it depended on having a plugin-side
sentry_init.py to call into).

Implements Tier 2A Task 2A.7 from spec 2026-05-06 §3.

* feat(plugin): add pre_gateway_dispatch hook (Tier 2A Task 2A.4)

Replaces the skip_user_authorization semantics that PR #20 removed.
Adds myah_hermes_plugin/myah_platform/pre_dispatch_hook.py with a
single hook callback that returns {'action': 'allow'} for Myah-platform
messages and None (passthrough) for everything else.

Architectural note: returning 'allow' does NOT bypass the gateway-level
user-allowlist check at gateway/run.py:3655. Auth bypass for Myah's
single-tenant deployment is handled by allow_all_env=MYAH_ALLOW_ALL_USERS
on the platform registration. The hook exists as a documented choke
point for future Myah-specific routing logic (rate limiting, silent
ingest, content rewrite) that must NOT live in upstream gateway/run.py.

Wired via ctx.register_hook('pre_gateway_dispatch', ...) in the plugin's
register() — guarded with hasattr(ctx, 'register_hook') so older
PluginContext versions don't fail.

Implements Tier 2A Task 2A.4 from spec 2026-05-06 §3.

* feat(plugin): vendor _dispatch_approval_notify + register_gateway_notify

Plugin owns its own copy of the notify-dispatch chain. Without it, the
plugin's vendored request_action_confirmation has no transport to the
platform adapter and the agent silently auto-approves.

Variadic dispatch (1/2/3-arity callbacks) matches upstream's
gateway/run.py:_dispatch_approval_notify behavior at lines 334-460.
Uses threading.Lock for the registry to match upstream's thread-safe
semantics in tools/approval.py.

Implements docs/superpowers/plans/2026-05-06-myah-oss-completion-2a-plugin-stock-upstream.md
Task 2A.2.0.

* feat(plugin): vendor request_action_confirmation + _action_queues registry

Plugin owns its own action confirmation primitives mirroring upstream's
tools/approval.py:request_action_confirmation (sync, threading.Event-based).

The sync contract is required because cronjob_tools._execute() runs in
a tool-runner threadpool and blocks on event.wait(timeout) — an asyncio
version would force rewriting the cron tool.

Two of the six tests intentionally fail until Task 2A.2.2 lands the
plugin's cron_tool.py shadow.

Implements docs/superpowers/plans/2026-05-06-myah-oss-completion-2a-plugin-stock-upstream.md
Task 2A.2.1.

* feat(plugin): shadow upstream cronjob_tools.py with plugin-owned cron tool

Verbatim copy of upstream's tools/cronjob_tools.py, with the single
change of importing request_action_confirmation from the
plugin-vendored myah_hermes_plugin.cron_approval. Tool registration
uses the same 'cronjob' name as upstream so last-writer-wins on import
order shadows upstream's handler with the plugin's.

All 6 tests in test_cron_approval.py now pass (the two cron_tool
sanity checks added in 2A.2.1 had been failing pending this commit).

Implements docs/superpowers/plans/2026-05-06-myah-oss-completion-2a-plugin-stock-upstream.md
Task 2A.2.2.

* refactor(plugin): adapter.py uses plugin-vendored approval primitives

Move 4 import sites in adapter.py from upstream's tools.approval to the
plugin's vendored modules:

- resolve_action_confirmation         → myah_hermes_plugin.cron_approval
- resolve_action_confirmation_by_session → myah_hermes_plugin.cron_approval
- unregister_gateway_notify (×2)      → myah_hermes_plugin.dispatcher

Update mock targets in test_myah_confirm_dispatch.py to follow the new
import paths. Comment block in send_action_confirmation updated to
reference the plugin's cron_approval module.

resolve_gateway_approval (legacy terminal-command approval) stays on
tools.approval — it's outside Tier 2A's cron-approval scope.

Implements docs/superpowers/plans/2026-05-06-myah-oss-completion-2a-plugin-stock-upstream.md
Task 2A.2.3.

* feat(plugin): adapter standalone aiohttp runner on MYAH_GATEWAY_PORT

Tier 2A Task 2A.3 — collapse the previous hosted/standalone split. The
plugin's MyahAdapter now ALWAYS owns its own aiohttp AppRunner +
TCPSite via the new MyahStandaloneRunner helper, removing the
dependency on upstream's gateway/platforms/api_server.py
register_pre_setup_hook + get_shared_app.

This is a one-way door for hosted Myah: once shipped, hosted Myah uses
the standalone runner forever. See spec
docs/superpowers/specs/2026-05-06-myah-oss-completion-design.md §3
Task 2A.3 for the rationale.

Default port is now 8643 via MYAH_GATEWAY_PORT (was 8642 hard-coded).
Port resolution still respects config.extra.port and MYAH_ADAPTER_PORT
overrides for backward compat with existing yaml.

Tests cover ephemeral binding (port=0 → OS-assigned), live HTTP route
dispatch, and env-var resolution fallback. Zero new failures in the
gateway/tools test directories vs baseline (55 pre-existing failures
unchanged).

* test: fix Tier 2A regressions surfaced by full CI suite

The PR's local verification only ran tests/gateway/ + tests/tools/ —
three tests in tests/cron/, tests/plugins/, and tests/tools/test_resolve_path.py
broke under the full CI run.

1. tests/cron/test_origin_capture_log.py — DELETED.
   Task 2A.1.5 dropped the WARNING that this entire 'Bug E' file was
   designed to verify. The file's purpose is gone with the warning;
   keeping the happy-path companion test would just duplicate coverage
   that lives elsewhere.

2. tests/plugins/test_myah_admin_providers.py::test_list_providers_against_real_catalog —
   import switched to pytest.importorskip(). Task 2A.6 moved
   myah_overrides.py from hermes_cli/ into the plugin distribution,
   but the hermes Tests workflow only runs 'uv pip install -e ".[all,dev]"'
   against the core package — it never installs the plugin. Local dev
   runs (where the plugin is editable-installed) still exercise the
   real-catalog assertion; CI now skips it cleanly. The plugin's own
   pytest suite covers the catalog path independently.

3. tests/tools/test_resolve_path.py — added isolated_resolve_state
   fixture that clears tools.file_tools._file_ops_cache and
   tools.terminal_tool._active_environments before the two tests that
   rely on falling through to TERMINAL_CWD. Adding the new plugin
   modules to the suite shifted xdist worker assignment in a way that
   exposed pre-existing pollution from other tests leaving a 'default'
   task_id entry in those module-level dicts. The fixture is per-test
   (monkeypatch.setattr) so it doesn't disrupt other tests.

* test: gentler test_resolve_path isolation (don't rebind module attrs)

Previous fix used monkeypatch.setattr to swap _file_ops_cache and
_active_environments with empty dicts. This unintentionally caused
test_file_state_registry tests to fail on Linux CI under xdist —
rebinding the module attribute discarded entries other tests had
written under non-'default' keys, and the cross-worker timing
exposed the leak.

The replacement fixture mutates the existing dicts in place — only
pops the 'default' key (which is what _resolve_path reads when no
explicit task_id is passed) and restores the prior value after the
test. No module-attribute rebinding, no impact on entries other
tests created.

* ci: trigger re-run to confirm test flakiness

Recent runs showed 3 new tests intermittently failing on Linux CI
that pass locally and pass in isolation:
- test_concurrent_inserts_settle_at_cap (30s timeout)
- test_modal_sandbox_fixes::TestToolResolution::test_terminal_tool_present
- test_modal_sandbox_fixes::TestToolResolution::test_terminal_and_file_toolsets_resolve_all_tools

This empty commit re-runs CI to confirm whether they're deterministic
regressions or pre-existing test-ordering flakes exposed by the changed
test count from this PR.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.