Skip to content

fix(cron): multiplex ticks isolate per profile, fire the bound port, deliver satellites fail-closed (#74878 #100489 #101113 #86519, salvage #70747 #84755 #96944) - #101245

Merged
teknium1 merged 7 commits into
mainfrom
salvage/mux-cron
Sep 2, 2026
Merged

teknium1 merged 7 commits into
mainfrom
salvage/mux-cron

Conversation

@teknium1

@teknium1 teknium1 commented Sep 2, 2026

Copy link
Copy Markdown
Collaborator

Summary

Five multiplex cron bugs fixed in one branch: one profile's broken store killed every profile's ticker; dashboard fires targeted an unbound port; the desktop ticker delivered secondaries with the default identity; credentialless profile_routes satellites couldn't deliver at all; notepad/suggestions paths were frozen at import. Root causes: missing per-profile except in the ticker loops, target-profile port lookup under a shared listener, an unscoped ThreadPoolExecutor + unconditional desktop ticking, an empty adapter map with no route-exact transport resolver, and import-time get_hermes_home().

Changes

  • cron/scheduler_provider.py: per-profile except BaseException in recovery + tick loops (CronTickYielded/_profile_errors + gateway: cron scheduler permanently stalls after EMFILE while heartbeat keeps reporting healthy #87644 backoff preserved); optional profile_gate; SharedRouteAdapters view for map-less secondaries.
  • cron/scheduler.py: copy_context() on the asyncio.run fallback pool; _primary_profile_routes_for_current_home() shared by preflight + delivery; SharedRouteAdapters (route-exact, fail-closed) used per target.
  • hermes_cli/web_server.py: multiplex fire URL resolves port from the default listener (env override parity, logged fallback); desktop ticker stands down for profiles whose own gateway runs.
  • cron/notepad.py, cron/suggestions.py: call-time path resolution (fix(cron): resolve profile store paths per call #96944).
  • Tests: 9 new tests pinning each invariant; docs paragraph for routed-profile cron delivery.

Validation

Fix Before (origin/main) After
Ticker isolation alpha's corrupt DB killed thread; no profile ticked default+beta tick, alpha error recorded
Fire port :8701/p/worker_alpha/… → connection refused bound default port
Fallback pool scope UnscopedSecretError, default home ops home + OPS-TOKEN
Desktop gate ticked ['default','ops'] beside ops' live gateway ticked ['default']
Satellite delivery "DISCORD_BOT_TOKEN is not set" on routed chat primary bot sends routed chat; unrouted never uses it
Store paths notepad/suggestions written to default home written to scoped profile

Credits

@Cyber-Yichen (#70747, authored), @OYLFLMH (#74888), @webtecnica (#74952), @bergusdz (#84755, authored), @wanliqin + @honor2030 (#96944, cherry-picked), @tachyon-r (#74017, first on suggestions half), @TheBlueHouse75 (#93043, independent identification), reporters @minkiboo (#101113), @herovivian-collab (#100489), @wanliqin (#86519).

Fixes #100489
Fixes #101113
Fixes #86519
Fixes #74878

Infographic

mux-cron

@github-actions

github-actions Bot commented Sep 2, 2026 •

Copy link
Copy Markdown

૮ >ﻌ< ა ci review

ran on 803c683 — docs(profiles): cron delivery for routed profiles rides the

⚠️ Warnings

OSV vulnerability scan · View job

13 known vulnerabilities found in pinned dependencies.

How to fix:

Review the findings in the Security tab. Update the affected dependencies if a patched version is available.


debug info

CI timings

CI timings · View report · View job

Wall time 5m39s vs 6m13s (-9.1%). 4 job(s) slower, 10 faster, 1 unchanged.

  • Check no case-colliding filenames / check-case-collisions: +62.0s
  • Python tests / e2e: -51.0s
  • Python tests / Run tests: -41.0s
  • Python lints / ruff enforcement (blocking): +40.0s
  • OS-specific tests / Windows-only tests: -39.0s

@alt-glitch alt-glitch added type/bug Something isn't working P2 Medium — degraded but workaround exists comp/cron Cron scheduler and job management comp/cli CLI entry point, hermes_cli/, setup wizard area/config Config system, migrations, profiles area/profiles Multi-profile isolation, HERMES_HOME scoping sweeper:risk-message-delivery Sweeper risk: may drop, duplicate, misroute, or suppress messages labels Sep 2, 2026
@webtecnica

Copy link
Copy Markdown

Thanks for the credit on #74952 — glad the per-profile ticker isolation landed.

@teknium1
teknium1 force-pushed the salvage/mux-cron branch 4 times, most recently from 69b633f to 74e9597 Compare September 2, 2026 13:08
wanliqin and others added 7 commits September 2, 2026 06:19
Resolve notepad and suggestion paths at transaction time so multiplexed profile ticks cannot write into the import-time home. Preserve explicit test overrides and cover writes after a profile context switch.

Co-authored-by: 이민재 <19909783+honor2030@users.noreply.github.com>
One profile's broken cron store no longer takes the whole multiplex ticker
down with it:

- startup recovery loop: a per-profile exception (e.g. an unreadable
  executions.db raising sqlite3.DatabaseError) was uncaught and killed the
  ticker thread before its first tick — no profile ever fired.
- tick loop: only CronTickYielded was caught per profile; any other
  exception escaped to the cycle-wide handler, skipping every remaining
  profile that cycle and marking all of them failed.

Both loops now catch per profile, record the failure into THAT profile's
ticker_last_error (`hermes cron status`), and keep ticking the siblings.
The existing CronTickYielded/_profile_errors semantics and the #87644
EMFILE reclaim/backoff are preserved (backoff is applied once per cycle
from the worst per-profile failure).

Salvaged from PR #70747 (@Cyber-Yichen); the recovery test's real
sqlite3.OperationalError shape is from PR #74888 (@OYLFLMH). Same class
also reported in PR #74952 (@webtecnica).

Co-authored-by: OYLFLMH <95945448+OYLFLMH@users.noreply.github.com>
Co-authored-by: webtecnica <75556242+webtecnica@users.noreply.github.com>
…stener port

Under gateway.multiplex_profiles only the DEFAULT profile's api_server is
bound; secondaries share it via /p/<profile>/ mirrors. _gateway_fire_endpoint
read the port from the TARGET profile's config.yaml/.env and then prefixed
the mirror path, so a secondary with its own API_SERVER_PORT produced a URL
nothing listens on (connection refused on every Chronos fire).

Multiplex is now detected first (config.yaml + the GATEWAY_MULTIPLEX_PROFILES
override via gateway.config._env_multiplex_profiles_override — same
semantics as the gateway loader), and in that mode the port is resolved
from the default root's config/.env with the fallback logged. Per-profile
gateway topology is unchanged.

Salvaged from PR #84755 (@bergusdz), with env-override parity restored and
the silent except replaced by a debug log.
…p ticker stands down for profiles with their own gateway (#100489)

Two mechanisms let the desktop multiplex ticker deliver a secondary
profile's cron output through the default profile's identity:

1. _deliver_result's `asyncio.run` ThreadPoolExecutor fallback (taken when
   the caller already has a running loop — the desktop dashboard shape) ran
   the standalone sender on a fresh thread with NO profile ContextVars: the
   home override and secret scope were gone, so the sender resolved the
   process default's home/token (or, fail-closed under multiplex, raised
   UnscopedSecretError). Wrap the submit in copy_context().run like the
   session-db (:6562), heartbeat (:4650) and parallel-pool (:8314) workers.

2. _start_desktop_cron_ticker ticked EVERY local profile, including ones
   whose own gateway (with live adapters) is running; winning the tick-lock
   race meant the adapter-less desktop ticker delivered standalone. The
   multiplex loop gains an optional per-cycle `profile_gate(name, home)`;
   the desktop wires it to `_check_gateway_running(home)` so such profiles
   are neither ticked nor heartbeated by the dashboard while their gateway
   is alive (re-evaluated every cycle, no restart needed).

Fixes #100489
…he primary adapter for exact profile_routes targets (#101113)

Under gateway.multiplex_profiles a shared-token satellite profile (routed
via gateway.profile_routes, no bot credential of its own) got an empty
adapter map from the multiplex ticker, so _deliver_result fell through to
the standalone sender under the satellite's secret scope and failed with
"DISCORD_BOT_TOKEN is not set" — even though the primary adapter owns the
exact routed channel and had delivered the same target before. Preflight
already rescued this topology (#97476); the delivery half did not.

- cron/scheduler.py: factor the preflight's primary-config route loader into
  `_primary_profile_routes_for_current_home()` (one owner for both halves,
  so route semantics cannot drift) and add `SharedRouteAdapters`, a
  read-only view over the primary adapter map that resolves an adapter for
  a (platform, target) ONLY when an enabled primary route with a
  chat_id/thread_id maps that exact target to the current profile —
  using the same `ProfileRoute.matches` predicate as inbound routing.
  `_deliver_result` resolves the transport per target from it; everything
  else (unmatched chat, disabled route, route for another profile, no
  primary adapter, guild-only route) is a miss and never uses the primary
  bot. Execution stays scoped to the satellite; no credential is copied.
- cron/scheduler_provider.py: a secondary with no adapter map of its own
  gets the SharedRouteAdapters view instead of `{}`. This is NOT a default
  fallback: with no matching route the view is falsy and delivers nothing.

Fixes #101113
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/config Config system, migrations, profiles area/profiles Multi-profile isolation, HERMES_HOME scoping comp/cli CLI entry point, hermes_cli/, setup wizard comp/cron Cron scheduler and job management P2 Medium — degraded but workaround exists sweeper:risk-message-delivery Sweeper risk: may drop, duplicate, misroute, or suppress messages type/bug Something isn't working

Projects

None yet

7 participants