Skip to content

fix(cron): multiplexed cron fires and delivers for every served profile — restart-safe scope, guild-scoped shared routes, Desktop ticker allowlist, idle-exit (#107399, #89302, salvage #107413/#104979/#103844/#103742) - #108428

Merged
teknium1 merged 7 commits into
mainfrom
fix/mux-cron-multiplex
Sep 11, 2026

Conversation

@teknium1

Copy link
Copy Markdown
Collaborator

Under gateway.multiplex_profiles, cron jobs now actually fire and deliver for every served profile: the systemd restart-safe handoff no longer dies on UnscopedSecretError, routed satellites deliver through the shared bot for the documented guild_id + chat_id route shape (and without a platforms: block), the Desktop/serve ticker honours the allowlist and yields to the multiplexer, and the SSH-isolated idle-exit no longer kills a running cron job.

Changes

Root cause (one line each)

F3: the handoff ran before _run_one_job_body installs the fire scope. F5: route.matches(...) was called without guild_id, so any route declaring one returned False. F10/#89302: _resolve_target_transport applied the satellite's native enabled gate to a transport the primary authorized. F6: profiles_to_serve(multiplex=True) without the allowlist + a pid-only gate. F7: the idle probe read only the dashboard session table. Named: profiles_to_serve yields default + allowlist, never the -p launch profile.

Live repro

Script: /tmp/mux_audit/fix-cron-multiplex/repro.py (temp HERMES_HOME, real imports, profiles/worker, profiles/guest, allowlist [worker], routes {guild G1, chat C1} and {chat C2}, env_passthrough: [BW_SESSION]).

Before (origin/main ad03f20dd61):

F3     FIRE: UnscopedSecretError: get_secret('BW_SESSION') called with no profile secret scope active while multiplexing is
F5     FIRE: guild+chat route -> None; chat-only route -> adapter
F10    FIRE: platform 'discord' not configured/enabled
89302  FIRE: platform 'discord' not configured/enabled
F6     FIRE: ticked profiles=['default', 'guest', 'worker']; gate(worker served by multiplexer) -> True
F7     FIRE: turn_in_flight() with a running cron job -> False
named  FIRE: -p guest multiplexer ticks ['default', 'worker']

After:

F3     clean: launch -> True, BW_SESSION in worker env: 'profile-value'
F5     clean: guild+chat route -> adapter; chat-only route -> adapter
F10    clean: shared transport resolved
89302  clean: native transport resolved
F6     clean: ticked profiles=['default', 'worker']; gate(worker served by multiplexer) -> False
F7     clean: turn_in_flight() with a running cron job -> True
named  clean: -p guest multiplexer ticks ['default', 'worker', 'guest']

E2E (e2e_delivery.py, real config loader + DeliveryRouter + _deliver_result, satellite with NO platforms: block): before routed C1 -> "platform 'discord' not configured/enabled"; after routed C1 -> error=None, primary sent ['C1'], unrouted C9 still fails closed (never the primary bot).

Validation

Check Result
scripts/run_tests.sh tests/cron/ 101 files, 1238 passed, 0 failed
scripts/run_tests.sh tests/hermes_cli/ 893 files, 9443 passed, 1 failed — test_update_head_moved_gate.py::test_update_success_when_head_moves, fails identically on unmodified origin/main (pre-existing, unrelated hermes update test)
scripts/run_tests.sh tests/gateway/ 824 files, 8383 passed, 0 failed
New tests red on base 8/8 red on origin/main (tests_red_on_base.txt)
ruff / windows-footguns / compat pointers / subprocess-stdin / git diff --check clean

Tests added: test_guild_scoped_route_authorizes_cron_target_even_when_satellite_has_no_platform_block, test_live_native_adapter_without_platform_block_is_not_treated_as_disabled, test_turn_probe_counts_in_flight_cron_execution, test_desktop_ticker_honours_allowlist_and_yields_to_default_multiplexer, test_cron_shared_adapter_owner_is_the_launch_profile (+ salvaged test_cron_tick_homes_include_active_named_host, test_single_profile_ticks_only_without_gateway, test_launch_external_worker_uses_restart_safe_scope_and_acknowledges extension).

Credits

Not done / design call

Fixes #107399
Fixes #89302
Addresses #107485
Addresses #94590

Infographic

infographic

fangliquanflq and others added 7 commits September 11, 2026 09:27
…ild-scoped routes and profiles without a platforms block

SharedRouteAdapters.get called ProfileRoute.matches without guild_id, so the
documented Discord route shape (guild_id + chat_id) never authorized a cron
target and the satellite fell to standalone delivery ("DISCORD_BOT_TOKEN is
not set" every fire). A cron target has no inbound guild anchor; the route's
own guild_id is passed so its target-exact discriminators decide.

_resolve_target_transport then vetoed the authorized shared transport on the
SATELLITE's platforms.<p>.enabled (absent block or enabled: false), although
that block describes a connector the satellite never runs. The shared hit now
builds the transport directly (keeping the satellite's non-credential platform
settings), and a live native adapter with no config block is no longer read as
"disabled" (#89302) — same normalization the relay path already had.

Fixes #89302
Co-authored-by: web3blind <264741654+web3blind@users.noreply.github.com>
…s down for served satellites

The Desktop/serve backend ticked every installed profile (ignoring
gateway.multiplex_profile_allowlist) and gated only on the profile's OWN
gateway.pid. A satellite served by the default multiplexer has no pid file, so
both tickers raced its fires and the Desktop one won nondeterministically —
adapter-less standalone delivery, and the environment behind #107485.

Homes now come from profiles_to_serve with the default profile's allowlist
(the multiplexer's served set); the per-tick gate also consults
named_profile_served_by_running_multiplexer.

Addresses #107485, #94590
Co-authored-by: fangliquan <fangliquan@qq.com>
…on job runs

turn_in_flight read only the dashboard session table; an in-process cron run
never registers there, so the watchdog reported "no running turn" and exited
mid-job (tool calls then failed with "cannot schedule new futures after
interpreter shutdown", the execution was marked unknown, the slot lost). The
probe now also consults cron.scheduler.get_running_job_ids — the ledger the
gateway shutdown drain already uses.

Addresses #107485
…d owns the shared adapters

profiles_to_serve(multiplex=True) yields default + allowlist, so a gateway run
as `hermes -p <name>` with multiplex on never ticked its own profile's jobs
unless allowlisted (which would start a second adapter on the same token). The
ticker's home list now unions the active profile. The shared-adapter owner
passed to the ticker is the runner's launch profile instead of the literal
"default", so that profile's jobs reuse its live adapters rather than the
fail-closed empty map.

Co-authored-by: Paul Pincente <101599379+pincente@users.noreply.github.com>
Co-authored-by: r3x443 <325334945+r3x443@users.noreply.github.com>
@github-actions

github-actions Bot commented Sep 11, 2026 •

Copy link
Copy Markdown

૮ >ﻌ< ა ci review

ran on 121c624 — docs(multiplex): cron ticker allowlist, named multiplexer, g

⚠️ Warnings

OSV vulnerability scan · View job

80 known vulnerabilities found in pinned dependencies.

How to fix:

Review the findings in the Security tab. Update the affected dependencies if a patched version is available.


debug info

CI timings

CI timings · View report · View job

Wall time 5m57s vs 4m51s (+22.7%). 11 job(s) slower, 3 faster, 1 unchanged.

  • OS-specific tests / Windows-only tests: +67.0s
  • Docs Site / docs-site-checks: +63.0s
  • Python lints / Windows footguns (blocking): +16.0s
  • Python tests / Run tests: +10.0s
  • OS-specific tests / macOS-only tests: -10.0s

@alt-glitch alt-glitch added type/bug Something isn't working P2 Medium — degraded but workaround exists comp/cron Cron scheduler and job management comp/gateway Gateway runner, session dispatch, delivery comp/cli CLI entry point, hermes_cli/, setup wizard area/profiles Multi-profile isolation, HERMES_HOME scoping sweeper:risk-message-delivery Sweeper risk: may drop, duplicate, misroute, or suppress messages labels Sep 11, 2026
@teknium1
teknium1 merged commit d807a34 into main Sep 11, 2026
40 checks passed
@teknium1
teknium1 deleted the fix/mux-cron-multiplex branch September 11, 2026 22:28
teknium1 added a commit that referenced this pull request Sep 12, 2026
… unified init path

The unified runtime copied _init_registries_and_clocks and the cron bootstrap from a
pre-fix snapshot: HookRegistry() (one process-wide registry, so served secondaries'
hooks/ never load and fire the launch profile's handlers) and default_profile="default"
with _multiplex_profile_homes (a --profile <name> multiplexer's own jobs treated as a
secondary's). Restore ProfileHookRegistries and _cron_tick_profile_homes /
runner._primary_profile_name from main (#108453, #108428).
teknium1 added a commit that referenced this pull request Sep 12, 2026
…ery profile

The multiplexing default gateway now serves default + every live named profile
under profiles/. profiles_to_serve(multiplex=True) is a pure directory read
(tombstoned profiles skipped, never mkdir); every reader — gateway served set,
/p/<profile>/ prefixes for api_server + webhook, the named-profile standalone
guard, the Desktop cron ticker (its #108428 standdown for a profile owned by a
running gateway is unchanged) — drops the allowlist parameter.

Config v43 migration deletes the key from user config.yaml; DEFAULT_CONFIG,
GatewayConfig and the top-level yaml bridge no longer carry it.

BREAKING: anyone who set an allowlist now has their excluded profiles served.
Archive or delete a profile you do not want served (Teknium approved).
teknium1 added a commit that referenced this pull request Sep 12, 2026
…ery profile

The multiplexing default gateway now serves default + every live named profile
under profiles/. profiles_to_serve(multiplex=True) is a pure directory read
(tombstoned profiles skipped, never mkdir); every reader — gateway served set,
/p/<profile>/ prefixes for api_server + webhook, the named-profile standalone
guard, the Desktop cron ticker (its #108428 standdown for a profile owned by a
running gateway is unchanged) — drops the allowlist parameter.

Config v43 migration deletes the key from user config.yaml; DEFAULT_CONFIG,
GatewayConfig and the top-level yaml bridge no longer carry it.

BREAKING: anyone who set an allowlist now has their excluded profiles served.
Archive or delete a profile you do not want served (Teknium approved).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/profiles Multi-profile isolation, HERMES_HOME scoping comp/cli CLI entry point, hermes_cli/, setup wizard comp/cron Cron scheduler and job management comp/gateway Gateway runner, session dispatch, delivery P2 Medium — degraded but workaround exists sweeper:risk-message-delivery Sweeper risk: may drop, duplicate, misroute, or suppress messages type/bug Something isn't working

Projects

None yet

4 participants