Skip to content

fix(cron): preserve active turns across gateway restarts (salvage #101877) - #101940

Merged
kshitijk4poor merged 5 commits into
NousResearch:mainfrom
kshitijk4poor:salvage/101877-cron-restart-durability
Sep 3, 2026
Merged

kshitijk4poor merged 5 commits into
NousResearch:mainfrom
kshitijk4poor:salvage/101877-cron-restart-durability

Conversation

@kshitijk4poor

Copy link
Copy Markdown

Summary

Cron jobs fired by a systemd-managed Linux gateway now survive systemctl restart / hermes gateway restart: the fire is handed to a detached systemd-run --scope worker that owns the execution durably, and its final delivery is queued for whichever gateway instance is live.

Salvage of #101877 by @jayleaton — the three commits are authored by Brooklyn Nicholson (@OutThisLife) and cherry-picked with authorship preserved. Two follow-up commits fold in the review findings.

Root cause: #60711 made the shutdown drain see in-flight cron work, but a forced restart still kills the gateway cgroup, and the cron agent turn with it. start_new_session alone does not escape a systemd service cgroup.

Changes

  • cron/scheduler.py: run_one_job hands a managed-gateway fire to python -m cron.scheduler --external-worker-file … via restart_safe_gateway_child_argv; parent waits on the worker process (1s wait(timeout)), never re-runs the job in-process on an uncertain handoff; worker adopts the execution row (CAS claimed+handoff_pending → running) before any side effect; restart-safe waiters are excluded from forced-shutdown interruption; stranded payload/ack files reaped once terminal; worker entrypoint sets up hermes logging.
  • cron/executions.py: handoff_pending/handoff_started_at columns (via sqlite_util.add_column_if_missing), mark_execution_handoff_pending, adopt_claimed_execution, owner-fenced running/finish/recovery with a 30s adoption grace.
  • cron/delivery_queue.py (new): profile-local sqlite queue; gateway claims a row before transport I/O, dead claimant → unknown never replayed; unclaimed rows stay queued across a long gateway outage (they are certainly unsent, not uncertain) and are reported as deferred, not delivery_failed; WAL via apply_wal_with_fallback; prune only where terminal rows are created.
  • gateway/run.py: housekeeping thread drains the queue per profile scope (gated on the queue file existing, so non-systemd gateways never open it).
  • tools/process_registry.py: restart_safe_gateway_child_argv — fail closed in managed topology when a user scope can't be created; passthrough everywhere else.
  • hermes_cli/kanban_db.py: Kanban worker spawn uses the same scope wrapper and build_subprocess_env.
  • tools/cronjob_tools.py: cronjob(run) reports the exact execution outcome instead of stale last_status.
  • _deliver_result queues only when the worker marker matches the delivering job's own execution_id (a nested in-process dispatch inside a worker must not be keyed under the outer attempt).

Validation

Before After
systemd gateway restarted mid cron run agent turn killed, false interrupted worker survives, side effect once, delivered once (E2E: parent SIGKILLed while worker running)
gateway down > worker wait budget (300s) delivery marked failed, never sent row stays pending, next gateway drains it
parent ledger polling during a run 20 sqlite opens/s for the whole job ~1 read/s, wakeup on process exit
macOS / Windows / launchd / Docker gateways — unchanged: wrapper returns before any I/O; cron run CLI untouched
tests/cron + housekeeping + kanban handoff + cronjob_tools 1228 passed; 5 test_monitor_kind failures are pre-existing on main (env)
ruff / ty (diff vs main) clean / 0 new diagnostics

Regression guards added: nested-dispatch queue keying (mutation-checked), pending-stays-queued, worker exit-then-recheck; test_lost_fire_claim_stops_stale_delivery had become vacuous under the new ownership CAS and is fixed (mutation-checked).

Related

Closes #101877

OutThisLife and others added 5 commits September 3, 2026 12:30
Follow-up to the salvaged restart-safe worker (NousResearch#101877):

- delivery_queue: a row still `pending` at the worker's wait timeout was
  marked `failed` and never drained, so any gateway outage longer than the
  300s budget (e.g. a restart that runs `hermes update`) silently lost the
  delivery. Unclaimed rows are certainly unsent, not uncertain — leave them
  queued for the next gateway; only mid-send rows are fenced `unknown`.
- delivery_queue: stop running the full-table prune UPDATE+COUNT inside
  every transaction (each `get_status` poll paid for it; terminalizing
  paths already prune explicitly); poll at 1s instead of 250ms.
- delivery_queue/executions: use `hermes_state.apply_wal_with_fallback`
  (bare `journal_mode=WAL` raises on NFS/SMB homes) and the race-safe
  `hermes_cli.sqlite_util.add_column_if_missing`; drop the copied
  owner-liveness helpers in favour of the ones in cron.executions.
- scheduler: the parent waited on the worker by re-opening the executions
  ledger every 50ms for the whole run (~20 opens/s, hours). Wait on the
  process with a 1s timeout instead — the worker commits its terminal row
  before exiting — and reap stranded payload/ack files once terminal.
- scheduler: skip the housekeeping drain until a worker has actually
  created deliveries.db, so non-systemd gateways never open it.
- scheduler: set up hermes logging in the detached worker entrypoint; it
  runs with stdout/stderr on DEVNULL and previously logged nowhere.
- tests: test_lost_fire_claim_stops_stale_delivery still mocked
  `mark_execution_running -> None`, which now means "ownership lost, return
  before run_job" — the test passed without ever reaching the path it
  names. Mocking `{}` restores it (mutation-checked).
…end is not a failure

Review fold-in on the salvage of NousResearch#101877:

- `_deliver_result` routed to the durable queue whenever the worker's
  `_HERMES_CRON_EXTERNAL_WORKER` marker was set, regardless of WHICH job was
  delivering. A worker whose script dispatches another job in-process
  (`hermes cron run <other>`) inherits that env and would have queued the
  nested job's message under the outer execution id — `INSERT OR IGNORE`
  then drops it silently. Match the marker against the delivering job's own
  `execution_id`, as `run_one_job` already does. Regression test added
  (mutation-checked: fails with the guard removed).
- A `pending` row left queued at the worker's wait timeout was still
  reported as a delivery error, so `mark_job_run` recorded
  `last_status=delivery_failed` for a message the next gateway's drain goes
  on to send, and nothing ever corrects the job record. Log and return
  success instead; the deliveries row is the authority for the send.
- Reuse `cron.executions._TERMINAL_STATES` in the parent wait loop instead
  of a second hardcoded terminal set.
@kshitijk4poor
kshitijk4poor enabled auto-merge (rebase) September 3, 2026 07:03
@kshitijk4poor
kshitijk4poor merged commit c3e9b28 into NousResearch:main Sep 3, 2026
37 checks passed
@alt-glitch alt-glitch added type/bug Something isn't working comp/cron Cron scheduler and job management comp/gateway Gateway runner, session dispatch, delivery P1 High — major feature broken, no workaround sweeper:risk-session-state Sweeper risk: may lose/corrupt/mis-associate session or context state sweeper:risk-message-delivery Sweeper risk: may drop, duplicate, misroute, or suppress messages labels Sep 3, 2026
@jayleaton

Copy link
Copy Markdown

Live managed-restart follow-up: the original exact head's process isolation worked, but the end-to-end acceptance did not.

At 16:35 local time I restarted the sole default-profile managed gateway while cron execution eb4ef99dccd440deb5e2e88140cea024 (job c8e80ee31ad4) was active:

  • gateway PID changed 908682 → 1059241;
  • cron worker PID 1021707 stayed alive and running in hermes-worker-cron-c8e80ee31ad4-exec-eb4ef99dccd440deb5e2e88140cea024.scope;
  • exactly one default gateway remained active;
  • the replacement gateway delivered the successful result once (for_failure=0, one delivery row, status=delivered, owner PID 1059241);
  • there was one execution row and one same-job execution — no duplicate fire or delivery.

However, immediately after delivery the worker terminalized that same execution as:

failed: Interrupted by shutdown before terminal completion.

The job likewise recorded last_status=error. So the scoped worker and successful delivery survived, but durable terminal ownership still produced a false shutdown failure.

The existing live test passes but misses this topology: test_managed_gateway_restart_preserves_active_worker_and_single_side_effect terminates a harness parent and manually drains through an in-process replacement loop; it does not run the real gateway's 30-second graceful drain and replacement ticker. On current origin/main, the focused file remains green (18 passed, 1 skipped), confirming a coverage gap rather than an existing red test.

This live run was on the original exact head 74d90c3198; afterward I deployed current origin/main (80fae22bf5, containing merge c3e9b28a42) through the same worktree and verified one active default gateway at that SHA. I have not repeated the production restart. QA follow-up is tracked as Kanban t_8d12dfaf to isolate whether the external worker's fire-claim heartbeat sets lost_ownership during the graceful restart/replacement window or another path rewrites the claim. No second PR was opened.

kiwipaulrob added a commit to kiwipaulrob/hermes-agent that referenced this pull request Sep 6, 2026
A systemd-supervised gateway (INVOCATION_ID set) with no user D-Bus
session — containers, minimal LXCs, macOS-style supervisors — fails
EVERY scheduled job at dispatch: restart_safe_gateway_child_argv()
raises and run_one_job() records a failure. The only symptom is
silently skipped executions (a missed nightly backup, dead watchdogs).

Degrade to a direct external subprocess with a warning instead of
raising, unless HERMES_GATEWAY_CHILD_REQUIRE_SCOPE=1 re-enables
fail-closed. Degraded jobs keep process separation and the full
NousResearch#101940 ownership handoff; only cgroup isolation is lost.

The dispatch is a GatewayChildDispatch (in_process / scoped /
degraded) so the degraded case can never collapse into the
'not managed, stay in-process' sentinel — the exact failure mode
that would recreate the restart-interruption edge NousResearch#101940 closed.

Also drops OOMPolicy=kill from scope argv: it is a service-unit
property that transient scopes reject ('Unknown assignment'), so the
probe always failed where scopes actually work (NousResearch#102357, by gkd2323c,
incorporated here with credit).

Co-authored-by: gkd2323c (probe fix from NousResearch#102357)
kiwipaulrob added a commit to kiwipaulrob/hermes-agent that referenced this pull request Sep 8, 2026
A systemd-supervised gateway (INVOCATION_ID set) with no user D-Bus
session (containers, minimal LXCs, supervisors without linger) fails
EVERY scheduled job at dispatch: restart_safe_gateway_child_argv()
raises, run_one_job() records a failure, and the only symptom is
silently skipped executions (a missed nightly backup, dead watchdogs,
no alert).

Cron now degrades to a direct external subprocess with a
once-per-process warning instead of raising, unless
cron.require_restart_safe_scope=true (config.yaml, default false)
restores fail-closed. Degraded jobs keep process separation and the
full NousResearch#101940 ownership handoff - only cgroup isolation is lost, so a
mid-job gateway restart kills the worker and the execution ledger
records exactly that.

The dispatch is a GatewayChildDispatch NamedTuple (in_process /
scoped / degraded) so the degraded case can never collapse into the
"not managed, stay in-process" sentinel - the failure mode that would
recreate the restart-interruption edge NousResearch#101940 closed.

Kanban stays fail-closed (require_restart_safe_scope=True at its call
sites): its workers are long-lived agentic runs, so the degrade policy
is limited to bounded cron jobs in this PR.

Addresses the NousResearch#102431 review: the env-var flag became a config key per
AGENTS.md (no new HERMES_* non-secret vars), Kanban keeps fail-closed
instead of updating its tests to a degraded contract, main's
enable-linger remedy message is preserved, and the degrade warning
fires once per process.
kiwipaulrob added a commit to kiwipaulrob/hermes-agent that referenced this pull request Sep 11, 2026
A systemd-supervised gateway (INVOCATION_ID set) with no user D-Bus
session (containers, minimal LXCs, supervisors without linger) fails
EVERY scheduled job at dispatch: restart_safe_gateway_child_argv()
raises, run_one_job() records a failure, and the only symptom is
silently skipped executions (a missed nightly backup, dead watchdogs,
no alert).

Cron now degrades to a direct external subprocess with a
once-per-process warning instead of raising, unless
cron.require_restart_safe_scope=true (config.yaml, default false)
restores fail-closed. Degraded jobs keep process separation and the
full NousResearch#101940 ownership handoff - only cgroup isolation is lost, so a
mid-job gateway restart kills the worker and the execution ledger
records exactly that.

The dispatch is a GatewayChildDispatch NamedTuple (in_process /
scoped / degraded) so the degraded case can never collapse into the
"not managed, stay in-process" sentinel - the failure mode that would
recreate the restart-interruption edge NousResearch#101940 closed.

Kanban stays fail-closed (require_restart_safe_scope=True at its call
sites): its workers are long-lived agentic runs, so the degrade policy
is limited to bounded cron jobs in this PR.

Addresses the NousResearch#102431 review: the env-var flag became a config key per
AGENTS.md (no new HERMES_* non-secret vars), Kanban keeps fail-closed
instead of updating its tests to a degraded contract, main's
enable-linger remedy message is preserved, and the degrade warning
fires once per process.
kiwipaulrob added a commit to kiwipaulrob/hermes-agent that referenced this pull request Sep 11, 2026
A systemd-supervised gateway (INVOCATION_ID set) with no user D-Bus
session (containers, minimal LXCs, supervisors without linger) fails
EVERY scheduled job at dispatch: restart_safe_gateway_child_argv()
raises, run_one_job() records a failure, and the only symptom is
silently skipped executions (a missed nightly backup, dead watchdogs,
no alert).

Cron now degrades to a direct external subprocess with a
once-per-process warning instead of raising, unless
cron.require_restart_safe_scope=true (config.yaml, default false)
restores fail-closed. Degraded jobs keep process separation and the
full NousResearch#101940 ownership handoff - only cgroup isolation is lost, so a
mid-job gateway restart kills the worker and the execution ledger
records exactly that.

The dispatch is a GatewayChildDispatch NamedTuple (in_process /
scoped / degraded) so the degraded case can never collapse into the
"not managed, stay in-process" sentinel - the failure mode that would
recreate the restart-interruption edge NousResearch#101940 closed.

Kanban stays fail-closed (require_restart_safe_scope=True at its call
sites): its workers are long-lived agentic runs, so the degrade policy
is limited to bounded cron jobs in this PR.

Addresses the NousResearch#102431 review: the env-var flag became a config key per
AGENTS.md (no new HERMES_* non-secret vars), Kanban keeps fail-closed
instead of updating its tests to a degraded contract, main's
enable-linger remedy message is preserved, and the degrade warning
fires once per process.
kshitijk4poor pushed a commit to kshitijk4poor/hermes-agent that referenced this pull request Sep 12, 2026
A systemd-supervised gateway (INVOCATION_ID set) with no user D-Bus
session (containers, minimal LXCs, supervisors without linger) fails
EVERY scheduled job at dispatch: restart_safe_gateway_child_argv()
raises, run_one_job() records a failure, and the only symptom is
silently skipped executions (a missed nightly backup, dead watchdogs,
no alert).

Cron now degrades to a direct external subprocess with a
once-per-process warning instead of raising, unless
cron.require_restart_safe_scope=true (config.yaml, default false)
restores fail-closed. Degraded jobs keep process separation and the
full NousResearch#101940 ownership handoff - only cgroup isolation is lost, so a
mid-job gateway restart kills the worker and the execution ledger
records exactly that.

The dispatch is a GatewayChildDispatch NamedTuple (in_process /
scoped / degraded) so the degraded case can never collapse into the
"not managed, stay in-process" sentinel - the failure mode that would
recreate the restart-interruption edge NousResearch#101940 closed.

Kanban stays fail-closed (require_restart_safe_scope=True at its call
sites): its workers are long-lived agentic runs, so the degrade policy
is limited to bounded cron jobs in this PR.

Addresses the NousResearch#102431 review: the env-var flag became a config key per
AGENTS.md (no new HERMES_* non-secret vars), Kanban keeps fail-closed
instead of updating its tests to a degraded contract, main's
enable-linger remedy message is preserved, and the degrade warning
fires once per process.
kshitijk4poor pushed a commit that referenced this pull request Sep 12, 2026
A systemd-supervised gateway (INVOCATION_ID set) with no user D-Bus
session (containers, minimal LXCs, supervisors without linger) fails
EVERY scheduled job at dispatch: restart_safe_gateway_child_argv()
raises, run_one_job() records a failure, and the only symptom is
silently skipped executions (a missed nightly backup, dead watchdogs,
no alert).

Cron now degrades to a direct external subprocess with a
once-per-process warning instead of raising, unless
cron.require_restart_safe_scope=true (config.yaml, default false)
restores fail-closed. Degraded jobs keep process separation and the
full #101940 ownership handoff - only cgroup isolation is lost, so a
mid-job gateway restart kills the worker and the execution ledger
records exactly that.

The dispatch is a GatewayChildDispatch NamedTuple (in_process /
scoped / degraded) so the degraded case can never collapse into the
"not managed, stay in-process" sentinel - the failure mode that would
recreate the restart-interruption edge #101940 closed.

Kanban stays fail-closed (require_restart_safe_scope=True at its call
sites): its workers are long-lived agentic runs, so the degrade policy
is limited to bounded cron jobs in this PR.

Addresses the #102431 review: the env-var flag became a config key per
AGENTS.md (no new HERMES_* non-secret vars), Kanban keeps fail-closed
instead of updating its tests to a degraded contract, main's
enable-linger remedy message is preserved, and the degrade warning
fires once per process.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/cron Cron scheduler and job management comp/gateway Gateway runner, session dispatch, delivery P1 High — major feature broken, no workaround sweeper:risk-message-delivery Sweeper risk: may drop, duplicate, misroute, or suppress messages sweeper:risk-session-state Sweeper risk: may lose/corrupt/mis-associate session or context state type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants