Repository navigation
fix(cron): preserve active turns across gateway restarts (salvage #101877) - #101940
kshitijk4poor merged 5 commits into
Conversation
Follow-up to the salvaged restart-safe worker (NousResearch#101877): - delivery_queue: a row still `pending` at the worker's wait timeout was marked `failed` and never drained, so any gateway outage longer than the 300s budget (e.g. a restart that runs `hermes update`) silently lost the delivery. Unclaimed rows are certainly unsent, not uncertain — leave them queued for the next gateway; only mid-send rows are fenced `unknown`. - delivery_queue: stop running the full-table prune UPDATE+COUNT inside every transaction (each `get_status` poll paid for it; terminalizing paths already prune explicitly); poll at 1s instead of 250ms. - delivery_queue/executions: use `hermes_state.apply_wal_with_fallback` (bare `journal_mode=WAL` raises on NFS/SMB homes) and the race-safe `hermes_cli.sqlite_util.add_column_if_missing`; drop the copied owner-liveness helpers in favour of the ones in cron.executions. - scheduler: the parent waited on the worker by re-opening the executions ledger every 50ms for the whole run (~20 opens/s, hours). Wait on the process with a 1s timeout instead — the worker commits its terminal row before exiting — and reap stranded payload/ack files once terminal. - scheduler: skip the housekeeping drain until a worker has actually created deliveries.db, so non-systemd gateways never open it. - scheduler: set up hermes logging in the detached worker entrypoint; it runs with stdout/stderr on DEVNULL and previously logged nowhere. - tests: test_lost_fire_claim_stops_stale_delivery still mocked `mark_execution_running -> None`, which now means "ownership lost, return before run_job" — the test passed without ever reaching the path it names. Mocking `{}` restores it (mutation-checked).
…end is not a failure Review fold-in on the salvage of NousResearch#101877: - `_deliver_result` routed to the durable queue whenever the worker's `_HERMES_CRON_EXTERNAL_WORKER` marker was set, regardless of WHICH job was delivering. A worker whose script dispatches another job in-process (`hermes cron run <other>`) inherits that env and would have queued the nested job's message under the outer execution id — `INSERT OR IGNORE` then drops it silently. Match the marker against the delivering job's own `execution_id`, as `run_one_job` already does. Regression test added (mutation-checked: fails with the guard removed). - A `pending` row left queued at the worker's wait timeout was still reported as a delivery error, so `mark_job_run` recorded `last_status=delivery_failed` for a message the next gateway's drain goes on to send, and nothing ever corrects the job record. Log and return success instead; the deliveries row is the authority for the send. - Reuse `cron.executions._TERMINAL_STATES` in the parent wait loop instead of a second hardcoded terminal set.
|
Live managed-restart follow-up: the original exact head's process isolation worked, but the end-to-end acceptance did not. At 16:35 local time I restarted the sole default-profile managed gateway while cron execution
However, immediately after delivery the worker terminalized that same execution as:
The job likewise recorded The existing live test passes but misses this topology: This live run was on the original exact head |
A systemd-supervised gateway (INVOCATION_ID set) with no user D-Bus session — containers, minimal LXCs, macOS-style supervisors — fails EVERY scheduled job at dispatch: restart_safe_gateway_child_argv() raises and run_one_job() records a failure. The only symptom is silently skipped executions (a missed nightly backup, dead watchdogs). Degrade to a direct external subprocess with a warning instead of raising, unless HERMES_GATEWAY_CHILD_REQUIRE_SCOPE=1 re-enables fail-closed. Degraded jobs keep process separation and the full NousResearch#101940 ownership handoff; only cgroup isolation is lost. The dispatch is a GatewayChildDispatch (in_process / scoped / degraded) so the degraded case can never collapse into the 'not managed, stay in-process' sentinel — the exact failure mode that would recreate the restart-interruption edge NousResearch#101940 closed. Also drops OOMPolicy=kill from scope argv: it is a service-unit property that transient scopes reject ('Unknown assignment'), so the probe always failed where scopes actually work (NousResearch#102357, by gkd2323c, incorporated here with credit). Co-authored-by: gkd2323c (probe fix from NousResearch#102357)
A systemd-supervised gateway (INVOCATION_ID set) with no user D-Bus session (containers, minimal LXCs, supervisors without linger) fails EVERY scheduled job at dispatch: restart_safe_gateway_child_argv() raises, run_one_job() records a failure, and the only symptom is silently skipped executions (a missed nightly backup, dead watchdogs, no alert). Cron now degrades to a direct external subprocess with a once-per-process warning instead of raising, unless cron.require_restart_safe_scope=true (config.yaml, default false) restores fail-closed. Degraded jobs keep process separation and the full NousResearch#101940 ownership handoff - only cgroup isolation is lost, so a mid-job gateway restart kills the worker and the execution ledger records exactly that. The dispatch is a GatewayChildDispatch NamedTuple (in_process / scoped / degraded) so the degraded case can never collapse into the "not managed, stay in-process" sentinel - the failure mode that would recreate the restart-interruption edge NousResearch#101940 closed. Kanban stays fail-closed (require_restart_safe_scope=True at its call sites): its workers are long-lived agentic runs, so the degrade policy is limited to bounded cron jobs in this PR. Addresses the NousResearch#102431 review: the env-var flag became a config key per AGENTS.md (no new HERMES_* non-secret vars), Kanban keeps fail-closed instead of updating its tests to a degraded contract, main's enable-linger remedy message is preserved, and the degrade warning fires once per process.
A systemd-supervised gateway (INVOCATION_ID set) with no user D-Bus session (containers, minimal LXCs, supervisors without linger) fails EVERY scheduled job at dispatch: restart_safe_gateway_child_argv() raises, run_one_job() records a failure, and the only symptom is silently skipped executions (a missed nightly backup, dead watchdogs, no alert). Cron now degrades to a direct external subprocess with a once-per-process warning instead of raising, unless cron.require_restart_safe_scope=true (config.yaml, default false) restores fail-closed. Degraded jobs keep process separation and the full NousResearch#101940 ownership handoff - only cgroup isolation is lost, so a mid-job gateway restart kills the worker and the execution ledger records exactly that. The dispatch is a GatewayChildDispatch NamedTuple (in_process / scoped / degraded) so the degraded case can never collapse into the "not managed, stay in-process" sentinel - the failure mode that would recreate the restart-interruption edge NousResearch#101940 closed. Kanban stays fail-closed (require_restart_safe_scope=True at its call sites): its workers are long-lived agentic runs, so the degrade policy is limited to bounded cron jobs in this PR. Addresses the NousResearch#102431 review: the env-var flag became a config key per AGENTS.md (no new HERMES_* non-secret vars), Kanban keeps fail-closed instead of updating its tests to a degraded contract, main's enable-linger remedy message is preserved, and the degrade warning fires once per process.
A systemd-supervised gateway (INVOCATION_ID set) with no user D-Bus session (containers, minimal LXCs, supervisors without linger) fails EVERY scheduled job at dispatch: restart_safe_gateway_child_argv() raises, run_one_job() records a failure, and the only symptom is silently skipped executions (a missed nightly backup, dead watchdogs, no alert). Cron now degrades to a direct external subprocess with a once-per-process warning instead of raising, unless cron.require_restart_safe_scope=true (config.yaml, default false) restores fail-closed. Degraded jobs keep process separation and the full NousResearch#101940 ownership handoff - only cgroup isolation is lost, so a mid-job gateway restart kills the worker and the execution ledger records exactly that. The dispatch is a GatewayChildDispatch NamedTuple (in_process / scoped / degraded) so the degraded case can never collapse into the "not managed, stay in-process" sentinel - the failure mode that would recreate the restart-interruption edge NousResearch#101940 closed. Kanban stays fail-closed (require_restart_safe_scope=True at its call sites): its workers are long-lived agentic runs, so the degrade policy is limited to bounded cron jobs in this PR. Addresses the NousResearch#102431 review: the env-var flag became a config key per AGENTS.md (no new HERMES_* non-secret vars), Kanban keeps fail-closed instead of updating its tests to a degraded contract, main's enable-linger remedy message is preserved, and the degrade warning fires once per process.
A systemd-supervised gateway (INVOCATION_ID set) with no user D-Bus session (containers, minimal LXCs, supervisors without linger) fails EVERY scheduled job at dispatch: restart_safe_gateway_child_argv() raises, run_one_job() records a failure, and the only symptom is silently skipped executions (a missed nightly backup, dead watchdogs, no alert). Cron now degrades to a direct external subprocess with a once-per-process warning instead of raising, unless cron.require_restart_safe_scope=true (config.yaml, default false) restores fail-closed. Degraded jobs keep process separation and the full NousResearch#101940 ownership handoff - only cgroup isolation is lost, so a mid-job gateway restart kills the worker and the execution ledger records exactly that. The dispatch is a GatewayChildDispatch NamedTuple (in_process / scoped / degraded) so the degraded case can never collapse into the "not managed, stay in-process" sentinel - the failure mode that would recreate the restart-interruption edge NousResearch#101940 closed. Kanban stays fail-closed (require_restart_safe_scope=True at its call sites): its workers are long-lived agentic runs, so the degrade policy is limited to bounded cron jobs in this PR. Addresses the NousResearch#102431 review: the env-var flag became a config key per AGENTS.md (no new HERMES_* non-secret vars), Kanban keeps fail-closed instead of updating its tests to a degraded contract, main's enable-linger remedy message is preserved, and the degrade warning fires once per process.
A systemd-supervised gateway (INVOCATION_ID set) with no user D-Bus session (containers, minimal LXCs, supervisors without linger) fails EVERY scheduled job at dispatch: restart_safe_gateway_child_argv() raises, run_one_job() records a failure, and the only symptom is silently skipped executions (a missed nightly backup, dead watchdogs, no alert). Cron now degrades to a direct external subprocess with a once-per-process warning instead of raising, unless cron.require_restart_safe_scope=true (config.yaml, default false) restores fail-closed. Degraded jobs keep process separation and the full #101940 ownership handoff - only cgroup isolation is lost, so a mid-job gateway restart kills the worker and the execution ledger records exactly that. The dispatch is a GatewayChildDispatch NamedTuple (in_process / scoped / degraded) so the degraded case can never collapse into the "not managed, stay in-process" sentinel - the failure mode that would recreate the restart-interruption edge #101940 closed. Kanban stays fail-closed (require_restart_safe_scope=True at its call sites): its workers are long-lived agentic runs, so the degrade policy is limited to bounded cron jobs in this PR. Addresses the #102431 review: the env-var flag became a config key per AGENTS.md (no new HERMES_* non-secret vars), Kanban keeps fail-closed instead of updating its tests to a degraded contract, main's enable-linger remedy message is preserved, and the degrade warning fires once per process.
Summary
Cron jobs fired by a systemd-managed Linux gateway now survive
systemctl restart/hermes gateway restart: the fire is handed to a detachedsystemd-run --scopeworker that owns the execution durably, and its final delivery is queued for whichever gateway instance is live.Salvage of #101877 by @jayleaton — the three commits are authored by Brooklyn Nicholson (@OutThisLife) and cherry-picked with authorship preserved. Two follow-up commits fold in the review findings.
Root cause: #60711 made the shutdown drain see in-flight cron work, but a forced restart still kills the gateway cgroup, and the cron agent turn with it.
start_new_sessionalone does not escape a systemd service cgroup.Changes
cron/scheduler.py:run_one_jobhands a managed-gateway fire topython -m cron.scheduler --external-worker-file …viarestart_safe_gateway_child_argv; parent waits on the worker process (1swait(timeout)), never re-runs the job in-process on an uncertain handoff; worker adopts the execution row (CASclaimed+handoff_pending→running) before any side effect; restart-safe waiters are excluded from forced-shutdown interruption; stranded payload/ack files reaped once terminal; worker entrypoint sets up hermes logging.cron/executions.py:handoff_pending/handoff_started_atcolumns (viasqlite_util.add_column_if_missing),mark_execution_handoff_pending,adopt_claimed_execution, owner-fencedrunning/finish/recovery with a 30s adoption grace.cron/delivery_queue.py(new): profile-local sqlite queue; gateway claims a row before transport I/O, dead claimant →unknownnever replayed; unclaimed rows stay queued across a long gateway outage (they are certainly unsent, not uncertain) and are reported as deferred, notdelivery_failed; WAL viaapply_wal_with_fallback; prune only where terminal rows are created.gateway/run.py: housekeeping thread drains the queue per profile scope (gated on the queue file existing, so non-systemd gateways never open it).tools/process_registry.py:restart_safe_gateway_child_argv— fail closed in managed topology when a user scope can't be created; passthrough everywhere else.hermes_cli/kanban_db.py: Kanban worker spawn uses the same scope wrapper andbuild_subprocess_env.tools/cronjob_tools.py:cronjob(run)reports the exact execution outcome instead of stalelast_status._deliver_resultqueues only when the worker marker matches the delivering job's ownexecution_id(a nested in-process dispatch inside a worker must not be keyed under the outer attempt).Validation
interruptedrunning)failed, never sentpending, next gateway drains itcron runCLI untouchedtests/cron+ housekeeping + kanban handoff + cronjob_toolstest_monitor_kindfailures are pre-existing onmain(env)Regression guards added: nested-dispatch queue keying (mutation-checked), pending-stays-queued, worker exit-then-recheck;
test_lost_fire_claim_stops_stale_deliveryhad become vacuous under the new ownership CAS and is fixed (mutation-checked).Related
process_registry/scheduler/run.py. This PR is user-scope-only and fails closed when a user scope is unavailable; if feat(gateway): isolate workers and queue under resource pressure #97739 lands first,restart_safe_gateway_child_argvshould consume its backend-neutral seam.Closes #101877