Repository navigation
fix(cron): stop an adopted external worker from re-exec'ing and losing its payload (#124827) - #124855
PRATHAMESH75 wants to merge 1 commit into
Conversation
…g its payload (NousResearch#124827) An external cron worker imports the agent lazily inside run_one_job, which pulls in hermes_bootstrap; its module-level prepare_launch() can decide the process must re-exec (a pending self-update or a different store interpreter) and os.execv() itself. In an already-adopted worker that is fatal: the one-shot payload named by --external-worker-file was consumed and deleted before adoption, so the re-exec'd process starts with a dangling path, cannot resume, and the gateway records the run 'unknown' with no output or delivery while the next occurrence advances normally. The gateway, not a job worker, owns updates. Spawn the external worker with HERMES_DISABLE_LAZY_INSTALLS=1 so prepare_launch() returns None and never re-execs; dependency activation still runs, only the lazy update/re-exec path is suppressed. Add a regression test asserting the spawned worker env carries the flag.
Setting One correction to the comment, since the flag is load-bearing beyond the re-exec:
Minor: Unverified / please confirm: I could not run the repo's pytest, so the above comes from executing the real |
|
Fixed on main by de116d8 (#134358), salvaged from #133720 (thanks @mzyas, authorship kept). The external cron worker now runs Thanks @PRATHAMESH75 for this alternative. Closing in favour of #134358. If you still see #124827 on a build that includes de116d8, please comment and we'll reopen. |
What
A scheduled cron job can vanish after the restart-safe external worker acknowledges durable ownership: the execution is recorded
unknown(notfailed), with no output or delivery, and the next occurrence advances normally.Fixes #124827.
Root cause
The external worker imports the agent lazily inside
run_one_job, which pulls inhermes_bootstrap. That module's top-levelprepare_launch()(hermes_cli/venv_sync.py) can decide the process must restart — a pending self-update, or a different store interpreter — andos.execv()itself. Inside an already-adopted worker this is fatal:_run_external_worker_payloadconsumed and deleted the one-shot--external-worker-filepayload before adoption, so the re-exec'd process starts with a dangling path, cannot resume, and the gateway — which already received the acknowledgement — records an indeterminateunknownrun with no automatic retry.This is distinct from #122222 (missing deps before acknowledgement, reported
failed) and #120328 (worker death after acknowledgement).Fix
The gateway, not a job worker, owns updates. Spawn the external worker with
HERMES_DISABLE_LAZY_INSTALLS=1in its environment.prepare_launch()already honors this flag and returnsNone(no update completion, no re-exec), so the adopted worker can never lose its handoff to anos.execv(). Dependency activation still runs — only the lazy update/re-exec path is suppressed, which is exactly what an unattended worker should never perform.This is a one-line env addition alongside the existing worker-env hardening (presence-var scrubbing, PYTHONPATH pinning) in
_launch_external_cron_worker, matching the mechanism the boot path already exposes.Tests
tests/cron/test_restart_safe_worker.py::test_launch_external_worker_disables_lazy_reexec— the spawned worker env carriesHERMES_DISABLE_LAZY_INSTALLS=1.While in this file the Windows-footgun gate flagged four pre-existing lines (BOM-intolerant
read_text(encoding="utf-8"), a bareread_text(), and a baresignal.SIGKILL); fixed them per the repo policy (utf-8-sigreads,getattr(signal, "SIGKILL", signal.SIGTERM)) so the touched file passes the gate.The affected cron suite passes.