Skip to content

fix(cron): renew fire claims without contending with delivery fence - #104254

Closed
jerejones wants to merge 2 commits into
NousResearch:mainfrom
jerejones:fix/cron-heartbeat-delivery-fence
Closed

jerejones wants to merge 2 commits into
NousResearch:mainfrom
jerejones:fix/cron-heartbeat-delivery-fence

Conversation

@jerejones

Copy link
Copy Markdown

Problem

Cron delivery holds the per-job fire fence across network I/O. The concurrent claim heartbeat attempts that same fence and fails closed after 30 seconds of contention. The scheduler mistakes this for ownership loss, marking successfully executed/delivered runs interrupted.

Production evidence: output saved 06:32:03, local fence timeout 06:32:49, confirmed Discord delivery 06:32:53, then an interrupted execution record. The actual gateway restart was later, at 06:40:48.

Fix

  • Renew the current owner's lease under the jobs-file lock without taking its delivery fence.
  • Retain fencing for owner changes, terminal writes, and external side effects.
  • Require an acquired cross-process jobs lock for renewal; do not renew from a degraded, unsynchronized snapshot. Lock errors use the existing heartbeat transient-error grace.
  • Cover delivery contention, actual ownership replacement, and strict-lock failure paths.

Verification

  • Independent real-thread probe, temporary HERMES_HOME, unmodified 30-second timeout: original code returned false after 30 seconds; patched code renewed immediately and rejected a stale owner.
  • Final patch: scripts/run_tests.sh tests/cron/ -j 8 --file-retries 0: 1,227 passed, 0 failed, 2 skipped across 96 files. Independent contention probe again passed with the strict-lock follow-up.
  • Installed original fix and gracefully restarted gateway; gateway_state.json confirmed deployed SHA. A real Gmail cron run completed, produced triage output, and recorded no execution or delivery error.
  • Full repository suite was exercised, but is not green: 89 failures in 20 files outside cron. Two failures (test_hermes_state.py and test_tui_gateway_server.py) were independently reproduced on unchanged base 089bb32. Remaining broad-suite failures are not yet baseline-classified; no claim that the full suite passes.

@alt-glitch alt-glitch added type/bug Something isn't working P2 Medium — degraded but workaround exists comp/cron Cron scheduler and job management sweeper:risk-message-delivery Sweeper risk: may drop, duplicate, misroute, or suppress messages labels Sep 6, 2026
@alt-glitch

Copy link
Copy Markdown

This was generated by AI during triage.

Related: #100418 and #102627 are competing fixes for the same #100401 fence-contention bug. This PR renews the lease without taking the delivery fence; #100418 treats a busy fence as non-loss; #102627 tri-states the heartbeat result. Linking so reviewers can compare approaches.

@jerejones jerejones closed this Sep 6, 2026
@jerejones
jerejones deleted the fix/cron-heartbeat-delivery-fence branch September 6, 2026 12:18
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/cron Cron scheduler and job management P2 Medium — degraded but workaround exists sweeper:risk-message-delivery Sweeper risk: may drop, duplicate, misroute, or suppress messages type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants