Skip to content

fix(cron): don't treat a busy fire fence as lost claim ownership - #100418

Closed
oheckmann74 wants to merge 2 commits into
NousResearch:mainfrom
oheckmann74:fix/cron-heartbeat-fence-deadlock
Closed

oheckmann74 wants to merge 2 commits into
NousResearch:mainfrom
oheckmann74:fix/cron-heartbeat-fence-deadlock

Conversation

@oheckmann74

Copy link
Copy Markdown

Fixes #100401.

Problem

heartbeat_fire_claim() takes the per-job fire fence. The run thread already holds that fence around its side effects — save_job_output() and _deliver_result() (_side_effect_fence()). The heartbeat runs on a different thread, so _fire_job_lock's thread-local reentrancy bookkeeping (_fire_fence_lock_state = threading.local()) does not apply and the RLock is genuinely contended.

So a delivery slower than _JOBS_LOCK_TIMEOUT_SECONDS (30s) makes the heartbeat block, time out, and return False. The scheduler cannot tell that apart from a real takeover:

if not heartbeat_fire_claim(job_id, expected_owner=owner):
    lost_ownership.set()
    logger.warning("Job '%s': fire claim ownership lost; interrupting stale run", job_id)

and interrupts a perfectly healthy run, which operators see as:

failed     Interrupted by shutdown before terminal completion.

Nothing has shut down, and the work itself already succeeded — in my case a nightly backup that had committed and pushed before being reported as failed.

Production timeline (deliver: bot-chat:default), the two constants adding to exactly 90s:

03:45:00.49  run starts
03:46:30.55  Timed out waiting for local fire fence <dir>::<job_id>; failing closed   (t+90s)
             Job '<job_id>': fire claim ownership lost; interrupting stale run
03:47:33.32  Job '<job_id>': delivered to Bot Chat of profile 'default'               (t+153s)

Delivery was still in flight, holding the fence, when the heartbeat tried to renew.

Fix

A heartbeat is not an external side effect. It only compare-and-swaps fire_claim["at"] when fire_claim["by"] still equals expected_owner, and _heartbeat_fire_claim_locked() already performs that CAS under _jobs_lock() — which holds both an in-process lock and a cross-process flock on .jobs.lock. The fence adds nothing to its correctness.

So: prefer the fence to keep the common path serialized with other owner mutations, but when it is unavailable, fall back to the CAS instead of reporting a loss of ownership that was never observed. A genuine takeover still returns False, because by no longer matches.

No change to the scheduler, no change to the constants, no new public API.

Scope

Only jobs whose delivery is slow enough to span a heartbeat tick are affected. --deliver local never reproduces it — I confirmed with a 115s job that completes on both patched and unpatched builds. Every job I have seen fail delivers to bot-chat, which may make this worth a look alongside #95307 (also bot-chat, also a spurious lost fire claim).

Testing

  • New regression test reproduces the production log line and fails without the fix:
    tests/cron/test_claim_job_for_fire.py::test_heartbeat_survives_fence_held_by_its_own_run
  • It also asserts that a real takeover is still detected while the fence is busy, so the fix does not weaken the ownership guarantee.
  • Full tests/cron/ suite: 1081 passed, 1 skipped.
  • One pre-existing unrelated failure, test_repeated_heartbeat_errors_cancel_after_bounded_grace, fails identically with and without this change (5/5 runs each) — it asserts calls >= 3 against a 0.01s/0.03s timing window.

Verified against main @ 18a76be12; both touched files are byte-identical between that revision and current main.

heartbeat_fire_claim() took the per-job fire fence, which the run thread
already holds across its side effects (save_job_output, _deliver_result).
The heartbeat runs on a different thread, so the fence's thread-local
reentrancy does not apply: a delivery slower than _JOBS_LOCK_TIMEOUT_SECONDS
made the heartbeat time out and return False, which the scheduler reads as
lost ownership and uses to interrupt a healthy run -- surfaced to operators
as "Interrupted by shutdown before terminal completion."

A heartbeat is not an external side effect. It only compare-and-swaps
fire_claim["at"] when fire_claim["by"] still matches expected_owner, and
_heartbeat_fire_claim_locked already runs that CAS under _jobs_lock(), which
holds both an in-process lock and a cross-process flock. Prefer the fence,
but fall back to the CAS rather than reporting a loss never observed. A real
takeover still returns False, because "by" no longer matches.
Fails without the fix with the production symptom:
  Timed out waiting for local fire fence <dir>::<job_id>; failing closed
@alt-glitch alt-glitch added type/bug Something isn't working comp/cron Cron scheduler and job management P2 Medium — degraded but workaround exists sweeper:risk-message-delivery Sweeper risk: may drop, duplicate, misroute, or suppress messages labels Sep 1, 2026
@alt-glitch

Copy link
Copy Markdown

This was generated by AI during triage.

Related: #95432 and #97565 address the same false ownership-loss family using tri-state heartbeat handling. This PR instead retries the authoritative locked compare-and-swap after fence contention; maintainer choice is needed.

@kshitijk4poor

Copy link
Copy Markdown

The fix for this bug landed on main via #109310 (merge 9a60a7f): heartbeat_fire_claim now refreshes under _jobs_lock only and no longer waits on the per-thread fire fence its own run holds, so a delivery/agent turn longer than the 30s fence timeout is no longer misread as ownership loss. _refresh_claim still CASes on claim["by"], so a real takeover keeps returning False (pinned by test).

Thanks @oheckmann74 — earliest working fix in the cluster with a red-on-base test. The landed change is the fence-free variant of your fallback (skips the fence rather than trying it first), and the merged regression test follows your test's shape (fence held by a worker thread, heartbeat True on the calling thread, foreign owner still False); you are credited as Co-author on that commit. Closing with credit.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/cron Cron scheduler and job management P2 Medium — degraded but workaround exists sweeper:risk-message-delivery Sweeper risk: may drop, duplicate, misroute, or suppress messages type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: cron — fire-claim heartbeat deadlocks on its own run's fence, killing every job that runs >60s as "Interrupted by shutdown"

3 participants