Conversation
… fire_claim Multi-process schedulers sharing one jobs store (gateway + Desktop serve tabs) re-armed a job every tick while a long run in another process was still heartbeating its fire_claim, producing claim-fight churn and killing the live run (brain, 2026-09-02). Treat a fresh fire_claim as 'running elsewhere'.
Summary: Stale-error recovery no longer re-arms a job whose Findings (Non-blocking):
Verdict: Non-blocking. Correct incident-driven fix; no issues. |
|
Merged via #104518 as a cherry-pick of your commit (authorship preserved) — main What landed: your fresh- Thanks — closing in favour of the merged salvage. |
Symptom
With several scheduler processes sharing one jobs store (messaging gateway + Desktop profile tabs is enough), a recurring job that persisted an error state gets re-armed by every OTHER ticker while it is actively running in one of them. Live incident on our deployment: 171 spurious re-arms, ~175 junk "Fire claim lost" execution rows, and the live run was ultimately killed.
Where it manifests
cron/jobs.py::_job_is_stale_error_recurringchecks_job_running_in_this_process(...)— in-process liveness only. A run owned by another process heartbeatsfire_claim(cadence fromclaim_job_fire), but this predicate never consults it, so from any sibling process the job looks wedged and gets re-armed.Fix
Treat a fresh
fire_claimas cross-process liveness, using the existing_claim_is_live()helper that the due-scan and dispatch paths already use for exactly this purpose (4 existing call sites). TTL matches the 300s claim TTL used at the siblingclaim_job_firesite.Tests
Invariant test added: a job with a live foreign
fire_claimis not classified stale-error; an expired claim still is. Related reports: #100946, #97565, #102174 describe neighboring fire-claim ownership confusion — this PR fixes the stale-error re-arm lane specifically.