Skip to content

fix(cron): skip stale-error re-arm while another process holds a live fire_claim - #103740

Closed
r3x443 wants to merge 1 commit into
NousResearch:mainfrom
r3x443:fix/cron-stale-error-rearm-live-claim
Closed

r3x443 wants to merge 1 commit into
NousResearch:mainfrom
r3x443:fix/cron-stale-error-rearm-live-claim

Conversation

@r3x443

@r3x443 r3x443 commented Sep 5, 2026

Copy link
Copy Markdown

Symptom

With several scheduler processes sharing one jobs store (messaging gateway + Desktop profile tabs is enough), a recurring job that persisted an error state gets re-armed by every OTHER ticker while it is actively running in one of them. Live incident on our deployment: 171 spurious re-arms, ~175 junk "Fire claim lost" execution rows, and the live run was ultimately killed.

Where it manifests

cron/jobs.py::_job_is_stale_error_recurring checks _job_running_in_this_process(...) — in-process liveness only. A run owned by another process heartbeats fire_claim (cadence from claim_job_fire), but this predicate never consults it, so from any sibling process the job looks wedged and gets re-armed.

Fix

Treat a fresh fire_claim as cross-process liveness, using the existing _claim_is_live() helper that the due-scan and dispatch paths already use for exactly this purpose (4 existing call sites). TTL matches the 300s claim TTL used at the sibling claim_job_fire site.

Tests

Invariant test added: a job with a live foreign fire_claim is not classified stale-error; an expired claim still is. Related reports: #100946, #97565, #102174 describe neighboring fire-claim ownership confusion — this PR fixes the stale-error re-arm lane specifically.

… fire_claim

Multi-process schedulers sharing one jobs store (gateway + Desktop serve tabs)
re-armed a job every tick while a long run in another process was still
heartbeating its fire_claim, producing claim-fight churn and killing the live
run (brain, 2026-09-02). Treat a fresh fire_claim as 'running elsewhere'.
@alt-glitch alt-glitch added type/bug Something isn't working P1 High — major feature broken, no workaround comp/cron Cron scheduler and job management labels Sep 5, 2026
@Enough1122

Copy link
Copy Markdown

AI code review — automated review for reference; please use your judgment.

Summary: Stale-error recovery no longer re-arms a job whose fire_claim is fresh (<300s, matching claim TTL vs 60s heartbeat), fixing cross-process claim fights where one ticker's re-arms killed another process's live run (171 re-arms in the cited incident). Expired claims resume recovery normally.

Findings (Non-blocking):

  • cron/jobs.py:26 — liveness is inferred from fire_claim.at timestamps written by other hosts. Cross-host clock skew could stretch or shrink the 300s window; with a 5x margin over the 60s heartbeat, skew would need to exceed ~4 minutes to cause a false re-arm. Acceptable; just noting the implicit NTP assumption.
  • The test pins both directions (live claim skipped, expired claim recovered). Good.

Verdict: Non-blocking. Correct incident-driven fix; no issues.

@kshitijk4poor

Copy link
Copy Markdown

Merged via #104518 as a cherry-pick of your commit (authorship preserved) — main 193f05dec5.

What landed: your fresh-fire_claim liveness check in _job_is_stale_error_recurring via the existing _claim_is_live(). One change of mine: the 300s TTL was three separate literals (claim_job_for_fire default, rearm_oneshot, your guard), now one FIRE_CLAIM_TTL_SECONDS; the incident narrative in the comment was cut to the WHY. Your invariant test is the one that goes red on main.

Thanks — closing in favour of the merged salvage.

Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

comp/cron Cron scheduler and job management P1 High — major feature broken, no workaround type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants