Repository navigation
Conversation
…wnership loss (c-027) heartbeat_fire_claim() collapsed fire-fence acquisition failure into False. During slow fenced delivery the worker's own heartbeat thread cannot acquire the same process-local RLock, so a successfully delivering run was misclassified as ownership loss and terminalled as 'Interrupted by shutdown before terminal completion.' (TrustMRR exec 4219b48d while RSI recovery exec f502e929 was settling). Tri-state contract: - True = renewed/confirmed owner - False = authoritatively inspected: owner mismatch / absent claim - None = fence unavailable/unconfirmed (contention or store error) - initial validation stays fail-closed: run starts only on exactly True - heartbeat loop: None rides the existing grace window (180s prod vs 30s fence wait); immediate cancel only on explicit False - post-run probes (_fire_claim_ownership_lost, both interrupted blocks, terminal owner-CAS read): None can never adjudicate a confirmed loss; uncertain outcomes get distinct ledger errors - mark_job_run also tri-state: None = fence unavailable (CAS never ran), False stays reserved for confirmed owner mismatch - _side_effect_fence unchanged: still the exactly-once save/delivery barrier Tests (all written red-first, verified failing on unfixed code): slow_owned_delivery (parent handoff), sustained fence-contention grace, pre/post-delivery None probes, grace-exhausted None probe, terminal write None, mark_job_run fence-unavailable None, RSI recovery dispatch fails closed during original's fenced delivery, TrustMRR slow-owned delivery terminal success. Managed-gateway restart E2E (systemd scope) passes on host: one gateway, single side effect, single delivery.
Supersedes your own open #100965 (same fence-timeout-vs-ownership-loss fix, now post-#101940). Should #100965 be closed in favor of this one? Also competing with the open tri-state repairs #95432, #97565 and #100418 for the same bug; a maintainer needs to pick one mechanism. |
andrexibiza
left a comment
There was a problem hiding this comment.
Reviewed exact head 9b26dcb9fb1f5425b198495b655d7ef54304da3f against current main / merge-base 63279301bcbdc185c1b07b98a9312eb0c862f26d (1 commit, 4 changed files, branch not behind main). I traced the fire-claim contract through cron/jobs.py, the heartbeat wrapper, both post-run ownership probes, both side-effect fences, terminal mark_job_run, the new regression coverage, and the overlapping open implementations.
The core semantic correction is right: False must mean an owner mismatch that was actually observed, while fence acquisition failure is None / unconfirmed. The initial is True gate and the terminal marked is False vs marked is None split are both materially better than main’s current conflation.
There is still one merge-blocking hole in the exact defect class, though.
Blocker: self-held fence contention can still age through the grace window and demote a successfully delivered run
In _run_with_fire_claim_heartbeat(), renewed is None feeds the same _FIRE_CLAIM_HEARTBEAT_GRACE_SECONDS clock as a store/heartbeat failure. Once that clock expires, the heartbeat thread sets lost_ownership even though the PR itself identifies the common None case as this execution’s own _side_effect_fence being held across delivery.
That event is sticky. _run_one_job_body._fire_claim_ownership_lost() checks fire_claim_lost.is_set() first and immediately returns True, before a fresh owner inspection. After the slow delivery finally releases _side_effect_fence, the post-delivery interruption branch probes heartbeat_fire_claim() again; because the claim was never stolen, that probe is now True, and the branch records:
Interrupted by shutdown before terminal completion.
via mark_job_run(..., False, expected_fire_owner=fire_owner).
So the current patch fixes “own fence busy for less than grace” but still reproduces the same false terminal demotion when a legitimate delivery itself holds the fence longer than grace. No takeover is required.
The new tests leave exactly this boundary uncovered:
test_slow_owned_delivery_does_not_false_interrupt_completed_runshortens the heartbeat/fence timeout but leaves the production grace far above the delivery duration.test_heartbeat_lock_contention_is_unconfirmed_not_ownership_lossexplicitly keeps contention “well under the grace window”.test_interrupted_path_none_probe_does_not_claim_confirmed_lossdoes cross the grace window, but only asserts that the ledger wording is not “ownership lost”; it accepts cancellation of an otherwise still-owned run.
This is also a supersession regression relative to prior work. #100965 — from the same contributor and called out by triage as superseded by this PR — already carried test_successful_long_delivery_is_not_demoted_after_heartbeat_grace plus a heartbeat_uncertain/post-delivery recovery path specifically for this case. That test holds the real delivery fence beyond a shortened grace window and requires last_status == "ok" and the execution ledger to remain completed. The protection did not make it into this successor.
There are two other prior mechanisms worth preserving as design evidence rather than silently discarding:
- #95432 (BrunoBza) introduced the same tri-state shape but deliberately did not burn exception grace on
Nonefence-busy beats, because the busy fence is itself excluding replacement owners while the legitimate side effect runs. - #97565 (sycamoregroupltd) uses the same “
Noneconsumes grace” policy as this PR, so it has the same residual boundary. - #100418 (oheckmann74) takes the alternate route of falling back to the owner CAS under
_jobs_lockwhen the fire fence is busy, keeping true owner replacement detectable without treating fence contention as loss.
I am not prescribing which of those mechanisms to transplant post-#101940; the invariant is the important part: this execution holding its own side-effect fence must never, by duration alone, manufacture a cancellation verdict. Store/I/O uncertainty still needs a bounded failure policy, and a real external cancel or confirmed owner replacement must continue to fail closed.
Required regression before merge: use the real _side_effect_fence, shorten heartbeat/fence/grace constants, keep _deliver_result inside the fence longer than grace, and prove one delivery + last_status == "ok" + ledger completed. Pair it with the existing real-cancel / genuine-takeover assertion so the repair cannot weaken at-most-once settlement.
Interlock / attribution
This is the right current-main carrier for the post-#101940 shape, but the PR description should explicitly record the lineage before the older branches are closed: #95432, #97565, #100418, and #100965 are not unrelated duplicates. They are earlier competing implementations of the same fence-contention/ownership-verdict defect, and #100965 contains the long-delivery-beyond-grace acceptance case that this successor currently loses. Preserve that contributor credit and mark supersession/alternative-mechanism status in both directions rather than erasing the history.
FILE-LIST confirms the collision is direct: this PR changes cron/jobs.py, cron/scheduler.py, tests/cron/test_claim_job_for_fire.py, and tests/cron/test_script_claim_heartbeat.py; the predecessor PRs overlap those same authority and test seams. Merge-order should therefore be one carrier only, not additive merges.
CI gate
Exact-head CI is not green yet. For 9b26dcb9fb1f5425b198495b655d7ef54304da3f, GitHub currently reports 0 check runs; CI, Docker, and Nix workflow runs are action_required with no jobs executed. The local 1224/787 receipts are useful evidence but cannot substitute for exact-head repository CI. Because this PR is one commit, the every-commit gate reduces to this exact SHA: it still needs an actual green run before merge.
Aside from the grace-boundary regression above, the True / False / None separation itself is coherent with the current owner-CAS and side-effect-fence architecture. Fix that residual “other side of the shape,” restore the dropped long-delivery acceptance case, preserve the predecessor lineage, and this becomes a much stronger settlement repair.
|
Exact-head follow-up on |
|
@andrexibiza Exact-head correction is now pushed at The beyond-grace #100965 invariant now uses the real side-effect fence and passes with one delivery, |
|
Exact-head live deployment acceptance for @andrexibiza Please perform the requested exact-head re-review. GitHub CI remains externally blocked: CI, Docker, and Nix runs are still |
andrexibiza
left a comment
There was a problem hiding this comment.
Exact-head re-review: d74a8d157a5a4e3fabac716df2d9c14420ec34ae against current main / merge base 63279301bcbdc185c1b07b98a9312eb0c862f26d (4 commits, 6 changed files, 0 behind, mergeable).
Prior semantic blocker: resolved
The correction now preserves the three distinct facts that the original path collapsed:
True: the fence was acquired and this owner was renewed/confirmed.False: the fence was acquired and owner absence/replacement was authoritatively observed.None: the fence is known to be contended, so ownership was not inspected and takeover remains excluded by that same fence.
_fire_job_lock(..., raise_unavailable=True) now separates ordinary lock contention from lock-backend/I/O failure. _run_with_fire_claim_heartbeat() gives raised backend/store failures their own bounded grace clock, while either a confirmed renewal or known contention resets that clock. A worker holding its own _side_effect_fence can therefore remain there beyond the old grace window without manufacturing lost_ownership; explicit owner replacement still interrupts immediately, initial validation remains exactly-True, and external cancellation and unavailable terminal persistence remain fail-closed.
The dropped #100965 acceptance boundary is now restored with the real side-effect fence: delivery remains inside the fence beyond shortened heartbeat grace and must settle exactly once with last_status == "ok" and ledger completed. The paired external-cancel and confirmed-takeover controls still settle failed without leaking the stale result. The pending _interrupting_job_ids versus confirmed _interrupted_job_ids split also closes the shutdown race: a pending shutdown suppresses a plausible normal delivery, but only a successfully persisted shutdown terminal lets the worker skip its own terminal write.
I found no remaining code-level logic or security blocker in this six-file exact-head implementation. The linked live receipt is also coherent with the repaired state machine: exact code SHA, one surviving external worker across one gateway restart, one side effect, completed execution, and last_status=ok. That is strong runtime evidence, while still remaining an author-produced receipt rather than repository CI.
Remaining merge gates
1. Exact-head and every-commit CI are not green. For d74a8d157a5a4e3fabac716df2d9c14420ec34ae, CI 33841575617, Docker 33841574994, and Nix 33841574929 are all action_required; GitHub reports zero check runs and those workflows executed zero jobs. The preceding three commits likewise have only action_required runs. The focused/full-cron/live receipts do not replace exact-object repository CI, so this branch still has no green commit train.
2. The decomposition/landing interlock is not closed. Open #102117 directly overlaps cron/jobs.py, cron/scheduler.py, and the heartbeat tests from the same base. Its current head 2c6c645803055d213aad179a94054de925ce0542 still carries the pre-fix binary heartbeat_fire_claim() -> bool contract and the old _fire_job_lock() behavior that does not distinguish backend failure from contention. Any landing order that allows that head to overwrite this repair silently reintroduces the defect. Whichever carrier lands second must explicitly consume this exact tri-state postcondition and the beyond-grace/cancel/takeover acceptance cases.
This branch also grows the already-over-limit owners by net +74 lines in cron/jobs.py and +156 in cron/scheduler.py; both files remain far beyond the 2K boundary. The accepted landing topology therefore needs the bounded post-#102117 owner rather than re-entrenching the godfiles or letting the later extraction erase the semantic repair.
3. Canonical publication state needs reconciliation. The PR body still says no production gateway restart was performed, while the latest acceptance records a default managed-gateway restart. Clarify whether that host was explicitly non-production or update the statement. The duplicate label also conflicts with the body’s declaration that this is the selected current-main carrier; remove it or explain the intended canonical ownership before older alternatives are retired.
Verdict: the previous implementation blocker is resolved on this exact head. Merge remains blocked on a green exact-head/every-commit train, the #102117/2K landing interlock, and a single current public record of the acceptance and carrier state.
…nges to be re-applied next)
…wnership loss (c-027) heartbeat_fire_claim() collapsed fire-fence acquisition failure into False. During slow fenced delivery the worker's own heartbeat thread cannot acquire the same process-local RLock, so a successfully delivering run was misclassified as ownership loss and terminalled as 'Interrupted by shutdown before terminal completion.' (TrustMRR exec 4219b48d while RSI recovery exec f502e929 was settling). Tri-state contract: - True = renewed/confirmed owner - False = authoritatively inspected: owner mismatch / absent claim - None = fence unavailable/unconfirmed (contention or store error) - initial validation stays fail-closed: run starts only on exactly True - heartbeat loop: None rides the existing grace window (180s prod vs 30s fence wait); immediate cancel only on explicit False - post-run probes (_fire_claim_ownership_lost, both interrupted blocks, terminal owner-CAS read): None can never adjudicate a confirmed loss; uncertain outcomes get distinct ledger errors - mark_job_run also tri-state: None = fence unavailable (CAS never ran), False stays reserved for confirmed owner mismatch - _side_effect_fence unchanged: still the exactly-once save/delivery barrier Tests (all written red-first, verified failing on unfixed code): slow_owned_delivery (parent handoff), sustained fence-contention grace, pre/post-delivery None probes, grace-exhausted None probe, terminal write None, mark_job_run fence-unavailable None, RSI recovery dispatch fails closed during original's fenced delivery, TrustMRR slow-owned delivery terminal success. Managed-gateway restart E2E (systemd scope) passes on host: one gateway, single side effect, single delivery.
… paths Forward-port completion on top of the main-side cron refactor (scheduler phase helpers): the exception-path failure notification and the fence-local interrupted recheck before _deliver_result were lost in the refactor's _deliver_crash_failure extraction. Restored so a stale/uncertain worker can never emit a failure alert the replacement run will also send, and a shutdown that begins during the delivery fence can never leak the plausible final response as a success.
|
Forward-ported onto current main (cron refactor); conflict blocker resolved; exact-head verification + tests green. Re-review requested at this exact head. What changed on the branch (head
Verification at this exact head:
No merge performed; no production gateway restart. Exact-head maintainer re-review requested (fork-author permissions prevent formal reviewer assignment). |
|
The fix for this bug landed on main via #109310 (merge 9a60a7f): Thanks @jayleaton — the exact-head verification and live deployment acceptance here were thorough. The landed fix is the minimal form (heartbeat no longer takes the fence), which removes the contention class the tri-state machinery classifies. Re the review by @andrexibiza: the self-held long-delivery invariant is covered by the merged test (heartbeat True while a worker thread holds |
Problem
Follow-up to #101940 (merged): live acceptance still failed. An independently surviving cron worker held
_side_effect_fenceacross slow delivery; its heartbeat thread could not acquire the same process-local RLock, andheartbeat_fire_claim()collapsed that lock-contention timeout intoFalse. Scheduler code treated everyFalseas authoritative ownership loss, terminaling a successfully delivered run asInterrupted by shutdown before terminal completion.The first correction made the heartbeat tri-state, but exact-head review found a residual boundary: returned
Nonestill consumed the same grace budget as genuine store/I/O errors. A legitimate delivery held under its own side-effect fence longer than grace was therefore still demoted toerrorafter delivering once.Root cause
cron/jobs.py: heartbeat lock contention and lock-backend/I/O unavailability lacked a durable semantic distinction.cron/scheduler.py: returnedNone(known fence contention) and raised store/backend errors shared one grace clock and one sticky cancellation event.Fix
True: renewed / confirmed owner.False: fence acquired and owner mismatch/absence confirmed.None: known fence contention; ownership was not inspected, but takeover is excluded while the holder owns the fence.EACCES,EAGAIN, andEWOULDBLOCK, and Windows uses bounded nonblocking acquisition.Noneno longer consumes error grace. Consecutive backend/store exceptions have their own bounded grace clock, reset by either a successful renewal or known contention.Truefail-closed. Confirmed takeover, external cancellation, store/backend uncertainty beyond grace, and unavailable terminal persistence remain fail-closed._side_effect_fenceremains the exactly-once save/delivery authority.Tests
Ported and adapted the real #100965 acceptance case to the current execution-row contract:
last_status == "ok", and execution ledgercompleted;Verification on exact head:
tests/cron/: 1234 passed, 0 failed, 2 platform skips;Lineage and contributor credit
This PR is the current-main carrier after #101940, not an unrelated duplicate. It preserves and reconciles earlier competing work on the same fence-contention/ownership-verdict defect:
These overlap the same authority and test seams; they are alternative/superseded mechanisms, not additive merge candidates. One carrier should land.
No model/provider spend changes and no production gateway restart were performed for this correction.