Repository navigation
fix(cron): keep a delivered run's success through a transient post-delivery claim blip (#105861) - #105882
fix(cron): keep a delivered run's success through a transient post-delivery claim blip (#105861)#105882PRATHAMESH75 wants to merge 1 commit into
Conversation
…livery claim blip (NousResearch#105861) Agent-mode cron jobs hold a durable fire claim a heartbeat re-validates every 60s. `_run_one_job_body` short-circuited to `_record_fire_ownership_lost` whenever `d.side_effect_ownership_lost or _fire_claim_ownership_lost()` was truthy *after* the save/compose/deliver phase. `_fire_claim_ownership_lost()` (`fence.lost()`) is a sampled flag that flips permanently on a single transient missed tick (disk latency, a concurrent jobs.json rewrite, a brief spawn spike). So a run that had already delivered its output got `last_status` overwritten with "Interrupted by shutdown before terminal completion.", turning a healthy job's history into an error and firing false watchdog alerts on every later tick — while the message had in fact been sent (reporter saw it 3x in one day, all verified delivered). The post-delivery interruption decision is now `_delivery_phase_interrupted(d)`: True only when the claim was lost *during* the side effect (`_save_compose_deliver` raised `_FireClaimLostDuringSideEffect`, delivery incomplete). A completed delivery falls through to `_finish_completed_run`, whose owner-fenced `mark_job_run(expected_fire_owner=...)` is the authoritative check — it records a genuine ownership loss atomically at mark time and the real success when the claim is still held. Gateway shutdown is signaled separately via `_consume_interrupted_flag`, not this sampled fence, so that path is unaffected. Fixes NousResearch#105861
Related: this joins the open fix cluster for #100401 (same |
|
Heads-up: the root cause of the 'sampled lost flag flips on a transient tick' behaviour landed on main via #109310 (merge 9a60a7f). The tick that flipped |
|
Closing this as superseded — thanks @kshitijk4poor for the heads-up, and for landing the actual root-cause fix in #109310 ( I re-checked #105861 against current That matches your read: the fix belongs at the ownership-check source, not as a downstream exception in the bookkeeping tail. Maintainers may want to keep #105861 open until they've confirmed the P1 symptom is gone on main, but there's nothing left for this PR to add. Closing. |
Problem (#105861)
Agent-mode cron jobs hold a durable fire claim a heartbeat re-validates every 60s. After the save/compose/deliver phase,
_run_one_job_bodyshort-circuited to_record_fire_ownership_lostwhenever:_fire_claim_ownership_lost()(fence.lost()) is a sampled flag that flips permanently on a single transient missed heartbeat tick (disk latency, a concurrentjobs.jsonrewrite, a brief spawn spike). So a run that had already delivered its output gotlast_statusoverwritten withInterrupted by shutdown before terminal completion.— turning a healthy job's history intoerrorand firing false watchdog alerts on every subsequent tick, while the message had in fact been sent.The reporter observed this 3× in one day (~12% of that day's agent runs), all independently verified delivered, zero duplicates.
Confirmed in
_record_fire_ownership_lostitself: it only writes the errorif ... heartbeat_fire_claim(job_id, expected_owner=fire_owner)returns True — i.e. it wrote the interruption because the claim was still held. Thefence.lost()sample and the atomic heartbeat disagreed; the sample was the stale one.Fix
The post-delivery interruption decision is now
_delivery_phase_interrupted(d), True only when the claim was lost during the side effect (_save_compose_deliverraised_FireClaimLostDuringSideEffect, so delivery did not complete under a held claim). A completed delivery falls through to_finish_completed_run, whose owner-fencedmark_job_run(expected_fire_owner=...)is the authoritative check:mark_job_runreturns False → records the real ownership loss;last_status: ok.This removes the redundant, sampled post-delivery check without weakening any guarantee:
_save_compose_deliver(side_effect_fence()→ raises_FireClaimLostDuringSideEffect), which is unchanged._consume_interrupted_flag(a separate per-execution flag), not this fence, so the shutdown path is unaffected._fire_claim_ownership_lost()is still used for the pre-run ownership check earlier in the body.Tests
tests/gateway/test_cron_delivered_success_survives_claim_blip.py: a completed delivery is not an interruption (the regression), a loss during the side effect still is, a normal delivery error is not an ownership interruption, the predicate takes no sampled-ownership input, and the wrapper routes through_delivery_phase_interrupted(guards against re-adding the disjunction).Fixes #105861