Repository navigation
Conversation
…oss (NousResearch#95307) heartbeat_fire_claim() returned False both for a definitive claim takeover and for 'could not acquire the per-job fire fence within the lock budget'. Those are different facts: the fence that blocks the refresh is the very lock that excludes replacement owners while the legitimate runner performs a long fenced side effect (a bot-chat delivery runs a full child-agent turn holding it). Callers classified the contention as ownership loss and discarded their own successful result ('Fire claim ownership lost; stale result was discarded.'), which deterministic repros show for finite jobs delivering to bot-chat. heartbeat_fire_claim() is now tri-state: True (owned + refreshed), False (definitive loss: job/claim gone or owner replaced), None (fence busy — ownership unknown). Scheduler consumers updated: - pre-run validation treats None as unverifiable and closes the row with the existing 'could not be validated' record instead of 'lost' - heartbeat beats skip on None without burning the exception grace budget; refreshes resume once the fence frees up - post-run probes treat None as not-lost and proceed; the CAS mark_job_run(expected_fire_owner=...) remains the authoritative guard against completing a stolen claim - the two shutdown-interrupt re-probes only take the fenced-interrupted branch when ownership is positively confirmed (is True) False still means loss everywhere, so real takeovers keep interrupting stale runs exactly as before. Regression tests cover the tri-state store contract, the busy-fence startup path, terminal bookkeeping under a busy probe, and assert definitive takeaways are still detected.
The tri-state design itself is right: contention-as-None matches the fence's actual semantics (it's the same lock that already excludes replacement owners), the old truthiness call sites now use explicit |
|
The fix for this bug landed on main via #109310 (merge 9a60a7f): Thank you for the earliest report and fix of this bug (#95307). The landed change takes the smaller route — dropping the fence from the heartbeat instead of a tri-state result — which also resolves the startup-branch concern raised in the review here (a busy fence can no longer make the pre-run heartbeat return |
What does this PR do?
Fixes #95307 — finite cron jobs (once / repeat-limited) delivering to
bot-chatcomplete their agent turn and then get recorded as failed withFire claim ownership lost; stale result was discarded., losing the result entirely. The root cause is a semantic conflation in the fire-claim ownership probe; this PR separates "definitively lost" from "could not verify" so the runner stops discarding its own successful result.Bug cause
heartbeat_fire_claim()returnedFalsefor two different facts:\_JOBS_LOCK_TIMEOUT_SECONDS— pure lock contention, with no information about who owns the claim.Those are not interchangeable. The fence that blocks the refresh is the very lock that excludes replacement owners while the legitimate runner performs a fenced side effect. A
bot-chatdelivery runs an entire child-agent turn (hermes chat -Q ...) while holding that fence, so heartbeat/probe calls issued during delivery contend with the runner's own fence. Reading that as ownership loss makes the runner discard its own completed result. Deliveries tolocal/ ordinary platforms finish in well under a second and never straddle the probe window, which matches the reporter's isolation matrix exactly.On the current main,
mark_job_run(expected_fire_owner=...)already CAS-guards terminal writes and the post-run probes run before finalization — so the reporter's proposed fixes #1/#2 are largely landed upstream. What remains broken is exactly the conflatedFalse: any caller that cannot acquire the fence is told "you lost the claim" when the truth is "I don't know".Fix
heartbeat_fire_claim()becomes tri-state:Trueexpected_owner, refreshedFalseNoneScheduler consumers updated accordingly:
Nonecloses the execution row with the existing "Fire claim ownership could not be validated before execution started." record (same treatment as a transport exception) instead of reporting a loss.Nonebeat is skipped without consuming the exception grace budget; refreshes resume once the fence frees up. Only a definitiveFalsesets the ownership-lost event.\_fire_claim_ownership_lost):Noneis treated as not lost, and the run proceeds to its owner-fenced terminal write —mark_job_run(expected_fire_owner=...)remains the authoritative last-line guard against completing on a stolen claim.is True).Falsestill means loss everywhere: a real takeover keeps interrupting stale runs and suppressing stale deliveries exactly as before.Testing
New regression suite
tests/cron/test_fire_claim_contention_95307.py(behavior-contract, real store under a tempHERMES_HOME, no snapshot tests):None; uncontended ⇒True; wrong owner ⇒FalseAll six fail on pristine main (three of them reproduce the exact reported ledger error string); all pass with the fix.
Neighbor suites:
test_claim_job_for_fire.py,test_script_claim_heartbeat.py,test_shutdown_interrupt.py,test_run_one_job.py,test_cron_bot_chat_delivery.py,test_dead_owner_claim_reclaim.py,test_inflight_stale_guard.py,test_cron_run_stale_claim_reap_86721.py— 116 passed. Fulltests/cron/: 971 passed, 1 skipped, plus one pre-existing environment failure intest_cron_created_delivery.pythat fails identically on pristine main (unrelated to this change).tests/gateway/test_restart_resume_pending.py: 36 passed.Related