Skip to content

fix(bin): bound the startup-network worker's lock waits by its budget - #5528

Merged
kunchenguid merged 2 commits into
kunchenguid:mainfrom
karotkriss:fm/fm-up-5377-startup-network-lock-waits
Sep 24, 2026
Merged

kunchenguid merged 2 commits into
kunchenguid:mainfrom
karotkriss:fm/fm-up-5377-startup-network-lock-waits

Conversation

@karotkriss

Copy link
Copy Markdown
Contributor

Intent

Fixes #5377

The deferred startup network worker bounds its sweep work with stage budgets but still calls an unbounded lock wait for publication and delivery, so a live holder of the publish lock can keep the detached worker alive for hours past its timeout, burning CPU with its output discarded.
Make every wait in that worker respect its budget: when the lock cannot be taken in time the worker stops within budget and leaves a durable diagnose-for-rerun result instead of spinning.
This closes upstream issue #5377.

What Changed

  • Route every lock the deferred startup-network worker takes (stage, sweep-lease, publication, delivery, and harvest) through a new bounded take_lock helper, replacing the unbounded fm_lock_acquire_wait calls so a live lock holder can no longer keep the detached worker spinning past its timeout with its output discarded.
  • Add stage_deadline/DELIVERY_DEADLINE budgets and a publish_lock_held path: when a lock cannot be acquired in time, the worker writes a durable failed/rerun NETWORK_CHECKS record naming the holder (FM_LOCK_HELD_PID) and rerun command, queues a wake, and exits non-zero; cmd_run now propagates that exit code and await_delivery loops on the delivery deadline instead of a fixed iteration cap.
  • Refactor publish into record_result (file writing) plus lock/deadline handling, centralize cleanup in run_cleanup, update docs/configuration.md to document the bounded lock waits, and add a test that holds the publish lock from a live process and asserts a failed-rerun record instead of a hang.

Risk Assessment

✅ Low: Well-bounded shell change that routes every worker lock wait through a budget-bounded acquire, records a durable failed-rerun result on refusal, and is covered by a behavior-based regression test; the minimal fix-round commit correctly propagates publish's exit code and introduces no new defect.

Testing

I ran the full fm-startup-network behavior suite against the fixed worker: all 21 assertions pass, including the new regression that drives the real fm-startup-network.sh run/start worker under a live publish-lock holder held from a separate process. To prove the regression is real I swapped in the base (pre-fix) worker script and re-ran: the new test fails with the worker still waiting on the held publish lock 15s past a 2s budget, then restored the fixed script and confirmed a clean worktree. The test asserts observable behavior end-to-end - the worker gives up in under 6s on a 2s budget, records state=failed, the report output names the holding pid and the exact rerun command, the pre-sweep case runs no sweeps, the post-sweep case preserves the sweep output (PROBE_RAN), and both queue a startup-network wake. The docs/configuration.md line is a one-line comment with no live surface.

  • Live validation: ✅ go - 3 of 4 scenarios driven live against the product
Scenario Result Live Evidence
Publish lock held before the worker registers -> worker stops within budget, records failed-rerun naming the holder, runs no sweeps, and wakes ✅ pass live test_a_held_publish_lock_cannot_keep_the_worker_alive_past_its_budget (pre-sweep half): live worker gives up in under 6s on a 2s budget (rc!=124, rc!=0), state=failed, report contains 'still held by p…
Publish lock held after sweeps finish -> worker stops within delivery budget, preserves sweep output, records failed-rerun naming the holder, and wakes ✅ pass live test_a_held_publish_lock_cannot_keep_the_worker_alive_past_its_budget (post-sweep half): detached worker sweeps freely then finds the lock held; exits within a 2s delivery budget, state=failed, report…
Regression reproduces the reported unbounded wait (fail-before/pass-after) ✅ pass live Base worker from d4f3b78 makes the new test fail with 'still waiting on the held publish lock 15s past a 2s budget'; fixed worker makes it pass - captured in regression-before-after.txt
docs/configuration.md FM_STARTUP_NETWORK_TIMEOUT comment update ⏸️ untested no Pure docs edit; no live-executable product surface, so no tool or credential would enable a live check.
Evidence: fm-startup-network bounded lock-wait: fail-before / pass-after

Source: fm-startup-network bounded lock-wait: fail-before / pass-after

## After fix (HEAD=75b9a11): ok - fm-startup-network: a held publish lock ends the worker inside its budget with a failed-rerun record ## Before fix (base script from d4f3b78): not ok - the worker was still waiting on the held publish lock 15s past a 2s budget

=== fm-startup-network bounded lock-wait regression ===

## After fix (HEAD=75b9a11) - new regression test:
ok - fm-startup-network: a held publish lock ends the worker inside its budget with a failed-rerun record

## Before fix (base script from d4f3b78) - same test:
not ok - the worker was still waiting on the held publish lock 15s past a 2s budget

worktree restored: 0 modified files

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

🔧 **Review** - 1 issue found → auto-fixed ✅
  • ℹ️ bin/fm-startup-network.sh:596 - The new doc line (bin/fm-startup-network.sh:64-65) promises run "Exits non-zero when the stage was refused or could not publish, including a lock a live process still held at its deadline." The two explicit refusal paths honor this (registration lock line 501, lease lock line 558 both return 1). But the terminal publish paths (line 596/600/605) ignore publish's return value and fall through to return 0 (line 609), so when the publication lock is held past DELIVERY_DEADLINE, publish -> publish_lock_held writes the failed-rerun record yet cmd_run exits 0. A manual run from the lock-owning harness checking $? sees success despite a failed record. Remedy is trivial (propagate publish's rc, or narrow the doc claim); a durable record is still written either way, so this is cosmetic.

🔧 Fix applied.
✅ Re-checked - no issues remain.

✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 3 of 4 scenarios driven live against the product
Scenario Result Live Evidence
Publish lock held before the worker registers -> worker stops within budget, records failed-rerun naming the holder, runs no sweeps, and wakes ✅ pass live test_a_held_publish_lock_cannot_keep_the_worker_alive_past_its_budget (pre-sweep half): live worker gives up in under 6s on a 2s budget (rc!=124, rc!=0), state=failed, report contains 'still held by p…
Publish lock held after sweeps finish -> worker stops within delivery budget, preserves sweep output, records failed-rerun naming the holder, and wakes ✅ pass live test_a_held_publish_lock_cannot_keep_the_worker_alive_past_its_budget (post-sweep half): detached worker sweeps freely then finds the lock held; exits within a 2s delivery budget, state=failed, report…
Regression reproduces the reported unbounded wait (fail-before/pass-after) ✅ pass live Base worker from d4f3b78 makes the new test fail with 'still waiting on the held publish lock 15s past a 2s budget'; fixed worker makes it pass - captured in regression-before-after.txt
docs/configuration.md FM_STARTUP_NETWORK_TIMEOUT comment update ⏸️ untested no Pure docs edit; no live-executable product surface, so no tool or credential would enable a live check.
  • bash tests/fm-startup-network.test.sh (fixed worker) - all 21 assertions pass including test_a_held_publish_lock_cannot_keep_the_worker_alive_past_its_budget
  • Swapped base worker (git show d4f3b78:bin/fm-startup-network.sh) and re-ran the suite: new regression fails with 'still waiting on the held publish lock 15s past a 2s budget'
  • git checkout restored fixed worker; git status confirms clean worktree at HEAD 75b9a11
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

Fixes kunchenguid#5377

The deferred startup network worker bounded its sweeps with a stage budget but
took the publish lock and the fleet-lock lease with an unbounded wait, so a live
holder of that lock kept the detached worker alive for hours past its timeout
with its output discarded at the end. Every wait now goes through the bounded
acquire and shares the remaining stage or delivery budget; a lock a live process
still holds at the deadline ends the worker with a failed record naming the
holder and the rerun command, and a wake so the result surfaces.
@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate:

Contract-class: restore — tip vs main d4f3b78: main's deferred fm-startup-network.sh worker still calls unbounded fm_lock_acquire_wait on publication/delivery/lease while stage/delivery budgets already exist; tip routes every wait through take_lock → fm_lock_acquire_wait_bounded, records durable failed-rerun naming the holder, and propagates publish exit. Unconfigured path: same timeouts, no new default-on surface. Restores the already-specified budget path that was broken (issue #5377). Not new-default (FM-LEARN-4627).

VISION.md (each rule)

  • One captain, one interface: aligns — unbounded CPU with discarded output after a deferred check is hidden mechanics that erodes looking-away confidence.
  • Authority is explicit: aligns — no new grant; budgets already promised.
  • Scripts own the mechanics: aligns — lock waits and deadlines are exact script concerns.
  • A restart is a non-event: aligns — worker stops within budget; durable failed-rerun + wake survive.
  • Delegation with a spine: aligns — no new task shape; strengthens an existing refusal/bound.
  • The fleet outlives any vendor: aligns — lock/timeout path is harness-agnostic.
  • Scope: aligns — command-layer startup-network worker lifecycle.

aligns when / resist when: Aligns — strengthens a refusal path and survives lock-contention. Does not widen into wake-dedupe (#5378).

Closes check: Author Fixes #5377 verified against issue body (25h+ worker past budget on unbounded publish lock) and tip diff. Issue labeled ready-for-pr.

CI/NM: all SUCCESS. Attestation MATCH on HEAD 75b9a11. MERGEABLE/CLEAN. Safe review; no security tip. Auto-merge as restore.

@kunchenguid
kunchenguid merged commit 5842d42 into kunchenguid:main Sep 24, 2026
19 checks passed
@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate: this is merged. Thank you @karotkriss — really appreciate you taking the time on this.

mituso89 pushed a commit to mituso89/firstmate that referenced this pull request Sep 26, 2026
…kunchenguid#5528)

* fix(bin): bound the startup-network worker's lock waits by its budget

Fixes kunchenguid#5377

The deferred startup network worker bounded its sweeps with a stage budget but
took the publish lock and the fleet-lock lease with an unbounded wait, so a live
holder of that lock kept the detached worker alive for hours past its timeout
with its output discarded at the end. Every wait now goes through the bounded
acquire and shares the remaining stage or delivery budget; a lock a live process
still holds at the deadline ends the worker with a failed record naming the
holder and the rerun command, and a wake so the result surfaces.

* no-mistakes(review): propagate publish exit code from cmd_run terminal paths
mehulbhagwani pushed a commit to mehulbhagwani/firstmate that referenced this pull request Sep 26, 2026
…kunchenguid#5528)

* fix(bin): bound the startup-network worker's lock waits by its budget

Fixes kunchenguid#5377

The deferred startup network worker bounded its sweeps with a stage budget but
took the publish lock and the fleet-lock lease with an unbounded wait, so a live
holder of that lock kept the detached worker alive for hours past its timeout
with its output discarded at the end. Every wait now goes through the bounded
acquire and shares the remaining stage or delivery budget; a lock a live process
still holds at the deadline ends the worker with a failed record naming the
holder and the rerun command, and a wake so the result surfaces.

* no-mistakes(review): propagate publish exit code from cmd_run terminal paths
mehulbhagwani pushed a commit to mehulbhagwani/firstmate that referenced this pull request Sep 26, 2026
…kunchenguid#5528)

* fix(bin): bound the startup-network worker's lock waits by its budget

Fixes kunchenguid#5377

The deferred startup network worker bounded its sweeps with a stage budget but
took the publish lock and the fleet-lock lease with an unbounded wait, so a live
holder of that lock kept the detached worker alive for hours past its timeout
with its output discarded at the end. Every wait now goes through the bounded
acquire and shares the remaining stage or delivery budget; a lock a live process
still holds at the deadline ends the worker with a failed record naming the
holder and the rerun command, and a wake so the result surfaces.

* no-mistakes(review): propagate publish exit code from cmd_run terminal paths
mehulbhagwani pushed a commit to mehulbhagwani/firstmate that referenced this pull request Sep 27, 2026
…kunchenguid#5528)

* fix(bin): bound the startup-network worker's lock waits by its budget

Fixes kunchenguid#5377

The deferred startup network worker bounded its sweeps with a stage budget but
took the publish lock and the fleet-lock lease with an unbounded wait, so a live
holder of that lock kept the detached worker alive for hours past its timeout
with its output discarded at the end. Every wait now goes through the bounded
acquire and shares the remaining stage or delivery budget; a lock a live process
still holds at the deadline ends the worker with a failed record naming the
holder and the rerun command, and a wake so the result surfaces.

* no-mistakes(review): propagate publish exit code from cmd_run terminal paths
mehulbhagwani pushed a commit to mehulbhagwani/firstmate that referenced this pull request Sep 27, 2026
…kunchenguid#5528)

* fix(bin): bound the startup-network worker's lock waits by its budget

Fixes kunchenguid#5377

The deferred startup network worker bounded its sweeps with a stage budget but
took the publish lock and the fleet-lock lease with an unbounded wait, so a live
holder of that lock kept the detached worker alive for hours past its timeout
with its output discarded at the end. Every wait now goes through the bounded
acquire and shares the remaining stage or delivery budget; a lock a live process
still holds at the deadline ends the worker with a failed record naming the
holder and the rerun command, and a wake so the result surfaces.

* no-mistakes(review): propagate publish exit code from cmd_run terminal paths
mehulbhagwani pushed a commit to mehulbhagwani/firstmate that referenced this pull request Sep 27, 2026
…kunchenguid#5528)

* fix(bin): bound the startup-network worker's lock waits by its budget

Fixes kunchenguid#5377

The deferred startup network worker bounded its sweeps with a stage budget but
took the publish lock and the fleet-lock lease with an unbounded wait, so a live
holder of that lock kept the detached worker alive for hours past its timeout
with its output discarded at the end. Every wait now goes through the bounded
acquire and shares the remaining stage or delivery budget; a lock a live process
still holds at the deadline ends the worker with a failed record naming the
holder and the rerun command, and a wake so the result surfaces.

* no-mistakes(review): propagate publish exit code from cmd_run terminal paths
RooseveltAdvisors pushed a commit to RooseveltAdvisors/firstmate that referenced this pull request Sep 29, 2026
…kunchenguid#5528)

* fix(bin): bound the startup-network worker's lock waits by its budget

Fixes kunchenguid#5377

The deferred startup network worker bounded its sweeps with a stage budget but
took the publish lock and the fleet-lock lease with an unbounded wait, so a live
holder of that lock kept the detached worker alive for hours past its timeout
with its output discarded at the end. Every wait now goes through the bounded
acquire and shares the remaining stage or delivery budget; a lock a live process
still holds at the deadline ends the worker with a failed record naming the
holder and the rerun command, and a wake so the result surfaces.

* no-mistakes(review): propagate publish exit code from cmd_run terminal paths
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Bug: deferred startup network worker can outlive its timeout

2 participants