Skip to content

fix(bin): keep process identity stable across host clock steps - #45

Merged
knowttl merged 2 commits into
mainfrom
fm/fm-lstart-identity-siblings
Sep 23, 2026
Merged

knowttl merged 2 commits into
mainfrom
fm/fm-lstart-identity-siblings

Conversation

@knowttl

@knowttl knowttl commented Sep 23, 2026

Copy link
Copy Markdown
Owner

Intent

can you explicitly verify that it is an actual bug against the source code then implement a fix so we can also create a pull request for it in upstream?

also ensure that before you create a pull request on the upstream repo please also ensure we properly follow the upstream repo contribution rules if there are any and the standards it has defined?

also ensure that the fixes generalize across the firstmate repo.

two pull requests

Context the ask refers to: a scout confirmed that firstmate's Linux remote-job worker identified processes by the text of ps -o lstart=, which on Linux is the kernel's wall-clock-derived boot time plus start ticks, so every host clock step (WSL2 steps about every 30 seconds) changes every running process's recorded start and a live process stops matching its recorded identity.
The fix for the remote-job worker itself is the first of the two pull requests and is a separate task.
To make the fix generalize, the scout swept bin/ for the same flaw and found two more vulnerable sites (code-read verdicts, not reproduced): bin/fm-pending-reply-lib.sh (fm_pending_reply_pid_identity / fm_pending_reply_sender_alive, around lines 979-993), where after a clock step a live recovery sender reads as dead; and the Herdr lab viewer/launcher ownership check in bin/fm-herdr-lab.sh (around lines 202-231) and bin/fm-herdr-lab-viewer.py (around line 83), where after a step the pair reads as not owned.
Sites already safe for reference: fm_pid_identity in bin/fm-wake-lib.sh (the July watcher fix, upstream kunchenguid#752), task_process_identity in bin/fm-teardown.sh, and pidIdentity in bin/fm-extension.mjs, which all use /proc/<pid>/stat start ticks.
This task is the second pull request: fix those two remaining sites the same way, with compatibility for records already written in the old lstart form.
The scout report's section 9 D recommends exactly this: switch those two sites to /proc/<pid>/stat start-tick identity where readable, keep ps elsewhere, keep recognizing legacy lstart records, and pin each with a regression test that changes boot time but not start ticks.

What Changed

  • Use /proc start ticks for recovery sender identity while preserving command identity and a ps fallback.
  • Use start ticks for Herdr lab viewer ownership checks and continue recognizing existing lstart records.
  • Add regression tests for clock steps, legacy records, and sender PID reuse.

Risk Assessment

✅ Low: The change is limited to the two reported identity checks, preserves comparison of legacy records, and adds behavioral regressions for clock steps.

Testing

Both focused test scripts passed, covering simulated clock steps, PID reuse, and legacy records. A real named Herdr lab viewer attached with start-tick identities, stopped, and was torn down. An initial lab attempt hit the helper’s duplicate-tripwire guard because provision already performs prepare; the fresh provision-only run passed. No real host clock was changed.

  • Live validation: ⚠️ inconclusive - 1 of 4 scenarios driven live against the product
Scenario Result Live Evidence
A recovery sender remains in progress when boot time changes but its start ticks do not ⏸️ untested no The prior payload cited a regression test but did not establish this behavior against the live product.
Recovery rejects a reused PID and recognizes an existing lstart record ⏸️ untested no The prior payload cited test assertions but did not establish this behavior against the live product.
Viewer ownership survives a simulated clock step while rejecting an unowned process and accepting legacy records ⏸️ untested no The prior payload cited simulated test assertions but did not establish these behaviors against the live product.
An operator starts and stops a viewer in a named Herdr lab; its recorded process identities use start ticks and its ownership record clears on stop ✅ pass live Real Herdr lab viewer lifecycle
Evidence: Real Herdr lab viewer lifecycle
viewer attached to fm-lab-identity-633913-12057 (pid 634764)
launcher_start=proc-starttime=11953650
viewer_start=proc-starttime=11953656
viewer stopped; record present=no
- Outcome: ⚠️ 1 warning across 1 run (3m34s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

✅ **Review** - passed

✅ No issues found.

⚠️ **Test** - 1 warning
  • ⚠️ live validation verdict: inconclusive (1 of 4 scenarios were driven live against the product); untested: A recovery sender remains in progress when boot time changes but its start ticks do not, Recovery rejects a reused PID and recognizes an existing lstart record, Viewer ownership survives a simulated clock step while rejecting an unowned process and accepting legacy records
  • Live validation: ⚠️ inconclusive - 1 of 4 scenarios driven live against the product
Scenario Result Live Evidence
A recovery sender remains in progress when boot time changes but its start ticks do not ⏸️ untested no The prior payload cited a regression test but did not establish this behavior against the live product.
Recovery rejects a reused PID and recognizes an existing lstart record ⏸️ untested no The prior payload cited test assertions but did not establish this behavior against the live product.
Viewer ownership survives a simulated clock step while rejecting an unowned process and accepting legacy records ⏸️ untested no The prior payload cited simulated test assertions but did not establish these behaviors against the live product.
An operator starts and stops a viewer in a named Herdr lab; its recorded process identities use start ticks and its ownership record clears on stop ✅ pass live Real Herdr lab viewer lifecycle
  • tests/fm-pending-reply.test.sh
  • tests/fm-herdr-lab.test.sh
  • bin/fm-herdr-lab.sh provision <named-lab>
  • bin/fm-herdr-lab.sh viewer start <named-lab>
  • bin/fm-herdr-lab.sh viewer stop <named-lab>
  • bin/fm-herdr-lab.sh teardown <named-lab>
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

…ross clock steps

The pending-reply recovery sender check and the Herdr lab viewer ownership
check identified processes by ps lstart text, which on Linux is the
wall-clock-derived boot time plus start ticks. A host clock step (WSL2 steps
about every 30 seconds) re-renders it, so a live recovery sender read as dead
and a running lab viewer pair read as not owned.

Both now use /proc/<pid>/stat start ticks where readable, like fm_pid_identity
and task_process_identity, and keep the ps form elsewhere. Records already
written in the legacy lstart form are still compared through ps, so an
upgrade does not strand an in-flight recovery or a running viewer.
@knowttl
knowttl merged commit 29f8b97 into main Sep 23, 2026
36 of 37 checks passed
knowttl added a commit that referenced this pull request Sep 30, 2026
* fix(bin): keep process identity stable across host clock steps (#45)

* fix(bin): keep pending-reply sender and lab viewer identity stable across clock steps

The pending-reply recovery sender check and the Herdr lab viewer ownership
check identified processes by ps lstart text, which on Linux is the
wall-clock-derived boot time plus start ticks. A host clock step (WSL2 steps
about every 30 seconds) re-renders it, so a live recovery sender read as dead
and a running lab viewer pair read as not owned.

Both now use /proc/<pid>/stat start ticks where readable, like fm_pid_identity
and task_process_identity, and keep the ps form elsewhere. Records already
written in the legacy lstart form are still compared through ps, so an
upgrade does not strand an in-flight recovery or a running viewer.

* no-mistakes(document): Document clock-stable process identity and legacy records

* fix(bin): prevent Linux remote job worker pileups after clock changes (#46)

* fix(bin): keep Linux remote job worker identity stable across clock steps

The remote job worker identified its own processes (lock owner, staging
owner, job claims, lanes, command groups) by `ps -o lstart=` text. On Linux,
procps renders lstart from the current boot time, which moves whenever the
wall clock is stepped (NTP, VM or WSL2 time sync, resume). After a step a
healthy worker no longer matched its own lock record, so every remote call
started another detached supervisor beside it, the losers restarted for
minutes, the serving loop blocked on live lanes it thought had exited, a
competing worker reclaimed the live lock, and running jobs were published as
"remote job worker stopped before this job completed".

- Record Linux process identity as starttime=<stat field 22>, which no clock
  step moves; Darwin keeps ps lstart, unchanged.
- Keep records written by earlier workers comparable: an lstart record is
  compared as lstart, and a Linux lock owner still recorded as lstart is
  identified by pid and exact command, so an update replaces it in place and
  drains supervisors already piled beside it instead of stranding it.
- A serving worker that has lost its ownership lock now stops its own active
  execution and exits on a stop signal instead of re-arming, and never writes
  quarantine into a lock it does not own.

* no-mistakes(document): Document Linux remote worker identity and shutdown behavior
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant