Skip to content

test: stabilize watcher lock race fixtures - #6887

Merged
kunchenguid merged 2 commits into
kunchenguid:mainfrom
ironerumi:fm/fm-6455-upstream-pr
Oct 8, 2026
Merged

kunchenguid merged 2 commits into
kunchenguid:mainfrom
ironerumi:fm/fm-6455-upstream-pr

Conversation

@ironerumi

@ironerumi ironerumi commented Oct 8, 2026 •

Copy link
Copy Markdown
Contributor

Summary

Closes #6455.

Only tests/fm-watcher-lock.test.sh changes; the lock implementation and other tests remain unchanged.

The stale-lock concurrency fixture now combines 39 nonblocking contenders with one fm_lock_acquire_wait contender. A FIFO reports the actual winner, and a second FIFO holds that winner alive until the other contenders finish, replacing the timing-dependent one-second hold.

The competing-reaper fixture writes a verified-dead PID after creating its abandoned lock, allowing the tested steal-mutex API to enter its reaping path. This ensures PID reuse cannot bypass the intended race hook without introducing a duplicate direct reap call.

Evidence

Validation used upstream main at b062eb94c50a68dd4467fba25346a7d7b4083288 on macOS. Full-suite repetitions ran in groups of four alongside two CPU burners. Both reported failures were also reproduced with forced interleavings in temporary copies of the original and corrected cases:

  • A live steal-mutex holder remained in place until every nonblocking contender had completed its attempt, then released the mutex. The original fixture had no winner; the corrected waiting contender acquired it after release.
  • The killed fixture owner's PID was replaced with a known-live PID to model reuse before the next liveness probe. The original race hook never fired; the corrected fixture overwrote that PID with a verified-dead one and exercised the hook.
Validation Original Corrected
Full watcher-lock suite under identical CPU load, before removal of the redundant reap call 20/20 passed 20/20 passed
Additional corrected full-suite runs during concurrent forced-case stress Not run 19/20 passed; both changed cases passed 20/20
Forced steal-mutex contention under CPU load 20/20 failed with expected exactly one stale-lock stealer, got 0 20/20 passed
Forced owner-PID reuse under CPU load 20/20 failed with reap race hook never fired 20/20 passed

After review removed the redundant direct reap call, both forced races passed another 20/20 repetitions each against the final head d834b26519de4d7420cbf19fde73d25023bdeb83 with two persistent CPU burners. Independent validation also passed the two changed cases 20/20 under that load and the full suite once without load; exploratory loaded full-suite repetitions encountered existing timing failures in unchanged arm/child-cleanup cases, which are not claimed as passing evidence.

Unforced loaded repetitions did not reproduce either reported intermittent failure on this machine; the forced interleavings established the before/after difference. One corrected full-suite repetition failed the unchanged dead-nested-steal-chain case (rc=8) while the separate forced-interleaving batch was also running. That case is outside this test-only fix's scope; a subsequent independent batch with only the original CPU load passed all 20 full-suite repetitions. The canonical targeted runner also passed (total=1 failed=0, 176.6 seconds). Bash syntax checking, git diff --check, and the pinned ShellCheck/actionlint checks also passed.

Published validation logs

Pipeline

Updates from git push no-mistakes

✅ **Intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

🔧 **Review** - 1 issue found → auto-fixed ✅
  • ⚠️ tests/fm-watcher-lock.test.sh:550 - Simplification: the new fm_lock_reap_dead_link "$2" call at tests/fm-watcher-lock.test.sh:550 duplicates the same reaping step already performed by fm_lock_try_acquire_steal_mutex for this argument (bin/fm-wake-lib.sh:1159; the nested .steal.steal check at line 1155 does not change that). With $2 equal to .contend.lock.steal, removing line 550 still drives the fake rm hook through the tested API while preserving the verified-dead PID fix. The direct call is not required by the stated intent and adds an unnecessary parallel state transition; remove it as the smallest simplification.

🔧 Fix applied.
✅ Re-checked - no issues remain.

⚠️ **Test** - 1 warning
  • ⚠️ tests/fm-watcher-lock.test.sh:1 - An exploratory loaded full-suite loop encountered pre-existing timing failures in unchanged arm/child-cleanup cases before the loop timed out; the two changed cases and the forced stale-steal interleaving remained green in isolated 20/20 runs.
  • Live validation: ✅ go - 3 of 3 scenarios driven live against the product
Scenario Result Live Evidence
Verified-dead owner PID enables successor-safe steal reaping ✅ pass live Changed watcher-lock cases: 20/20 under CPU load
Waiting contender wins a forced stale-steal interleaving ✅ pass live Forced stale-steal interleaving: 20/20 under CPU load
Watcher-lock suite remains green ✅ pass live Canonical watcher-lock suite
  • timeout 600 tests/fm-watcher-lock.test.sh
  • Changed watcher-lock cases invoked 20 times with two persistent yes >/dev/null CPU burners
  • Forced real-lock interleaving invoked 20 times with two persistent yes >/dev/null CPU burners
  • git diff --check b062eb94c50a68dd4467fba25346a7d7b4083288 d834b26519de4d7420cbf19fde73d25023bdeb83 -- tests/fm-watcher-lock.test.sh
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

Issues

Closes #6455

@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate (OSS triage, not the captain).

I read the thread, issue #6455, and the diff against main b062eb94. I did not take the PR's own claims on trust.

What the tip does (my read): only tests/fm-watcher-lock.test.sh changes. test_lock_steal_reap_cannot_remove_successor now writes a verified-dead PID (dead_pid) into the abandoned steal owner, so PID reuse can't send the race down the liveness path. test_lock_stale_steal_single_winner_under_concurrency swaps one of its 40 one-shot fm_lock_try_acquire contenders for a waiting fm_lock_acquire_wait contender. A FIFO reports the winner, and a second FIFO keeps it alive until the other contenders finish, replacing the old sleep 1. The lock implementation is unchanged. This matches the done criteria in #6455.

Contract-class: restore. These are existing default CI fixtures for the existing lock contract, and they were failing intermittently because of how the fixtures were built. The tip makes their steps run in an explicit order and adds no product surface.

VISION.md, rule by rule

  • One captain, one interface: n/a.
  • Authority is explicit: n/a.
  • Scripts own the mechanics: aligns. Lock behavior stays covered by deterministic tests.
  • A restart is a non-event: n/a.
  • Delegation with a spine: aligns. Flaky gates weaken validation, and fixing the fixture strengthens it.
  • The fleet outlives any vendor: n/a.
  • Scope: aligns. This is test-only, and it turns a field/CI incident into stable regression coverage.

Gate: attestation MATCH d834b265. no-mistakes is green, and every CI job on this head is green. MERGEABLE/CLEAN. Security: none (test-only, no workflow or secrets changes). Fork CI didn't need approval on this pass.

Non-blocking note: if a future regression ever let two contenders win, the second winner would block on the gate FIFO and the case would hang until the shard timeout rather than fail fast. That's acceptable for this fix.

Merging as restore.

@kunchenguid
kunchenguid merged commit 19fcbbd into kunchenguid:main Oct 8, 2026
20 checks passed
@kunchenguid

Copy link
Copy Markdown
Owner

Speaking as Kun's firstmate: this is merged. Thank you @ironerumi — really appreciate you taking the time on this.

thechrisfischer added a commit to thechrisfischer/firstmate that referenced this pull request Oct 9, 2026
…up pull (#2)

* test: stabilize watcher lock race fixtures (kunchenguid#6887)

* test: order stale watcher-lock fixture races explicitly

* no-mistakes(review): Removed duplicate stale-steal reap call

* Add guarded primary startup pull coordination

* no-mistakes(review): Fix strict inventory, guard-write scope, profile launch_env parity

* no-mistakes(document): Refresh startup-pull verification record for review fixes

---------

Co-authored-by: YifuGu <39033099+ironerumi@users.noreply.github.com>
nathanjgaul-agile added a commit to nathanjgaul-agile/firstmate that referenced this pull request Oct 9, 2026
* feat(bin): defer spawns beyond a declared per-project capacity, opt-in (kunchenguid#5343)

* refactor(bin): share the local Firstmate home walk from the wake library

Teardown's walk over the root home and its registered local secondmate homes
moves into bin/fm-wake-lib.sh as fm_local_firstmate_state_dirs, next to
fm_firstmate_root_home, so a second consumer can count task records across
this machine's homes without a copy. Teardown keeps its exact refusal wording
through a thin wrapper.

* feat(bin): defer spawns beyond a project's declared machine capacity

A project whose machine-local resource only serves a few workers at once had
no way to tell Firstmate so: every queued item was launched, and the surplus
workers spent full-context turns retrying the resource.

config/project-capacity in the root home now declares how many workers each
named project admits at once on this machine. bin/fm-spawn.sh counts the ship
and scout records on the same project origin across the root and its local
secondmate homes, skipping ones whose ready PR is recorded, while holding the
shared project lock through publication. A spawn with every place held exits 75
before any brief render, endpoint, worktree, record, or backlog move, so the
item stays queued; batches report it as deferred. Undeclared projects keep
today's uncapped dispatch, and an unreadable declaration refuses rather than
guessing the limit.

Refs kunchenguid#4237

* no-mistakes(review): Document that capacity matches the clone directory name

* no-mistakes(document): Rewrap stale fm-wake-lib root-home doc comment

* no-mistakes(review): Dedupe local state dirs by identity to avoid double-counting

* no-mistakes(document): Rewrap fm_local_firstmate_state_dirs error doc comment

* no-mistakes(ci): I fixed all four Greptile findings. All 14 tests in tests/fm-project-capacity.test.sh pass, and shellcheck at warning level is clean on the changed files. Each new test failed against the old code and passes now. - **ci-1 (spaced names):** a declaration line must give a name its capacity whenever the name is a valid clone directory name. `fm_project_capacity_lookup` now trims each line, skips blank lines and lines whose first non-blank character is `#`, and takes the last field as the capacity. Everything before that field is the name, so it may contain spaces. The old error cases still refuse: a single field is rejected, and trailing text leaves a last field that is not an integer. The library header and docs/configuration.md now say a name starting with `#` cannot be declared. New test `test_spaced_project_name_is_declared` declares `my heavy project 1` next to an indented comment line and gets a deferral. - **ci-2 (unreadable records):** the holder count must never silently leave out a holder. `fm_project_capacity_occupants` now refuses when a local home's state directory exists but cannot be read or listed, or when a `.meta` file cannot be read. The error names the path, and `fm-spawn.sh` shows it in its existing refusal message. New test `test_unreadable_holders_refuse_admission` covers an unreadable record in the root home and an unreadable state directory in a registered local secondmate home, then checks that the spawn is admitted once both are readable. The test is skipped when run as root. - **ci-3 (Orca lock):** any spawn that can become a holder for a capped project must take that project's lock. The lookup now also reports whether the declaration caps any project at all, and an Orca spawn takes the per-origin lock whenever it does. This covers every capped same-origin clone. It also covers some cases where no same-origin clone is capped, because a spawn cannot find clones under other directory names without searching for them. With no declaration file, Orca still skips the lock. The comments in the library and in the `fm-spawn.sh` header are updated. The Orca test now clones the origin as `project-2`, which has no declaration, and checks that its Orca spawn refuses while the lock is held and publishes no record. - **ci-4 (worktrees):** `assert_nothing_created` now also compares the project's `git worktree list` from before and after a deferred spawn. Both tests that call it take that snapshot first. Files changed: bin/fm-project-capacity-lib.sh, bin/fm-spawn.sh, docs/configuration.md, tests/fm-project-capacity.test.sh

* fix(bin): declare capacity for a project name that begins with #

A clone directory whose name begins with # was skipped as a comment, so that project stayed uncapped. A line is a declaration when the # is written against the rest of the name and the line ends with a capacity; a # followed by whitespace stays a comment.

* no-mistakes(document): Rewrap project-capacity library header comment

* no-mistakes(ci): Lint 2 fails because this PR's code pushes ShellCheck past its memory cap. ShellCheck ran out of memory analyzing bin/fm-teardown.sh in CI (reason=memory, rc=251, peak about 8.39 GB). On current main the same file passes at about 7.29 GB. **Cause:** the new `fm_local_firstmate_state_dirs` function in bin/fm-wake-lib.sh had a conditional `. fm-secondmate-registry-lib.sh` with a `# shellcheck source=` directive inside the function. ShellCheck followed that source again, inside a function scope, wherever fm-wake-lib.sh is sourced, and bin/fm-teardown.sh is the heaviest root that sources it. Measured locally with `shellcheck --norc --external-sources bin/fm-teardown.sh`: - current main (fd325b1): 7.29 GB - main merged with this PR: 7.86 GB - the same merge without the in-function source: 7.27 GB **Rule this restores:** this change must not make any lint root heavier than it is on main. That function holds the only new nested source in the change. **Fix:** I removed the in-function source, which no caller needs. Both callers already load the registry library at top level before calling the function: - bin/fm-teardown.sh sources it directly. - bin/fm-spawn.sh, the only user of bin/fm-project-capacity-lib.sh, gets it through bin/fm-ff-lib.sh. I also documented the requirement in the function's comment and in the "Requires" note in bin/fm-project-capacity-lib.sh. No behaviour changes. **Verification:** - ShellCheck on head: bin/fm-teardown.sh peaks at 7.12 GB and bin/fm-spawn.sh at 6.68 GB, both with rc=0. bin/fm-wake-lib.sh and bin/fm-project-capacity-lib.sh lint clean. - tests/fm-project-capacity.test.sh, tests/fm-teardown.test.sh (102 ok) and tests/fm-teardown-endpoint-safety.test.sh all pass. Files changed: bin/fm-wake-lib.sh, bin/fm-project-capacity-lib.sh

* fix(bin): release the Herdr session lock when reclaim finishes

A concurrent resume in another home waits five seconds for that lock.
Reclaim is the last presentation change on the recovery path, so holding
the lock through the launch tail made the waiter time out. The contributions
arm check also freezes its one-second clock, the same way the budget tests
do, because an unfrozen clock can tick past before the first forge read.

* no-mistakes(review): Keep Herdr session lock through launch handoff after reclaim

* no-mistakes(review): Skip the spawning task's own record in capacity count

* no-mistakes(review): Restore release test comment above its test

* docs: scope PR-ready re-evaluation to a declared project capacity

A ready pull request frees a place only when that project declares capacity, so the always-loaded backlog contract should re-evaluate on that handoff only in that case.

* fix(bin): name the Lavish read message count by the same label as its section (kunchenguid#6785)

fm-procevent-lavish.sh read labels tag=message rows SESSION-ENDING MESSAGE
only when session_ended is true and CAPTAIN MESSAGE otherwise, but the
count line always said session_ending_message_count. Several composer
messages on a still-open board were therefore counted as session-ending.

The count line now follows the same session_ended switch:
session_ending_message_count once the session ended, captain_message_count
otherwise. Message rows stay out of the annotation count, per triage.

Fixes kunchenguid#6743

* fix(herdr): make the presentation lock namespace per OS account (kunchenguid#6780)

The Herdr presentation lock namespace was the fixed machine-global
/tmp/firstmate-herdr-presentation, so on a host where two OS users run
Firstmate on Herdr the first account to create it owned it and every
teardown from the other account was refused with no way to clear it.

Suffix the namespace with the account uid. The owner-uid and mode-700
checks are unchanged, so a foreign-owned or wrong-mode name at this
account's path is still refused and never adopted, chowned, or removed.

Fixes kunchenguid#4716.

* Fix OpenCode arm plugin to decide with the shared supervision predicate (kunchenguid#6809)

The OpenCode session plugin's shouldArm kept its own copy of the need
test that only looked for in-flight task records, while the turn-end
guard decides with fm_supervision_needed in bin/fm-supervision-lib.sh,
which also counts registered process-event sources and trusted custom
checks. With an empty fleet but any registered source or check, the
guard blocked every turn end while the plugin declined to arm - a loop
the guard's own repair line could not resolve because it names the
plugin as the fix.

The plugin now delegates the decision to the shared predicate through
bash, keeping the local away-record decline and the x-mode.env arm
override. OpenCode plugin test fixtures now carry the real predicate
their arming path sources, and the arm suite gains six cases asserting
the plugin's decision against the shared verdict over the same
synthetic state directories.

Co-authored-by: Mia Sun <mia@Bigs-Mac-mini.localdomain>

* fix(bin): resolve a pending reply only from its own task's status line (kunchenguid#6792)

* fix(bin): resolve a pending reply only from its own task's status line

Remote reply ingestion handed every corr= token in a mate's payload to
fm_pending_reply_try_resolve together with that mate's own status log,
so one mate echoing another mate's token resolved the other request.
Honor a status-file override only when it is the record's own
parent_status, and match the corr= token as a whole word.

Fixes kunchenguid#6538

* no-mistakes(document): docs: scope remote reply settlement to the asked mate

* docs(skills): index the six missing agent-only skill triggers (kunchenguid#6784)

agent-skill-trigger-index claims to be the complete agent-only trigger
index but omitted operational-home-layout, session-start-recovery,
validation-supervision, ship-landing, scout-completion, and
away-quiet-supervision. Add each with its own description's trigger,
placed beside the related entries.

The decision-hold-lifecycle redirect stub stays out, per triage.

Fixes kunchenguid#6503

* test: make agent process fixtures compatible with multicall sleep (kunchenguid#6814)

* test: share a rename-safe agent stand-in across liveness suites

On Ubuntu 26.04, `sleep` is the uutils multicall binary, which refuses to
run when invoked through a symlink named after another utility. The Herdr
descendant process-walk tests built their agent-named process as a `pi`
symlink to the host `sleep`, so the process exited at once, its parent shell
was gone before the walk ran, and both cases read `unknown unreadable` and
failed on that host. The suite stops at its first failure, so every later
case went unrun. The Herdr control smoke test's `claude` symlink has the
same construction.

The tmux liveness suite already solved this with a host-compiled spinner and
a survival-checked `sleep` fallback. That builder moves into tests/lib.sh as
fm_agent_standin, and the tmux suite, both Herdr descendant cases, and the
Herdr control smoke test now use it. When no stand-in can survive a foreign
name, a case skips with the reason instead of failing.

tests/fm-test-fixtures.test.sh gains a portable regression with a fake
multicall `sleep`, so it bites on hosts whose own `sleep` is single-purpose.

* no-mistakes(document): Correct Herdr verification fixture reference

* ci: retrigger cancelled shard

* fix(bin): keep the supervision host's successor watcher alive after the Stop hook's group is torn down (kunchenguid#6787)

* fix(bin): keep the supervision host's pass-through successor out of the hook's process group

The successor a main-only pass-through leaves for main shared the Stop hook's
process group, so the harness tearing that group down after the exit-2 rewake
stopped it. The stop published downtime and the next park's first cycle
announced an empty check: rearm-resurface, which woke main again in a loop.
Start that successor in a process group of its own, as the hook's own
handling successor already is.

* no-mistakes(review): Give the at-turn successor left for main its own group

* no-mistakes(document): Document own-group successor for turn-start hand-back too

* no-mistakes(ci): I made the change you asked for: both new teardown tests in tests/fm-supervision-host.test.sh now call the existing `stop_home_processes "$home"` just before `pass`. The tests are `test_successor_left_at_the_turn_survives_the_hook_process_group_teardown` and `test_pass_through_successor_survives_the_hook_process_group_teardown`. No production code and no other tests changed. The rule broken was that a test must not leave a home's watcher or arm processes running after it passes. These two were the only cases in the changed area that broke it. The other host+hook tests already stop their home, and `test_successor_close_during_main_turn_is_delivered_at_the_next_turn_end` leaves its watcher behind too, but it is an older test you said not to touch. The only reason anything was left over is that the successor's arm now sits in its own process group, outside the hook's teardown. `stop_home_processes` kills the watcher by the pid in its lock file, which stops it no matter which group it is in. **Checks run:** - I ran just these two tests from a scratch copy of the suite (since deleted). Both pass in about 13 seconds. - After each test, a process listing filtered to that test's home directory came back empty once the processes had about a second to exit after TERM. - `bash -n` on the test file passes. - `shellcheck` is not installed here, so I did not lint the file. - I did not run the full serial-2 suite locally. Whether it now finishes under its 30-minute limit will only show on the next CI run

* test: make tmux Claude readiness checks independent of permission footers (kunchenguid#6823)

* test: use idle composer readiness for Claude tmux guards

* no-mistakes(test): Fix attended supervision test expectations and isolate worker state

* no-mistakes(document): Correct live guard coverage and readiness documentation

* no-mistakes(ci): Captain, fixed SC2100 by quoting the cursor-agent assignment in tests/fm-host-mirror-live-e2e.test.sh. Reproduced the failure before editing; pinned ShellCheck lint on both PR test files, bash syntax checks, and git diff --check now pass

* test: preserve attended successor close assertions

* no-mistakes(test): Fix attended live test watcher takeover expectations

* no-mistakes(document): Correct stale attended guard documentation

* Revert "no-mistakes(document): Correct stale attended guard documentation"

This reverts commit 8e59d89.

* Revert "no-mistakes(test): Fix attended live test watcher takeover expectations"

This reverts commit c0b8510.

* test(herdr): accept safe exit refusals and clean lab once (kunchenguid#6818)

* test: stabilize watcher lock race fixtures (kunchenguid#6887)

* test: order stale watcher-lock fixture races explicitly

* no-mistakes(review): Removed duplicate stale-steal reap call

* Rerun CI after the merge

---------

Co-authored-by: Tiago <tiagop@hey.com>
Co-authored-by: yairtech <39274208+falkoro@users.noreply.github.com>
Co-authored-by: Asser AboElkhair <asser.aboelkhair@gmail.com>
Co-authored-by: dubiousenvelope <greg.ecklin@gmail.com>
Co-authored-by: Mia Sun <mia@Bigs-Mac-mini.localdomain>
Co-authored-by: Pedro Guimarães <21346846+0x7067@users.noreply.github.com>
Co-authored-by: menidi <menidi@users.noreply.github.com>
Co-authored-by: YifuGu <39033099+ironerumi@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

Flaky fm-watcher-lock stale-steal tests: one-shot contenders can all lose, and the dead-owner fixture PID can be reused

2 participants