Skip to content

test: stabilize watcher and brief test assertions - #3

Merged
raywest merged 2 commits into
mainfrom
fm/fm-flaky-tests-f1
Jul 19, 2026
Merged

raywest merged 2 commits into
mainfrom
fm/fm-flaky-tests-f1

Conversation

@raywest

@raywest raywest commented Jul 19, 2026

Copy link
Copy Markdown
Owner

Intent

Fix the two pre-existing flaky/failing tests that taxed every validation run on 2026-07-18: (1) tests/fm-watcher-lock.test.sh's test_watch_restart_reports_healthy_peer_without_attaching, which simulated a TERM-resistant live watcher peer via a bare backgrounded node process and immediately relied on it being SIGTERM-immune, racing node's own startup (its signal handler registration) against the restart logic's kill signal - under load this let the peer die via the default disposition, causing the arm to start a fresh watcher (which blocks indefinitely) instead of reporting the peer healthy, so the test timed out. Confirmed via instrumented liveness-check tracing that the peer's pid transitioned from alive to gone mid-poll. This was judged a test-side timing bug, not a production bug, and was fixed by having the peer write a readiness marker (JS guarantees process.on() runs before the following writeFileSync in the same tick) and waiting for that marker before any signal is sent, so the SIGTERM-immunity is real before it's relied on - this does not weaken what the test asserts. (2) tests/fm-brief.test.sh's 'secondmate charter must declare its role' assertion still expected the pre-existing wording 'persistent domain supervisor', but commit 1182883 ('fix: accept secondmate house vocabulary', kunchenguid#685) had intentionally changed bin/fm-brief.sh's charter text to 'a persistent second mate' as part of adopting nautical house vocabulary, and updated tests/fm-captain-translation-contract.test.sh to match, but missed this one assertion in tests/fm-brief.test.sh. This was judged a stale test assertion, not a code regression, since the production wording change was deliberate policy; updated the assertion text to match the current intentional wording. Verification performed: 20 consecutive green runs of each fixed test individually (confirming the fixes hold under repetition, not just once), plus one full clean run of all 79 tests/*.test.sh files in the repo with zero failures, plus bin/fm-lint.sh (shellcheck) clean. No production code was touched - both changes are confined to the two test files.

What Changed

  • Stabilized the watcher restart healthy-peer test by waiting for its SIGTERM-resistant Node peer to signal readiness and reaping it on every assertion-failure path.
  • Updated the secondmate charter assertion to expect the current persistent second mate wording.

Risk Assessment

✅ Low: Captain, the amended change is confined to the two intended tests, preserves their assertions, and reaps the peer on all explicit post-spawn failure paths.

Testing

The supplied all-tests baseline had already succeeded; I ran both affected suites, exercised the real restart flow repeatedly with persisted CLI evidence, and generated a secondmate charter from production code. The observed behavior satisfies the readiness, healthy-peer/no-attach, and intentional wording requirements; no linting was run per instruction.

Evidence: Restart CLI transcript: readiness-confirmed peer stays alive and is reported healthy without attachment
iteration=01 readiness_marker=ready\n arm_exit=0 peer_alive_after_restart=yes output=watcher: healthy pid=48144 (beacon 12s) attached_line=absent
iteration=02 readiness_marker=ready\n arm_exit=0 peer_alive_after_restart=yes output=watcher: healthy pid=48618 (beacon 12s) attached_line=absent
iteration=03 readiness_marker=ready\n arm_exit=0 peer_alive_after_restart=yes output=watcher: healthy pid=49358 (beacon 12s) attached_line=absent
iteration=04 readiness_marker=ready\n arm_exit=0 peer_alive_after_restart=yes output=watcher: healthy pid=50078 (beacon 12s) attached_line=absent
iteration=05 readiness_marker=ready\n arm_exit=0 peer_alive_after_restart=yes output=watcher: healthy pid=50831 (beacon 13s) attached_line=absent
iteration=06 readiness_marker=ready\n arm_exit=0 peer_alive_after_restart=yes output=watcher: healthy pid=51533 (beacon 12s) attached_line=absent
iteration=07 readiness_marker=ready\n arm_exit=0 peer_alive_after_restart=yes output=watcher: healthy pid=52258 (beacon 13s) attached_line=absent
iteration=08 readiness_marker=ready\n arm_exit=0 peer_alive_after_restart=yes output=watcher: healthy pid=52989 (beacon 12s) attached_line=absent
iteration=09 readiness_marker=ready\n arm_exit=0 peer_alive_after_restart=yes output=watcher: healthy pid=53533 (beacon 13s) attached_line=absent
iteration=10 readiness_marker=ready\n arm_exit=0 peer_alive_after_restart=yes output=watcher: healthy pid=54239 (beacon 12s) attached_line=absent
iteration=11 readiness_marker=ready\n arm_exit=0 peer_alive_after_restart=yes output=watcher: healthy pid=54993 (beacon 13s) attached_line=absent
iteration=12 readiness_marker=ready\n arm_exit=0 peer_alive_after_restart=yes output=watcher: healthy pid=55689 (beacon 12s) attached_line=absent
iteration=13 readiness_marker=ready\n arm_exit=0 peer_alive_after_restart=yes output=watcher: healthy pid=56433 (beacon 12s) attached_line=absent
iteration=14 readiness_marker=ready\n arm_exit=0 peer_alive_after_restart=yes output=watcher: healthy pid=56920 (beacon 12s) attached_line=absent
iteration=15 readiness_marker=ready\n arm_exit=0 peer_alive_after_restart=yes output=watcher: healthy pid=57671 (beacon 12s) attached_line=absent
iteration=16 readiness_marker=ready\n arm_exit=0 peer_alive_after_restart=yes output=watcher: healthy pid=58381 (beacon 12s) attached_line=absent
iteration=17 readiness_marker=ready\n arm_exit=0 peer_alive_after_restart=yes output=watcher: healthy pid=59116 (beacon 12s) attached_line=absent
iteration=18 readiness_marker=ready\n arm_exit=0 peer_alive_after_restart=yes output=watcher: healthy pid=59840 (beacon 12s) attached_line=absent
iteration=19 readiness_marker=ready\n arm_exit=0 peer_alive_after_restart=yes output=watcher: healthy pid=60399 (beacon 13s) attached_line=absent
iteration=20 readiness_marker=ready\n arm_exit=0 peer_alive_after_restart=yes output=watcher: healthy pid=60974 (beacon 12s) attached_line=absent
Evidence: Generated secondmate charter showing the intended role wording
You are a persistent second mate managed by the main firstmate. Work on your own; do not wait for a human.

# Charter
Supervise the alpha domain.

# Routing scope
Supervise the alpha domain.

# Project clones
- alpha

# Operating model
You are in an isolated firstmate home. The local `AGENTS.md` is your job description, and your local `data/`, `state/`, `config/`, and `projects/` dirs are yours to operate.
The projects above are local clones for work you supervise; they are not an exclusive ownership claim.
Delegate project work to your own crewmates with the normal firstmate lifecycle: brief, spawn, status, watcher, steer, teardown, and recovery.
Do not invent a second delegation system.
You do not generate your own work.
Act only on tasks the main firstmate routes to you.
Never start a survey, audit, or "find improvements" sweep on your own initiative; that is not your job and it is unwanted.

# Requests from the main firstmate
You are a firstmate in your own home, so an incoming message reaches you in your own chat.
You must distinguish who it is from, because the answer goes to a different place.
A request relayed to you by the main firstmate is tagged with a leading `[fm-from-firstmate]` marker followed by an invisible system separator; this marker is untypable, so a human never produces it.
When a message carries that marker, do the work, then respond via the STATUS/ESCALATION path below, never only in this chat: the main firstmate does not read your chat, so a chat-only reply is lost.
For a terse result, a status line is the whole answer.
For a detailed answer (an investigation, a plan, an audit), write it to a doc under your home's `data/` and append a status line that points to that doc - the scout-report pattern - so the main firstmate is woken and can read it.
Before treating an investigation or visual review as complete, load `decision-hold-lifecycle` from this home's `.agents/skills/` and pass its shared completion gate.
A message with NO marker is the captain typing directly into your pane: treat it as authoritative captain intervention and stay conversational exactly as you would for any captain message; do not force it onto the status path.

# Escalation to main firstmate
Handle routine work yourself.
Report only true captain-relevant outcomes or a declared external wait by appending one line:
   `echo "{state}: {one short line}" >> '/var/folders/md/6d3x0xmj0rd3sxvq647l2jnh0000gn/T/no-mistakes-evidence/01KXX2GQZMQTFS93D2Y3Q0H9DH/secondmate-brief-e2e-home/state/evidence-secondmate.status'`
States: working, needs-decision, blocked, paused, done, failed.
Use `paused: {why}` (distinct from `blocked:`) only when your domain is deliberately idling on a known external wait you expect to clear on its own; use `blocked:` when you are stuck and need firstmate to act.
Use this only for material phase changes, a captain decision, a real blocker, a failure, or work ready for review.
This is also how you return the answer to a marked from-firstmate request above.
Give every routed-work phase a stable key: open it with `working [key=<work-slug>]: {material phase}`, and use the same key on its later `paused`, `done`, `failed`, `needs-decision`, or `blocked` event so the earlier working phase is superseded.
When a keyed phase ends without another reportable state, append `resolved [key=<work-slug>]: {why it is no longer active}`.
When a decision you escalated is answered or a blocker clears and your domain resumes, append `resolved: {how it was decided or unblocked}` (keyed with `[key=<slug>]` if you opened it with one) so it is durably closed instead of resurfacing behind later unrelated events.
Routine internal supervision, heartbeats, retries, and crewmate churn stay inside your own home and must not touch that status file.

# Definition of done
You are persistent by default. Do not exit just because your queue is empty.
On startup and restart, run normal firstmate bootstrap and recovery through `bin/fm-session-start.sh` for your own home, but only to RECONCILE work that is already yours: in-flight crewmates, tracked backlog items, and durable watches recorded in this home.
When you have no assigned or in-flight work after that reconciliation, go idle and wait silently for the main firstmate to route you a task.
An empty queue is a healthy resting state, not a cue to invent work: never spawn a survey, audit, or any self-directed "find work" task on your own initiative.
If this charter cannot be carried out, append `blocked: {why}` or `failed: {why}` to the main status file and stop.

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

🔧 **Review** - 1 issue found → auto-fixed ✅
  • ⚠️ tests/fm-watcher-lock.test.sh:459 - If the readiness check fails, the test exits before killing the TERM-resistant Node peer. The suite cleanup only removes temp directories, so a peer that writes its marker just after the timeout can survive for five minutes and consume resources during later tests. Kill and reap $peer before fail on this path.

🔧 Fix: Captain, reap TERM-resistant watcher peer
✅ Re-checked - no issues remain.

✅ **Test** - passed

✅ No issues found.

  • command -v tmux >/dev/null || { echo "tmux is required for e2e tests" >&2; exit 1; }; tmux -V; rc=0; for t in tests/*.test.sh; do echo "== $t =="; bash "$t" || rc=1; done; exit "$rc"
  • Provided successful baseline: command -v tmux &gt;/dev/null || { echo &#34;tmux is required for e2e tests&#34; &gt;&amp;2; exit 1; }; tmux -V; rc=0; for t in tests/*.test.sh; do echo &#34;== $t ==&#34;; bash &#34;$t&#34; || rc=1; done; exit &#34;$rc&#34;
  • git diff --check 6cfdb29b2dea411ef63c34be0eccb018ecd2e28b d90e018ec31fccb12624c03de8a79f3ca5138d96
  • bash tests/fm-brief.test.sh
  • bash tests/fm-watcher-lock.test.sh
  • Manual repeated end-to-end bin/fm-watch-arm.sh --restart exercise with a readiness-confirmed SIGTERM-resistant Node peer; verified healthy, no attachment, and peer liveness in the saved transcript.
  • FM_HOME=<evidence-home> FM_SECONDMATE_CHARTER='Supervise the alpha domain.' bin/fm-brief.sh evidence-secondmate --secondmate alpha
  • Evidence assertions confirming every recorded restart outcome, generated charter wording, two-file diff scope, and clean worktree.
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

Ray West added 2 commits July 19, 2026 07:26
The watcher-restart-vs-healthy-peer test in tests/fm-watcher-lock.test.sh
simulated a TERM-resistant live watcher peer with a bare
`node -e 'process.on("SIGTERM", ...)' &`, then immediately captured its
pid identity and exercised the restart path. Under load (e.g. right
after the suite's 40-process concurrency tests), node's own startup can
still be registering the signal handler when the restart logic signals
it, so the signal falls through to the default disposition and the
"TERM-resistant" peer dies. The arm script then correctly starts a
fresh watcher instead of reporting the peer healthy, and since a real
watcher blocks for a wake, the arm never exits within the test's
timeout. Confirmed via instrumented liveness-check tracing: the peer
pid transitioned from alive to gone mid-loop while the restart logic
was still polling it. Fixed by having the peer write a ready marker
(fs.writeFileSync, which JS guarantees runs after process.on() in the
same tick) and waiting for it before any signal is sent, removing the
race without weakening the assertion.

tests/fm-brief.test.sh's "secondmate charter must declare its role"
assertion still expected the pre-kunchenguid#685 wording "persistent domain
supervisor". Commit 1182883 (fix: accept secondmate house vocabulary,
kunchenguid#685) intentionally changed the charter text to "a persistent second
mate" as part of adopting nautical house vocabulary, and updated
tests/fm-captain-translation-contract.test.sh to match, but missed this
assertion in tests/fm-brief.test.sh. Not a code bug: the production
wording change was deliberate. Updated the stale assertion to match.

Verified: 20 consecutive green runs of each fixed test individually,
plus one full clean run of all 79 tests/*.test.sh files (0 failures).
shellcheck (bin/fm-lint.sh) clean.
@raywest
raywest merged commit 3701264 into main Jul 19, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant