Skip to content

test: stop leaked supervision-host cycles per case and widen the remote-reply recapture wait - #27

Merged
cloud-practitioner merged 1 commit into
mainfrom
fm/fm-remote-reply-supervision-host-tests
Oct 1, 2026
Merged

cloud-practitioner merged 1 commit into
mainfrom
fm/fm-remote-reply-supervision-host-tests

Conversation

@cloud-practitioner

@cloud-practitioner cloud-practitioner commented Oct 1, 2026 •

Copy link
Copy Markdown
Owner

Intent

Fix the test failures in tests/fm-remote-reply.test.sh and the supervision-host tests (tests/fm-supervision-host.test.sh and, if they also fail, tests/fm-supervision-host-live-e2e.test.sh and tests/fm-supervision-host-attended-live-e2e.test.sh) in the fork https://github.com/cloud-practitioner/firstmate. Context: while validating the upstream-merge branch, these tests failed both on the fork's unmerged main and on upstream https://github.com/kunchenguid/firstmate main in this environment, so they are pre-existing failures rather than merge breakage; one supervision-host run hit its 900-second timeout. A standing rule says test failures found along the way get fixed. Establish the cause first, fix it in the fork, and note whether the fix belongs upstream as well.

What Changed

  • tests/fm-supervision-host.test.sh: adds a run_case wrapper that runs each test function and then calls stop_home_processes for every home that case added to $HOMES_FILE. All 65 cases now run through it. Before this, a case that finished with its host parked or a successor watcher still running left that FM_POLL=1 cycle alive until the suite's EXIT trap. Those cycles piled up from case to case and pushed later cases past their fixed wait_until budgets. In an instrumented run the suite took 1558s, well past the 900s cap.
  • tests/fm-remote-reply.test.sh: raises the await_reply_result polling bound from 800 to 2400 iterations, so the 40s budget becomes 120s. This makes room for the cursor-loss whole-log recapture, which re-fetches every document one remote job at a time and runs slower since upstream fix: reduce remote-job and supervision polling churn kunchenguid/firstmate#6255. A healthy wait still returns as soon as the capture is applied. A new comment records why the bound is this size.
  • Only test files change; product scripts under bin/ are untouched. Both fixes also apply upstream. When upstream is merged in, the call-list hunk will conflict with upstream's added test_park_exit_probe_uses_half_second_child_sleeps line. Resolve it by prefixing that line with run_case.

🤖 Generated with Claude Code

Causes and evidence

Both failures come from timing assumptions in the tests that a busy host exposes.
Neither is a product bug in bin/ or a missing tool: tmux and ruby are absent here, but neither suite uses them, and both live e2e variants are opt-in and skip without their opt-in variables.
Upstream CI at f593060 passes both suites (supervision-host 683s, remote-reply 184s) on runners with no other work.

Supervision-host: leaked hosts and watchers.
Most cases end with their host parked (the default park bound is 27000s) or with a successor watcher still polling at FM_POLL=1, and only the EXIT trap stopped them.
An instrumented run of the unfixed suite counted 17 live watchers and 8 live hosts by the last case.
Over the run, load average rose from 1.5 to 25.8 on this 4-core host, cases ran about 3x slower (latch 35s -> 131s, attended latch 37s -> 159s), and the suite took 1558s, past the 900s cap.
When other lanes added load, closes ran past the fixed 15-25s wait_until budgets.
Observed failures were "latch: the later failed probe did not hand the wake back" (fork main fc795e5, 1482s), "away grok: the wake was not handled on the engine" (upstream-merge branch), and "the watcher's downtime resurface did not reach main".
With run_case and the same conditions, no watchers leaked between cases, load stayed between 2 and 9, and the suite passed in 558s.
It also passed in 659s through bin/fm-test-run.sh on fork main with other tests running, and in 731s on the upstream f593060 tree with the same change.

Remote-reply: the recapture wait after upstream kunchenguid#6255.
After a cursor loss, the whole-log recapture reads the log once and then fetches 13 documents, one remote job at a time.
Upstream kunchenguid#6255 (549e07f) made each sequential remote job slower to be picked up and sampled:

  • The dispatcher burst dropped from 20 passes to 4 before the 1s quiet scan.
  • Active and result sampling went from 0.05s to 0.25s.
  • Delta polling went from 0.2s to 0.5s.

Two runs side by side under the same load timed that step at 36s on f593060 (about 2.5s per fetch).
With 549e07f's bin/ changes reverted, it took 20s (about 1.1s per fetch).
The old await_reply_result budget was 40s.
"The replay-identity whole-log recapture was not captured" failed only on trees that include kunchenguid#6255.
Fork main, which does not have kunchenguid#6255 yet, still passes, but that step already took 30s of its 40s under load.

Upstream: both fixes belong upstream as well.
Upstream main has the same leaked processes and the same 40s budget after kunchenguid#6255, and its CI passes only because nothing else runs on its runners.

Risk Assessment

✅ Low: The change only touches test code. It adds a per-case cleanup wrapper that reuses the existing idempotent stop_home_processes over just the homes each case registered, and it widens one success-only polling bound. No product code changed, no caller of await_reply_result expects a timeout, and no case depends on processes left running by an earlier case.

Testing

I ran the changed suites directly on this 4-core host while other work was also using it. The supervision-host suite passed 65/65 in 599s with at most 3 hosts and 1 watcher alive at once. Base fc795e5 reproduced the pile-up: after 450s it had finished only 36 cases, with 11 hosts, 12 watchers and load average 17.5. The remote-reply suite passed 33/33 in 166s, and an instrumented copy showed the two whole-log recapture waits took 19.6s and 21.9s against the new 120s bound. Both live e2e suites skip cleanly without their opt-in variables. Separately, I found an existing product temp-directory leak, which is reported as a finding. There is no UI surface, so I captured no screenshots. The evidence is CLI logs and per-10s process-count tables.

  • Live validation: ✅ go - 4 of 4 scenarios driven live against the product
Scenario Result Live Evidence
Full tests/fm-supervision-host.test.sh run at HEAD passes every case within the 900s per-command cap ✅ pass live supervision-host-head.log: 65 'ok' lines, 0 'not ok', rc=0 elapsed=599s
Adversarial: hosts and watchers left behind by one case no longer pile up across later cases (base reproduces the pile-up) ✅ pass live supervision-host-head-procs.tsv: peaks of 3 hosts, 1 watcher, load 5.59 over the whole run. supervision-host-base-procs.tsv: base fc795e5 reached 11 hosts, 12 watchers, load 17.49 and finished only 36…
tests/fm-remote-reply.test.sh passes at HEAD, and the whole-log recapture finishes well inside the widened await bound ✅ pass live remote-reply-head.log: ALL TESTS PASSED, rc=0, 166s. remote-reply-await-timings.tsv: replay-identity recapture 19.6s, cursor-loss recapture 21.9s, all others <=5.6s, bound now 120s
The live e2e supervision-host variants skip cleanly without their opt-in variables, so they are not a source of the failures ✅ pass live live-e2e-gate.log: both print 'skip: live: opt-in; set FM_SUPERVISION_HOST_..._LIVE_E2E=1 to run' and exit 0; neither file changed between base and HEAD
Evidence: Evidence summary (HEAD vs base table, recapture timings)

Source: Evidence summary (HEAD vs base table, recapture timings)

# Test-phase evidence: supervision-host leak fix and remote-reply recapture bound

Host: 4 cores, WSL2, other lanes active (load average 3-7 when the runs started).

## tests/fm-supervision-host.test.sh
| tree | cases | wall time | peak live fixture hosts | peak live fixture watchers | peak load1 |
|---|---|---|---|---|---|
| HEAD e389c56 (run_case per-case stop) | 65/65 ok, rc=0 | 599s (< 900s cap) | 3 | 1 | 5.59 |
| base fc795e5 (stopped after 450s) | 36/65 done in 450s | would exceed 900s at that pace | 11 | 12 | 17.49 |

Per-10s samples: supervision-host-head-procs.tsv, supervision-host-base-procs.tsv
(only processes whose FM_HOME is under the suite's fm-supervision-host.* fixture root).

## tests/fm-remote-reply.test.sh (HEAD)
33/33 ok, ALL TESTS PASSED, 166s. Instrumented copy timed every await_reply_result:
the two whole-log recapture waits took 19.6s (replay-identity, orig line 678) and
21.9s (cursor-loss recapture, orig line 942); all others <= 5.6s. New bound 120s.

## Live e2e variants
Both gate-skip without their opt-in env and exit 0 (live-e2e-gate.log); unchanged by this change.
Evidence: supervision-host HEAD full run transcript (65 ok, rc=0, 599s)

Source: supervision-host HEAD full run transcript (65 ok, rc=0, 599s)

ok - report surface: only the branch actor's current turn may report, and only on the tasks its wake names
ok - report surface: visible late outcomes queue a relay, while silent outcomes remain stored without a wake or note
ok - dispatch entry: the host reads branch eligibility, the offer rule, and the wake prompt from the Pi branch's own owner
ok - drain: BRANCH OUTCOMES runs on a Claude home by default and on another primary with the file, never with off, and never on Pi
ok - drain: captain outcomes come first, and routine overflow collapses into a count one drain clears
ok - drain: repeated captain outcomes collapse per task, and the byte cap presents only the run its acknowledgement covers
ok - drain: a long away window costs one short drain, captain outcomes collapsed per task and routine overflow counted, and nothing from it is shown again
ok - drain: the BRANCH OUTCOMES budgets count bytes, cutting multibyte summaries by whole characters in any locale
ok - drain: branch outcomes stay unread when a projection of the store fails
ok - drain: branch outcomes stay unread and the drain fails when jq is missing
ok - drain: branch outcomes stay unread when the drain cannot print them
ok - drain: a legacy backlog is presented with each outcome's age and a check-first instruction, never adopted
ok - drain: a keyed decision survives acknowledgement through a newer outcome for its task
ok - drain: an outcome carried across a switch off Pi comes back with its age, never adopted
ok - drain: an outcome nothing has shown is presented until acknowledged, and a repeated acknowledgement changes nothing
ok - drain: an outcome the host drain presented but main never acknowledged stays unprocessed across a switch to Pi
ok - drain: an outcome the host drain presented but main never acknowledged survives an outcome index repair
ok - host: an attended wake the branch may take is handled on the engine, and its routine outcome never wakes main
ok - host: an attended captain outcome wakes main once and stays in its drain until main acknowledges it
ok - host: a captain outcome recorded after the captain left waits for the return, then reaches main's drain
ok - host: a quiet record without its daemon is a present captain, so outcomes and decisions reach main
ok - host: an attended decision close stays main's exactly as the plain arm delivers it
ok - host: an off written while the host is parked sends the next attended close to main, naming the opt-out
ok - host: a main-only pass-through leaves the successor watcher running and the close undelivered for main
ok - host: an attended close whose main session cannot be identified reaches main and runs no engine turn
ok - host: a decision close accepted away whose turn starts attended still reaches main unchanged
ok - host: an attended close whose task turns main-only before its turn still reaches main unchanged
ok - host+hook: an attended main-only pass-through rewakes main and keeps its successor watcher
ok - host+hook: a captain outcome beside a quiet record rewakes the present captain with no away note
ok - host+hook: a Claude home without config/supervision-host runs the host at the default engine, and an off file restores the plain arm
ok - host+hook: a close that turns main-only at its turn rewakes main and keeps its successor watcher
ok - host+hook: failed at-turn downtime write notifies main despite a healthy successor
ok - host+hook: a successor close that lands during main's turn is delivered at the next turn end
ok - host: a primary with no verified dialog mirror keeps every attended close on main, and its away posture still runs
ok - host: each wake carries the captain's dialog since the last wake, without operational input
ok - host: the dialog mirror, its feed, and the wake file are owner-only
ok - host: dialog a turn never completed with its report (a park boundary, a stopped turn, no report) reaches the next turn
ok - host: an attended wake whose mirror is missing, cannot be read (at the feed or at the prompt), or holds a malformed entry reaches main before any engine turn, and the cursor stays put
ok - host: an away wake is handled on the engine through the branch contract and never reaches main
ok - host: an engine turn that records no outcome hands its durable wake to main
ok - host: a captain return during an engine turn hands that turn's outcomes to main
ok - host: silent outcomes are excluded from both captain-return handoff paths
ok - host: an early visible outcome survives more than 1,000 same-turn receipts
ok - host: an outcome lookup failure forces a visible main handoff
ok - host: an outcome recorded after the return reaches main even when its host dies at the turn's end
ok - host: a killed predecessor's engine is reaped and the next host removes its turn files
ok - host: a turn that reports but leaves its granted rows queued hands the wake to main
ok - host: a captain return during a failed engine turn still hands that turn's outcomes to main
ok - host: an engine turn whose result is incomplete hands its wake to main
ok - host: two engine errors latch the session, main keeps every away wake in the cooldown, a failed probe doubles it up to its cap, and a report recovers it silently
ok - host: an attended close in a latched session reaches main as the arm printed it and leaves the latch as it was, and a home that opted out with off never reads it
ok - host: attended, two engine errors latch the session, main keeps every close unchanged in the cooldown, a failed probe doubles it up to its cap, and a routine probe's recovery stays in the ledger, off main
ok - host: an engine turn is bounded, and tool processes outside its process group are reaped
ok - host: a restarted host stops, by recorded identity, the cycle a killed predecessor left running
ok - host: the park ends itself with a boundary wake and a stopped watcher
ok - host: waiting closes cannot carry the park past its boundary
ok - host: a close whose margin runs out while the successor starts reaches main at the boundary without a turn
ok - host: the park's test clock is inert without the test marker
ok - host: a park at or beyond the Stop-hook registration falls back to the default boundary
ok - host: an owner's later park limit lets a turn outlive the boundary, and no other limit does
ok - host: the first cycle's status streams once, --restart replaces a stale watcher, and an owner predecessor makes a handling successor
ok - host: an unchanged held outcome reaches the captain once across cadences, and a later decision on the task still does
ok - host: a home naming an unverified engine hands every away wake to main with the reason
ok - host: a host outside the session-lock owner stands down without arming
ok - host: a host under a superseded auto-arm generation stands down without touching the owner
rc=0 elapsed=599s
Evidence: HEAD per-10s live fixture hosts/watchers/load

Source: HEAD per-10s live fixture hosts/watchers/load

elapsed_s	load1	hosts	watchers
2	3.18	0	0
12	3.23	0	0
23	3.61	0	0
33	3.45	2	1
44	3.08	3	1
56	4.28	3	1
66	4.15	0	0
77	4.52	2	0
88	4.74	2	0
99	4.94	0	1
110	4.75	1	1
121	4.50	0	1
132	4.43	0	0
143	4.21	2	1
154	4.79	2	1
165	4.59	2	1
175	5.50	2	1
186	5.19	2	0
197	5.16	1	0
207	4.81	1	0
219	4.86	1	0
229	5.13	3	1
240	5.27	1	0
251	5.58	2	1
262	5.17	3	1
273	5.47	2	1
284	5.10	2	0
295	5.00	2	1
306	5.59	0	0
316	5.51	2	0
327	5.42	0	0
338	5.13	2	1
349	5.03	2	0
360	4.81	2	0
370	5.12	2	1
381	5.22	2	1
392	5.13	2	0
403	5.42	2	1
414	5.10	2	0
425	4.63	1	0
435	4.62	1	0
446	4.51	1	0
457	4.58	2	1
468	4.41	2	0
478	4.19	2	1
489	3.86	3	1
500	4.48	0	0
511	4.97	2	1
521	5.05	2	0
532	4.73	2	1
543	4.70	0	0
555	4.77	0	0
565	4.80	2	1
576	4.26	2	1
Evidence: base fc795e5 partial run transcript (36 cases in 450s)

Source: base fc795e5 partial run transcript (36 cases in 450s)

ok - report surface: only the branch actor's current turn may report, and only on the tasks its wake names
ok - report surface: visible late outcomes queue a relay, while silent outcomes remain stored without a wake or note
ok - dispatch entry: the host reads branch eligibility, the offer rule, and the wake prompt from the Pi branch's own owner
ok - drain: BRANCH OUTCOMES runs on a Claude home by default and on another primary with the file, never with off, and never on Pi
ok - drain: captain outcomes come first, and routine overflow collapses into a count one drain clears
ok - drain: repeated captain outcomes collapse per task, and the byte cap presents only the run its acknowledgement covers
ok - drain: a long away window costs one short drain, captain outcomes collapsed per task and routine overflow counted, and nothing from it is shown again
ok - drain: the BRANCH OUTCOMES budgets count bytes, cutting multibyte summaries by whole characters in any locale
ok - drain: branch outcomes stay unread when a projection of the store fails
ok - drain: branch outcomes stay unread and the drain fails when jq is missing
ok - drain: branch outcomes stay unread when the drain cannot print them
ok - drain: a legacy backlog is presented with each outcome's age and a check-first instruction, never adopted
ok - drain: a keyed decision survives acknowledgement through a newer outcome for its task
ok - drain: an outcome carried across a switch off Pi comes back with its age, never adopted
ok - drain: an outcome nothing has shown is presented until acknowledged, and a repeated acknowledgement changes nothing
ok - drain: an outcome the host drain presented but main never acknowledged stays unprocessed across a switch to Pi
ok - drain: an outcome the host drain presented but main never acknowledged survives an outcome index repair
ok - host: an attended wake the branch may take is handled on the engine, and its routine outcome never wakes main
ok - host: an attended captain outcome wakes main once and stays in its drain until main acknowledges it
ok - host: a captain outcome recorded after the captain left waits for the return, then reaches main's drain
ok - host: a quiet record without its daemon is a present captain, so outcomes and decisions reach main
ok - host: an attended decision close stays main's exactly as the plain arm delivers it
ok - host: an off written while the host is parked sends the next attended close to main, naming the opt-out
ok - host: a main-only pass-through leaves the successor watcher running and the close undelivered for main
ok - host: an attended close whose main session cannot be identified reaches main and runs no engine turn
ok - host: a decision close accepted away whose turn starts attended still reaches main unchanged
ok - host: an attended close whose task turns main-only before its turn still reaches main unchanged
ok - host+hook: an attended main-only pass-through rewakes main and keeps its successor watcher
ok - host+hook: a captain outcome beside a quiet record rewakes the present captain with no away note
ok - host+hook: a Claude home without config/supervision-host runs the host at the default engine, and an off file restores the plain arm
ok - host+hook: a close that turns main-only at its turn rewakes main and keeps its successor watcher
ok - host+hook: failed at-turn downtime write notifies main despite a healthy successor
ok - host+hook: a successor close that lands during main's turn is delivered at the next turn end
ok - host: a primary with no verified dialog mirror keeps every attended close on main, and its away posture still runs
ok - host: each wake carries the captain's dialog since the last wake, without operational input
ok - host: the dialog mirror, its feed, and the wake file are owner-only
Evidence: base per-10s live fixture hosts/watchers/load (peaks 11/12/17.5)

Source: base per-10s live fixture hosts/watchers/load (peaks 11/12/17.5)

elapsed_s	load1	hosts	watchers
1	3.41	0	0
11	3.49	0	0
22	3.70	0	0
33	3.59	0	0
44	3.89	0	0
55	3.93	4	1
66	4.24	4	2
78	4.59	6	2
89	4.34	4	3
100	5.26	5	4
111	5.69	6	5
123	6.83	6	6
136	6.92	4	7
148	6.71	4	7
160	7.58	5	7
174	7.95	4	8
186	9.12	4	8
200	9.29	4	9
212	9.48	5	9
224	11.46	4	9
237	12.08	6	10
250	12.77	6	9
262	13.73	4	9
277	13.50	6	10
289	14.35	6	9
302	14.20	6	9
316	14.46	8	10
330	14.45	9	11
344	14.82	10	11
359	15.75	10	12
373	15.51	9	12
387	16.96	10	11
400	17.37	10	11
414	17.45	10	12
428	17.49	11	12
442	17.03	10	11
458	17.19	9	11
Evidence: remote-reply HEAD timestamped transcript (ALL TESTS PASSED, 166s)

Source: remote-reply HEAD timestamped transcript (ALL TESTS PASSED, 166s)

[+5s] ok - a blocking non-destructive remote delta reaches durable process-event capture
[+5s] ok - a captured delta is applied, acknowledged, and re-armed without a handler
[+6s] ok - ingest appends one validated line, fetches its document, and advances the cursor
[+7s] ok - replayed capture has one deduplicated append and one durable handling identity
[+14s] ok - later generations cannot invalidate an unacknowledged ingested result
[+17s] ok - the remote status and decision model mirrors and the cursor advances
[+17s] ok - a remote mate's new decision folds open exactly as a local mate's does
[+17s] ok - a replayed mirrored delta is idempotent in both the stream and the cursor
[+19s] ok - transported control bytes are normalized in place and never stop the stream
[+22s] ok - payload protocol-field names cannot collide with transport metadata
[+25s] ok - NUL bytes are normalized in place before shell line processing
[+31s] ok - local document storage failures remain retryable until delivery succeeds
[+34s] ok - a structured pointer must end at its token boundary
[+39s] ok - adjacent pointers are fetched while malformed tokens remain unchanged
[+43s] ok - a rejected candidate never gives the following text a false leading boundary
[+48s] ok - nested remote reports relay while an undeliverable foreign pointer fails open
[+59s] ok - the reported incident raises no standing decision and still delivers the report
[+70s] ok - an undeliverable structured offer fails open with one note and never a decision
[+75s] ok - a remote refusal surfaces its own reason without opening a decision
[+81s] ok - a failed pointer extraction never commits a partial delta
[+88s] ok - a failed mirror write never drops status content or advances the cursor
[+111s] ok - source-line identity survives commit failure and cursor-loss recapture
[+115s] ok - a mirrored reserved-key line cannot squat or clear the parent's own decision
[+118s] ok - a reply that arrives after escalation resolves it and clears the open decision
[+124s] ok - a remote reply listener stays owned across empty waits and a delta
[+126s] ok - a failed remote read exits instead of relistening
[+132s] ok - persistent ingestion failure leaves exactly one durable capture
[+136s] ok - a preempted reply poll reports a closed window without publishing channel freshness
[+141s] ok - a preempted reply poll keeps its listener and reconcile launches nothing
[+143s] ok - a quiet reply window publishes the caught-up watermark the reply guard reads
[+161s] ok - a cursor-loss whole-log recapture is acknowledged quietly with no duplicate wake
[+165s] ok - truncation is detected, escalated once, and not silently rebased
[+166s] ok - remote reply retirement quiesces and refuses unhandled captured results
[+166s] ALL TESTS PASSED
rc=0 elapsed=166s
Evidence: remote-reply await_reply_result per-call durations

Source: remote-reply await_reply_result per-call durations

220	3.3	rc=0	remote-reply-ios.2.result
236	2.8	rc=0	remote-reply-ios.3.result
271	2.4	rc=0	remote-reply-ios.4.result
331	1.4	rc=0	remote-reply-ios.5.result
348	2.7	rc=0	remote-reply-ios.6.result
361	2.9	rc=0	remote-reply-ios.7.result
431	2.9	rc=0	remote-reply-ios.9.result
431	5.6	rc=0	remote-reply-ios.10.result
431	2.8	rc=0	remote-reply-ios.11.result
431	5.0	rc=0	remote-reply-ios.12.result
431	3.6	rc=0	remote-reply-ios.13.result
431	2.6	rc=0	remote-reply-ios.14.result
431	4.4	rc=0	remote-reply-ios.15.result
431	4.0	rc=0	remote-reply-ios.16.result
431	4.1	rc=0	remote-reply-ios.17.result
431	5.4	rc=0	remote-reply-ios.18.result
681	19.6	rc=0	remote-reply-ios.22.result
724	2.7	rc=0	remote-reply-ios.23.result
743	4.1	rc=0	remote-reply-ios.24.result
945	21.9	rc=0	remote-reply-ios.27.result
Evidence: live e2e suites skip without their opt-in variables

Source: live e2e suites skip without their opt-in variables

== fm-supervision-host-live-e2e
skip: live: opt-in; set FM_SUPERVISION_HOST_LIVE_E2E=1 to run
rc=0
== fm-supervision-host-attended-live-e2e
skip: live: opt-in; set FM_SUPERVISION_HOST_ATTENDED_LIVE_E2E=1 to run
rc=0
Evidence: Process monitor used for the samples

Source: Process monitor used for the samples

#!/usr/bin/env bash
# Samples every 10s the live hosts and watchers belonging to the suite's fixture
# homes (FM_HOME under an fm-supervision-host.* temp root), plus load average.
out=$1; pidf=$2
printf 'elapsed_s\tload1\thosts\twatchers\n' > "$out"
start=$(date +%s)
while kill -0 "$(cat "$pidf")" 2>/dev/null; do
  h=0; w=0
  for p in /proc/[0-9]*; do
    tr '\0' '\n' < "$p/environ" 2>/dev/null | grep -q '^FM_HOME=.*/fm-supervision-host\.' || continue
    a=$(tr '\0' ' ' < "$p/cmdline" 2>/dev/null)
    case "$a" in *bin/fm-supervision-host.sh*) h=$((h+1)) ;; *bin/fm-watch.sh*) w=$((w+1)) ;; esac
  done
  printf '%s\t%s\t%s\t%s\n' $(( $(date +%s) - start )) "$(cut -d' ' -f1 /proc/loadavg)" "$h" "$w" >> "$out"
  sleep 10
done
Evidence: HEAD vs base supervision-host comparison
HEAD e389c56: 65/65 ok rc=0 in 599s; peak live fixture hosts=3 watchers=1 load1=5.59
base fc795e5: 36/65 after 450s (stopped); peak live fixture hosts=11 watchers=12 load1=17.49
- Outcome: ⚠️ 1 info across 1 run (27m33s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

✅ **Review** - passed

✅ No issues found.

⚠️ **Test** - 1 info
  • ℹ️ bin/fm-procevent-remote-reply.sh:557 - Existing product leak, unrelated to this change. cmd_ingest's continuity-broken branch does return 3 without removing its staging directory (created at line 508). The cleanup is an EXIT trap (trap &#39;rm -rf -- &#34;$tmp&#34;&#39; EXIT), but tmp is local to cmd_ingest, so by the time the script exits the variable is out of scope and nothing is removed. Every tests/fm-remote-reply.test.sh run leaves 4 /tmp/fm-remote-reply-ingest.* directories, each holding an empty payload and normalized-payload. Dozens of these from earlier runs are on this host. The fix would be trap - EXIT; rm -rf -- &#34;$tmp&#34; before return 3, mirroring the success path at lines 636-637. I did not change it because product code is out of scope for this test phase. It likely affects upstream too.
  • Live validation: ✅ go - 4 of 4 scenarios driven live against the product
Scenario Result Live Evidence
Full tests/fm-supervision-host.test.sh run at HEAD passes every case within the 900s per-command cap ✅ pass live supervision-host-head.log: 65 'ok' lines, 0 'not ok', rc=0 elapsed=599s
Adversarial: hosts and watchers left behind by one case no longer pile up across later cases (base reproduces the pile-up) ✅ pass live supervision-host-head-procs.tsv: peaks of 3 hosts, 1 watcher, load 5.59 over the whole run. supervision-host-base-procs.tsv: base fc795e5 reached 11 hosts, 12 watchers, load 17.49 and finished only 36…
tests/fm-remote-reply.test.sh passes at HEAD, and the whole-log recapture finishes well inside the widened await bound ✅ pass live remote-reply-head.log: ALL TESTS PASSED, rc=0, 166s. remote-reply-await-timings.tsv: replay-identity recapture 19.6s, cursor-loss recapture 21.9s, all others <=5.6s, bound now 120s
The live e2e supervision-host variants skip cleanly without their opt-in variables, so they are not a source of the failures ✅ pass live live-e2e-gate.log: both print 'skip: live: opt-in; set FM_SUPERVISION_HOST_..._LIVE_E2E=1 to run' and exit 0; neither file changed between base and HEAD
  • bash tests/fm-supervision-host.test.sh at HEAD, with a 10s monitor counting live fixture hosts and watchers (FM_HOME under the fm-supervision-host.* fixture root) and load average
  • git show fc795e5:tests/fm-supervision-host.test.sh run from a temporary copy in tests/ for 450s with the same monitor, then stopped (before/after comparison of the leak)
  • bash tests/fm-remote-reply.test.sh at HEAD, with each output line timestamped
  • Temporary instrumented copy of tests/fm-remote-reply.test.sh that wraps await_reply_result to log each call's duration
  • bash tests/fm-supervision-host-live-e2e.test.sh and bash tests/fm-supervision-host-attended-live-e2e.test.sh with the opt-in variables unset (checks that both skip)
  • Cleanup check: confirmed no leftover fixture processes after the runs, removed the temporary test copies and my /tmp ingest leftovers, and confirmed the worktree is clean
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

…te-reply recapture wait

Both failures are test-side timing assumptions that a busy host exposes.
They are not product bugs, and they do not come from missing tools: tmux
and ruby are absent here but neither suite needs them (supervision-host
stubs tmux), and the live e2e variants are opt-in and gate-skip unless
FM_SUPERVISION_HOST_LIVE_E2E / FM_SUPERVISION_HOST_ATTENDED_LIVE_E2E is set.
Upstream CI at f593060 passes both suites (supervision-host 683s,
remote-reply 184s) on uncontended runners.

tests/fm-supervision-host.test.sh
Most cases end with their host parked (default park bound 27000s) or with a
pass-through's successor watcher still running at FM_POLL=1. Before this
change only the EXIT trap stopped them, so they piled up across the 65 cases.
An instrumented run on a quiet 4-core host counted 17 live watchers and 8
live hosts by the last case. Load average rose from 1.5 to 25.8, per-case
time roughly tripled (latch case 35s -> 131s, attended latch 37s -> 159s),
and the suite took 1558s, past any 900s per-command cap. When concurrent
lanes added load, the pile-up pushed closes past the suite's fixed 15-25s
wait_until budgets, with failures that varied between runs:
"latch: the later failed probe did not hand the wake back" (fork main
fc795e5, 1482s), "away grok: the wake was not handled on the engine"
(upstream-merge branch), and "the watcher's downtime resurface did not reach
main". run_case now stops every home a case made after the case returns.
With that change no watchers leak between cases, load stayed at 2-9, and the
suite passed in 558s under the same conditions. It also passed in 659s
under concurrent load from fm-test-run, and in 731s on the upstream f593060
tree with the same change applied.

tests/fm-remote-reply.test.sh
After a cursor loss, the whole-log recapture fetches again every document
that the whole log offers, one remote job at a time: one delta read and 13
document fetches. Upstream kunchenguid#6255 (549e07f) made each sequential remote job
slower. It cut the dispatcher burst from 20 passes to 4 before the 1s quiet
scan, raised active/result sampling from 0.05s to 0.25s, and raised delta
polling from 0.2s to 0.5s. Paired runs under the same load measured that step
at 36s on f593060 (about 2.5s per fetch). With 549e07f's bin changes reverted
it took 20s (about 1.1s per fetch). Either way it has to fit inside
await_reply_result's 40s budget (800 x 0.05s polls). The failures ("the
replay-identity whole-log recapture was not captured") appeared only on trees
that include kunchenguid#6255. Fork main without it passes, but the same step already
used 30s of the 40s under load. The bound is now 2400 polls (120s). A healthy
wait still returns as soon as the capture is applied.

Upstream: both changes belong upstream too. The leak and the post-kunchenguid#6255 budget
exist the same way on upstream main, and upstream CI passes only because its
runners are uncontended. When upstream is merged into the fork, the call-list
hunk conflicts trivially with upstream's added
test_park_exit_probe_uses_half_second_child_sleeps line. Resolve it by
prefixing that line with run_case.
@cloud-practitioner
cloud-practitioner merged commit c8de201 into main Oct 1, 2026
20 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant