Skip to content

feat: add shared CPU pass pool and host load reporting - #53

Merged
MrGTV-love merged 12 commits into
mainfrom
fm/fm-cpu-pool-gate-overlay
Oct 8, 2026
Merged

MrGTV-love merged 12 commits into
mainfrom
fm/fm-cpu-pool-gate-overlay

Conversation

@MrGTV-love

@MrGTV-love MrGTV-love commented Oct 8, 2026 •

Copy link
Copy Markdown
Owner

What Changed

  • Add a host-wide, CPU-count-sized pass pool with exact multi-pass reservations, inherited lock ownership, nested-run handling, status reporting, and degraded execution when the pool is unavailable. Integrate per-script reservations into fm-test-run.sh, keeping queue waits outside script timeouts and limiting nested concurrency to inherited passes.
  • Add load recording and text/JSON reports covering pool usage, pipeline agent durations, failures, timeout errors, and load/convergence verdicts for a fixed first-ten review-reaching run cohort.
  • Add pool and reporting regression coverage, isolate runner tests with private pools, document the cross-repository protocol and measurement recipe, and refresh portable serial-shard duration hints.

Risk Assessment

✅ Low: The Firstmate-side change is bounded, incorporates the recorded decisions, and has no substantiated material source defects; documentation correctly conditions Phase 1 completion and measurement on Vernant participation.

Testing

Both focused public-command scenario scripts passed, followed by manual live checks of the real host-sized pool, nested and queued runner execution, recorder output, and convergence transitions. CLI transcripts, host samples, and JSON reports were preserved; disposable fixtures were removed. Private F6 operation and the post-Vernant full-fleet measurement were not validated in this phase.

  • Live validation: ✅ go - 12 of 13 scenarios driven live against the product
Scenario Result Live Evidence
Read the host-sized pool and receive the requested reservation ✅ pass live Host-sized CPU-pool CLI transcript; focused CPU-pool script exercised child markers and exit-status propagation.
Queue behind occupied passes and identify the current partial collector ✅ pass live Host-sized CPU-pool CLI transcript; tests/fm-cpu-pass.test.sh partial-collection scenario.
Cancel queued work without starting its command ✅ pass live Host-sized CPU-pool CLI transcript.
Retain a running child's reservation after wrapper death ✅ pass live Host-sized CPU-pool CLI transcript; focused CPU-pool signal scenarios.
Refuse invalid reservations and nested over-requests before work starts ✅ pass live Host-sized CPU-pool CLI transcript; tests/fm-cpu-pass.test.sh.
Announce degraded execution on default stderr while preserving workload output and status ✅ pass live Host-sized CPU-pool CLI transcript.
Reject malformed inherited markers in Python, no-Python, and runner execution ✅ pass live Host-sized CPU-pool CLI transcript; tests/fm-cpu-pass.test.sh inherited-marker scenarios.
Request two nested runner jobs under one pass without exceeding the inherited budget ✅ pass live Nested and queued runner transcript, including persisted start/end events.
Wait for a pass outside the script timeout while retaining the timeout for hung work ✅ pass live Nested and queued runner transcript.
Record real host load and pool occupancy with record and watch ✅ pass live Actual host-load samples; recorder and convergence reporter transcript.
Keep convergence pending until earlier live runs settle, then preserve the first-ten verdict ✅ pass live Pending, settled unsuccessful, and unchanged-later-successes JSON artifacts; reporter transcript.
Report load and convergence failures without allowing a smaller acceptance cohort ✅ pass live Above-budget load report; reporter transcript; tests/fm-load-report.test.sh.
Validate the coordinated Phase 1 rollout and sustained fleet acceptance ⏸️ untested no This change delivers the Firstmate participant and explicitly requires the separate Vernant-repository participant before measurement begins. Disposable checks validated Firstmate's runtime and report…
Evidence: Host-sized CPU-pool CLI transcript

Source: Host-sized CPU-pool CLI transcript

$ ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/bin/fm-cpu-pass.sh size
18
exit=0

$ ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/bin/fm-cpu-pass.sh run --passes 18 --label finished-earlier -- /bin/true
fm-cpu-pass: command not found: /bin/true
exit=127

$ [background] ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/bin/fm-cpu-pass.sh run --passes 17 --label active-holder -- /bin/bash -c touch "$1/started"; while [ ! -e "$1/release" ]; do sleep .05; done _ ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/.cpu-live-6osl5vs9

$ [background] ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/bin/fm-cpu-pass.sh run --passes 2 --label current-collector -- /bin/bash -c echo "collector-started held=$FM_CPU_PASS_HELD"

$ ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/bin/fm-cpu-pass.sh status --json
{"available": true, "dir": "~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/.cpu-live-6osl5vs9/pool", "free": 0, "held": 18, "holders": [{"holder": "pid=87250 passes=17 since=1791458081 label=active-holder", "slot": 0}, {"holder": "pid=87250 passes=17 since=1791458081 label=active-holder", "slot": 1}, {"holder": "pid=87250 passes=17 since=1791458081 label=active-holder", "slot": 2}, {"holder": "pid=87250 passes=17 since=1791458081 label=active-holder", "slot": 3}, {"holder": "pid=87250 passes=17 since=1791458081 label=active-holder", "slot": 4}, {"holder": "pid=87250 passes=17 since=1791458081 label=active-holder", "slot": 5}, {"holder": "pid=88674 passes=2 since=1791458082 label=current-collector", "slot": 6}, {"holder": "pid=87250 passes=17 since=1791458081 label=active-holder", "slot": 7}, {"holder": "pid=87250 passes=17 since=1791458081 label=active-holder", "slot": 8}, {"holder": "pid=87250 passes=17 since=1791458081 label=active-holder", "slot": 9}, {"holder": "pid=87250 passes=17 since=1791458081 label=active-holder", "slot": 10}, {"holder": "pid=87250 passes=17 since=1791458081 label=active-holder", "slot": 11}, {"holder": "pid=87250 passes=17 since=1791458081 label=active-holder", "slot": 12}, {"holder": "pid=87250 passes=17 since=1791458081 label=active-holder", "slot": 13}, {"holder": "pid=87250 passes=17 since=1791458081 label=active-holder", "slot": 14}, {"holder": "pid=87250 passes=17 since=1791458081 label=active-holder", "slot": 15}, {"holder": "pid=87250 passes=17 since=1791458081 label=active-holder", "slot": 16}, {"holder": "pid=87250 passes=17 since=1791458081 label=active-holder", "slot": 17}], "size": 18}
exit=0

$ [background] ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/bin/fm-cpu-pass.sh run --label queued-waiter -- /bin/bash -c echo SHOULD-NOT-RUN

collector stderr while waiting:
fm-cpu-pass: waiting 2s for 2 CPU pass(es) for current-collector: all passes in use; pool size 18; holders: pid=87250 passes=17 since=1791458081 label=active-holder; pid=87250 passes=17 since=1791458081 label=active-holder; pid=87250 passes=17 since=1791458081 label=active-holder; pid=87250 passes=17 since=1791458081 label=active-holder; pid=87250 passes=17 since=1791458081 label=active-holder; and 13 more

$ ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/bin/fm-cpu-pass.sh status --json
{"available": true, "dir": "~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/.cpu-live-6osl5vs9/pool", "free": 0, "held": 18, "holders": [{"holder": "pid=87250 passes=17 since=1791458081 label=active-holder", "slot": 0}, {"holder": "pid=87250 passes=17 since=1791458081 label=active-holder", "slot": 1}, {"holder": "pid=87250 passes=17 since=1791458081 label=active-holder", "slot": 2}, {"holder": "pid=87250 passes=17 since=1791458081 label=active-holder", "slot": 3}, {"holder": "pid=87250 passes=17 since=1791458081 label=active-holder", "slot": 4}, {"holder": "pid=87250 passes=17 since=1791458081 label=active-holder", "slot": 5}, {"holder": "pid=88674 passes=2 since=1791458082 label=current-collector", "slot": 6}, {"holder": "pid=87250 passes=17 since=1791458081 label=active-holder", "slot": 7}, {"holder": "pid=87250 passes=17 since=1791458081 label=active-holder", "slot": 8}, {"holder": "pid=87250 passes=17 since=1791458081 label=active-holder", "slot": 9}, {"holder": "pid=87250 passes=17 since=1791458081 label=active-holder", "slot": 10}, {"holder": "pid=87250 passes=17 since=1791458081 label=active-holder", "slot": 11}, {"holder": "pid=87250 passes=17 since=1791458081 label=active-holder", "slot": 12}, {"holder": "pid=87250 passes=17 since=1791458081 label=active-holder", "slot": 13}, {"holder": "pid=87250 passes=17 since=1791458081 label=active-holder", "slot": 14}, {"holder": "pid=87250 passes=17 since=1791458081 label=active-holder", "slot": 15}, {"holder": "pid=87250 passes=17 since=1791458081 label=active-holder", "slot": 16}, {"holder": "pid=87250 passes=17 since=1791458081 label=active-holder", "slot": 17}], "size": 18}
exit=0

collector stderr while waiting:
fm-cpu-pass: waiting 2s for 2 CPU pass(es) for current-collector: all passes in use; pool size 18; holders: pid=87250 passes=17 since=1791458081 label=active-holder; pid=87250 passes=17 since=1791458081 label=active-holder; pid=87250 passes=17 since=1791458081 label=active-holder; pid=87250 passes=17 since=1791458081 label=active-holder; pid=87250 passes=17 since=1791458081 label=active-holder; and 13 more

Driver correction: wait summaries show at most five slot records plus an omitted count; require no stale holder, and use status for exhaustive partial-holder attribution.

collector stderr while waiting:
fm-cpu-pass: waiting 2s for 2 CPU pass(es) for current-collector: all passes in use; pool size 18; holders: pid=87250 passes=17 since=1791458081 label=active-holder; pid=87250 passes=17 since=1791458081 label=active-holder; pid=87250 passes=17 since=1791458081 label=active-holder; pid=87250 passes=17 since=1791458081 label=active-holder; pid=87250 passes=17 since=1791458081 label=active-holder; and 13 more

waiter stderr while waiting:
fm-cpu-pass: waiting 2s for 1 CPU pass(es) for queued-waiter: another request is collecting passes; pool size 18; holders: pid=87250 passes=17 since=1791458081 label=active-holder; pid=87250 passes=17 since=1791458081 label=active-holder; pid=87250 passes=17 since=1791458081 label=active-holder; pid=87250 passes=17 since=1791458081 label=active-holder; pid=87250 passes=17 since=1791458081 label=active-holder; and 13 more

Cancelled waiter raw subprocess returncode=-15 (shell exit 143 for SIGTERM); work output=''

collector stdout after release:
collector-started held=2

$ ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/bin/fm-cpu-pass.sh status --json
{"available": true, "dir": "~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/.cpu-live-6osl5vs9/pool", "free": 18, "held": 0, "holders": [], "size": 18}
exit=0

$ [background] ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/bin/fm-cpu-pass.sh run --label surviving-child -- /bin/bash -c echo $$ > "$1/child.pid"; exec sleep 30 _ ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/.cpu-live-6osl5vs9

$ ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/bin/fm-cpu-pass.sh status --json
{"available": true, "dir": "~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/.cpu-live-6osl5vs9/pool", "free": 17, "held": 1, "holders": [{"holder": "pid=68410 passes=1 since=1791458160 label=surviving-child", "slot": 5}], "size": 18}
exit=0

$ ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/bin/fm-cpu-pass.sh status --json
{"available": true, "dir": "~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/.cpu-live-6osl5vs9/pool", "free": 18, "held": 0, "holders": [], "size": 18}
exit=0

$ ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/bin/fm-cpu-pass.sh run --passes 0 -- /usr/bin/touch ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/.cpu-live-6osl5vs9/invalid-ran
fm-cpu-pass: --passes must be between 1 and the pool size (18)
exit=125

$ ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/bin/fm-cpu-pass.sh run --passes -1 -- /usr/bin/touch ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/.cpu-live-6osl5vs9/invalid-ran
fm-cpu-pass: --passes must be between 1 and the pool size (18)
exit=125

$ ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/bin/fm-cpu-pass.sh run --passes 19 -- /usr/bin/touch ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/.cpu-live-6osl5vs9/invalid-ran
fm-cpu-pass: --passes must be between 1 and the pool size (18)
exit=125

$ ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/bin/fm-cpu-pass.sh run --passes invalid -- /usr/bin/touch ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/.cpu-live-6osl5vs9/invalid-ran
usage: fm-cpu-pass.sh run [-h] [--passes PASSES] [--label LABEL]
                          [--log-fd LOG_FD]
                          ...
fm-cpu-pass.sh run: error: argument --passes: invalid int value: 'invalid'
exit=125

$ ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/bin/fm-cpu-pass.sh run --passes 2 -- /usr/bin/touch ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/.cpu-live-6osl5vs9/invalid-ran
fm-cpu-pass: --passes must not exceed FM_CPU_PASS_HELD (1)
exit=125

$ ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/bin/fm-cpu-pass.sh run -- /bin/bash -c echo "held=$FM_CPU_PASS_HELD"; exit 7
held=0
fm-cpu-pass: running bash without a CPU pass: pool directory ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/.cpu-live-6osl5vs9/not-a-directory is unusable: [Errno 17] File exists: '~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/.cpu-live-6osl5vs9/not-a-directory'
exit=7

$ ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/bin/fm-cpu-pass.sh run -- bash -c echo "held=$FM_CPU_PASS_HELD"; echo "child stderr" >&2; exit 5
held=0
fm-cpu-pass: running bash without a CPU pass: python3 not found
child stderr
exit=5

$ ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/bin/fm-cpu-pass.sh run -- /usr/bin/touch ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/.cpu-live-6osl5vs9/invalid-ran
fm-cpu-pass: FM_CPU_PASS_HELD must be a nonnegative decimal integer
exit=125

$ ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/bin/fm-cpu-pass.sh run -- /usr/bin/touch ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/.cpu-live-6osl5vs9/invalid-ran
fm-cpu-pass: FM_CPU_PASS_HELD must be a nonnegative decimal integer
exit=125
Evidence: Nested and queued runner transcript

Source: Nested and queued runner transcript

$ ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/bin/fm-cpu-pass.sh run -- ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/.cpu-live-6osl5vs9/runner-repo/bin/fm-test-run.sh --jobs 2 tests/fm-brief.test.sh tests/fm-composer-lib.test.sh --json ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/.cpu-live-6osl5vs9/nested-timing.json
FM_TEST_BEGIN 2026-10-08T11:16:38Z tests/fm-brief.test.sh family=pure-contract-unit expected_gate_skip=none
ok - held=1
FM_TEST_END 2026-10-08T11:16:40Z tests/fm-brief.test.sh exit=0 duration_ms=1504 gate_skip=false
FM_TEST_BEGIN 2026-10-08T11:16:40Z tests/fm-composer-lib.test.sh family=pure-contract-unit expected_gate_skip=none
ok - held=1
FM_TEST_END 2026-10-08T11:16:42Z tests/fm-composer-lib.test.sh exit=0 duration_ms=1405 gate_skip=false
FM_TEST_SUMMARY total=2 failed=0 skipped_gate=0 duration_ms=3432
FM_TEST_SUMMARY_FAMILY family=pure-contract-unit count=2 duration_ms=2909 failed=0
FM_TEST_SLOWEST rank=1 script=tests/fm-brief.test.sh duration_ms=1504
FM_TEST_SLOWEST rank=2 script=tests/fm-composer-lib.test.sh duration_ms=1405
fm-test-run: reducing --jobs 2 to inherited FM_CPU_PASS_HELD=1
fm-test-run: wrote timing artifact: ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/.cpu-live-6osl5vs9/nested-timing.json
exit=0

Persisted nested execution events:
start tests/fm-brief.test.sh held=1
end tests/fm-brief.test.sh
start tests/fm-composer-lib.test.sh held=1
end tests/fm-composer-lib.test.sh

$ [background] ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/bin/fm-cpu-pass.sh run --passes 18 --label outside-burst -- /bin/bash -c touch "$1/full-started"; while [ ! -e "$1/full-release" ]; do sleep .05; done _ ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/.cpu-live-6osl5vs9

$ [background] ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/.cpu-live-6osl5vs9/runner-repo/bin/fm-test-run.sh --per-script-timeout-secs 3 ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/.cpu-live-6osl5vs9/runner-repo/tests/probe.test.sh

Queued runner out:
FM_TEST_BEGIN 2026-10-08T11:16:42Z ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/.cpu-live-6osl5vs9/runner-repo/tests/probe.test.sh family=unclassified expected_gate_skip=none
skip: disposable probe (held=1)
FM_TEST_END 2026-10-08T11:16:49Z ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/.cpu-live-6osl5vs9/runner-repo/tests/probe.test.sh exit=0 duration_ms=6358 gate_skip=true
FM_TEST_SUMMARY total=1 failed=0 skipped_gate=1 duration_ms=6788
FM_TEST_SUMMARY_FAMILY family=unclassified count=1 duration_ms=6358 failed=0
FM_TEST_SLOWEST rank=1 script=~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/.cpu-live-6osl5vs9/runner-repo/tests/probe.test.sh duration_ms=6358

Queued runner err:
fm-cpu-pass: waiting 2s for 1 CPU pass(es) for fm-test-run ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/.cpu-live-6osl5vs9/runner-repo/tests/probe.test.sh: all passes in use; pool size 18; holders: pid=35024 passes=18 since=1791458202 label=outside-burst; pid=35024 passes=18 since=1791458202 label=outside-burst; pid=35024 passes=18 since=1791458202 label=outside-burst; pid=35024 passes=18 since=1791458202 label=outside-burst; pid=35024 passes=18 since=1791458202 label=outside-burst; and 13 more
fm-cpu-pass: got 1 CPU pass(es) for fm-test-run ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/.cpu-live-6osl5vs9/runner-repo/tests/probe.test.sh after 4s
fm-test-run: gate skip: ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/.cpu-live-6osl5vs9/runner-repo/tests/probe.test.sh: disposable probe (held=1)

$ ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/.cpu-live-6osl5vs9/runner-repo/bin/fm-test-run.sh --per-script-timeout-secs 1 tests/hang.test.sh
FM_TEST_BEGIN 2026-10-08T11:16:49Z tests/hang.test.sh family=unclassified expected_gate_skip=none
not ok - tests/hang.test.sh exceeded the per-script bound of 1s and was terminated
FM_TEST_END 2026-10-08T11:16:50Z tests/hang.test.sh exit=124 duration_ms=1310 gate_skip=false
FM_TEST_SUMMARY total=1 failed=1 skipped_gate=0 duration_ms=1706
FM_TEST_SUMMARY_FAMILY family=unclassified count=1 duration_ms=1310 failed=1
FM_TEST_SLOWEST rank=1 script=tests/hang.test.sh duration_ms=1310
exit=1

$ ~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/bin/fm-cpu-pass.sh status --json
{"available": true, "dir": "~/.no-mistakes/worktrees/32d18ed9638d/01M4DDA7E1CT1X3J8R87EWF7BN/.cpu-live-6osl5vs9/pool", "free": 18, "held": 0, "holders": [], "size": 18}
exit=0
Evidence: Actual host-load samples

Source: Actual host-load samples

1791458258	63.99	60.81	53.03	18	18	0
1791458259	63.99	60.81	53.03	18	18	0
1791458259	63.99	60.81	53.03	18	18	0
Evidence: Pending convergence cohort

Source: Pending convergence cohort

{"load": {"cpus": 4, "first": 1100, "last": 1200, "load1_max": 7.0, "load1_p50": 6.0, "load1_p95": 7.0, "load_within_2x_cpus": true, "pool_held_max": 3, "pool_held_mean": 2.5, "samples": 2, "share_above_2x_cpus": 0.0}, "pipeline": {"agent_failures": [], "agent_minutes_by_purpose": {"review": {"count": 10, "minutes_p50": 10.0, "minutes_p95": 10.0}, "review-fix": {"count": 9, "minutes_p50": 10.0, "minutes_p95": 10.0}}, "cohort_runs": [{"agent_minutes": 10.0, "converged": true, "created_at": 2000, "review_fix_rounds": 0, "run": "run-00", "status": "completed", "test_fix_rounds": 0, "timed_out": false}, {"agent_minutes": 20.0, "converged": true, "created_at": 2001, "review_fix_rounds": 1, "run": "run-01", "status": "completed", "test_fix_rounds": 0, "timed_out": false}, {"agent_minutes": 30.0, "converged": true, "created_at": 2002, "review_fix_rounds": 2, "run": "run-02", "status": "completed", "test_fix_rounds": 0, "timed_out": false}, {"agent_minutes": 10.0, "converged": true, "created_at": 2003, "review_fix_rounds": 0, "run": "run-03", "status": "completed", "test_fix_rounds": 0, "timed_out": false}, {"agent_minutes": 20.0, "converged": true, "created_at": 2004, "review_fix_rounds": 1, "run": "run-04", "status": "completed", "test_fix_rounds": 0, "timed_out": false}, {"agent_minutes": 30.0, "converged": true, "created_at": 2005, "review_fix_rounds": 2, "run": "run-05", "status": "completed", "test_fix_rounds": 0, "timed_out": false}, {"agent_minutes": 10.0, "converged": true, "created_at": 2006, "review_fix_rounds": 0, "run": "run-06", "status": "completed", "test_fix_rounds": 0, "timed_out": false}, {"agent_minutes": 20.0, "converged": true, "created_at": 2007, "review_fix_rounds": 1, "run": "run-07", "status": "completed", "test_fix_rounds": 0, "timed_out": false}, {"agent_minutes": 30.0, "converged": true, "created_at": 2008, "review_fix_rounds": 2, "run": "run-08", "status": "completed", "test_fix_rounds": 0, "timed_out": false}, {"agent_minutes": 10.0, "converged": true, "created_at": 2009, "review_fix_rounds": 0, "run": "run-09", "status": "completed", "test_fix_rounds": 0, "timed_out": false}], "converged_within_2_fix_rounds": null, "runs_considered": 10, "runs_wanted": 10, "timeout_class_run_errors": []}, "since": 1100}
Evidence: Settled unsuccessful convergence cohort

Source: Settled unsuccessful convergence cohort

{"load": {"cpus": 4, "first": 1100, "last": 1200, "load1_max": 7.0, "load1_p50": 6.0, "load1_p95": 7.0, "load_within_2x_cpus": true, "pool_held_max": 3, "pool_held_mean": 2.5, "samples": 2, "share_above_2x_cpus": 0.0}, "pipeline": {"agent_failures": [{"category": "timeout", "count": 1, "exit_status": "error"}], "agent_minutes_by_purpose": {"review": {"count": 11, "minutes_p50": 10.0, "minutes_p95": 10.0}, "review-fix": {"count": 9, "minutes_p50": 10.0, "minutes_p95": 10.0}}, "cohort_runs": [{"agent_minutes": 10.0, "converged": false, "created_at": 1999, "review_fix_rounds": 0, "run": "older-live", "status": "failed", "test_fix_rounds": 0, "timed_out": true}, {"agent_minutes": 10.0, "converged": true, "created_at": 2000, "review_fix_rounds": 0, "run": "run-00", "status": "completed", "test_fix_rounds": 0, "timed_out": false}, {"agent_minutes": 20.0, "converged": true, "created_at": 2001, "review_fix_rounds": 1, "run": "run-01", "status": "completed", "test_fix_rounds": 0, "timed_out": false}, {"agent_minutes": 30.0, "converged": true, "created_at": 2002, "review_fix_rounds": 2, "run": "run-02", "status": "completed", "test_fix_rounds": 0, "timed_out": false}, {"agent_minutes": 10.0, "converged": true, "created_at": 2003, "review_fix_rounds": 0, "run": "run-03", "status": "completed", "test_fix_rounds": 0, "timed_out": false}, {"agent_minutes": 20.0, "converged": true, "created_at": 2004, "review_fix_rounds": 1, "run": "run-04", "status": "completed", "test_fix_rounds": 0, "timed_out": false}, {"agent_minutes": 30.0, "converged": true, "created_at": 2005, "review_fix_rounds": 2, "run": "run-05", "status": "completed", "test_fix_rounds": 0, "timed_out": false}, {"agent_minutes": 10.0, "converged": true, "created_at": 2006, "review_fix_rounds": 0, "run": "run-06", "status": "completed", "test_fix_rounds": 0, "timed_out": false}, {"agent_minutes": 20.0, "converged": true, "created_at": 2007, "review_fix_rounds": 1, "run": "run-07", "status": "completed", "test_fix_rounds": 0, "timed_out": false}, {"agent_minutes": 30.0, "converged": true, "created_at": 2008, "review_fix_rounds": 2, "run": "run-08", "status": "completed", "test_fix_rounds": 0, "timed_out": false}], "converged_within_2_fix_rounds": false, "runs_considered": 10, "runs_wanted": 10, "timeout_class_run_errors": [{"class": "timed out", "run": "older-live"}]}, "since": 1100}
Evidence: Unchanged cohort after later successes

Source: Unchanged cohort after later successes

{"load": {"cpus": 4, "first": 1100, "last": 1200, "load1_max": 7.0, "load1_p50": 6.0, "load1_p95": 7.0, "load_within_2x_cpus": true, "pool_held_max": 3, "pool_held_mean": 2.5, "samples": 2, "share_above_2x_cpus": 0.0}, "pipeline": {"agent_failures": [{"category": "timeout", "count": 1, "exit_status": "error"}], "agent_minutes_by_purpose": {"review": {"count": 21, "minutes_p50": 10.0, "minutes_p95": 10.0}, "review-fix": {"count": 9, "minutes_p50": 10.0, "minutes_p95": 10.0}}, "cohort_runs": [{"agent_minutes": 10.0, "converged": false, "created_at": 1999, "review_fix_rounds": 0, "run": "older-live", "status": "failed", "test_fix_rounds": 0, "timed_out": true}, {"agent_minutes": 10.0, "converged": true, "created_at": 2000, "review_fix_rounds": 0, "run": "run-00", "status": "completed", "test_fix_rounds": 0, "timed_out": false}, {"agent_minutes": 20.0, "converged": true, "created_at": 2001, "review_fix_rounds": 1, "run": "run-01", "status": "completed", "test_fix_rounds": 0, "timed_out": false}, {"agent_minutes": 30.0, "converged": true, "created_at": 2002, "review_fix_rounds": 2, "run": "run-02", "status": "completed", "test_fix_rounds": 0, "timed_out": false}, {"agent_minutes": 10.0, "converged": true, "created_at": 2003, "review_fix_rounds": 0, "run": "run-03", "status": "completed", "test_fix_rounds": 0, "timed_out": false}, {"agent_minutes": 20.0, "converged": true, "created_at": 2004, "review_fix_rounds": 1, "run": "run-04", "status": "completed", "test_fix_rounds": 0, "timed_out": false}, {"agent_minutes": 30.0, "converged": true, "created_at": 2005, "review_fix_rounds": 2, "run": "run-05", "status": "completed", "test_fix_rounds": 0, "timed_out": false}, {"agent_minutes": 10.0, "converged": true, "created_at": 2006, "review_fix_rounds": 0, "run": "run-06", "status": "completed", "test_fix_rounds": 0, "timed_out": false}, {"agent_minutes": 20.0, "converged": true, "created_at": 2007, "review_fix_rounds": 1, "run": "run-07", "status": "completed", "test_fix_rounds": 0, "timed_out": false}, {"agent_minutes": 30.0, "converged": true, "created_at": 2008, "review_fix_rounds": 2, "run": "run-08", "status": "completed", "test_fix_rounds": 0, "timed_out": false}], "converged_within_2_fix_rounds": false, "runs_considered": 10, "runs_wanted": 10, "timeout_class_run_errors": [{"class": "timed out", "run": "older-live"}]}, "since": 1100}
Evidence: Above-budget load report

Source: Above-budget load report

{"load": {"cpus": 4, "first": 1100, "last": 1300, "load1_max": 9.0, "load1_p50": 7.0, "load1_p95": 9.0, "load_within_2x_cpus": false, "pool_held_max": 3, "pool_held_mean": 1.67, "samples": 3, "share_above_2x_cpus": 0.3333}, "pipeline": {"agent_failures": [{"category": "timeout", "count": 1, "exit_status": "error"}], "agent_minutes_by_purpose": {"review": {"count": 21, "minutes_p50": 10.0, "minutes_p95": 10.0}, "review-fix": {"count": 12, "minutes_p50": 10.0, "minutes_p95": 10.0}}, "cohort_runs": [{"agent_minutes": 10.0, "converged": false, "created_at": 1999, "review_fix_rounds": 0, "run": "older-live", "status": "failed", "test_fix_rounds": 0, "timed_out": true}, {"agent_minutes": 40.0, "converged": false, "created_at": 2000, "review_fix_rounds": 3, "run": "run-00", "status": "completed", "test_fix_rounds": 0, "timed_out": false}, {"agent_minutes": 20.0, "converged": true, "created_at": 2001, "review_fix_rounds": 1, "run": "run-01", "status": "completed", "test_fix_rounds": 0, "timed_out": false}, {"agent_minutes": 30.0, "converged": true, "created_at": 2002, "review_fix_rounds": 2, "run": "run-02", "status": "completed", "test_fix_rounds": 0, "timed_out": false}, {"agent_minutes": 10.0, "converged": true, "created_at": 2003, "review_fix_rounds": 0, "run": "run-03", "status": "completed", "test_fix_rounds": 0, "timed_out": false}, {"agent_minutes": 20.0, "converged": true, "created_at": 2004, "review_fix_rounds": 1, "run": "run-04", "status": "completed", "test_fix_rounds": 0, "timed_out": false}, {"agent_minutes": 30.0, "converged": true, "created_at": 2005, "review_fix_rounds": 2, "run": "run-05", "status": "completed", "test_fix_rounds": 0, "timed_out": false}, {"agent_minutes": 10.0, "converged": true, "created_at": 2006, "review_fix_rounds": 0, "run": "run-06", "status": "completed", "test_fix_rounds": 0, "timed_out": false}, {"agent_minutes": 20.0, "converged": true, "created_at": 2007, "review_fix_rounds": 1, "run": "run-07", "status": "completed", "test_fix_rounds": 0, "timed_out": false}, {"agent_minutes": 30.0, "converged": true, "created_at": 2008, "review_fix_rounds": 2, "run": "run-08", "status": "completed", "test_fix_rounds": 0, "timed_out": false}], "converged_within_2_fix_rounds": false, "runs_considered": 10, "runs_wanted": 10, "timeout_class_run_errors": [{"class": "timed out", "run": "older-live"}]}, "since": 1100}
- Outcome: 🔧 2 issues found → auto-fixed ✅ across 2 runs (19m56s)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

🔧 **Rebase** - 1 issue found → auto-fixed ✅
  • ⚠️ docs/fm-test-portable-shards.md - merge conflict rebasing onto origin/main

🔧 Fix applied.
✅ Re-checked - no issues remain.

🔧 **Review** - 7 issues found → auto-fixed (5) ✅
  • 🚨 bin/fm-test-run.sh:2538 - The intent requires "Phase 1, which must land together: F2 ... and F6", including Vernant worker reservations, gate sub-agent concurrency 1, and advisor/eager-sub-agent disabling for pipeline agents only. This diff implements Firstmate reservations, while docs/cpu-pass-pool.md:53 merely tells other repositories how to participate. No companion integration or gate configuration is established by the change; commit 3a70ee8 explicitly says F6 is external and "this change does not depend on it", contradicting the required coordinated cutover. Establish the companion implementation and joint-cutover dependency, or obtain authorization to split Phase 1.
  • 🚨 bin/fm-cpu-pass.py:243 - Only the wrapper owns the locks, and close_fds=True prevents its command from retaining them. If the wrapper crashes or receives SIGKILL while CPU-heavy work continues, the kernel immediately frees its passes and another burst starts alongside the surviving work. The new test constructs exactly this state at tests/fm-cpu-pass.test.sh:169-173: it asserts zero held passes before explicitly killing the still-running child. This violates the reservation-until-work-ends invariant in docs/cpu-pass-pool.md:4 and bin/fm-cpu-pass.py:10. Related ownership sites are bin/fm-cpu-pass.py:116 (close-on-exec locks) and :323 (wrapper-owned release); bin/fm-test-run.sh:2538 uses this wrapper for every script. Make reservation ownership survive with executing work at the shared child-launch boundary, and correct the crash assertion.
  • 🚨 tests/fm-cpu-pass.test.sh:26 - Failure cleanup sends TERM only to recorded wrapper PIDs, but bin/fm-cpu-pass.py:239 deliberately ignores TERM while its child runs. An assertion failure while a holder is waiting for its release file therefore leaves both wrapper and workload alive; deleting the fixture removes the possibility of release. Their inherited output descriptors can also keep the runner's streaming tee waiting indefinitely. Affected holder launches are tests/fm-cpu-pass.test.sh:99, :139, :163, :183, :205, :231, :244 and :304; other tracked background launches are :110, :144, :210 and :309. The crash case's child cleanup at :172-173 occurs only on success. Track and terminate the fixture workloads or their isolated groups, then wait for termination before removing fixtures.
  • ⚠️ bin/fm-load-report.py:163 - A run that reached review, was cancelled or failed for a non-timeout reason, and had zero to two review-fix invocations is reported as converged. Ten such unsuccessful runs produce converged_within_2_fix_rounds=true. Additionally, the selection at :141 includes pending runs, which bin/fm-nm-run-lib.sh:125-128 explicitly classifies as live. Determine eligibility and successful convergence at pipeline_section's shared boundary: exclude live runs and do not mark failed/cancelled runs as successful. Related changed consumers are bin/fm-load-report.py:196 (aggregate verdict), :226 (text verdict), :27-34 (documented selection), docs/cpu-pass-pool.md:62 (measurement recipe), and tests/fm-load-report.test.sh:29, :76-77 (fixture statuses and verdict expectations).
  • ⚠️ bin/fm-cpu-pass.py:83 - Simplification: FM_CPU_POOL_SIZE introduces a production sizing authority beyond the required pool "with its size taken from hw.ncpu". For example, two participants using the same directory with sizes 18 and 36 can hold up to 36 distinct slots on an 18-core host; the turnstile does not reconcile their budgets. No intent requirement needs this override. Remove it and use the host CPU count consistently. Related sites are bin/fm-cpu-pass.py:297 and :332 (run/status sizing), docs/cpu-pass-pool.md:26-27 (override contract), and tests/fm-cpu-pass.test.sh:67-71 plus its per-case FM_CPU_POOL_SIZE assignments.
  • ⚠️ bin/fm-cpu-pass.py:293 - Simplification: FM_CPU_POOL=off adds an unconditional bypass of the host-wide budget, including extra case-insensitive and whitespace-normalized spellings. A heavy suite can execute immediately while every pass is held, as tests/fm-cpu-pass.test.sh:244-251 demonstrates. The intent exempts interactive agents, not CPU-heavy test bursts, and supplies no requirement for this opt-out. Remove the bypass while retaining the separately necessary nested-work handling. Related sites are bin/fm-test-run.sh:127 and :2402, docs/cpu-pass-pool.md:37 and :54, and tests/fm-cpu-pass.test.sh:249-255.
  • ⚠️ bin/fm-cpu-pass.py:301 - Simplification: silently clamping an oversized reservation is not required by the CPU-budget intent and can under-account the command's actual workers. On a two-core host, the documented run --passes 4 -- pytest -n 4 takes only two passes but still starts four workers; clamping the reservation does not resize pytest. Remove silent clamping and use the narrower exact-count contract: callers size their workers using size and reserve that same positive count. Related sites are bin/fm-cpu-pass.py:30 and :375, docs/cpu-pass-pool.md:34 and :48, and tests/fm-cpu-pass.test.sh:80-84, which currently endorses oversized clamping.

🔧 Fix applied.
3 warnings still open:

  • ⚠️ bin/fm-load-report.py:163 - A run that reached review, was cancelled or failed for a non-timeout reason, and had zero to two review-fix invocations is reported as converged. Ten such unsuccessful runs produce converged_within_2_fix_rounds=true. Additionally, the selection at :141 includes pending runs, which bin/fm-nm-run-lib.sh:125-128 explicitly classifies as live. Determine eligibility and successful convergence at pipeline_section's shared boundary: exclude live runs and do not mark failed/cancelled runs as successful. Related changed consumers are bin/fm-load-report.py:196 (aggregate verdict), :226 (text verdict), :27-34 (documented selection), docs/cpu-pass-pool.md:62 (measurement recipe), and tests/fm-load-report.test.sh:29, :76-77 (fixture statuses and verdict expectations).
  • ⚠️ bin/fm-cpu-pass.sh:14 - Round 1 fixed reservation validation in the Python engine but left the no-python sibling behind. This loop discards --passes without validating it, so with python3 unavailable, fm-cpu-pass.sh run --passes 0 -- COMMAND executes COMMAND and returns its status instead of refusing work with 125. Negative, malformed, and oversized counts likewise bypass validation, in both split and equals forms. Validate the count before the degraded exec, preserving the authorized fallback for valid requests. Related invariant sites: bin/fm-cpu-pass.sh:29 (unchecked execution), bin/fm-cpu-pass.py:290 (normal-path validation), docs/cpu-pass-pool.md:36 (invalid counts never start work), tests/fm-cpu-pass.test.sh:124 and :129 (normal and nested refusal cases), and tests/fm-cpu-pass.test.sh:327 (no-python case currently exercises only a valid request).
  • ⚠️ bin/fm-load-report.py:144 - The intent requires that "the next 10 runs converge in at most 2 fix rounds", but the added query uses ORDER BY created_at DESC LIMIT ?, measuring the latest ten instead. With twenty eligible runs after the recorded cutover epoch, a three-round failure in the first ten disappears when the last ten converge, and the report returns true despite the required cohort failing. Round 1's false-convergence fix corrected terminal-status eligibility but left this independent window-selection issue unchanged. Use the first ten eligible runs after --since for the acceptance verdict; retaining a separate rolling cohort would need authorization. Related sites: bin/fm-load-report.py:26 (most-recent selection contract), :196 (aggregate verdict), docs/cpu-pass-pool.md:69-71 (fixed cutover epoch paired with latest-ten reporting), and tests/fm-load-report.test.sh:75-83 (exactly ten eligible fixtures cannot distinguish the cohorts).

🔧 Fix applied.
4 warnings still open:

  • ⚠️ bin/fm-load-report.py:163 - A run that reached review, was cancelled or failed for a non-timeout reason, and had zero to two review-fix invocations is reported as converged. Ten such unsuccessful runs produce converged_within_2_fix_rounds=true. Additionally, the selection at :141 includes pending runs, which bin/fm-nm-run-lib.sh:125-128 explicitly classifies as live. Determine eligibility and successful convergence at pipeline_section's shared boundary: exclude live runs and do not mark failed/cancelled runs as successful. Related changed consumers are bin/fm-load-report.py:196 (aggregate verdict), :226 (text verdict), :27-34 (documented selection), docs/cpu-pass-pool.md:62 (measurement recipe), and tests/fm-load-report.test.sh:29, :76-77 (fixture statuses and verdict expectations).
  • ⚠️ bin/fm-load-report.py:144 - The intent requires that "the next 10 runs converge in at most 2 fix rounds", but the added query uses ORDER BY created_at DESC LIMIT ?, measuring the latest ten instead. With twenty eligible runs after the recorded cutover epoch, a three-round failure in the first ten disappears when the last ten converge, and the report returns true despite the required cohort failing. Round 1's false-convergence fix corrected terminal-status eligibility but left this independent window-selection issue unchanged. Use the first ten eligible runs after --since for the acceptance verdict; retaining a separate rolling cohort would need authorization. Related sites: bin/fm-load-report.py:26 (most-recent selection contract), :196 (aggregate verdict), docs/cpu-pass-pool.md:69-71 (fixed cutover epoch paired with latest-ten reporting), and tests/fm-load-report.test.sh:75-83 (exactly ten eligible fixtures cannot distinguish the cohorts).
  • ⚠️ bin/fm-load-report.py:146 - Round 2's oldest-first fix still recomputes the cohort from current eligibility. Concrete sequence: an older run remains running while ten newer runs complete successfully; the report returns true. When the older run subsequently fails after reaching review, it enters the oldest-first cohort, displaces a successful run, and changes the verdict to false. This contradicts the recorded requirement: "once 10 eligible runs exist the verdict is fixed." Related sites: bin/fm-load-report.py:143 (live-run exclusion), :197 (recomputed verdict), :202 (JSON cohort), :225 (text verdict), :35 (stability promise); docs/cpu-pass-pool.md:74 (same promise); tests/fm-load-report.test.sh:37 and :83 (older live fixtures never transition). Address this at pipeline_section's cohort-selection boundary. Freezing cohort membership and its verdict requires durable state keyed by cutover; that remedy, rather than the defect, needs authorization because the reporter currently promises read-only operation.
  • ⚠️ bin/fm-load-report.py:276 - Simplification: the new --runs option makes the acceptance cohort arbitrarily configurable, although the recorded decision requires "the first 10 eligible runs after --since." With one successful eligible run, report --runs 1 returns converged_within_2_fix_rounds=true without the required ten-run evidence. No stated intent requires an alternate acceptance cohort. Remove this option and use the narrower fixed-ten contract, rather than maintaining another verdict mode. Related sites: bin/fm-load-report.py:8 and :26 (public option contract), :146 (variable query limit), :197 and :201 (variable verdict threshold and output), :251 (argument forwarding), :297 (option validation).

🔧 Fix applied.
2 issues (1 error, 1 warning) still open:

  • 🚨 docs/cpu-pass-pool.md:64 - The intent requires Phase 1 to "land together" and explicitly includes a CPU pool "used by Vernant's pytest -n worker sizing ... and by Firstmate's fm-test-run.sh". This added line instead defers Vernant participation to a follow-up, while docs/cpu-pass-pool.md:63 declares the coordinated cutover satisfied by Firstmate plus F6 alone. Round 1 introduced this cutover section, but it leaves a required CPU-demand source outside the documented rollout. Coordinate the Vernant participant before declaring Phase 1 complete, or obtain explicit authorization for narrower Firstmate-only containment.
  • ⚠️ bin/fm-cpu-pass.py:198 - Collected slots retain their previous holder records until the entire reservation is acquired. With two slots, let an earlier run leave its label in slot 1, another run hold slot 0, and a two-pass request collect slot 1 while waiting for slot 0. Status and periodic wait notices now identify the finished earlier run as a current holder, rather than the waiting collector; this persists throughout the wait and misdirects timeout diagnosis. Write the current reservation metadata when each slot is acquired. Related sites: bin/fm-cpu-pass.py:215-220 (deferred metadata publication), :147 (reading stale records), :175 (wait-notice consumer), :335-344 (status consumers), and tests/fm-cpu-pass.test.sh:195 (partial-collection coverage currently checks only the held count).

🔧 Fix applied.
1 error still open:

  • 🚨 bin/fm-test-run.sh:2402 - Nested execution bypasses reservations without limiting concurrency to the inherited pass count. Concrete intended path: an outer runner reserves one pass for tests/fm-test-run.test.sh; that test invokes the real runner with --jobs 2 over fm-session-lock-ancestry.test.sh and fm-task-inbox.test.sh (tests/fm-test-run.test.sh:1628, executed at :2052). Both nested scripts run concurrently with FM_CPU_PASS_HELD=1 and take no additional passes. With the other host slots occupied, the fleet executes more scripts than its CPU budget while status reports a full, correctly sized pool. Round 1's reservation-validation fix left this sibling invariant unchecked: bin/fm-cpu-pass.py:289 validates against host size, but :292 permits a nested --passes 2 request under an inherited one-pass reservation. Related changed sites: bin/fm-test-run.sh:2537 (reservation bypass), docs/cpu-pass-pool.md:36-40 (exact worker-count and nesting contracts), and tests/fm-cpu-pass.test.sh:483-498 (nested coverage exercises only one worker). Enforce the inherited reservation at the shared nested-dispatch boundary: nested concurrency must not exceed its reserved count, and callers needing parallel work must reserve sufficient passes before starting the outer workload rather than acquiring more while holding an insufficient reservation.

🔧 Fix applied.
✅ Re-checked - no issues remain.

🔧 **Test** - 2 issues found → auto-fixed ✅
  • ⚠️ bin/fm-cpu-pass.sh:75 - Live execution with Python absent from a restricted PATH ran the workload with FM_CPU_PASS_HELD=0 but produced empty stderr instead of the promised degradation notice. The group's 2>/dev/null also redirects the notice when log_fd defaults to 2. An explicit --log-fd 1 emitted the notice, isolating this to default-stderr handling. This silently removes the CPU budget without telling the operator, undermining timeout diagnosis. Preserve the original notice destination while suppressing only write errors, and add a public-command regression that checks default stderr rather than only a separate log fd.
  • 🚨 live validation verdict: no-go (13 of 13 scenarios were driven live against the product); failed: Announce no-Python degradation on the default stderr notice fd
  • Live validation: ❌ no-go - 13 of 13 scenarios driven live against the product
Scenario Result Live Evidence
Size a test reservation from the actual logical CPU count ✅ pass live Live command transcript: size returned 18, matching os.cpu_count().
Queue competing bursts and report the current partial collector rather than a finished holder ✅ pass live Live command transcript: status identified active-holder and current-collector; both wait notices excluded finished-earlier; queued commands ran after acquisition, and unreserved work ran while the po…
Kill a reservation wrapper while its workload survives and retain its passes until workload exit ✅ pass live Live command transcript: held remained 18 after wrapper SIGKILL and became 0 after the surviving child exited.
Cancel a queued burst without starting its workload ✅ pass live Live command transcript: the waiter terminated on SIGTERM, its workload marker was absent, and the pool returned to zero held passes.
Request parallel nested scripts under one inherited pass and observe serialized execution ✅ pass live Live command transcript: explicit --jobs 2 and automatic concurrency each produced start/end/start/end events, one reduction notice on stderr, and timing selection jobs=1 while the other 17 passes wer…
Reject oversized reservations and malformed inherited counts before work starts ✅ pass live Live command transcript: zero passes, 19 passes, nested two-pass requests under one pass, and malformed engine/runner markers exited 125 without starting work; no-Python validation also refused invali…
Wait outside a script deadline while preserving skip output and bounding actual overruns ✅ pass live Live command transcript: a runner queued beyond its one-second script bound and then completed with exit=0 and gate_skip=true; an actual overrun produced script exit=124, runner exit=1, and released i…
Run top-level parallel workers with separate passes and propagate a worker failure ✅ pass live Parallel runner artifact: two workers overlapped with one pass each alongside a 16-pass holder; worker exits 0 and 7 produced runner exit 1.
Continue degraded work when the pool or Python is unavailable ✅ pass live Live command transcript: an unusable pool emitted a notice, exported held=0, and preserved workload exit 7; the real no-Python shell fallback ran with held=0 and emitted a notice when given an explici…
Announce no-Python degradation on the default stderr notice fd ❌ fail live Live command transcript: the default no-Python invocation returned degraded-held=0 with empty stderr; --log-fd 1 emitted the expected python3-not-found notice.
Record actual host load and current pool utilization ✅ pass live Actual host-load samples: record and watch appended real load values with cpus=18, pool_size=18, and pool_held=2.
Report matching tri-state convergence verdicts for the fixed first-ten cohort without modifying the database ✅ pass live Live transcript and report-boundary artifact: an earlier live nonreview run yielded pending, its review-reaching failure yielded false, later successes did not alter the settled cohort, ten successful…
Expose excessive load and unavailable measurement inputs without reporting false success ✅ pass live Reporting artifacts: below/above twice-CPU samples produced true/false load verdicts; --runs was rejected; an absent database produced exit 1 and pipeline unavailable while preserving load facts and c…
  • bash tests/fm-cpu-pass.test.sh with a worktree-local TMPDIR; passed.
  • bash tests/fm-load-report.test.sh with a worktree-local TMPDIR; passed.
  • python3 .live-validation/drive.py drove the real CPU-pool and reporting commands with an isolated pool at the actual 18-CPU host size; retained the failed default-notice observation and completed the remaining scenarios.
  • fm-test-run.sh --jobs 2 tests/fm-brief.test.sh tests/fm-composer-lib.test.sh in a disposable repository containing unchanged product scripts: observed two separate one-pass reservations while 16 other passes were held, concurrent execution, and propagation of a deliberate worker exit 7.
  • fm-load-report.sh report --samples <disposable.tsv> --since 1000 --nm-db <disposable.sqlite> in JSON and text modes: exercised pending, true, and false convergence, cohort stability, timeout classification, load thresholds, and byte-for-byte database read-only behavior.
  • fm-load-report.sh report --runs 1 refused the removed option; reporting against an absent database exited 1, retained load facts, and did not create the database.
  • Confirmed the isolated live pool had zero held passes, then removed all disposable drivers, fixture repositories, pools, databases, restricted-PATH tools, and temporary data.

🔧 Fix applied.
✅ Re-checked - no issues remain.

  • Live validation: ✅ go - 12 of 13 scenarios driven live against the product
Scenario Result Live Evidence
Read the host-sized pool and receive the requested reservation ✅ pass live Host-sized CPU-pool CLI transcript; focused CPU-pool script exercised child markers and exit-status propagation.
Queue behind occupied passes and identify the current partial collector ✅ pass live Host-sized CPU-pool CLI transcript; tests/fm-cpu-pass.test.sh partial-collection scenario.
Cancel queued work without starting its command ✅ pass live Host-sized CPU-pool CLI transcript.
Retain a running child's reservation after wrapper death ✅ pass live Host-sized CPU-pool CLI transcript; focused CPU-pool signal scenarios.
Refuse invalid reservations and nested over-requests before work starts ✅ pass live Host-sized CPU-pool CLI transcript; tests/fm-cpu-pass.test.sh.
Announce degraded execution on default stderr while preserving workload output and status ✅ pass live Host-sized CPU-pool CLI transcript.
Reject malformed inherited markers in Python, no-Python, and runner execution ✅ pass live Host-sized CPU-pool CLI transcript; tests/fm-cpu-pass.test.sh inherited-marker scenarios.
Request two nested runner jobs under one pass without exceeding the inherited budget ✅ pass live Nested and queued runner transcript, including persisted start/end events.
Wait for a pass outside the script timeout while retaining the timeout for hung work ✅ pass live Nested and queued runner transcript.
Record real host load and pool occupancy with record and watch ✅ pass live Actual host-load samples; recorder and convergence reporter transcript.
Keep convergence pending until earlier live runs settle, then preserve the first-ten verdict ✅ pass live Pending, settled unsuccessful, and unchanged-later-successes JSON artifacts; reporter transcript.
Report load and convergence failures without allowing a smaller acceptance cohort ✅ pass live Above-budget load report; reporter transcript; tests/fm-load-report.test.sh.
Validate the coordinated Phase 1 rollout and sustained fleet acceptance ⏸️ untested no This change delivers the Firstmate participant and explicitly requires the separate Vernant-repository participant before measurement begins. Disposable checks validated Firstmate's runtime and report…
  • TMPDIR=<worktree disposable tmp> bash tests/fm-cpu-pass.test.sh
  • TMPDIR=<worktree disposable tmp> bash tests/fm-load-report.test.sh
  • bin/fm-cpu-pass.sh size and status --json against a private pool: confirmed the actual 18-CPU host size.
  • Public CPU-pool commands with a 17-pass holder, a partially collecting two-pass request, and a queued waiter; inspected status and notices, cancelled the waiter, and released the holder.
  • Killed a CPU-pool wrapper while its child remained running; inspected the retained reservation, terminated the child, and observed release.
  • Exercised invalid counts, nested over-requests, malformed inherited markers, an unusable pool directory, and a restricted PATH without Python.
  • Ran the real runner in a disposable repository under a real one-pass reservation with --jobs 2; inspected execution events, notice output, and timing JSON.
  • Ran the real runner with --per-script-timeout-secs 3 while all 18 passes were occupied; released the pool after the script bound had elapsed, then separately exercised a hung script with a one-second bound.
  • bin/fm-load-report.sh record and watch --interval .2 against a private samples file; inspected actual host samples.
  • Ran text and JSON reports against disposable samples and SQLite state through pending, failed, and later-successful run transitions; checked database hashes for read-only behavior.
  • Exercised load-threshold verdicts, successful and excessive-fix cohorts, rejected --runs options, and an absent database.
  • Stopped validation processes, confirmed zero held passes, and removed all disposable worktree fixtures.
✅ **Document** - passed

✅ No issues found.

🔧 **Lint** - 1 issue found → auto-fixed ✅
  • ⚠️ linter found issues (exit code 1)

🔧 Fix applied.
✅ Re-checked - no issues remain.

✅ **Push** - passed

✅ No issues found.

Host CPU oversubscription (load 250-482 on 18 cores) was the shared
condition behind the 2026-10-07 agent timeouts: every worktree's test
runner sized itself to the whole host and nothing owned total demand.
This is Phase 1 F2 of that root-cause plan. It bounds CPU-heavy test
bursts, not agents: interactive agents, lanes and sessions never take a
pass and are never queued or capped.

- bin/fm-cpu-pass.sh (+ .py engine): `run [--passes K] -- CMD` holds K
  passes from one host-wide pool while CMD runs; `status`; `size`.
  Passes are fcntl.flock locks on slot files in $HOME/.cache/fm-cpu-pool,
  pool size = CPU count, so a killed holder never leaks a pass. A request
  queues and never fails; a broken pool degrades to running without a
  pass, with one notice.
- bin/fm-test-run.sh: each executed script takes one pass, outside its
  per-script bound, so waiting never trips that bound; notices go to the
  runner's stderr, never into captured script output. Runs nested inside
  a pass, with FM_CPU_POOL=off, without python3, or from a copy without
  the tool run directly.
- docs/cpu-pass-pool.md owns the cross-repository protocol, so Vernant's
  pytest -n sizing (d7) can join the same pool.
- bin/fm-load-report.sh (+ .py engine): records load samples and reports
  load p95 against 2x CPUs plus pipeline fix-round convergence and
  timeout-class run errors from the no-mistakes database (read-only), to
  judge Phase 1 after 24-48 h.
- Thirteen portable-serial duration hints from green CI run 37404365422
  keep the unhinted share under the coverage guard's 15% limit now that
  the serial lane gains the two new tests (provenance in
  docs/fm-test-portable-shards.md).

F6 (omp advisor and eager sub-agents off for pipeline agents only) is not
in this repository: the gate overlay is written by the no-mistakes binary.
It ships as a no-mistakes binary change applied by Main to the private
build; this change does not depend on it. No timeout was raised and the
shared daemon was not touched.

System check:
- Callers of bin/fm-test-run.sh that execute scripts: CI lanes in
  .github/workflows/ci.yml (all serial on their own runners, so one pass
  at a time from an uncontended pool), no-mistakes Test agents and
  workers running it locally (now take turns host-wide), and the
  runner's own fixture tests (copies without the tool run directly).
  Inspection callers (bin/fm-test-isolation-proof.sh --list modes,
  tests/fm-ci-workflow.test.sh --list-lanes) execute nothing and are
  unchanged.
- Blast radius: a local run may now wait for a pass when the host is
  busy; the wait is outside per-script bounds but inside a caller's own
  command limit and --max-wall-ms. --jobs above the pool size runs only
  pool-size scripts at once. Timing durations include any wait.
- Deadlock guards: nested runners inherit FM_CPU_PASS_HELD and take no
  pass; multi-pass requests collect behind one turnstile.
- Tests: tests/fm-cpu-pass.test.sh (exclusive passes, queued waiter,
  multi-pass, SIGKILL release, TERM keeps pass with running work, waiter
  TERM never runs, nested and opt-out, degraded and no-python paths,
  runner wait outside the bound, gate-skip detection unchanged, bound
  still fires, nested runner) and tests/fm-load-report.test.sh, plus the
  existing runner, fixture, isolation-proof, timeout-lib, documentation
  and CI-workflow suites.
Only an unset or blank FM_CPU_POOL_SIZE falls back to the CPU count;
any other non-positive-integer value is a usage error, as
bin/fm-cpu-pass.sh already enforces. Wording only.
…the CPU-pass wrapper's inherited slot-lock descriptor shifted a live extension claim handle onto fd 7, which the capture handoff then overwrote, preventing terminal source retirement. Reproduced the failure locally with fd 4 occupied. The shared Perl handoff now reserves destination descriptors before allocating input handles, fixing both result.silent and result.terminal paths. Added public-command regression coverage for terminal retirement and silent acknowledgement with inherited descriptors, and updated verification documentation. Verification passed: the complete fm-extension-binding suite through the CPU-pass-enabled fm-test-run.sh runner, tests/fm-cpu-pass.test.sh, shellcheck -x on the changed test, and Perl syntax checking. No timeouts or budgets changed. The remote CI check was not rerun; verification was local on macOS
…n bin/fm-test-run.sh: raised the stale fm-supervision-host duration hint from 41512 to 877426 ms and added fm-skill-pick at 120549 ms. Both hints remain sorted and unique. Updated docs/fm-test-portable-shards.md with one-sentence-per-line provenance citing supervision-host measurements from main runs 37772520915 (562966 ms), 37764027766 (853538 ms), and 37774962436 (877426 ms), plus PR runs 37774432736 (549 s) and 37778222434 (859832 ms); skill-pick measured 120549 ms in run 37778222434. These measurements and run IDs are included here for the outer executor's commit message. Verified the packer's public --list-scheduled interface over a disposable three-way merged snapshot of this branch and available origin/main (32fbe8e), including main's skill-pick test. Per-shard hinted totals in milliseconds, shards 1–9: 877426, 864666, 864675, 864667, 864672, 864654, 864669, 864670, 864637. These are 14.41–14.62 minutes, with a 12789 ms spread and at least 15m22.574s estimated headroom below 30 minutes. Supervision-host occupies shard 1 alone; skill-pick is on shard 3. The merged partition is complete and disjoint. These are packing estimates, not measured CI runtimes. Verification passed: full bash tests/fm-test-run.test.sh on the branch; --check-coverage on both branch and merged snapshot; bash -n bin/fm-test-run.sh. The initial suite invocation encountered a local Python-launcher issue before its first assertion; the complete suite passed using a resolved Python binary through a temporary worktree-local PATH entry. All verification scaffolding was removed. No test code, timeout, --max-wall-ms budget, or shard count changed. No remote CI or pipeline-control command was invoked
@MrGTV-love
MrGTV-love merged commit 6fe80c3 into main Oct 8, 2026
21 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant