Skip to content

test: remove read-only strip hooks and staged launch dirs in fixture teardown - #16

Merged
cloud-practitioner merged 8 commits into
mainfrom
fm/fm-e2e-tmp-hooks-cleanup
Oct 1, 2026
Merged

cloud-practitioner merged 8 commits into
mainfrom
fm/fm-e2e-tmp-hooks-cleanup

Conversation

@cloud-practitioner

@cloud-practitioner cloud-practitioner commented Sep 29, 2026 •

Copy link
Copy Markdown
Owner

Intent

Fix test hygiene and flakiness found along the way: tests/fm-backend-herdr-presentation-e2e.test.sh cleanup leaves read-only git-hooks directories in /tmp. The earlier fix cleans up read-only strip hooks and staged launch dirs in test teardown, already open as #16. That earlier fix (pull request 16) makes test teardown remove fixture trees even when they contain the read-only state/.git-hooks strip-hook directories that real spawns create, and also remove the launch directories those spawns staged outside the fixture tree, so test runs no longer leave those directories behind in /tmp.

What Changed

  • Added tests/fixture-tree-helpers.sh, sourced by both tests/lib.sh and tests/herdr-test-safety.sh. It provides fm_test_remove_tree, which used to live in lib.sh. It also adds fm_test_home_hash and fm_test_remove_spawn_launch_dirs. That last helper finds every Firstmate home under a fixture root and removes the /tmp/fm-<id>+<home-hash> launch directories that real spawns staged outside the tree. It returns 0 even when there is nothing to remove, so it is safe under set -e.

  • Shared cleanup in lib.sh (fm_test_cleanup and fm_test_reap_orphans) now removes those launch directories before deleting each tree. So do the Herdr suites: presentation e2e, launcher-workspace e2e, workspace-per-home e2e, autodetect smoke and control smoke. Previously they ran a plain rm -rf or an inline chmod, which left read-only state/<id>.git-hooks directories and launch dirs behind in /tmp. The presentation e2e now fails if cleanup leaves its fixture tree or any of its staged launch directories behind. Before the duplicate-live-agent check, it also starts a sleep 600 foreground process in the pane so Herdr does not expire that registration.

  • Added regression cases to tests/fm-test-fixtures.test.sh:

    • removing a tree that holds real read-only strip hooks, without changing a directory reached through a symlink
    • exit cleanup of a spawn that was never torn down and whose task record was deleted
    • set -e exit cleanup when no launch directory was staged

    Also mapped the new helper in bin/fm-test-run.sh's changed-path family lookup and listed it among the shared test helpers in CONTRIBUTING.md.

🤖 Generated with Claude Code

Risk Assessment

✅ Low: The change only touches test teardown helpers. Launch-dir removal uses the same physical-path sha256 token that fm-spawn.sh and fm-teardown.sh use (bin/fm-spawn.sh:5295), so it only reaches the fixture's own homes. Symlinks are never followed when restoring permissions, and the only problem found is an inaccurate comment.

Testing

I removed the scratch repro file as the operator asked. I ran the real-Herdr presentation E2E to completion on Herdr 0.9.3 in an isolated fm-lab session and every scenario passed. That includes the changed duplicate-live-agent step: with a real sleep 600 foreground process keeping the registration alive, live duplicate risk refused launch. The suite's final assertion also passed: cleanup removed the whole fixture tree, read-only strip-hook directories included, and every launch dir this run staged in /tmp. The fixture-cleanup test passed, covering the read-only removal, set -e safety and outside-target guards. Afterwards no lab Herdr session was left and the default session was untouched. The new /tmp entries in the before/after snapshot (fm-rl*, fm-sm*, fm-control-, fm-resume-) belong to suites I did not run, such as fm-captain-hold-lifecycle and fm-control-relaunch, so other runs on this machine likely created them. This change has no UI, so there are no screenshots.

  • Live validation: ✅ go - 3 of 5 scenarios driven live against the product
Scenario Result Live Evidence
Scratch repro file is removed and will not be committed ⏸️ untested no The prior payload recorded this as a worktree hygiene check (git status), not a scenario driven against the live product, so it did not establish a live result.
Duplicate-live-agent pane keeps its Herdr 0.9.3 registration because a real foreground process runs, so the live duplicate risk reliably refuses launch ✅ pass live presentation-e2e.log: ok - real Herdr lab: missing, renamed, and duplicate tokens trigger zero destructive or adoptive calls, and live duplicate risk refuses launch
Real-spawn E2E teardown removes the whole fixture tree, including read-only state/<id>.git-hooks, and its staged /tmp launch dirs ✅ pass live presentation-e2e.log: ok - cleanup removed the whole fixture tree, including read-only strip-hook directories, and its staged launch directories, rc=0
Lab teardown leaves no Herdr lab session behind and the default session untouched ✅ pass live herdr session list after the run lists only default; log reports default-session tripwire intact
Shared teardown helpers remove read-only strip-hook trees and staged launch dirs, finish under set -e, and leave outside targets alone ⏸️ untested no The prior payload ran only the fixture-level automated test with fake binaries (tests/fm-test-fixtures.test.sh), so it did not establish a live result. The live presentation E2E scenario above covers…
Evidence: Full real-Herdr presentation E2E transcript (Herdr 0.9.3, rc=0)

Source: Full real-Herdr presentation E2E transcript (Herdr 0.9.3, rc=0)

ok - real Herdr lab: missing, renamed, and duplicate tokens trigger zero destructive or adoptive calls, and live duplicate risk refuses launch ok - real Herdr lab validation completed on Herdr 0.9.3 with the default-session tripwire intact ok - cleanup removed the whole fixture tree, including read-only strip-hook directories, and its staged launch directories rc=0 elapsed=728s

ok - real Herdr lab: an opted-out spawn retains the Stage 1 Herdr command sequence with zero ordering calls
ok - real Herdr lab: a home that configured nothing is projected by default on herdr 0.9.3
ok - real Herdr lab: every projected create, task-tab create, seeded prune, and move preserves active workspace and tab
ok - real Herdr lab: persisted-focused seeded prune proceeds when no live client is attached
ok - real Herdr lab: bounded lock contention warns and falls back flat without projection or focus drift
ok - real Herdr lab: concurrent primary workers form one stable contiguous block without active workspace/tab drift
ok - real Herdr lab: forced workspace.move failure leaves a successful worker in default order with a warning and no cleanup
ok - real Herdr lab: concurrent post-create abort cleanup stays serialized with exact focus restoration
ok - real Herdr lab: Treehouse commands and metadata shape are byte-identical except for endpoint IDs and spawn incarnation
ok - real Herdr lab: exact task-pane close removes the projected workspace with no unrestored wrong-focus interval
ok - real Herdr lab: concurrent projected cleanup is serialized and leaves active workspace/tab unchanged
ok - real Herdr lab: three repeated concurrent create/order/cleanup waves have zero active workspace or tab drift
ok - real Herdr lab: the primary presentation setting inherits into real secondmate homes
ok - real Herdr lab: primary and two secondmate homes each own a top-level contiguous child block
ok - real Herdr lab: concurrent primary/A/B spawns preserve parent order and exact focus
ok - real Herdr lab: session lock contention from a secondmate home falls back flat with no journal
ok - real Herdr lab: Hi Bit and Wheelhouse-style same-identity restarts reclaim one nested space with exact focus and idempotence
ok - real Herdr lab: secondmate restart binding and reclaim stay isolated to the exact child home and parent
ok - real Herdr lab: concurrent cross-home recoveries replace exact husks under one session lock with no focus drift
ok - real Herdr lab: legacy projection labels and flat secondmate tabs are left unmigrated
ok - real Herdr lab: multi-home exact-pane teardowns restore captain focus without workspace close authority
warning: no exact herdr presentation token match for missing1; leaving any stale space untouched and spawning flat
warning: no exact herdr presentation token match for renamed1; leaving any stale space untouched and spawning flat
warning: 2 exact herdr presentation token matches for duplicate1 are quarantined; inspecting only for duplicate-agent risk
warning: quarantined herdr presentation for duplicate1 is dead or agent-free; exact bound reclaim may proceed, otherwise spawning flat
warning: 2 exact herdr presentation token matches for duplicate1 are quarantined; inspecting only for duplicate-agent risk
error: quarantined herdr presentation for duplicate1 has a live pane; refusing duplicate launch
ok - real Herdr lab: missing, renamed, and duplicate tokens trigger zero destructive or adoptive calls, and live duplicate risk refuses launch
ok - real Herdr lab validation completed on Herdr 0.9.3 with the default-session tripwire intact
ok - cleanup removed the whole fixture tree, including read-only strip-hook directories, and its staged launch directories
rc=0 elapsed=728s
Evidence: Fixture cleanup test log

Source: Fixture cleanup test log

ok - runner and shared helpers isolate host Git config and preserve explicit config and outside commits
ok - suite exit cleanup removes an untorn spawn's read-only strip hooks and staged launch directory
ok - a set -e suite with no staged launch directories finishes exit cleanup and exits 0
ok - fm_test_remove_tree removes a tree holding read-only strip hooks and leaves outside targets alone
ok - fm_touch_epoch preserves both epochs in the repeated DST hour
ok - a fixture commit starts no background maintenance that a local clone could race
ok - fake no-mistakes --version is the shared constant and overridable
ok - init/doctor no-mistakes stub touches markers and refuses other verbs
ok - fake gh authenticates and fake gh-axi reports the shared version
ok - spawn fakebin answers pane path, logs -l payloads, and installs extra tools
ok - send stubs log typed text and fake ssh records argv with a controllable exit
ok - spawn-home layout writes harness pin, beat, and brief
Evidence: /tmp fm-* snapshot before the run

Source: /tmp fm-* snapshot before the run

/tmp/fm-anchor
/tmp/fm-anchor+5e2aee5844c9c4415e4bc4bed5c159ea967445d3f4c3e70b6ba7400684c60f5f
/tmp/fm-anchor+ff2683e70e2969af369bf46edcf3bb0e0ef6b1d8d0ac8067c861548a95bbf5b7
/tmp/fm-base-path-sans.n5ZxQn
/tmp/fm-control-auth.ilwUHF
/tmp/fm-control-auth.mjYAz2
/tmp/fm-control-herdr-smoke-exit.filled.md
/tmp/fm-control-relaunch.H1Sqag
/tmp/fm-control-relaunch.H1Sqag
/tmp/fm-control-relaunch.QdXxID
/tmp/fm-control-relaunch.QdXxID
/tmp/fm-e2e-hooks-fm-backend-autodetect-smoke.log
/tmp/fm-e2e-hooks-fm-backend-herdr-launcher-workspace-e2e.log
/tmp/fm-e2e-hooks-fm-backend-herdr-launcher-workspace-e2e.log
/tmp/fm-e2e-hooks-fm-backend-herdr-workspace-per-home-e2e.log
/tmp/fm-e2e-hooks-fm-control-herdr-smoke.log
/tmp/fm-e2e-hooks-intent.txt
/tmp/fm-e2e-hooks-nm1.log
/tmp/fm-e2e-hooks-nm2.log
/tmp/fm-e2e-hooks-nm3.log
/tmp/fm-e2e-hooks-nm4.log
/tmp/fm-e2e-hooks-nm5.log
/tmp/fm-e2e-hooks-pres-main.log
/tmp/fm-e2e-hooks-pres.log
/tmp/fm-e2e-hooks-pres2.log
/tmp/fm-e2e-tmp-hooks-cleanup.filled.md
/tmp/fm-e2esm1
/tmp/fm-explicitbackendz4+450411b3e67c088c3a2d74a264543ce82c807234f22b72d9961423f94dea58ff
/tmp/fm-fdev+4733d433e5cb37786ca73e78c64ad23be1432fed8a53dee49235ef9aa3f1bc85
/tmp/fm-fdev+b26b9c5c19923cc9ab82a6231a96fc01c4c991675a393918367eec6e6269a930
/tmp/fm-fdev+f16937e38fe7869c9334ffe6356827813ff97a54f26c39856525c04c98348e73
/tmp/fm-fdev+fae625be605725dd01b2a654d126d312f702f8e813cde888baa62dcc408df14e
/tmp/fm-firstmate-maint
/tmp/fm-firstmate-maint+450411b3e67c088c3a2d74a264543ce82c807234f22b72d9961423f94dea58ff
/tmp/fm-fleet-tasks.rk01Yx
/tmp/fm-fm-e2e-tmp-hooks-cleanup
/tmp/fm-fm-e2e-tmp-hooks-cleanup+c1f7d6b63a9f67404bfca028eaec05edbbdbbdb76dc660f4bb2eac1ec5dc9bf0
/tmp/fm-fm-herdr-smoke-dup-label
/tmp/fm-fm-herdr-smoke-dup-label+c1f7d6b63a9f67404bfca028eaec05edbbdbbdb76dc660f4bb2eac1ec5dc9bf0
/tmp/fm-fm-spawn-tasktmp-home-scope
/tmp/fm-fm-spawn-tasktmp-home-scope+c1f7d6b63a9f67404bfca028eaec05edbbdbbdb76dc660f4bb2eac1ec5dc9bf0
/tmp/fm-fmtest-herdr-fhtZI6+a8189fbdb21cde1f8b92ccde6a2ccfbeed461819a2a39f575e468ad627d3e4f2
/tmp/fm-fmtest-sm-ZuPQ4f+1eba4c436fb562251c75f565f239bcf840bcfc41d400847a516f6485228d4fe7
/tmp/fm-fmtest-sm-ZuPQ4f+22035d7229d60e569ac0371de7d2ac2ff6d20cb13ea10eb3c564d821fb726911
/tmp/fm-fmtest-sm-ZuPQ4f+596d9f7c9e0f39b316b26c42e67d4375608a279990e312e9bfb5a49f68c98b36
/tmp/fm-fmtest-sm-ZuPQ4f+80256f9153c88c93e5f926584d17629f37ffc90f1efc1cf9b091d8e530215f77
/tmp/fm-fmtest-sm-ZuPQ4f+8120cfb37a0489871d61d2281c5e90985ee17f344cbf516763810f202944821a
/tmp/fm-fmtest-sm-fhtZI6+010362565dd75e4810ef79c5d476ab81daf1c4ff9d6356745f770a71f69aa7dc
/tmp/fm-fmtest-sm-fhtZI6+34fef571f6fef2724585779c6959513684720a6ac723c09c1391c58e893d48d1
/tmp/fm-fmtest-sm-fhtZI6+6423086d99ce3ae975e0d6bffb90892d8777561412cfb1c461cc800ede0f1506
/tmp/fm-fmtest-sm-fhtZI6+70413db028fc9c34e14fc79a5e27bdddaad3fb62d682636b2116cda08a3081fe
/tmp/fm-fmtest-sm-fhtZI6+8e11aade5dc3cf0e02bbdecfaa04f4afebca6d8917e84627372090d127f83785
/tmp/fm-herdr-lab-1000
/tmp/fm-herdr-presentation.bzrzuU
/tmp/fm-hsmoke
/tmp/fm-hsmoke+0ade4284487c9eadb41d2e9dde7b8451a81368732c49c12ad9e18fe21098143f
/tmp/fm-hsmoke+0feb30406ad0c638cd2b32d44a85042e18bc63deeef2aaa954b0a702e7837066
/tmp/fm-hsmoke+11b783723a3380026adcb505b2addc07e1f9f705faaa5719a6c48c1ea275b5bc
/tmp/fm-hsmoke+12c4fa710b4ebeabacf96550f6c413f3ba023b577590cdda5c40b5e1d95a206e
/tmp/fm-hsmoke+14d8c3dee8ca03faed395026aa863dbf3b6e0fa5dc35620657678aa5ba02df79
/tmp/fm-hsmoke+1a19f7664bf4b5e3e498ff518a617bc027fad376da40abc5f875aba0db4e3120
/tmp/fm-hsmoke+1a29f8a8fcd54ad430b234ae993144973e60314dc8bcf3f9f9eb800523f3c38d
/tmp/fm-hsmoke+23b49642bff33f76189dccf7574a71263a9df6e8837679132eb7e9301cc3ef35
/tmp/fm-hsmoke+24ede896fb7bf6e9edbf8f98fbd86091700ea09f3361cbd92544b72cdbed9b5c
/tmp/fm-hsmoke+29b43002738b37cf8d74d1ea6164258a557d16956bfb15d98d5975aaf089a084
/tmp/fm-hsmoke+2a8704320101b888bef39751621a8f9d54b8a9e1eb3da3498d3ae7c2f408fe47
/tmp/fm-hsmoke+2b16809a7474d4ddecd20abe34623139967e096f31e23ddf5bc8e468250a1248
/tmp/fm-hsmoke+468777b40f416be1f913fc32d041a91ee54ccb06fc64439d36534cd8f073d88a
/tmp/fm-hsmoke+48deb2ceb0908153debadce2a1be1bb3d2dc8befea5940087be3f271f6c0e15a
/tmp/fm-hsmoke+492f485f9de7cf2b8fd42bf2547ef8991ac8dd4760ac6c9ca51dde0b095a4b01
/tmp/fm-hsmoke+53cec20b49a6e7951afa65c53b9b74b72e1694e62a3b650f7c257157ee814fea
/tmp/fm-hsmoke+54230b07bd1b64e08a9173fd20d8d0e9e0804343dedc4d12cbac92ae7e03fa7b
/tmp/fm-hsmoke+59068412b67a0f7dd38aef2f1fd3ce242c14bba95058b666b4d060b4ffc92faf
/tmp/fm-hsmoke+59175b552c0fc538a1dc72befe4c210e14a92e51c552d8b7ab0abaa6b4e03ab4
/tmp/fm-hsmoke+5ac1aef8636fb4a4954640e50b09c204600e7467998b1db321387936135d6477
/tmp/fm-hsmoke+64ad50888f477837f34e85a6689cee4a4ee9ddc233ba67862f025fc336282ad1
/tmp/fm-hsmoke+680e0f2b2a634559c52f0a90d4c1146a0768235cffe9043b17a30052c18f618a
/tmp/fm-hsmoke+883ebbbd9443681d8501ddc87753f01408a5e9713a8e71dd33185a5c5348478e
/tmp/fm-hsmoke+8c9792b82c5ffc6fb218f1a08240778f6f157bbbd9fd245b99ec1778c35dc44a
/tmp/fm-hsmoke+970d42323a5eaf68025f5d08ed92e4ad1d74c270ed21bbec60ff5a4f6a1c5736
/tmp/fm-hsmoke+a5bd9e347640df885ae2118a25e82a47ac29b0f4134dc0a91f0489c206f5a80a
/tmp/fm-hsmoke+a96730a460bf4750fdcbc3c865e7a84ffea579b414db009da59e43b5ca910d19
/tmp/fm-hsmoke+ae298343481f366aa1ae091f98a3f77ed6f7dd43e1d1075a8382cd25822507d9
/tmp/fm-hsmoke+bf4b5b61c5ac8a08ddd1f5fc5bb1955887eea6dc00bf16618af1c7d56415b1d7
/tmp/fm-hsmoke+c149b503517eb0b956381f39140aaacf118f0b4ca6ebd4540c46292e21198f8e
/tmp/fm-hsmoke+e36ae758f882330854cbabcff52b1c8fe23bc08c7ceda1bf53f675e32bbdd4f3
/tmp/fm-hsmoke+ed141a9cbc2c35344bae6a0675e4bc60bfd09f8b4286191436152652d2d838b4
/tmp/fm-hsmoke+f0dfbc61323c95cf98da63f23c5989696ff18871026db112eb3e9dd9aa2d3de5
/tmp/fm-kimi-drop-z2+189bec0fb5a7526e53901d2cded19ed41e715ef90589fd05bef699a5a6ec6171
/tmp/fm-kimi-drop-z2+7418b9a50852d864a532ff1a3baf2119fcba2d744c08c1ee107720377be4ca60
/tmp/fm-kimi-drop-z2+a2fa192cb7b4c2ab5a5793ef8ea8ff65737aadcc545aa3f413b67e4e1fdaf5bc
/tmp/fm-kimi-fallback-z4+1d1d4e065b61d59a7463027145a9067ea448b084a44d56f4b789477b36ca75e7
/tmp/fm-kimi-fallback-z4+2e8b1e26ec0fb0348cc24df1c66c208823ff7af755cf1b4fbcca950aa25da043
/tmp/fm-kimi-fallback-z4+c8f13c2885b5ce061e74684ba83ea899eb845d34c455b225eccd2a369f28bcc7
/tmp/fm-kimi-hook-auth-z6+116c9e9a09f3a7b50d4b313a27d7194204aecf6d36d423f9b3a5ed63048845ee
/tmp/fm-kimi-hook-auth-z6+5a359263a53c3dedfa082a8d24b379cae6d3eefc48c0f58a9b52b8c5972877a9
/tmp/fm-kimi-hook-auth-z6+771281ec5ac29be8c990f0343fbc1e2dc6f78a9c5b95a5ca3903f98d68d493c5
/tmp/fm-kimi-not-ready-z3+3fb3277f5c994895b0efb4e899144cfd42dcd27d5ee30a924ef23aca1befb054
/tmp/fm-kimi-not-ready-z3+541ec83637194a2c8f8c5850400042dd4b37f175f0f62afd781c27fe0930f79e
/tmp/fm-kimi-not-ready-z3+effeaf091e7bdc0392e50b40d1a98ee8642180fe20c3a7d0cd32c680cd4b5f76
/tmp/fm-kimi-success-z1-1907832+c93c3662012571fae41fa133740b2350adae2280e24833ce56bc671ab1fb4928
/tmp/fm-kimi-success-z1-2963701+4397713589d248afa436dacb4273de62e401861208fdbb0932a6f36bf68f9129
/tmp/fm-kimi-success-z1-615273+0e560f7d4328f14946a53d1d3067eadb1faa8b841d654be03cd2ee41f5757ea4
/tmp/fm-kimi-trust-blank-y7+0efd5070fe74f25569dfa1f808aba3523937ffed3d61cf3f3b1c8c2733164ba9
/tmp/fm-kimi-trust-blank-y7+1525699d2e21b6d0cfff97daeca28135c96e83750bc1357167d7a13086dee817
/tmp/fm-kimi-trust-blank-y7+94f13b8b1d664ca004d7733997df82beea0a3cd30382ba69d37010c65e6d1b6f
/tmp/fm-kimi-trust-blink-z4+35a07f60ef27a234e39544e3b018d36282f1739ffff1a72f2964669178701cdd
/tmp/fm-kimi-trust-blink-z4+8cc659f5bfa105e812ce1cf2459c5d7a919a9be5c1c5ae299270ab1619edc229
/tmp/fm-kimi-trust-blink-z4+e95575af8a90bad4812bc2514cd1ffbf108440305016cf57440708f5e6359eb8
/tmp/fm-kimi-trust-decoy-y2+6c59b045d4e49853719209cd12677dfe957a43cf67d69aff1aa5da814eb1150d
/tmp/fm-kimi-trust-decoy-y2+e51e3e372adb54dd2a9e8ee6c6abecbb79322f6ab2c4cba04f5885486e9a6ddc
/tmp/fm-kimi-trust-decoy-y2+e9c916cb501e7fad8bb53f02481b4ec2d1f8524323099e497bede4f680d6a9cf
/tmp/fm-kimi-trust-history-y6+a0e434b7e4b7ca40a22cc35c57064c8dc4156bf5e3b61188e97f7f40678205b2
/tmp/fm-kimi-trust-history-y6+b886fba753607431e3ae249b182b07b11c7945287764f282a2c941a7c075e7ea
/tmp/fm-kimi-trust-history-y6+dfd6d7a233ffb63146087875a7c8ccdc56b381131bd80952f522ab444a950df0
/tmp/fm-kimi-trust-late-y5+694f49695eba04e481

... [5994 bytes truncated] ...

+80b46857e6088b318d6207521ebd08ef5a1c08fb0381c69a40567f4c8fd8624d
/tmp/fm-rl21+83a36cffe82c47ac436915ac46f816c3cbdbb99dea06449fbc2b232566b666b5
/tmp/fm-rl21+c24556eaffbdb395081e9be0fb6b80d99ac06c16446beb9ba1f9560182b31766
/tmp/fm-rl24+0fcb48fcc233b923047373d24b7dec1cec516aab5b21766c5bf14462df8fb23a
/tmp/fm-rl24+8a2980e30ebd32ee61a1f6794a200ab0ba58f869219650d4f10f9d919da63357
/tmp/fm-rl27+d944a162890312f34c9d01962300e215667e181f8cc23e316dc29f01db16c7e2
/tmp/fm-rl27+e51407c64abdc4ddc2c86c424a973c0093485030d0539dca82f2295577554cb5
/tmp/fm-rl28
/tmp/fm-rl28+504d6ae34b84497e591de7bde9f0e0c43351472ae11af5bcdfae96c0d8f18661
/tmp/fm-rl28+ad13e661df2b959f0bfad85e4a2e13801f862891b79275488e29ad822085b645
/tmp/fm-rl28+b8d7723af4ce528f73be5f8d7a9d7b499560c0748f8bb20eff09d56999be70e2
/tmp/fm-rl29
/tmp/fm-rl30+87613068b9821eb77051ff0ebb4042b2faafba7fdcd30e8074709ee4da34d091
/tmp/fm-rl30+a1d57908fa4b1e528bbfc41041ed78b90e0eb51d4bf19be81b7af6628cce5ce9
/tmp/fm-rl31+07c620fa26cb2b9449edb2dc20d870ca35b15256055e9bacfef60a6322a972eb
/tmp/fm-rl31+bc3e3219b145455b08a8ae41d6a31ae8c76d026f6c915fb5c9b8b7735d0d7043
/tmp/fm-rl32
/tmp/fm-rl32+381b67527395331434b2063e374bfc842edc0aff311015ce711ac1a22a301c14
/tmp/fm-rl32+6864d1a45dc5e26d460ebafdc6b6fcabf771f1072a067de1f42ad6b3b12ba3ec
/tmp/fm-rl32+98208b4a83aa91321b341b390636fd97bba5f72cb813304bff239a6353c487e1
/tmp/fm-rl33
/tmp/fm-rl33+97d55fef4af50551de737a6e6320cf5286c3aca47695443101bc3c1ca3dd6f85
/tmp/fm-rl33+f56b65d90bfa695c39f08b3ba83c1875dc6233ecb824130e930a5d9322a4e9cb
/tmp/fm-rl33+f861be859e07726f6b6af4c2200e7b57ecbff8c71cf8129d367c94b2bd76849c
/tmp/fm-rl35+54ac986ad24ed6bde57b1add323cfb50942cd08a89a54dbf122096b67890ece8
/tmp/fm-rl35+f3368d63eb4d4d99ab93de7d2d6242011509394aeb2df13378d15cb4a5ecdafd
/tmp/fm-rl38+8b7d42cb551239f7c1658e9889a9b40431cb651575157c5bce8fc4b5845b4c2f
/tmp/fm-rl38+d80029c43254eac75b87d13e9acb79d19392ffe12c7f52f397f5e727ca335b14
/tmp/fm-rl4
/tmp/fm-rl4+43397a932a16d96be943c24fd5f2395aafbef6a97dee84716cc5fb5aeed2faad
/tmp/fm-rl4+7a685ae380d602041a229c75d82968dd3ea67347b99c71e4483fb05dacbaf483
/tmp/fm-rl4+b50623d9fc8455a8e1d43fc2999726de7accb2195dc590983d9b727fb667bcc5
/tmp/fm-rl40+353c91d7e8a218ebb512d5a2aed5c8aaeb4c50e1df9f382be5f83f50e5d2b8fc
/tmp/fm-rl40+462b0a2dd0097d8192bd45d1ccaaacd9f705417f91416b80a66a6bd92d5f4dcc
/tmp/fm-rl41+349787ecdab43dbd48708c2193e981b8c9de4a254e63a0669bf0792ea9df615c
/tmp/fm-rl41+58aaf17ae390519c9723f39b15f97b8730fb41dd6cc445e4175264ac6d04d684
/tmp/fm-rl42
/tmp/fm-rl42+0a5c042c0c38f79629368ccfd3055b431c246c55a1bf18837b193c12f0788aa7
/tmp/fm-rl42+12a5e556e0ff7cececca52dc6e170f59e87b0b96f8e92aeaf792525114fa024f
/tmp/fm-rl42+2989cd3705fa567351e9b458e8520587fe553d6061b01b3751c61d927ad294c9
/tmp/fm-rl42+544109676e7a8fbd30025df001c096db1fdce90055011c5369d6c3bcf023790b
/tmp/fm-rl42+aaa5f88324657b987fd292908b94b3e0bc606ed41798d045a4bcc37f0bc2289c
/tmp/fm-rl42+df7006c2381bb3f968f82f4bf6432db672d6f6c0bb56d684b32052862b2e54c6
/tmp/fm-rl5
/tmp/fm-rl5+1ba48a5d0bb4ca95779b9245a4d5ac08faa627c053457f1fd692a577a0c137a2
/tmp/fm-rl5+48e02ea2ea0abbd0c35efce3d6d5476d9e939a09f9c732b21d978ac76a5c9393
/tmp/fm-rl5+618bcfbe45421237f33d09f15abee240c950379397ebfa0ad38340b292018a10
/tmp/fm-rl6
/tmp/fm-rl6+078276aa1af3e264aefc0240cbba40324e0f2f1550aa5b21e9cb8144f104631b
/tmp/fm-rl6+20caab99fe88329e54be9f97f69e9c762316746aacec5173a1c9fe9ca6412dc4
/tmp/fm-rl6+cabf1bf546159ca6bdf34ef7535fd7907b07e635e29e4b9d99d61ac568a61b24
/tmp/fm-rl68+6cf9a08b3d6180eb044c0ecbd1082d76c9db1025211869fca3d1267ece9e8743
/tmp/fm-rl68+b3a0d5df948d97d9cadbdc53d3456246c78667c62792ac77e430ea69e2495181
/tmp/fm-rl7
/tmp/fm-rl7+5e4c28fe3e1b6ac40fb808f86dce08cec6348729ecc359116c4e20e4158d1323
/tmp/fm-rl7+87c157714c20d62d676d5330ef2754233b2d14c8c96b8f3fb6d4f4eb6424ad26
/tmp/fm-rl7+f5b30267c50810e32a1cdaad0e8207a1e7efea5399f167ab4a556e821502c351
/tmp/fm-rl73+06c91d2f763b29500217bc50645be04bfa514bb97a14803e5fa9f3aaa58c0b1d
/tmp/fm-rl73+1af1063b81596513e9ae1436bb9ad1281ce3dd4afd1f0cfbe660e50a6c44c214
/tmp/fm-rl75+46005ea627c51ab5655b277f31cf18dd6760f9979f72bbfff7a779834c66211f
/tmp/fm-rl75+84a1a801c03e3b4f385ae76f1bc0dba8efa526e292d24432c087c8b02d9e98bd
/tmp/fm-rl9
/tmp/fm-rl9+0bd0c0ad2f1be097529d61c5707061284f59c4c57607924bcf23382e7136b952
/tmp/fm-rl9+a4395727a48f3dd7fa7cae6f866f8aaaca77a626ba96f25649923465447319c7
/tmp/fm-rl9+c3a256cf6e3fd2c7cf886650df6d0f6a9a34848d7137a99cf6c9adf6b5cea7cf
/tmp/fm-rl91
/tmp/fm-rl91+2aab734239408b9e20bcece7605515e8ea48356dfe6fea512a2fe8d759f439a9
/tmp/fm-rl91+7ce9c327288d34730e7ebb58b44479737fb18c11061c03ecac4c9d6f829e7df0
/tmp/fm-rl91+b4251f685853429cea6e7dc3f2aef17b8d5ee8c79932881efcf6f7fd3ea3f0a4
/tmp/fm-rl92+01d2b1cd39bdce3f397763026e066864b6c4f5f3e84e2eb3812f6781918066c2
/tmp/fm-rl92+364e11eaf5bb9bea10011e1a98b75b90e1a0d16e3747d97332ba0704e9bd3175
/tmp/fm-rl92+ad871839d4b9dd1e717f324df14e9d0c1209d037b5003ad7fa252de7c15bcecb
/tmp/fm-secondmate-safety.6gU9NR
/tmp/fm-session-start-stage.FhhxL3
/tmp/fm-session-start-stage.b3E5ai
/tmp/fm-session-start-tests.ZuPQ4f
/tmp/fm-shape
/tmp/fm-shape+5e2aee5844c9c4415e4bc4bed5c159ea967445d3f4c3e70b6ba7400684c60f5f
/tmp/fm-sm1+8c2fa8862c0d40899981ec5d6be68afd0c387f714e37ebc2370522c282c273bb
/tmp/fm-sm1+d77d869fed355256a4c1e827b7871d7b226c6859cc7776e724510979318dea94
/tmp/fm-sm3+3005cbd3d220f464e0e6ddc404c40c6ee07af8fa1ddb5393b115815c49c1cf1e
/tmp/fm-sm3+49094ad5bd5b84d956fc1f4b03d707c5d76e456ba0bb057e1c5fa5a8e8732c10
/tmp/fm-sm3+df587afbd30fae94ed95e2f1f6dd53140c57597768ab19d02c3de8947b45a0f4
/tmp/fm-sm4+5e1fa53c13f5ce47eaf24664abe9ddca2ac6f80885a045b9aa27f9821e13a812
/tmp/fm-sm4+73754f486d1f68544f4156830855115de14eb426ee4331c01f2e9c0ab584b302
/tmp/fm-sm4+ac8362763b9e78f2a952ed3948a072b39b053070210dd299052986a726c28834
/tmp/fm-sm6+019fedd3d7ed5d193575206706fdffc772a8f7754c870cfda943563ef4494a94
/tmp/fm-sm6+095be6e05e80da5b92d21c2cff2487884cc91129ab86dd3cc8d293a0c58eac3a
/tmp/fm-sm6+3a5e24527901dfca55e353b0428de0948315e1537571fdf42dc41b15f52897a9
/tmp/fm-sm6+9360dbc97c45f0a9ac249f763ccac9baa58cc5beb5a23c010cde3b7fe4a29784
/tmp/fm-sm6+bedaf84f18e29801c390e63ee0b86988d7fcbd4cb037493edba7df3239cf7050
/tmp/fm-smE
/tmp/fm-spawnsymlinklogical+0348f5e347c5573a299c3802dab29eb5561542f0de3f9760b55e16ec4f4ea98f
/tmp/fm-spawnsymlinklogical+1b920429340974921e1a0830ed3c61a0b939515546644b38a39fa2bfb9867119
/tmp/fm-spawnsymlinklogical+246ad6cd5c91f60595f540ff75b8db1156b6a232b1a4e85edd51cd06638a18c6
/tmp/fm-spawnsymlinklogical+3b7da2db1fb3e26c9100ccc82f1bee1d70b6f78c4510eeed8c2759ae9b4663cf
/tmp/fm-spawnsymlinklogical+793453f499dca68252cda06ddf0257b8a0142c573efcd6977c32c7476febe729
/tmp/fm-spawnsymlinklogical+7c08a82734d61e8d79724c4aef5b25b7c44283e4bfeea159aad9bddd40606ab0
/tmp/fm-spawnsymlinklogical+f7c45debaf0ebf378deb4f3e2da23e09aa2f526baea74307d3217eeca67634a7
/tmp/fm-spawnsymlinklogical+feb363d1c2fd7401e2622091e2f8195cb02e904daeca556ba12c90ae1362d1d3
/tmp/fm-spawnsymlinkphysical+0348f5e347c5573a299c3802dab29eb5561542f0de3f9760b55e16ec4f4ea98f
/tmp/fm-spawnsymlinkphysical+1b920429340974921e1a0830ed3c61a0b939515546644b38a39fa2bfb9867119
/tmp/fm-spawnsymlinkphysical+246ad6cd5c91f60595f540ff75b8db1156b6a232b1a4e85edd51cd06638a18c6
/tmp/fm-spawnsymlinkphysical+3b7da2db1fb3e26c9100ccc82f1bee1d70b6f78c4510eeed8c2759ae9b4663cf
/tmp/fm-spawnsymlinkphysical+793453f499dca68252cda06ddf0257b8a0142c573efcd6977c32c7476febe729
/tmp/fm-spawnsymlinkphysical+7c08a82734d61e8d79724c4aef5b25b7c44283e4bfeea159aad9bddd40606ab0
/tmp/fm-spawnsymlinkphysical+f7c45debaf0ebf378deb4f3e2da23e09aa2f526baea74307d3217eeca67634a7
/tmp/fm-spawnsymlinkphysical+feb363d1c2fd7401e2622091e2f8195cb02e904daeca556ba12c90ae1362d1d3
/tmp/fm-timeout-status.2kDugY
/tmp/fm-timeout-status.LRUJVT
/tmp/fm-timeout-status.W4Ixxm
/tmp/fm-timeout-status.WaFhSu
/tmp/fm-timeout-status.Wg06ud
/tmp/fm-timeout-status.X2jWaU
/tmp/fm-timeout-status.ZRvqG3
/tmp/fm-timeout-status.ZnGINi
/tmp/fm-timeout-status.hAB8mC
/tmp/fm-timeout-status.wNdZGV
/tmp/fm-timeout-status.xgH1qN
/tmp/fm-trustspawn+48b841f140add00df13845147c555c21877afbcc5f12139c4a6b0260b74854bb
/tmp/fm-trustspawn+fb56964249a1a38532f9f15050c293de0dd2df88006568f8b36cdce26393a77a
/tmp/fm-uniqA
/tmp/fm-uniqB
/tmp/fm-wake-tangle-root.vFRSEZ
/tmp/fm-wedge-rec.JoM3pK

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

⚠️ **Review** - 1 info
  • ℹ️ tests/fixture-tree-helpers.sh:53 - The new helper comment says the per-task /tmp/fm-<id> root is owned by "fm-spawn.sh's own cleanup", but fm-spawn.sh never removes TASK_TMP (bin/fm-spawn.sh:4378-4387 only creates it). Only fm-teardown.sh removes it (bin/fm-teardown.sh:3788-3790). So a task that is never torn down, such as the presentation suite's anchor or the untorn spawn in tests/fm-test-fixtures.test.sh, still leaves /tmp/fm-<id>/gotmp behind. The fixtures test handles this with its own rmdir at tests/fm-test-fixtures.test.sh:381. The intent only covers strip hooks and launch dirs, so leaving this root alone is acceptable. Fix: reword the comment to say fm-teardown.sh owns the per-task root and that a task that is never torn down keeps it, so readers aren't misled about who cleans it up.
🔧 **Test** - 2 issues found → no changes applied ✅
  • ⚠️ The Test agent did not finish within its invocation budget. Reported: agent run tests timed out after 30m0s: agent last produced output 3s ago (106 observed); agent reported: claude parse events: context deadline exceeded. This is a budget or provider-slowness cut, not a code failure. Re-running the same request costs another full budget, so no further attempt is made automatically. If this repository's targeted tests or evidence gathering routinely approach the default 30m0s, raise test_agent_timeout in global config. Respond with fix to spend another budget: a repair turn runs only for selected findings other than this budget cut, then validation re-runs. Or abort and retry after raising the budget.
  • 🚨 Approval is refused: the run worktree at ~/.no-mistakes/worktrees/450411b3e67c/01M3TG9RFM949G5A11TNACEFWR holds work no Test turn validated, and the steps after Test would commit and publish it. It holds uncommitted changes to tests/fm-backend-herdr-presentation-e2e.test.sh, tests/zz-base-presentation-repro.test.sh (inspect with git -C ~/.no-mistakes/worktrees/450411b3e67c/01M3TG9RFM949G5A11TNACEFWR status and git -C ~/.no-mistakes/worktrees/450411b3e67c/01M3TG9RFM949G5A11TNACEFWR diff). Respond with fix to validate it, or abort.

🔧 No changes applied.
✅ Re-checked - no issues remain.

  • Live validation: ✅ go - 3 of 5 scenarios driven live against the product
Scenario Result Live Evidence
Scratch repro file is removed and will not be committed ⏸️ untested no The prior payload recorded this as a worktree hygiene check (git status), not a scenario driven against the live product, so it did not establish a live result.
Duplicate-live-agent pane keeps its Herdr 0.9.3 registration because a real foreground process runs, so the live duplicate risk reliably refuses launch ✅ pass live presentation-e2e.log: ok - real Herdr lab: missing, renamed, and duplicate tokens trigger zero destructive or adoptive calls, and live duplicate risk refuses launch
Real-spawn E2E teardown removes the whole fixture tree, including read-only state/<id>.git-hooks, and its staged /tmp launch dirs ✅ pass live presentation-e2e.log: ok - cleanup removed the whole fixture tree, including read-only strip-hook directories, and its staged launch directories, rc=0
Lab teardown leaves no Herdr lab session behind and the default session untouched ✅ pass live herdr session list after the run lists only default; log reports default-session tripwire intact
Shared teardown helpers remove read-only strip-hook trees and staged launch dirs, finish under set -e, and leave outside targets alone ⏸️ untested no The prior payload ran only the fixture-level automated test with fake binaries (tests/fm-test-fixtures.test.sh), so it did not establish a live result. The live presentation E2E scenario above covers…
  • rm -f tests/zz-base-presentation-repro.test.sh (scratch file removed; git status --short shows only M tests/fm-backend-herdr-presentation-e2e.test.sh)
  • bash tests/fm-backend-herdr-presentation-e2e.test.sh against real herdr 0.9.3 in its own fm-lab named session (rc=0, 728s, every ok line, no failures)
  • bash tests/fm-test-fixtures.test.sh (rc=0)
  • Before/after snapshot of /tmp fm-* entries plus herdr session list after the run (only the untouched default session remains; no lab session left over)
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

cloud-practitioner and others added 7 commits October 1, 2026 01:12
…st cleanup

The Herdr presentation e2e suite removed its temp root with a plain rm -rf,
which fails on the read-only state/<id>.git-hooks directories that the real
spawn installs, so every run leaked its /tmp tree. Spawns the suite never
tears down also left their /tmp/fm-<id>+<home-hash> launch directories behind.

Move fm_test_remove_tree into tests/fixture-tree-helpers.sh, sourced by both
tests/lib.sh and tests/herdr-test-safety.sh, and add
fm_test_remove_spawn_launch_dirs, which removes only the launch directories
scoped to the fixture's own homes. Use them in the Herdr suites that spawn
for real, assert in the presentation suite that nothing is left behind, and
cover both helpers in tests/fm-test-fixtures.test.sh.
* fix(bin): run no repository hook when core.hooksPath is empty (kunchenguid#6216)

* fix(bin): run no repository hook when core.hooksPath is empty

The per-task hook wrapper refused every commit in a repository whose own
config sets core.hooksPath to the empty string, because git rev-parse
--git-path hooks fails on it. Plain git reads that setting as no hooks, so
the wrapper now runs none; every other lookup failure still refuses and
shows git's error.

Fixes kunchenguid#6171

* no-mistakes(review): Refuse commits when core.hooksPath is a valueless key

* no-mistakes(document): Document empty core.hooksPath handling in commit attribution docs

* no-mistakes(ci): When the wrapper refuses a commit, Git's hook-lookup error now shows up once instead of twice. That required changing one line in the wrapper, and the tests were extended so both bad-config cases would catch the duplicate. Invariant: when the wrapper refuses, Git's lookup error must appear exactly once. In the failure path, the only Git call besides the deliberate second lookup is the `git config --get --type=path core.hooksPath` check in `runtime_chain_body` (`bin/fm-git-strip-ai-trailers.sh:168`). That check prints the same error, so it was the one place to fix. I added `2>/dev/null` to it. Its exit status still decides the outcome: an empty value still runs no hook, and anything else goes on to the second lookup, which prints Git's error once, and the commit is refused. Tests (`tests/fm-git-strip-ai-trailers.test.sh`): - The unresolvable-path test (`~fm-no-such-user-6171/hooks`) now requires `failed to expand user dir` to appear exactly once in the refused commit's output. - The valueless-key test now requires `missing value for 'core.hookspath'` to appear exactly once. - Pre-existing bug in the unresolvable-path test: its `git add` ran after the bad config was set, so it failed silently (exit 128) and the "refused commit" had nothing staged. The test now stages the file before writing the config, the same way the valueless test does, so a real commit gets refused. - The empty-string test is unchanged and still passes, so an empty `core.hooksPath` still runs no hook. Verification: - With the wrapper change reverted, both new checks fail with `expected '1', got '2'`. With the change in place, the whole suite passes. - `bash -n` passes. shellcheck shows only an info-level SC1091 note about sourcing `lib.sh`, which was already there before this change. - `git status` lists only the two intended files

* fix(bin): let a stale record on a reassigned slot retire records-only (kunchenguid#6213)

* fix(bin): let a stale record on a reassigned slot retire records-only

When a pool slot's owner claim names another task, the stale record's
teardown touches nothing under the slot, so the exclusive-slot record scan
no longer refuses it. Full teardowns of a slot this task still claims, or
one with no claim, keep the refusal.

Fixes kunchenguid#6184

* no-mistakes(document): Note claim-over-record precedence for reassigned teardown slots

* fix(bin): keep the steering doorbell short under deep homes (kunchenguid#6240)

* fix(bin): keep the steering doorbell short under deep homes

The doorbell printed the task inbox's absolute path twice, so under a deep
home it grew to about 290 characters and a Herdr submit reported it never
reached the pane on every re-ring. It now names the inbox once by its short
<task>.inbox name and points at the full path the worker's brief already
gives, so its length no longer depends on the home's depth.

Fixes kunchenguid#6120

* no-mistakes(review): Export FM_TASK_INBOX at launch and name it in doorbell

* no-mistakes(ci): ci-1 (Behavior portable serial 9) was caused by this PR, and I fixed it in the test. tests/fm-claude-trust.test.sh failed with "the launch command did not carry a brief doorbell". Its claude_launch_doorbell helper stripped exactly two leading `export ...;` statements before reading the final prompt argument. This PR adds a third one (`export FM_TASK_INBOX=...`) to every launch, so the helper was reading the wrong command. The invariant: a test that parses the launch command must skip every leading export statement, however many there are. I checked every test that parses the launch this way. The only other ones are the two helpers in tests/fm-spawn-dispatch-profile.test.sh, and they already loop over all exports. The kimi and dispatch-profile exact-string checks were updated earlier in this PR. The fix makes claude_launch_doorbell use the same loop (`while [[ "$command" == export\ *\;* ]]; do command=${command#*; }; done`) and then take the last argument. The ordinary path still works: the claude spawn test and the secondmate-clone spawn test both resolve the brief record through the same helper. Verified locally: `bash tests/fm-claude-trust.test.sh` exits 0 with no failing cases. ci-2 (Behavior tests (Herdr)) was not caused by this change, and I made no code change for it. In tests/fm-backend-herdr-presentation-e2e.test.sh, the concurrent secondmate recovery failed with "herdr presentation recovery could not acquire its session lock; refusing a concurrent resume". Two reasons it is not this PR: - The same failure, in the same test and case, happened on run 36655209015 for the unrelated branch fm/fm-contributions-old-gh-compat about 14 hours earlier. - This PR's change cannot lengthen how long the lock is held. The launch is written to a file and sent to the pane as `. launch.N.sh`, so the extra export changes neither the pane submit nor the lock hold time. The cause is a race that was already there: spawn_herdr_presentation_order_lock_acquire gives up after 5 seconds, and a concurrent real-Herdr recovery can hold the lock longer. Fixing that means changing the product's lock timeout, which is outside this PR. It should be tracked separately, and a rerun of the Herdr job is expected to pass. The only file changed is tests/fm-claude-trust.test.sh

* feat(bin): add opt-in config/wait-no-turns so a waiting worker spends no turns (kunchenguid#4859)

* fix(dod): drive no-mistakes with one foreground call, not a background poll

The brief told workers to background the drive call and poll `axi status`
because one call "routinely outlives what your harness lets a single
command run". That advice contradicts the tool it drives: `no-mistakes
axi run --help` documents `--wait` with an 8m default, existing precisely
"so an agent harness with a 10-minute tool cap gets a structured return
instead of an unbounded hang".

Following the old text, a worker could never idle - a backgrounded call
returns in milliseconds, so it does not wait at all - and each attempt
leaked a live timer that later fired as a paid wake. Tell workers to make
one foreground call, let it block, and repeat it when it returns on
elapsed wait rather than on a gate or outcome.

Also drops the generalisation that told workers on any unestablished
harness to assume a command cap and use the same shape, which exported
the defect to harnesses with no such cap.

* fix(bin): let a waiting worker spend no turns until it is answered

A worker waiting on a decision, a pipeline gate, CI, or a heavy-test slot
kept taking model turns: the brief told it to list its inbox at any natural
checkpoint, and six automatic senders nudged secondmates whatever their open
decisions.

- The ship and scout briefs gain one Waiting section: end the turn after
  needs-decision or blocked, and hold an external wait inside ONE blocking
  command bounded by the harness's own command ceiling. The checkpoint clause
  is deleted. Forbidding the wrong shapes is not enough on its own, so the
  section also names the blocking foreground `until` loop as the wait a Claude
  Code worker may use, because that harness can refuse a sleep-then-check
  command while pointing at backgrounding, which is the one shape a waiting
  worker must not take.
- fm-send --automatic defers (exit 4, nothing written or rung) while the
  target has an open decision or blocker of its own; every automatic sender
  passes it and keeps its retry state, and the pending-reply recovery waits
  the same way.
- The two senders that report the result classified it by matching the text of
  the send's captured output against `deferred:*`. fm-send runs bin/fm-guard.sh
  as a supervision warning, and that guard prints its worktree-tangle banner
  whenever the primary checkout is on a feature branch, which is exactly what a
  CI pull-request checkout is. The banner lands ahead of the `deferred:` line,
  so the match fell through and a waiting mate was reported as a failed send,
  with the banner as the reason. Both senders now classify on fm-send's exit
  status, which is the contract the deferral is actually stated in, and select
  the `deferred:` line out of the output rather than assuming it came first.

The third root cause, a no-mistakes definition of done that backgrounded the
drive call and polled axi status, is fixed by this branch's parent commit
"drive no-mistakes with one foreground call, not a background poll"; this
commit takes that text as is and adds the regression test.

Upstream's spawn abort path no longer calls the lease-return helper at all,
so the fork's missing-helper guard and its pin-feature test line are moot
here and are not ported.

The command ceilings each harness enforces, and the probes behind the named
Claude Code wait, are recorded in docs/verification/runtime-backends.md.

* no-mistakes(review): Exempt captain holds, quiet deferred reconcile, clarify worker pauses

* no-mistakes(document): Document deferred automatic nudges, rereads, and reply recovery

* no-mistakes(document): Ring unlanded fire-and-forget steers exactly once more

* no-mistakes(ci): The failing check, "PR must be raised via no-mistakes", reads the pipeline's attestation record, which says document=skipped. No file in the repository can change that record, so I did not touch the check or the PR body. As you said, the no-mistakes rerun after this run finishes will re-execute the document step and record document=completed. The one change is the documentation sentence you ordered. It adds a line to docs/remote-secondmates.md, right after the line saying the remote host runs no re-ring ladder of its own: "A fire-and-forget record, such as a reconcile ask, gets its single retry ring only on the local plane: the remote steer leg owes no re-ring, so a swallowed remote doorbell for one waits for the next ring into that inbox, and a remote-side retry is known follow-up scope." No behavior changed. Checks: tests/fm-documentation-audiences.test.sh passes (4/4) and bin/fm-lint.sh is clean. The change is left uncommitted in the working tree for the pipeline to pick up

* no-mistakes(review): Hold automatic wakes until a mate's own decision closes

* no-mistakes(document): Document watcher delivery of deferred remote re-read nudges

* no-mistakes(review): Merge duplicate elapsed-wait reattach instructions in DOD

* no-mistakes(test): Resolve merged default decision in remote-reply recovery fixture

* no-mistakes(test): Source classify lib so config-push retry-deferred honors open decisions

* no-mistakes(ci): Fixed a flaky test that also fails on main. Neither this PR's bin/fm-brief.sh nor its bin/fm-dod-lib.sh change is involved: bin/fm-dispatch-resolve.sh sources neither file. Another branch (fm-attended-cutover-smoothing-s1, run 36343879084) failed the same shard 8 check the same way, on a different case ("a rule-criterion match prints one diagnostic line, got 2"). Root cause: `fm_quota_single_provider_for_harness` in bin/fm-quota-axi-lib.sh returned from its `while read` loop as soon as it found a match. That closed the pipe while `fm_quota_single_provider_table`'s `printf` was sometimes still writing. GitHub Actions runners ignore SIGPIPE, so bash printed `fm-quota-axi-lib.sh: line 138: printf: write error: Broken pipe` to the resolver's stderr. That is the extra line. I reproduced it locally by running the test with SIGPIPE ignored: 2 of 20 runs failed, one with the resolver's diagnostic line plus two broken-pipe lines. Invariant: looking up a harness in the provider table must never make the table writer fail. The only reader of that table is this function, and all of the resolver's lookups (line 208 without stderr redirected, line 222 with it) go through it. So the fix is in that one place: read the whole table, then print the match. The same file now shows it reads the full table first, like `fm_control_harness_supported` does. Return values and output are unchanged. Verification: with SIGPIPE ignored, tests/fm-dispatch-resolve.test.sh failed 0 of 30 runs after the fix (2 of 20 before). tests/fm-dispatch-resolve.test.sh, tests/fm-brief.test.sh, tests/fm-send-inbox.test.sh, tests/fm-quota-choose.test.sh and tests/fm-quota-array-dispatch-live-e2e.test.sh all pass, and shellcheck is clean. tests/fm-procevent-quota.test.sh fails locally with or without the change ("process-event state root is not a private directory"), so that failure comes from the local environment, not from this fix. No new test was added: the existing one-diagnostic-line assertions already catch this whenever SIGPIPE is ignored, as it is in CI

* Revert "no-mistakes(ci): Fixed a flaky test that also fails on main. Neither this PR's bin/fm-brief.sh nor its bin/fm-dod-lib.sh change is involved: bin/fm-dispatch-resolve.sh sources neither file. Another branch (fm-attended-cutover-smoothing-s1, run 36343879084) failed the same shard 8 check the same way, on a different case ("a rule-criterion match prints one diagnostic line, got 2"). Root cause: `fm_quota_single_provider_for_harness` in bin/fm-quota-axi-lib.sh returned from its `while read` loop as soon as it found a match. That closed the pipe while `fm_quota_single_provider_table`'s `printf` was sometimes still writing. GitHub Actions runners ignore SIGPIPE, so bash printed `fm-quota-axi-lib.sh: line 138: printf: write error: Broken pipe` to the resolver's stderr. That is the extra line. I reproduced it locally by running the test with SIGPIPE ignored: 2 of 20 runs failed, one with the resolver's diagnostic line plus two broken-pipe lines. Invariant: looking up a harness in the provider table must never make the table writer fail. The only reader of that table is this function, and all of the resolver's lookups (line 208 without stderr redirected, line 222 with it) go through it. So the fix is in that one place: read the whole table, then print the match. The same file now shows it reads the full table first, like `fm_control_harness_supported` does. Return values and output are unchanged. Verification: with SIGPIPE ignored, tests/fm-dispatch-resolve.test.sh failed 0 of 30 runs after the fix (2 of 20 before). tests/fm-dispatch-resolve.test.sh, tests/fm-brief.test.sh, tests/fm-send-inbox.test.sh, tests/fm-quota-choose.test.sh and tests/fm-quota-array-dispatch-live-e2e.test.sh all pass, and shellcheck is clean. tests/fm-procevent-quota.test.sh fails locally with or without the change ("process-event state root is not a private directory"), so that failure comes from the local environment, not from this fix. No new test was added: the existing one-diagnostic-line assertions already catch this whenever SIGPIPE is ignored, as it is in CI"

This reverts commit c719928.

* no-mistakes(review): Retry deferred local instruction nudges via the watcher

* no-mistakes(review): Document watcher retry for deferred local instruction nudges

* no-mistakes(ci): I fixed both review findings you selected (ci-1 and ci-3). I did not touch the deferral check in bin/fm-send.sh. ci-1 (bin/fm-config-push.sh, retry_deferred_rereads) - Rule that must hold: a deferred reread stays flagged until it is actually delivered. - Before the fix, the flag was removed before any of the steps that can skip a mate: the remote lock-path lookup, validate_secondmate_home, the local lock-path lookup, and the lock acquire. A skip at any of those dropped the flag, so the watcher lost track of the reread. - Now the flag is removed in one place only, when the send succeeds (rc 0). A skipped home, a busy lock, a deferred send (rc 4) or a failed send all leave it in place. The re-mark calls on a busy lock and on rc 4 were no longer needed, so I removed them. I updated the comment above the function to match. - Side effect: a send that keeps failing now stays flagged, so the watcher retries it on every poll and logs each failure. That follows your "don't clear until delivered" rule, but it replaces the old behaviour of leaving a failed send to the next config push or session start. - New test in tests/fm-secondmate-sync.test.sh: T8j "a deferred flag survives a skipped invalid home and is retried once it validates". It takes the home's marker away to make validation fail, checks that nothing is sent and the flag stays, then puts the marker back and checks that the nudge is delivered and both the flag and the retry marker are cleared. It fails on the old code and passes now. ci-3 (bin/fm-secondmate-restart.sh) - Rule that must hold: no automatic send wakes a mate that is waiting on its own open decision. - The two automatic sends in this script are the fallback reread nudge (fall_back_to_nudge) and the persist request. Both now pass --automatic. If a persist request is deferred, its correlation is discarded and the mate goes to the fallback nudge, which is also deferred, so the mate is reported as unreached. - New test in tests/fm-secondmate-restart.test.sh: T3b. It gives a mate an open needs-decision and runs a restart. It checks that both sends report as deferred, the mate's doorbell is never rung, its inbox gets no message, nothing is stopped, and the mate is reported as unreached with exit status 3. It fails on the old code and passes now. - The test marks the watcher as alive first. Without that, the watcher-down warning is printed first and becomes the reported reason instead of the deferral message. Verification - tests/fm-secondmate-sync.test.sh passes. - tests/fm-secondmate-restart.test.sh passes. - tests/fm-secondmate-harness.test.sh (the other test that exercises --retry-deferred) passes. - The fm-send-inbox test that covers automatic deferral passes. I only looked at the last lines of that run, not the whole file. - `shellcheck -x` on the four changed files is clean

* Pin autoarm supervision model in secondmate restart T3b

The fresh watcher beat the test writes proves a live watcher only under the
autoarm model; on CI hosts with no detected harness the persistent model
demands a lock-holding watcher, so the watcher-down banner became the
reported reason and the deferral assertion failed.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Keep deferred secondmate nudges retryable under the inheritance lock.

A bootstrap instruction nudge could write its deferral flag outside the lock the watcher retry holds, so a concurrent retry could delete a flag that had just been set. A restart fallback that is deferred now records the same marker and flag, so the watcher delivers it once the decision closes.

* no-mistakes(document): Document watcher retry of deferred restart re-read nudges

* Send secondmate reread and restart nudges immediately again.

Deferring those nudges let a later config push drop an incomplete transfer once the decision closed. They now send as they do on main.

* Make the no-turn wait opt-in behind config/wait-no-turns.

Homes that do not create the file keep the previous briefs, drive text, and sends.

* no-mistakes(document): Document wait-no-turns inbox wording change in configuration

* no-mistakes(review): Keep checkpoint inbox check; forbid only polling while waiting

* no-mistakes(ci): Fixed ci-2 (Greptile: a concurrent retry marker gets lost). The rule that was broken: the watcher may remove only the `.retry-ring` mark for the record it just processed. A newer mark written in the meantime is owed its own retry. `fm_task_inbox_clear_retry` is the one shared function that removes the mark, and I fixed it there. In `bin/fm-task-inbox-lib.sh` it now takes the record path. It compares the mark's content with that record's name and removes the mark only when they match. When the mark names a different record it returns success and leaves the mark alone. It still fails only when the processed record's own mark can't be removed. Both callers in `bin/fm-watch.sh` now pass `"$rec"`: the dead or missing pane path and the path after a retry ring. So the fix holds at both removal sites. Tests, in `tests/fm-task-inbox.test.sh`: - I added an optional `FM_RING_MARKS_RETRY` hook to the fake tmux. It writes a newer record's mark while the doorbell is being typed, which reproduces the race deterministically. - I added `test_watcher_retry_keeps_a_newer_mark`. The owed retry rings once, the newer mark survives, and a later check rings the newer record once and then clears its mark. The test fails without the fix ("the spent retry removed a newer record's mark written during its ring") and passes with it. - I updated the direct `clear_retry` call in the existing unit test to pass the record. Results: `tests/fm-task-inbox.test.sh` passes in full and `tests/fm-send-inbox.test.sh` passes 15/15. Shellcheck reports only SC1091 "not following sourced file" notices. As instructed, I didn't change the brief inbox wording

* no-mistakes(document): Fix stale wait-no-turns inbox wording in inbox lib comment

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>

---------

Co-authored-by: Christopher McKay <101884182+karotkriss@users.noreply.github.com>
Co-authored-by: Tiago <tiagop@hey.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Cursor Agent <cursoragent@cursor.com>
Co-authored-by: Kun Chen <kunchenguid@users.noreply.github.com>
@cloud-practitioner
cloud-practitioner force-pushed the fm/fm-e2e-tmp-hooks-cleanup branch from 2cfdba0 to 4483c14 Compare October 1, 2026 02:14
@cloud-practitioner cloud-practitioner changed the title test: clean up read-only strip hooks and staged launch dirs in fixture teardown test: remove read-only strip hooks and staged launch dirs in fixture teardown Oct 1, 2026
@cloud-practitioner
cloud-practitioner merged commit 07c549b into main Oct 1, 2026
20 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant