refactor(tests): give the launch-delivery pane one verifying owner - #73
Merged
Merged
Conversation
Seventeen fake terminals across thirteen suites each carried their own copy of fm-spawn's launch-delivery protocol. Every copy answered the staged-launch check by echoing back the marker it parsed out of the submitted line, so it reported success without ever running the checksum it was meant to verify: corrupt staging would have passed unnoticed in each of them. Add fm_fake_pane_shell to tests/lib.sh as the single owner of a healthy pane's answer, and compose it into every one of those fakes. It executes the submitted check against the bytes the pane actually accumulated, so all thirteen suites now run the real verification. Each fake keeps ownership of everything else it models - its own logging, window inventory, kimi readiness state, and pane content - and a fake for another backend feeds the shared shell directly, the way the herdr fixture answers `pane run` and `pane read`. tests/fm-spawn-launch-delivery.test.sh keeps its own purpose-built fake: it owns the truncation, wrapping, and retry-injection contract, which a healthy pane deliberately does not model. tests/fm-backend-orca.test.sh already ran the real check and is left alone. Net 171 fewer lines of duplicated emulation.
quinnbot-ai
force-pushed
the
fm/fm-main-red-spawn-tests
branch
from
August 2, 2026 01:38
23f6424 to
a0de4ac
Compare
The portable serial cap was a hang tripwire set at 20 minutes against a stale "~13 min wall" note. The lane now measures 16.1 minutes over 78 scripts on one hosted runner and 17.9 over 69 on another, while two same-revision runs exceeded 20 and were killed: runner speed moves this lane by more than a third, so a cap near the typical wall sits inside the variance and reports a healthy suite as red. Both main and every PR were red for that reason, with no failing assertion. Raise it to 40 minutes, roughly 2.5x the measured wall and the value the Herdr lane already uses, and replace the stale comment with the real measurements. Record those measurements in docs/fm-test-portable-shards.md under "Measured serial wall", with the script count and run id behind each one. The lane grows whenever a new script is neither proven-isolated nor Herdr-gated, so the next person to grow it can see the current wall and the runner spread instead of rediscovering both from a cancelled job. No test is weakened, skipped, or reassigned to another lane.
quinnbot-ai
added a commit
that referenced
this pull request
Aug 10, 2026
* refactor(tests): give the launch-delivery pane one verifying owner Seventeen fake terminals across thirteen suites each carried their own copy of fm-spawn's launch-delivery protocol. Every copy answered the staged-launch check by echoing back the marker it parsed out of the submitted line, so it reported success without ever running the checksum it was meant to verify: corrupt staging would have passed unnoticed in each of them. Add fm_fake_pane_shell to tests/lib.sh as the single owner of a healthy pane's answer, and compose it into every one of those fakes. It executes the submitted check against the bytes the pane actually accumulated, so all thirteen suites now run the real verification. Each fake keeps ownership of everything else it models - its own logging, window inventory, kimi readiness state, and pane content - and a fake for another backend feeds the shared shell directly, the way the herdr fixture answers `pane run` and `pane read`. tests/fm-spawn-launch-delivery.test.sh keeps its own purpose-built fake: it owns the truncation, wrapping, and retry-injection contract, which a healthy pane deliberately does not model. tests/fm-backend-orca.test.sh already ran the real check and is left alone. Net 171 fewer lines of duplicated emulation. * ci: size the serial lane timeout to its measured wall The portable serial cap was a hang tripwire set at 20 minutes against a stale "~13 min wall" note. The lane now measures 16.1 minutes over 78 scripts on one hosted runner and 17.9 over 69 on another, while two same-revision runs exceeded 20 and were killed: runner speed moves this lane by more than a third, so a cap near the typical wall sits inside the variance and reports a healthy suite as red. Both main and every PR were red for that reason, with no failing assertion. Raise it to 40 minutes, roughly 2.5x the measured wall and the value the Herdr lane already uses, and replace the stale comment with the real measurements. Record those measurements in docs/fm-test-portable-shards.md under "Measured serial wall", with the script count and run id behind each one. The lane grows whenever a new script is neither proven-isolated nor Herdr-gated, so the next person to grow it can see the current wall and the runner spread instead of rediscovering both from a cancelled job. No test is weakened, skipped, or reassigned to another lane. --------- Co-authored-by: QuinnBot <quinnbot@proton.me>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this changes
Seventeen fake terminals across thirteen test suites each carried their own copy of
bin/fm-spawn.sh's launch-delivery protocol.Every copy answered the staged-launch check the same weak way - parsing the token out of the submitted line and echoing the
LAUNCH_OKmarker back - so it reported success without ever executing the checksum it existed to verify.Corrupt staging would have passed unnoticed in all of them.
This adds
fm_fake_pane_shelltotests/lib.shas the single owner of a healthy pane's answer, and composes it into every one of those fakes.It runs the submitted check against the bytes the pane actually accumulated, so all thirteen suites now perform the real verification.
Scope notes
Each fake keeps ownership of everything else it models: its own logging, window inventory, kimi readiness state, and pane content.
A fake for a non-tmux backend feeds the shared shell directly through
fm_fake_pane_line, the way the herdr fixture answerspane runandpane read.Two files deliberately keep their own emulation:
tests/fm-spawn-launch-delivery.test.showns the truncation, wrapping, and retry-injection contract, which a healthy pane deliberately does not model.tests/fm-backend-orca.test.shalready executed the real check.Net 171 fewer lines of duplicated emulation.
History
This branch started as a fix for the eight scripts that went red on main after #71.
#72 landed first and independently restored those scripts by adding the per-suite copies described above, so the red-main framing no longer applies and this PR is now purely the consolidation plus the verification fix.
Verification
All three portable CI lanes on this branch, 102 scripts, zero failures:
bin/fm-lint.shclean,bin/fm-test-run.sh --check-coverageok (total=113),bin/fm-doc-audience-check.shok.