Skip to content

refactor(tests): give the launch-delivery pane one verifying owner - #73

Merged
quinnbot-ai merged 2 commits into
mainfrom
fm/fm-main-red-spawn-tests
Aug 2, 2026
Merged

quinnbot-ai merged 2 commits into
mainfrom
fm/fm-main-red-spawn-tests

Conversation

@quinnbot-ai

@quinnbot-ai quinnbot-ai commented Aug 2, 2026 •

Copy link
Copy Markdown
Owner

What this changes

Seventeen fake terminals across thirteen test suites each carried their own copy of bin/fm-spawn.sh's launch-delivery protocol.
Every copy answered the staged-launch check the same weak way - parsing the token out of the submitted line and echoing the LAUNCH_OK marker back - so it reported success without ever executing the checksum it existed to verify.
Corrupt staging would have passed unnoticed in all of them.

This adds fm_fake_pane_shell to tests/lib.sh as the single owner of a healthy pane's answer, and composes it into every one of those fakes.
It runs the submitted check against the bytes the pane actually accumulated, so all thirteen suites now perform the real verification.

Scope notes

Each fake keeps ownership of everything else it models: its own logging, window inventory, kimi readiness state, and pane content.
A fake for a non-tmux backend feeds the shared shell directly through fm_fake_pane_line, the way the herdr fixture answers pane run and pane read.

Two files deliberately keep their own emulation:

  • tests/fm-spawn-launch-delivery.test.sh owns the truncation, wrapping, and retry-injection contract, which a healthy pane deliberately does not model.
  • tests/fm-backend-orca.test.sh already executed the real check.

Net 171 fewer lines of duplicated emulation.

History

This branch started as a fix for the eight scripts that went red on main after #71.
#72 landed first and independently restored those scripts by adding the per-suite copies described above, so the red-main framing no longer applies and this PR is now purely the consolidation plus the verification fix.

Verification

All three portable CI lanes on this branch, 102 scripts, zero failures:

portable-parallel-1  total=11 failed=0
portable-parallel-2  total=13 failed=0
portable-serial      total=78 failed=0

bin/fm-lint.sh clean, bin/fm-test-run.sh --check-coverage ok (total=113), bin/fm-doc-audience-check.sh ok.

Seventeen fake terminals across thirteen suites each carried their own copy
of fm-spawn's launch-delivery protocol. Every copy answered the staged-launch
check by echoing back the marker it parsed out of the submitted line, so it
reported success without ever running the checksum it was meant to verify:
corrupt staging would have passed unnoticed in each of them.

Add fm_fake_pane_shell to tests/lib.sh as the single owner of a healthy pane's
answer, and compose it into every one of those fakes. It executes the submitted
check against the bytes the pane actually accumulated, so all thirteen suites
now run the real verification. Each fake keeps ownership of everything else it
models - its own logging, window inventory, kimi readiness state, and pane
content - and a fake for another backend feeds the shared shell directly, the
way the herdr fixture answers `pane run` and `pane read`.

tests/fm-spawn-launch-delivery.test.sh keeps its own purpose-built fake: it
owns the truncation, wrapping, and retry-injection contract, which a healthy
pane deliberately does not model. tests/fm-backend-orca.test.sh already ran the
real check and is left alone.

Net 171 fewer lines of duplicated emulation.
@quinnbot-ai
quinnbot-ai force-pushed the fm/fm-main-red-spawn-tests branch from 23f6424 to a0de4ac Compare August 2, 2026 01:38
@quinnbot-ai quinnbot-ai changed the title fix(tests): restore green main - shared pane-shell launch-delivery helper refactor(tests): give the launch-delivery pane one verifying owner Aug 2, 2026
The portable serial cap was a hang tripwire set at 20 minutes against a stale
"~13 min wall" note. The lane now measures 16.1 minutes over 78 scripts on one
hosted runner and 17.9 over 69 on another, while two same-revision runs
exceeded 20 and were killed: runner speed moves this lane by more than a third,
so a cap near the typical wall sits inside the variance and reports a healthy
suite as red. Both main and every PR were red for that reason, with no failing
assertion.

Raise it to 40 minutes, roughly 2.5x the measured wall and the value the Herdr
lane already uses, and replace the stale comment with the real measurements.

Record those measurements in docs/fm-test-portable-shards.md under "Measured
serial wall", with the script count and run id behind each one. The lane grows
whenever a new script is neither proven-isolated nor Herdr-gated, so the next
person to grow it can see the current wall and the runner spread instead of
rediscovering both from a cancelled job.

No test is weakened, skipped, or reassigned to another lane.
@quinnbot-ai
quinnbot-ai merged commit 48c0c2d into main Aug 2, 2026
10 of 11 checks passed
quinnbot-ai added a commit that referenced this pull request Aug 10, 2026
* refactor(tests): give the launch-delivery pane one verifying owner

Seventeen fake terminals across thirteen suites each carried their own copy
of fm-spawn's launch-delivery protocol. Every copy answered the staged-launch
check by echoing back the marker it parsed out of the submitted line, so it
reported success without ever running the checksum it was meant to verify:
corrupt staging would have passed unnoticed in each of them.

Add fm_fake_pane_shell to tests/lib.sh as the single owner of a healthy pane's
answer, and compose it into every one of those fakes. It executes the submitted
check against the bytes the pane actually accumulated, so all thirteen suites
now run the real verification. Each fake keeps ownership of everything else it
models - its own logging, window inventory, kimi readiness state, and pane
content - and a fake for another backend feeds the shared shell directly, the
way the herdr fixture answers `pane run` and `pane read`.

tests/fm-spawn-launch-delivery.test.sh keeps its own purpose-built fake: it
owns the truncation, wrapping, and retry-injection contract, which a healthy
pane deliberately does not model. tests/fm-backend-orca.test.sh already ran the
real check and is left alone.

Net 171 fewer lines of duplicated emulation.

* ci: size the serial lane timeout to its measured wall

The portable serial cap was a hang tripwire set at 20 minutes against a stale
"~13 min wall" note. The lane now measures 16.1 minutes over 78 scripts on one
hosted runner and 17.9 over 69 on another, while two same-revision runs
exceeded 20 and were killed: runner speed moves this lane by more than a third,
so a cap near the typical wall sits inside the variance and reports a healthy
suite as red. Both main and every PR were red for that reason, with no failing
assertion.

Raise it to 40 minutes, roughly 2.5x the measured wall and the value the Herdr
lane already uses, and replace the stale comment with the real measurements.

Record those measurements in docs/fm-test-portable-shards.md under "Measured
serial wall", with the script count and run id behind each one. The lane grows
whenever a new script is neither proven-isolated nor Herdr-gated, so the next
person to grow it can see the current wall and the runner spread instead of
rediscovering both from a cancelled job.

No test is weakened, skipped, or reassigned to another lane.

---------

Co-authored-by: QuinnBot <quinnbot@proton.me>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant