Skip to content

[ci] Get the fork's main green: fix the four standing CI failures - #8

Merged
keenvc merged 4 commits into
mainfrom
fm/fm-fork-main-green
Sep 28, 2026
Merged

keenvc merged 4 commits into
mainfrom
fm/fm-fork-main-green

Conversation

@keenvc

@keenvc keenvc commented Sep 28, 2026 •

Copy link
Copy Markdown
Owner

Main on this fork has been red since 2026-09-20 (runs 35485411963, 35492815658 and 35683307194), so no PR against it can report green. This PR fixes the four failures those runs share and nothing else. Each one was reproduced first; the quoted CI strings turned out to describe three test defects and one documentation defect, and no guard was weakened to get there.

1. Stock macOS Bash snapshot compatibility: expected 18 snapshot/fleet-view tests, got 19

The nineteenth test is test_large_payloads_compose_through_files in tests/fm-fleet-snapshot-view.test.sh. It is a genuine member of that family: it drives bin/fm-fleet-snapshot.sh with a 200KB status fold and backlog and asserts the snapshot survives, which is the regression for the same commit's change to compose jq payloads through --rawfile and --slurpfile instead of argv (the kernel caps one argument at 128KB). Test and fix landed together in d9356ca, a commit that preserved uncommitted local edits during the fork merge, which is why the count in .github/workflows/ci.yml never followed. The macOS job already printed 19 ok lines under /bin/bash 3.2, so only the expectation moves. Fix: the count is now 19.

2. Behavior portable serial 3: the documented example passes opted-in resolution (missing: ' status: clear')

bin/fm-dispatch-resolve.sh refuses docs/examples/crew-dispatch.json before any request:

error: malformed rules file: ... - use profiles whose harness lacks one authoritative provider family require provider: cline

The example is the wrong side, not the resolver. The resolver's provider is the quota-axi provider family it ranks a candidate by. Cline's model id prefix (cline-pass/...) is the launch provider cline reads from -m, and the same prefix could just as well name anthropic or openrouter, so cline is multi-provider by construction, exactly like pi and omp, and docs/configuration.md deliberately keeps it out of the single-provider table. The Pi default in the same example already declares provider: claude for the same reason. The fork's cline commit added the four cline profiles without the field and also added a sentence to docs/configuration.md claiming no field is needed, which contradicts the table two paragraphs above it.

Fix: provider: cline-pass on the four cline profiles (cline-pass is the verified provider id in docs/verification/cline.md), the contradicting sentence corrected, and the test's canned Choice answer given a rule_4 probability because the example now has four rules and the resolver checks the answer's option set against the file. The resolver now returns status: clear on the example, with the cline candidate reported as eligible but unranked because the test fixture's quota snapshot carries no cline-pass row, which is the disclosed-uncertainty outcome AGENTS.md section 4 asks for.

3. Behavior portable serial 9: expected exit 5, got 0

This one deserved the most suspicion, and the guard is fine. tests/fm-claim.test.sh picks the record to corrupt with find "$FM_CLAIM_ROOT" -name '*.claim' | head -n 1 while three claims exist (o/r#7, repos/example#9, owner/repo#11). Directory order differs between filesystems, so on the runner it corrupted a different record and then asked status owner/repo#11, whose intact record answered held with exit 0. CI's stdout shows exactly that held line. On this host find also returns repos/example#9 first, and the same probe with the right record corrupted returns exit 5 with "unreadable or corrupt". Fix: the test selects the record by its documented key= line.

4. Behavior portable serial 3: a dead publication lock wedged publication

This one is timing, and it only bit the 2026-09-22 run: the 09-20 runs passed the same section in 9.1s, the failing run took 19.7s, which is the same 9s plus the test's full 10s wait. Nothing in the lock, watcher or writer path changed between those heads (git diff dce008a0 66063a50 --stat touches only fm-spawn.sh and its tests).

Mechanism: in the restart section the watcher runs with FM_HOME_SUMMARY_INTERVAL=999999, but age_of reports 999999 for a missing ledger, so home_summary_refresh_detached fires on every one-second poll. Those detached refreshes run with the test's FM_HOME_SUMMARY_TIMEOUT=2. Once the lock holder is killed, the watcher's next attempt steals the dead lock first, the test's own idle-only refresh sees the lock held and returns 0 without publishing, and on a slow runner the watcher's attempt is killed at 2s (the producer takes 0.64s on this 32-core box) before it can publish. The next poll repeats the same death, and the ledger never appears. The assertion is about dead-lock reclamation; the 2-second bound was test plumbing that encoded a speed assumption.

Fix: the three restart watchers use FM_HOME_SUMMARY_TIMEOUT=30, the same deadline the file already uses for its accumulated-home publication, with a comment explaining the race. The idle-only refresh, the dead-lock reclamation and the 10-second wait are unchanged.

Evidence pairs

Methodology for 2, 3 and 4: the named test script run with bash from this worktree on Linux, output captured to a file, compared before and after the fix. For 4 the run is additionally CPU-throttled with systemd-run --user --scope -p CPUQuota=30% bash <script> to emulate a slow runner, identically for both halves.

Failure Before After
1 CI: expected 18 snapshot/fleet-view tests, got 19 (macOS job) 19 ok lines locally under bash 5.2; the macOS assertion is exercised by this PR's own CI run
2 not ok - the documented example passes opted-in resolution (missing: ' status: clear'), resolver exit 2 with the error above 18 ok, # all fm-dispatch-resolve tests passed; resolver prints status: clear and profile: --harness 'codex' --model 'gpt-5.5' --effort 'medium' on the example
3 probe: corrupt a non-target record, status owner/repo#11 prints held ..., exit 0 (CI's line); corrupt the target record, exit 5 ok - fm-claim, exit 0
4 throttled: 19 ok then not ok - a dead publication lock wedged publication, exit 1 (CI's line) throttled: 21 ok including publication remains single-flight across watcher restart, exit 0

bin/fm-lint.sh passes (ShellCheck on the changed scripts, actionlint 1.7.12 on the workflow). bin/fm-doc-audience-check.sh passes. bin/fm-test-run.sh --check-coverage passes. tests/fm-ci-workflow.test.sh needs ruby, which this host lacks, so that guard runs in CI only.

What stays red on this PR

Require no-mistakes fails: this is a direct PR with no pipeline attestation, by instruction, because the no-mistakes pipeline rebases against upstream while pushing to the fork. Every other check is green, the first time since 09-20.

One note for the record. The first run's Behavior tests (Herdr) failed in tests/fm-backend-herdr-presentation-e2e.test.sh with "herdr presentation recovery could not acquire its session lock; refusing a concurrent resume". That step gives a concurrent resume 5 seconds (50 tries at 100 ms) to take the presentation session lock while the other home's resume holds it. Nothing in this PR touches that path, the same lane was green on main's 09-22 run and on #7, and a re-run of the job passed, so it is a timing flake in that lane rather than one of the four failures fixed here.

The fork's tests/fm-fleet-snapshot-view.test.sh carries
test_large_payloads_compose_through_files, the regression for
bin/fm-fleet-snapshot.sh composing payloads above 128KB through files
instead of argv. It landed in d9356ca together with that snapshot
change and passes under /bin/bash 3.2 on the macOS runner, which
already counted 19 ok lines, so the hard-coded 18 in the macOS job was
the only thing left behind.
The resolver refuses docs/examples/crew-dispatch.json with "profiles
whose harness lacks one authoritative provider family require
provider: cline", so the documented-example check in
tests/fm-dispatch-resolve.test.sh has failed since the cline profiles
were added to the example.

The example is the wrong side. The resolver's provider is the quota-axi
provider family it ranks candidates by, not the launch prefix cline
reads from its model id, and the single-provider table in
docs/configuration.md leaves cline out on purpose: the model prefix,
not the harness, decides who bills the run, exactly as for pi and omp.
The Pi default in the same example already declares provider: claude
for that reason. Add provider: cline-pass to the four cline profiles,
correct the one sentence in docs/configuration.md that claimed no field
was needed, and give the test's canned Choice answer the fourth rule
the example now has, since the resolver checks the answer against the
full option set.
…urns

The corrupt-record check picked its victim with find | head -n 1 while
three claims exist, so on a filesystem whose directory order differs
from the author's it corrupted o/r#7 or repos/example#9 and then asked
about owner/repo#11, whose intact record answered "held" with exit 0.
CI's stdout showed exactly that line. Select the record by its
documented key= line instead. The guard itself is intact: corrupting
the right record returns exit 5 on the same inputs.
In the watcher-restart section of tests/fm-home-summary-refresh.test.sh
the watcher runs with FM_HOME_SUMMARY_INTERVAL=999999, but age_of
reports 999999 for a missing ledger, so its detached refresh fires on
every one-second poll. After the lock holder is killed that refresh
steals the dead lock ahead of the test's idle-only refresh, which then
returns without publishing, and with the section's
FM_HOME_SUMMARY_TIMEOUT=2 a slow runner kills the watcher's attempt
before it publishes. The next poll dies the same way and the ledger
never appears, which is the "a dead publication lock wedged
publication" failure in CI run 35683307194.

Raise the three restart watchers' bound to 30 seconds, the deadline
this file already uses for its accumulated-home publication. Under a
30% CPU quota the unchanged test fails on exactly that line and the
changed one passes all 21 checks. The idle-only refresh, the dead-lock
reclamation and the 10-second wait are untouched.
@keenvc keenvc added do-not-auto-ready Keep PR in draft; do not mark ready for review automatically and removed do-not-auto-ready Keep PR in draft; do not mark ready for review automatically labels Sep 28, 2026
@keenvc
keenvc marked this pull request as ready for review September 28, 2026 23:44
@keenvc
keenvc merged commit e60c3c6 into main Sep 28, 2026
35 of 38 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant