[ci] Get the fork's main green: fix the four standing CI failures - #8
Merged
Merged
Conversation
The fork's tests/fm-fleet-snapshot-view.test.sh carries test_large_payloads_compose_through_files, the regression for bin/fm-fleet-snapshot.sh composing payloads above 128KB through files instead of argv. It landed in d9356ca together with that snapshot change and passes under /bin/bash 3.2 on the macOS runner, which already counted 19 ok lines, so the hard-coded 18 in the macOS job was the only thing left behind.
The resolver refuses docs/examples/crew-dispatch.json with "profiles whose harness lacks one authoritative provider family require provider: cline", so the documented-example check in tests/fm-dispatch-resolve.test.sh has failed since the cline profiles were added to the example. The example is the wrong side. The resolver's provider is the quota-axi provider family it ranks candidates by, not the launch prefix cline reads from its model id, and the single-provider table in docs/configuration.md leaves cline out on purpose: the model prefix, not the harness, decides who bills the run, exactly as for pi and omp. The Pi default in the same example already declares provider: claude for that reason. Add provider: cline-pass to the four cline profiles, correct the one sentence in docs/configuration.md that claimed no field was needed, and give the test's canned Choice answer the fourth rule the example now has, since the resolver checks the answer against the full option set.
…urns The corrupt-record check picked its victim with find | head -n 1 while three claims exist, so on a filesystem whose directory order differs from the author's it corrupted o/r#7 or repos/example#9 and then asked about owner/repo#11, whose intact record answered "held" with exit 0. CI's stdout showed exactly that line. Select the record by its documented key= line instead. The guard itself is intact: corrupting the right record returns exit 5 on the same inputs.
In the watcher-restart section of tests/fm-home-summary-refresh.test.sh the watcher runs with FM_HOME_SUMMARY_INTERVAL=999999, but age_of reports 999999 for a missing ledger, so its detached refresh fires on every one-second poll. After the lock holder is killed that refresh steals the dead lock ahead of the test's idle-only refresh, which then returns without publishing, and with the section's FM_HOME_SUMMARY_TIMEOUT=2 a slow runner kills the watcher's attempt before it publishes. The next poll dies the same way and the ledger never appears, which is the "a dead publication lock wedged publication" failure in CI run 35683307194. Raise the three restart watchers' bound to 30 seconds, the deadline this file already uses for its accumulated-home publication. Under a 30% CPU quota the unchanged test fails on exactly that line and the changed one passes all 21 checks. The idle-only refresh, the dead-lock reclamation and the 10-second wait are untouched.
keenvc
marked this pull request as ready for review
September 28, 2026 23:44
This was referenced Sep 30, 2026
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Main on this fork has been red since 2026-09-20 (runs 35485411963, 35492815658 and 35683307194), so no PR against it can report green. This PR fixes the four failures those runs share and nothing else. Each one was reproduced first; the quoted CI strings turned out to describe three test defects and one documentation defect, and no guard was weakened to get there.
1.
Stock macOS Bash snapshot compatibility: expected 18 snapshot/fleet-view tests, got 19The nineteenth test is
test_large_payloads_compose_through_filesintests/fm-fleet-snapshot-view.test.sh. It is a genuine member of that family: it drivesbin/fm-fleet-snapshot.shwith a 200KB status fold and backlog and asserts the snapshot survives, which is the regression for the same commit's change to compose jq payloads through--rawfileand--slurpfileinstead of argv (the kernel caps one argument at 128KB). Test and fix landed together in d9356ca, a commit that preserved uncommitted local edits during the fork merge, which is why the count in.github/workflows/ci.ymlnever followed. The macOS job already printed 19oklines under/bin/bash3.2, so only the expectation moves. Fix: the count is now 19.2.
Behavior portable serial 3: the documented example passes opted-in resolution (missing: ' status: clear')bin/fm-dispatch-resolve.shrefusesdocs/examples/crew-dispatch.jsonbefore any request:The example is the wrong side, not the resolver. The resolver's
provideris the quota-axi provider family it ranks a candidate by. Cline's model id prefix (cline-pass/...) is the launch provider cline reads from-m, and the same prefix could just as well name anthropic or openrouter, so cline is multi-provider by construction, exactly likepiandomp, anddocs/configuration.mddeliberately keeps it out of the single-provider table. The Pi default in the same example already declaresprovider: claudefor the same reason. The fork's cline commit added the four cline profiles without the field and also added a sentence todocs/configuration.mdclaiming no field is needed, which contradicts the table two paragraphs above it.Fix:
provider: cline-passon the four cline profiles (cline-passis the verified provider id indocs/verification/cline.md), the contradicting sentence corrected, and the test's canned Choice answer given arule_4probability because the example now has four rules and the resolver checks the answer's option set against the file. The resolver now returnsstatus: clearon the example, with the cline candidate reported as eligible but unranked because the test fixture's quota snapshot carries no cline-pass row, which is the disclosed-uncertainty outcome AGENTS.md section 4 asks for.3.
Behavior portable serial 9: expected exit 5, got 0This one deserved the most suspicion, and the guard is fine.
tests/fm-claim.test.shpicks the record to corrupt withfind "$FM_CLAIM_ROOT" -name '*.claim' | head -n 1while three claims exist (o/r#7,repos/example#9,owner/repo#11). Directory order differs between filesystems, so on the runner it corrupted a different record and then askedstatus owner/repo#11, whose intact record answeredheldwith exit 0. CI's stdout shows exactly that held line. On this hostfindalso returnsrepos/example#9first, and the same probe with the right record corrupted returns exit 5 with "unreadable or corrupt". Fix: the test selects the record by its documentedkey=line.4.
Behavior portable serial 3: a dead publication lock wedged publicationThis one is timing, and it only bit the 2026-09-22 run: the 09-20 runs passed the same section in 9.1s, the failing run took 19.7s, which is the same 9s plus the test's full 10s wait. Nothing in the lock, watcher or writer path changed between those heads (
git diff dce008a0 66063a50 --stattouches onlyfm-spawn.shand its tests).Mechanism: in the restart section the watcher runs with
FM_HOME_SUMMARY_INTERVAL=999999, butage_ofreports 999999 for a missing ledger, sohome_summary_refresh_detachedfires on every one-second poll. Those detached refreshes run with the test'sFM_HOME_SUMMARY_TIMEOUT=2. Once the lock holder is killed, the watcher's next attempt steals the dead lock first, the test's own idle-only refresh sees the lock held and returns 0 without publishing, and on a slow runner the watcher's attempt is killed at 2s (the producer takes 0.64s on this 32-core box) before it can publish. The next poll repeats the same death, and the ledger never appears. The assertion is about dead-lock reclamation; the 2-second bound was test plumbing that encoded a speed assumption.Fix: the three restart watchers use
FM_HOME_SUMMARY_TIMEOUT=30, the same deadline the file already uses for its accumulated-home publication, with a comment explaining the race. The idle-only refresh, the dead-lock reclamation and the 10-second wait are unchanged.Evidence pairs
Methodology for 2, 3 and 4: the named test script run with
bashfrom this worktree on Linux, output captured to a file, compared before and after the fix. For 4 the run is additionally CPU-throttled withsystemd-run --user --scope -p CPUQuota=30% bash <script>to emulate a slow runner, identically for both halves.expected 18 snapshot/fleet-view tests, got 19(macOS job)oklines locally under bash 5.2; the macOS assertion is exercised by this PR's own CI runnot ok - the documented example passes opted-in resolution (missing: ' status: clear'), resolver exit 2 with the error aboveok,# all fm-dispatch-resolve tests passed; resolver printsstatus: clearandprofile: --harness 'codex' --model 'gpt-5.5' --effort 'medium'on the examplestatus owner/repo#11printsheld ..., exit 0 (CI's line); corrupt the target record, exit 5ok - fm-claim, exit 0okthennot ok - a dead publication lock wedged publication, exit 1 (CI's line)okincludingpublication remains single-flight across watcher restart, exit 0bin/fm-lint.shpasses (ShellCheck on the changed scripts, actionlint 1.7.12 on the workflow).bin/fm-doc-audience-check.shpasses.bin/fm-test-run.sh --check-coveragepasses.tests/fm-ci-workflow.test.shneeds ruby, which this host lacks, so that guard runs in CI only.What stays red on this PR
Require no-mistakesfails: this is a direct PR with no pipeline attestation, by instruction, because the no-mistakes pipeline rebases against upstream while pushing to the fork. Every other check is green, the first time since 09-20.One note for the record. The first run's
Behavior tests (Herdr)failed intests/fm-backend-herdr-presentation-e2e.test.shwith "herdr presentation recovery could not acquire its session lock; refusing a concurrent resume". That step gives a concurrent resume 5 seconds (50 tries at 100 ms) to take the presentation session lock while the other home's resume holds it. Nothing in this PR touches that path, the same lane was green on main's 09-22 run and on #7, and a re-run of the job passed, so it is a timing flake in that lane rather than one of the four failures fixed here.