fix(bin): let a live rebased run outrank the previous failed one in crew state - #4590
dsantosg1103 wants to merge 7 commits into
Conversation
bin/fm-crew-state.sh reported "state: failed - source: run-step - run failed" for a task whose no-mistakes run was alive and parked at its gate, and the watcher turned that into a terminal-outcome wake for an outcome that never happened. The ledger reader ended its scan at any newest row whose head the head rule could not bind. That is correct for a terminal or unclassifiable row, but a LIVE row whose head the pipeline replayed onto an advanced upstream shares no ancestry with the worktree HEAD in either direction, so it was discarded exactly like a foreign run - leaving the previous run's terminal record standing as the present. A rebased head is exactly as unprovable as an unfetched one, so both now reach the same anchored pipeline-continuation recognition. The anchor itself is unchanged: the immediately older row for the same branch must still resolve to EXACTLY the worktree HEAD, so branch-name coincidence and other tasks' runs still never match. Observed 2026-09-07 on nutrifam-cerrar-allow-authenticated.
|
Speaking as Kun's firstmate: Outcome: contract-class: restore — crew-state / VISION: One captain/one interface — aligns (no false terminal wakes). Authority explicit — aligns (no new autonomy). Scripts own mechanics — aligns (ledger ancestry rules). Restart non-event — aligns (durable run ledger). Delegation spine — aligns (evidence honesty). Fleet outlives vendor — aligns. Scope — aligns. Security: clean (no workflows/secrets). Review note on teardown vs read-path asymmetry is informational / out of intent — not a merge gate. |
Intent
bin/fm-crew-state.shreportostate: failed - source: run-step - run failedpara una tarea cuya corrida de no-mistakes estaba VIVA y parada en su compuerta. Observado el 2026-09-07 en la tarea nutrifam-cerrar-allow-authenticated, copia~/.treehouse/Nutrifam-fb5f91/1/Nutrifam:01M1YR324J8R91FYXYZC36ZPEM, status running, awaiting_agent parked 23m57s, compuerta en el paso review con status fix_review y 2 hallazgos, uno ask-user.01M1Y8WFNKZVWKTT544KRS4GXQ, muerta en review por limite de cuota. Despues de esa, el trabajador recupero custodia y arranco la corrida nueva, y el clasificador se quedo leyendo el registro de la vieja.Por que importa: el vigilante emitio un wake
check: inactive-outcome ... state=failedque pide presentarle al capitan un resultado terminal que NO ocurrio. Es la misma clase de falso positivo que ya esta documentada endata/learnings.mdpara el falso "termino" al disparar /no-mistakes, y es peor que aquella, porque aquella venia de una linea de status ambigua y esta viene del clasificador que existe justamente para ser la fuente de verdad del estado actual. Un aviso equivocado que llega hasta el capitan erosiona la confianza en los avisos de verdad.Criterio de aceptacion: con una corrida viva parada en compuerta y una corrida anterior fallida en la misma copia de trabajo,
bin/fm-crew-state.shdebe reportar el estado de la corrida VIVA, y no debe emitirse un wake de resultado terminal. Con cobertura de prueba de ese caso concreto, porque es reproducible.What Changed
fm_nm_runs_status_for_worktreeinbin/fm-nm-run-lib.shnow treats a newest ledger row whose head diverges from the worktree HEAD the same as one that never resolved locally: if the row is LIVE it reaches the pipeline-continuation anchor (the immediately older row for the same branch must resolve to exactly the worktree HEAD), so a run the pipeline rebased onto an advanced upstream is recognized instead of falling through to the previous failed run's status.tests/fm-crew-state.test.shadds three cases built on real git fixtures: the incident shape (rebased live row + terminal row at the worktree commit, pinning the wholestate: failed · source: run-step · run failedline as a regression literal), a negative control where the anchor row is a descendant rather than the exact commit and nothing binds, and a strict-ancestor live newest row that the anchor must not rescue. Header comments inbin/fm-crew-state.sh,bin/fm-nm-run-lib.sh, anddocs/architecture.mdare updated to state the widened unbindable-head rule and its limits.Risk Assessment
✅ Low: The change is tightly scoped and conservative - it widens exactly one ledger shape (live row with a diverged head) into the existing exact-head anchor, adds a strict-ancestor guard that closes the round-4 regression, leaves the sibling scan, the anchor rule, and teardown's call path untouched, and pins all four shapes with behavioral tests that drive the real script.
Testing
Reproduced the 2026-09-07 false-failed report against the real product before testing the fix: with the dead run bound by
axi statusat the worktree's exact commit and the live rebased run newest on the runs ledger, the base-commit binaries print the incident's exact line and the target-commit binaries print the live run instead, with the real watcher scan queueing nothing. A no-live-successor control on the same fixture still reports failed and still queues the terminal-outcome wake, so the silence is specific to the incident shape rather than a disabled wake path. Three adversarial shapes aimed at the widened anchor - a strict-ancestor live newest row, a diverged terminal newest row, and a rebased live row anchored only by a descendant - each refuse to bind a run and fall to the pane, and the deliberately reverted sibling rule still lets a terminal row at the worktree commit stand. Evidence is CLI transcripts because this change has no rendered UI surface;bin/fm-crew-state.sh's single emitted line is the end-user surface and is pinned verbatim in the transcripts. One unrelated bounded-scan failure in the inactive-reconcile suite reproduces identically at the base commit on this host and is reported informationally.state: failedstate: failed · source: run-step · run failed(the incident's own line), target-commit tree prints `state: working…bin/fm-inactive-reconcile.sh scanwith the real crew-state; state/.wake-queue and state/terminal-outcomes stay empty (incident-transcript.txt)state: failed · source: run-step · run failedand queuesinactive-outcome:… child=nutrifamcerrar state=failed(incident-t…state: working · source: pane · harness busy (claude-hook)- neversource: run-step(guards-transcript.txt)source: pane, with nostate: failedand nosource: run-step(guards-transcript.txt)state: working · source: pane · harness busy (claude-hook), so the anchor stayed exact-equality only (guards-transcript.txt)state: failed · source: run-step · run failed, matching the round-4 revert decision (guards-transcript.txt)bash tests/fm-crew-state.test.sh- all cases pass, including the four new ledger/ancestry cases (crew-state-targeted-cases.txt)Evidence: Incident end-to-end transcript (pre-fix vs fixed crew-state, watcher scan, control)
Source: Incident end-to-end transcript (pre-fix vs fixed crew-state, watcher scan, control)
worktree HEAD (submitted, where the DEAD run died): 042a5e4... live run head (rebased onto advanced upstream): 56d9ecf... neither commit is an ancestor of the other: confirmed === S1: the incident shape, PRE-FIX product (base bdcacb9) === state: failed · source: run-step · run failed === S1: the incident shape, FIXED product (HEAD) === state: working · source: run-step · validating (background run) === S2: the watcher (bin/fm-inactive-reconcile.sh scan) over the same shape === scan exit: 0 wake/outcome records emitted: [] === S3 control: the SAME dead run with no live successor still reports and wakes === state: failed · source: run-step · run failed actionable: inactive terminal outcome awaiting captain presentation: child=nutrifamcerrar state=failed scan exit: 0 wake/outcome records emitted: [inactive-outcome:dd0bb0d1524ff1478b40c4f872232aa6 inactive state=failed ]Evidence: Adversarial guard transcript (ancestor / diverged-terminal / unanchored / reverted sibling)
Source: Adversarial guard transcript (ancestor / diverged-terminal / unanchored / reverted sibling)
=== G1 strict-ancestor live newest row (head bfd3c97 is an ancestor of HEAD 6683856) === pre-fix : state: working · source: pane · harness busy (claude-hook) fixed : state: working · source: pane · harness busy (claude-hook) === G2 diverged TERMINAL newest row + perfect exact-head anchor behind it === fixed : state: working · source: pane · harness busy (claude-hook) === G3 rebased live row whose anchor is a DESCENDANT, not the exact commit === fixed : state: working · source: pane · harness busy (claude-hook) === G4 terminal row at the worktree commit, diverged live sibling behind it === pre-fix : state: failed · source: run-step · run failed fixed : state: failed · source: run-step · run failedEvidence: Incident driver script (reproducible)
Source: Incident driver script (reproducible)
Evidence: Adversarial guard driver script (reproducible)
Source: Adversarial guard driver script (reproducible)
Evidence: Targeted crew-state suite cases for this change
Source: Targeted crew-state suite cases for this change
ok - an unfetched live sibling outranks a terminal row at the worktree's exact commit ok - a live run whose head was rebased outranks a terminal row at the worktree's commit ok - a rebased live row without the exact anchor still binds nothing ok - a diverged terminal newest row is never anchored by the row behind it ok - a strict-ancestor live newest row is never anchored ok - a genuinely failed run with no later run is not hidden ok - active fix round with an unfetched pipeline head reads working all fm-crew-state tests passedPipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
bin/fm-teardown.sh:1778- Informational, pre-existing and outside this change's intent - no action expected here. This change taught the READ path (crew-state) that a rebased, resolvable-but-diverged run head is as unprovable as an unfetched one, but teardown's WRITE path still gates the shared ledger proof on the head being absent entirely:[ -z "$(fm_nm_resolve_commit "$wt" "$run_head")" ] || return 1runs beforefm_nm_runs_status_for_worktreeis ever consulted. So for the incident's own shape - a live run parked at its gate whose rebased head the task copy CAN resolve (the premise tests/fm-crew-state.test.sh:1263 states) - crew-state now correctly reports the live run, whiletask_status_is_own_parked_runstill returns 1 and teardown orphans that parked run instead of concluding it. The library header at bin/fm-nm-run-lib.sh:8-13 names both callers as sharing one recognition rule, so the two paths now disagree about what "cannot be bound" means. This is unchanged behavior, not a regression introduced by the diff, and closing it would extend the change beyond the stated intent (which is about fm-crew-state.sh reporting only); recorded here only so the asymmetry is a deliberate choice rather than an oversight.tests/fm-inactive-reconcile.test.sh:735- Pre-existing, unrelated test failure on this host:tests/fm-inactive-reconcile.test.shfails attest_stalled_state_read_is_bounded_and_scan_progresses("next bounded scan did not resume with the following child", tests/fm-inactive-reconcile.test.sh:735). It fails identically when run from a clean tree at the base commit bdcacb9, so this change did not cause it. This machine has no GNUtimeout/gtimeout, sofm_run_timed(bin/fm-timeout-lib.sh:134) takes the perl fork/setpgrp fallback; a manual replay of the bounded world shows the durable scan cursor advancing a -> b across the two budgeted scans while nochild=b state=donewake is queued. Left unfixed because it is outside this change's surface and a fix would touch the bounded-scan path; remote CI should confirm whether it also fails wheretimeoutexists.state: failedstate: failed · source: run-step · run failed(the incident's own line), target-commit tree prints `state: working…bin/fm-inactive-reconcile.sh scanwith the real crew-state; state/.wake-queue and state/terminal-outcomes stay empty (incident-transcript.txt)state: failed · source: run-step · run failedand queuesinactive-outcome:… child=nutrifamcerrar state=failed(incident-t…state: working · source: pane · harness busy (claude-hook)- neversource: run-step(guards-transcript.txt)source: pane, with nostate: failedand nosource: run-step(guards-transcript.txt)state: working · source: pane · harness busy (claude-hook), so the anchor stayed exact-equality only (guards-transcript.txt)state: failed · source: run-step · run failed, matching the round-4 revert decision (guards-transcript.txt)bash tests/fm-crew-state.test.sh- all cases pass, including the four new ledger/ancestry cases (crew-state-targeted-cases.txt)bash tests/fm-crew-state.test.sh(full targeted suite for the changed classifier, all cases pass, including the four new ones)bash tests/fm-gotmp.test.sh(other consumer of bin/fm-nm-run-lib.sh)bash tests/fm-inactive-reconcile.test.sh(consumer of the wake path; 1 pre-existing failure, reproduced at base commit bdcacb9 from a cleangit archivetree)Live incident driver:BASE_COMMIT=bdcacb9 bash drive-incident.sh <repo> <scratch>- builds a real git task copy with a rebased (bidirectionally non-ancestral) live head, serves the dead previous run through a fakeno-mistakes axi statusand the live row throughno-mistakes runs, then runs the realbin/fm-crew-state.shunder pre-fix and fixed bin trees and the realbin/fm-inactive-reconcile.sh scanLive guard driver:BASE_COMMIT=bdcacb9 bash drive-guards.sh <repo> <scratch>- drives the realbin/fm-crew-state.shover four adversarial ledger/ancestry shapes (strict-ancestor live newest row, diverged terminal newest row, rebased live row with descendant-only anchor, reverted diverged live sibling)Manual replay of the bounded-scan case to characterise the pre-existingfm-inactive-reconcilefailure (cursor advances a -> b, no wake queued; host has no GNUtimeout)✅ **Document** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.
Checks are waiting on a maintainer to authorize workflow runs - this PR comes from a fork, so GitHub reports both runs as
action_requiredwith zero jobs started rather than failing them. Local validation is complete: review passed with no gate, the test step reproduced the 2026-09-07 incident end-to-end against the realbin/fm-inactive-reconcile.shwatcher path and confirmed no terminal-outcome wake is emitted (with a control proving that path still fires when a run genuinely failed), documentation passed, and ShellCheck 0.11.0 is clean on all three changed shell files.