Skip to content

fix(bin): report a live run parked at its gate as parked, not working - #4600

Open
dsantosg1103 wants to merge 12 commits into
kunchenguid:mainfrom
dsantosg1103:fm/fm-crew-state-corrida-parada-se-lee-como-trabajando
Open

dsantosg1103 wants to merge 12 commits into
kunchenguid:mainfrom
dsantosg1103:fm/fm-crew-state-corrida-parada-se-lee-como-trabajando

Conversation

@dsantosg1103

@dsantosg1103 dsantosg1103 commented Sep 16, 2026 •

Copy link
Copy Markdown

Intent

Cuando bin/fm-crew-state.sh no puede atar la consulta directa a la corrida viva, reporta state: working - source: run-step - validating en vez de decir que esa corrida esta PARADA en su compuerta esperando una decision humana. En el incidente que lo destapo, el estado real era fix_review, parada 23m57s, con 2 hallazgos y uno de ellos ask-user.

La consecuencia es la que importa, y esta verificada: bin/fm-fleet-snapshot.sh clasifica working como trabajo activo y manda a la lista de esperas solo lo que dice parked, paused o blocked. Asi que un trabajador sentado en una compuerta que nadie contesta se le presenta al capitan como trabajo avanzando, y puede quedarse ahi indefinidamente sin que nadie lo note.

Es la misma familia de reporte falso que el incidente de la corrida vieja, ya corregido, y la misma que costo dos relanzamientos a ciegas la noche del 2026-09-15 con un trabajador que nunca habia arrancado y se reportaba working.

El capitan aprobo construirlo el 2026-09-15, con una condicion que se verifico y se cumple: que el defecto siguiera vivo en la version de arriba. Lo esta - ningun commit reciente de upstream toca bin/fm-crew-state.sh, bin/fm-nm-run-lib.sh ni bin/fm-fleet-snapshot.sh, y en upstream fm-crew-state.sh sigue degradando a RUN_SOURCE=coarse cuando la corrida atada no esta activa y el registro barato dice live.

Y acepto expresamente el costo que el remedio implica: el registro barato no guarda el id de la corrida, asi que para saber su compuerta hay que conseguir ese id y hacer una segunda consulta por corrida, en un camino que se ejecuta en cada supervision. Sus palabras: "si arreglalo".

Criterio de aceptacion: con una corrida viva PARADA en su compuerta que la consulta directa no puede atar, bin/fm-crew-state.sh debe reportar que esta parada y en que paso, no working, y la vista de flota debe clasificarla como espera y no como trabajo activo. Con cobertura de prueba de ese caso concreto.

What Changed

  • bin/fm-crew-state.sh no longer stops at the ledger's bare running word when the direct axi status query cannot be bound to the live run: nm_inspect_live_run recovers that run's id (from the recent-runs table already in hand, else from one home-view call) and re-queries axi status --run <id>, so a run sitting at an unanswered gate reports parked with its step and findings instead of working · validating. The enrichment proves the id names the attributed row, on this crew's branch, and still live, and keeps its own fixed 3s bound (NM_INSPECT_TIMEOUT) separate from FM_CREW_STATE_NM_TIMEOUT; anything unproven leaves the coarse word exactly as the ledger decided it.
  • bin/fm-nm-run-lib.sh gains fm_nm_runs_decision_for_worktree, which returns "<status> <row-head>" so callers can name the row's head, with fm_nm_runs_status_for_worktree kept as the status-word-only wrapper. Its unbindable-head rule now also covers a newest head that resolves but diverges from the worktree's line of history (a pipeline-replayed branch), while a head the worktree has already advanced past is treated as superseded history and ends the scan. New fm_nm_home_view_run_id parses the home view's recent-runs TOON table by its own header columns, validating branch, live status, head prefix, and id charset.
  • Tests cover the parked-gate case end to end: tests/fm-crew-state.test.sh adds cases for the parked live sibling and other-branch answers reporting their gate, rejection of home-view rows with a wrong/absent head, foreign-branch or terminal inspection results never displacing the live word, and the rebased/ancestor/diverged anchoring rules; tests/fm-fleet-snapshot-view.test.sh asserts the parked run lands on the fleet view's waiting list rather than counting as active work. docs/architecture.md and docs/configuration.md document the inspection, its ten-most-recent-runs horizon, and the separate timeout.

Risk Assessment

✅ Low: Every new path is fail-closed - id lookup, branch/id/liveness guards, and both enrichment calls degrade to exactly the pre-change coarse word - and I verified the new parsers and guards against real installed-CLI output for the incident's own parked run, with behavioral end-to-end coverage in both crew-state and fleet-snapshot; the only residuals (the ten-run id horizon and the enrichment latency) were explicitly accepted by the captain in this run and are now documented in source and in docs/architecture.md.

Testing

I stood up an isolated FM_HOME with a real git repo and a stand-in no-mistakes CLI returning the exact shapes the incident produced, then drove the real firstmate commands against both the base and target trees. The change flips the reported state from working to parked at review: 2 finding(s) (ask-user: authority decision) and moves the crew out of the fleet view's active work into the waiting list, which is precisely the acceptance criterion. Every failure path I could think to attack - a recycled id naming another branch, an inspection that comes back terminal, a table row carrying a different or empty head, a run past the ten-row id horizon, and both enrichment calls hanging - degrades to exactly the coarse word the tool already had, never to a wrong verdict, and the bounded inspection keeps a single crew read at ~5.1s under the snapshot's 10s per-crew budget. The new regression tests fail on the base implementation with the incident's own line and pass on the target. The only stand-in is the third-party no-mistakes binary (I cannot create a real parked pipeline run, and this phase may not invoke the pipeline); the product under test - the firstmate shell commands, the git worktree, the state files, the JSON contract and the rendered view - is entirely real. No screenshots: both surfaces are terminal CLIs, so the captured transcripts are the end-user surface itself.

  • Live validation: ✅ go - 10 of 10 scenarios driven live against the product
Scenario Result Live Evidence
A crew whose live run is parked at an unanswered gate reports the gate, not working ✅ pass live bin/fm-crew-state.sh parkedlive printed state: parked · source: run-step · parked at review: 2 finding(s) (ask-user: authority decision) on the target tree vs `state: working · source: run-step ·…
The captain's fleet view counts that crew as a wait, not as active work ✅ pass live bin/fm-fleet-snapshot.sh --secondmate-home-summary moved parkedlive from active_children (counts.active_children 1) into holds with reason parked at review... and source child-state; `bin/…
The same gate is reached when another crew's run is what the direct query answered ✅ pass live test_other_branch_answer_still_reports_the_live_gate in tests/fm-crew-state.test.sh passes on target and fails on base with the incident's working line
Adversarial: an inspected run naming another branch never answers for this worktree ✅ pass live Drove bin/fm-crew-state.sh with axi status --run returning branch fm/some-other-crew; output stayed state: working · source: run-step · validating (background run) - adversarial-guards.txt
Adversarial: an inspected run that comes back terminal never manufactures a failure verdict ✅ pass live Drove bin/fm-crew-state.sh with the inspection answering status: failed; the proven live coarse word stood and no state: failed was emitted - adversarial-guards.txt
Adversarial: an id-bearing row whose head is not the attributed run's is never inspected ✅ pass live Drove bin/fm-crew-state.sh with the home-view row carrying head 9999abc instead of the ledger's 0123abc; the gate detail was not adopted and the coarse word stood - adversarial-guards.txt
Accepted limitation: a run past the ten id-bearing rows degrades to the pre-change word, not a wrong one ✅ pass live Drove bin/fm-crew-state.sh with a home view of ten newer foreign runs (count: 10 of 60 total); output was state: working · source: run-step · validating (background run), exit 0 - adversarial-gu…
Adversarial: hung enrichment calls stay inside the fleet snapshot's per-crew budget ✅ pass live With both the home view and axi status --run sleeping 30s, a single bin/fm-crew-state.sh read measured 5.09s/5.18s/5.10s against FM_SNAPSHOT_CREW_STATE_TIMEOUT=10, and the snapshot reported the cr…
A live run the pipeline rebased ahead of a dead one is not reported as that dead run ✅ pass live test_rebased_live_run_outranks_terminal_row_at_worktree_head passes on target; on base bdcacb9 it fails with state: failed · source: run-step · run failed
The narrowing guards around that widening still reject rows they must not bind ✅ pass live tests/fm-crew-state.test.sh cases test_rebased_live_run_without_exact_anchor_binds_nothing, test_diverged_terminal_newest_row_is_never_anchored, test_ancestor_live_newest_row_is_never_anchored…
Evidence: Before/after live drive: crew-state line, snapshot classification, rendered fleet view

Source: Before/after live drive: crew-state line, snapshot classification, rendered fleet view

# Live drive: a crew whose no-mistakes run is PARKED at an unanswered gate

Isolated FM_HOME with a real git repo on fm/feat-parkedlive and a stand-in
no-mistakes CLI reproducing the 2026-09-15 incident shapes:
  - bare 'axi status'  -> the PREVIOUS run, failed, at this worktree's exact commit
  - 'runs' ledger      -> newest row: running, head 0123abc (never fetched here), no id column
  - 'axi' home view    -> the id-bearing surface: 01LIVE for this branch at 0123abc
  - 'axi status --run 01LIVE' -> running, awaiting_agent: parked 23m57s, gate step review,
                                 2 findings, one of them ask-user

$ no-mistakes runs --limit 200
  failed     fm/feat-parkedlive 31e6c7e  2026-09-15 11:20
  running    fm/feat-parkedlive 0123abc  2026-09-15 10:05

$ no-mistakes axi status --run 01LIVE
run:
  id: "01LIVE"
  branch: fm/feat-parkedlive
  status: running
  awaiting_agent: parked 23m57s
  head: 0123abc
  pr: ""
  findings: "2 awaiting, 1 auto-fix, 1 info"
  steps[3]{step,status,findings,duration_ms}:
    intent,completed,0,10
    review,fix_review,2,1437000
    test,pending,0,0
gate:
  step: review
  status: fix_review
  summary: "two findings await a decision"
  findings[2]{id,severity,file,action,description}:
    f1,warning,a.go,auto-fix,"ignored error"
    f2,error,b.go,ask-user,"changes product behavior"

================ BEFORE the change (base bdcacb9) ================
$ bin/fm-crew-state.sh parkedlive
state: working · source: run-step · validating (background run)

$ bin/fm-fleet-snapshot.sh --secondmate-home-summary   (captain waiting-list classification)
{
  "holds": [],
  "active_children": [
    {
      "id": "parkedlive",
      "kind": "ship",
      "state": "working",
      "repo": "alpha",
      "name": "Crew validating a change",
      "source": "run-step",
      "doing": "validating (background run)"
    }
  ],
  "active_count": 1
}

$ bin/fm-fleet-view.sh
# Fleet View

Schema: fm-fleet-snapshot.v1
Home: /tmp/fmE2E/home-after

## Under Way
| ID | Current | Kind | Repo/Project | Backend | Endpoint | Artifact | Path | Watch / return channel |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| parkedlive | working / run-step | ship | alpha | tmux | present | - | /tmp/fmE2E/home-after/projects/parkedlive | bin/fm-peek.sh fm-parkedlive |

## Queued
No queued backlog records found.

================ AFTER the change (target 4c232d7) ================
$ bin/fm-crew-state.sh parkedlive
state: parked · source: run-step · parked at review: 2 finding(s) (ask-user: authority decision)

$ bin/fm-fleet-snapshot.sh --secondmate-home-summary   (captain waiting-list classification)
{
  "holds": [
    {
      "id": "parkedlive",
      "title": "parkedlive",
      "blocked_by": null,
      "blocked_by_ids": [],
      "unresolved_blocker_ids": [],
      "reason": "parked at review: 2 finding(s) (ask-user: authority decision)",
      "source": "child-state"
    }
  ],
  "active_children": [],
  "active_count": 0
}

$ bin/fm-fleet-view.sh
# Fleet View

Schema: fm-fleet-snapshot.v1
Home: /tmp/fmE2E/home-after

## Under Way
| ID | Current | Kind | Repo/Project | Backend | Endpoint | Artifact | Path | Watch / return channel |
| --- | --- | --- | --- | --- | --- | --- | --- | --- |
| parkedlive | parked / run-step | ship | alpha | tmux | present | - | /tmp/fmE2E/home-after/projects/parkedlive | bin/fm-peek.sh fm-parkedlive |

## Queued
No queued backlog records found.
Evidence: Adversarial guards and enrichment timing

Source: Adversarial guards and enrichment timing

# Adversarial drives: every path that cannot prove the gate must keep the
# coarse ledger word the change already had, never a wrong one.

## inspected run names ANOTHER branch (recycled id)
$ bin/fm-crew-state.sh parkedlive
state: working · source: run-step · validating (background run)

## inspected run answers TERMINAL (finished between the two reads)
$ bin/fm-crew-state.sh parkedlive
state: working · source: run-step · validating (background run)

## the id-bearing row carries a DIFFERENT head than the attributed row
$ bin/fm-crew-state.sh parkedlive
state: working · source: run-step · validating (background run)

## accepted limitation: the live run sits past the ten id-bearing rows
$ bin/fm-crew-state.sh parkedlive
state: working · source: run-step · validating (background run)

## both enrichment calls hang 30s (NM_INSPECT_TIMEOUT=3 each)
   per-crew snapshot budget FM_SNAPSHOT_CREW_STATE_TIMEOUT is 10s;
   the crew must degrade to the coarse word, not to 'child current state unavailable'.
$ bin/fm-fleet-snapshot.sh --secondmate-home-summary
{"holds":[],"active":[{"id":"parkedlive","state":"working","doing":"validating (background run)"}]}

$ time bin/fm-crew-state.sh parkedlive        # the read the 10s per-crew bound applies to
single crew-state read: 5.09s
single crew-state read: 5.18s
single crew-state read: 5.10s
-> both enrichment calls hung: 5.1s per crew read, inside the 10s per-crew budget,
   and the crew keeps the coarse ledger word instead of folding to
   "child current state unavailable".

Normal-path cost of the two extra calls, same fixture, 3 reads averaged:
  base bdcacb9 : 1.28s per read
  target 4c232d7: 1.91s per read
Evidence: Core before/after, driven live
=== BEFORE (base bdcacb9) ===
state: working · source: run-step · validating (background run)
{"holds": [], "active_children": [{"id":"parkedlive","state":"working","doing":"validating (background run)"}], "active_count": 1}
| parkedlive | working / run-step | ship | alpha | ... |

=== AFTER (target 4c232d7) ===
state: parked · source: run-step · parked at review: 2 finding(s) (ask-user: authority decision)
{"holds": [{"id":"parkedlive","reason":"parked at review: 2 finding(s) (ask-user: authority decision)","source":"child-state"}], "active_children": [], "active_count": 0}
| parkedlive | parked / run-step | ship | alpha | ... |

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

⚠️ **Rebase** - 1 warning
  • ⚠️ bin/fm-crew-state.sh - branch carries 7 commit(s) that exist on your local main branch but were never pushed to origin/main; these may be unintended bundled work (proposed PR changes 5 file(s)):
  • 878de07 no-mistakes(document): scope unbindable-run-head claim to the newest ledger row
  • 5290f64 no-mistakes(document): document superseded-ancestor run heads in architecture
  • 62039d0 no-mistakes(review): revert sibling widening, reject strict-ancestor live heads
  • d796106 no-mistakes(review): cover the live-class guard against diverged terminal rows
  • 0bedfd8 no-mistakes(review): scope header no-widening claim, cover unanchored live sibling
  • 5ba4e9f no-mistakes(review): bind rebased live sibling rows through the exact-head anchor
  • d77216b fix(crew-state): let a rebased live run outrank the previous failed one

Confirm these commits belong in this PR before approving, or manually separate the intended work onto origin/main before gating.

🔧 **Review** - 3 issues found → auto-fixed (3) ✅
  • ⚠️ bin/fm-crew-state.sh:595 - The new axi status --run &lt;id&gt; inspection runs under nm_run, i.e. the full NM_TIMEOUT (default 10s), while bin/fm-fleet-snapshot.sh:146 bounds the ENTIRE per-crew crew-state read at FM_SNAPSHOT_CREW_STATE_TIMEOUT (default 10s) and folds a read that hits the bound to state: unknown. That single call can therefore consume the whole snapshot budget on its own. Concrete sequence: loaded daemon, foreign-branch axi status answer (the shape the code's own comment at bin/fm-crew-state.sh:653 calls routine once several crews validate the same repo) -> axi status + runs --limit 200 + home view (3s bound) + axi status --run (10s bound); the outer 10s bound trips and the parked crew reports unknown instead of parked - the same invisibility class this change fixes, just with a different word. Note the author applied exactly this reasoning to the home-view call at bin/fm-crew-state.sh:585-590 ("a slow enrichment would slow every supervision read that reaches for it") but classified axi status --run as an authoritative read and left it on NM_TIMEOUT; since failing it only keeps the coarse word, it is as much pure enrichment as the home view. This is flagged ask-user rather than auto-fix because which of the two calls counts as authoritative is the author's deliberate call, and the alternative remedy (budgeting inner bounds against the caller's remaining budget) would extend the change beyond its intent.
  • ⚠️ bin/fm-crew-state.sh:584 - Simplification pass: nm_inspect_live_run introduces TWO acceptance paths for the same value - first reading the recent-runs table out of the $RUN_OUT answer already in hand (line 584), then falling back to a dedicated home-view call (lines 592-593). The intent requires only that the id be obtained ("hay que conseguir ese id y hacer una segunda consulta por corrida") and names no source; the home-view path alone satisfies it in every case the first path covers, and it carries its own dedicated test (test_status_answer_carrying_the_runs_table_needs_no_home_view). The narrower form is the single home-view lookup. Stated fairly so it can be dismissed quickly: this is a real optimization, not dead weight - dropping it adds one no-mistakes axi subprocess (~0.7s warm per the code's own measurement) per supervision on the path where the status answer already carries the table, which is precisely the per-heartbeat cost the intent singles out. Raising it only because the pass requires every unrequired acceptance path to be named.
  • ℹ️ bin/fm-nm-run-lib.sh:354 - In fm_nm_home_view_run_id, the head guard index(row_head, want_head) != 1 &amp;&amp; index(want_head, row_head) != 1 silently accepts any row whose head column is empty: awk returns 1 for index(s, &#34;&#34;) on this platform (verified: awk &#39;BEGIN{print index(&#34;abc&#34;,&#34;&#34;)}&#39; prints 1), so the second half of the condition is false and the row is taken regardless of its head - printing the id of a run the ledger never attributed. Unlike the ledger parser in fm_nm_runs_decision_for_worktree, which charset- and length-validates sha (lines 227, 231), nothing here validates row_head, branch, or status - only the extracted id is checked, after the fact, in shell at line 359. The function header at lines 325-327 nonetheless claims "every extracted field is charset-validated, so a reshaped or malformed table yields nothing instead of a wrong id", which is not what the code does. Marked info, not warning: I could not construct a reachable production path, because runs.head_sha is NOT NULL in the CLI's own schema, so a real row never has an empty head. Remedy is one line - if (row_head == &#34;&#34;) next, or hex/length-validate row_head the way the ledger parser validates sha - plus correcting the header claim to describe the validation that actually exists.

🔧 Fix applied.
4 issues (2 warnings, 2 infos) still open:

  • ⚠️ bin/fm-crew-state.sh:584 - The gate recovery only works while the attributed live run is among the CLI's 10 most-recently-CREATED runs for the repo, but the ledger that proves the run scans 200 rows - so past that horizon the change silently returns to the exact false working report the intent forbids. Verified against the installed CLI v1.72.0, not inferred: BOTH id-bearing surfaces are hard-capped at 10 and ordered by created_at DESC (no-mistakes axi status and no-mistakes axi in ~/Dev/Nutrifam both print count: 10 of 60 total / runs[10]{id,branch,status,head,pr}:, and those 10 ids match select id from runs where repo_id=&#39;1014fcb4de41&#39; order by created_at desc limit 10 exactly), while no-mistakes runs --limit 200 returns all 60. no-mistakes runs --help exposes only --limit and emits no id column; axi status --help exposes only --run &lt;id&gt;; there is no id-bearing listing with a limit. Concrete reachable sequence: crew on branch B, ~8 runs/day created in that repo (measured: 6 runs in the 18.7h between created_at 1789360271 and 1789427746), bare axi status answers another crew's run so RUN_SOURCE=coarse, the ledger proves B's run running at head H -> nm_inspect_live_run reads the status answer's runs table (10 rows) and then the home view (10 rows), B's run is row 11+ because it was created a day earlier and has been sitting at an unanswered gate the whole time -> no id -> return 1 -> COARSE_STATUS=running -> state: working - source: run-step - validating (background run) and fm-fleet-snapshot.sh counts it as active work again. This is not hypothetical scale: right now run 01M2H391VCGZ6TX3ZKQKBHDBZ4 in that repo is awaiting_agent: parked 1d2h, i.e. exactly the >1-day park the intent calls out ("puede quedarse ahi indefinidamente"), and the file's own comment at bin/fm-crew-state.sh:126-131 already justifies 200 ledger rows because "a busy multi-crew fleet" pushes a branch's own run deep - the same condition that pushes it out of the 10-row id surface. docs/architecture.md:89 states the outcome unconditionally ("a run parked at a gate nobody has answered reports that gate instead of reading as work in progress"), which this horizon falsifies. Flagged ask-user, and the REMEDY is what needs authorization rather than the defect: the change already degrades to the pre-change word, and closing the gap needs an upstream CLI surface that lists run ids with a caller-chosen limit (or reading ~/.no-mistakes/state.sqlite directly) - both well beyond this change - so the alternative is to accept the bound explicitly and scope the architecture claim to it.
  • ⚠️ bin/fm-crew-state.sh:595 - Re-raising round 1's unresolved enrichment-call-uses-authoritative-timeout with new evidence that the worst case is larger than reported there. The new axi status --run at line 595 runs under NM_TIMEOUT (10s), and because a successful inspection sets RUN_SOURCE=full (line 606), the full path's ci branch becomes reachable from the coarse entry point for the first time: an inspected run whose steps table carries ci,running,... reaches nm_effective_ci_step_status -> nm_ci_checks_state (bin/fm-crew-state.sh:518) -> a FIFTH bounded call, nm_run axi logs --step ci --run &lt;id&gt;, also at 10s. Serial worst-case bounds on that path are now 10 (axi status) + 10 (runs --limit 200) + 3 (home view) + 10 (axi status --run) + 10 (axi logs) = 43s, against fm-fleet-snapshot.sh:146's FM_SNAPSHOT_CREW_STATE_TIMEOUT of 10s for the ENTIRE per-crew read, which folds to state: unknown (bin/fm-fleet-snapshot.sh:307-324). Before the change the same path bounded at 20s and could not reach the ci-log call at all. Stated fairly so it can be dismissed quickly: I measured the warm cost at 0.19-0.20s per invocation (three runs of no-mistakes runs --limit 1), so the normal added cost is ~0.4s, not the ~0.7s the comment at line 587-591 assumes; the tail only opens when the CLI's own network update check stalls, which is precisely the condition that comment cites as "unbounded by anything local when the network is slow". The author applied that reasoning to the home view (3s) but classified axi status --run as authoritative; since failing it only keeps the coarse word, it is as much pure enrichment as the home view. ask-user because which call is authoritative is the author's deliberate call, and the alternative remedy (budgeting each inner bound against the caller's remaining budget) would extend the change beyond its intent.
  • ℹ️ tests/fm-crew-state.test.sh:389 - run_parked_live_sibling - the fixture every acceptance test for this change depends on - renders a parked run as status: fix_review plus a top-level gate: block with step: review. The installed CLI does not emit that shape for a live parked run. Captured from a real parked crew (no-mistakes axi status in ~/.treehouse/Nutrifam-fb5f91/1/Nutrifam, run 01M2H391VCGZ6TX3ZKQKBHDBZ4): status: running, awaiting_agent: parked 1d2h, findings: &#34;3 awaiting, 1 auto-fix, 2 info&#34;, the gate is only visible as a ci,fix_review,3,349921 row inside steps[9]{step,status,findings,duration_ms}, and there is NO top-level gate: block and no findings[N]{...} table. Not a defect: I traced that exact text through the full path and it still lands correctly - nm_field status=running, awaiting non-empty so the parked elif fires, has_gate=0 so nm_gate_name falls to nm_gate_step_row and returns ci, nm_findings_count finds no findings[N] so the step row's 3 is used, and no ask-user substring is present - giving state: parked - source: run-step - parked at ci: 3 finding(s). The gap is coverage: the production shape exercises the awaiting_agent-only, no-gate-block, step-row-findings path, and no test in this change covers it, so a refactor that broke only that path would pass. Remedy is mechanical and non-user-visible: add one fixture variant carrying the captured shape above (or switch the existing one to it) so the acceptance test asserts against what the CLI returns.
  • ℹ️ bin/fm-crew-state.sh:592 - Informational, to retire a still-pending round 1 finding rather than to request work. Round 1's dual-id-source-not-required-by-intent claimed "the home-view path alone satisfies it in every case the first path covers", making the status-answer table path an unrequired second acceptance path. Real CLI output refutes both halves. (1) no-mistakes axi status run where the branch HAS a run prints only run: and branch_sync: with no runs[N] table at all (captured in ~/.treehouse/Nutrifam-fb5f91/1/Nutrifam), which is exactly the first call site's state (bare status bound this branch's TERMINAL run, ledger proved a live sibling) - so there the home-view call is mandatory, not optional. (2) no-mistakes axi status where the branch has NO run does carry runs[10]{id,branch,status,head,pr}: with ids (captured in ~/Dev/firstmate and ~/Dev/Nutrifam), which is the second call site's routine state - so there the table path saves a real subprocess. Neither path subsumes the other; both are required. No action needed on this one.

🔧 Fix applied.
2 warnings still open:

  • ⚠️ bin/fm-crew-state.sh:584 - The gate recovery only works while the attributed live run is among the CLI's 10 most-recently-CREATED runs for the repo, but the ledger that proves the run scans 200 rows - so past that horizon the change silently returns to the exact false working report the intent forbids. Re-verified today against the installed CLI v1.72.0: both id-bearing surfaces are hard-capped at 10 and ordered by created_at DESC (no-mistakes axi and no-mistakes axi status in ~/Dev/Nutrifam both print count: 10 of 60 total / runs[10]{id,branch,status,head,pr}:), no-mistakes axi --help exposes no limit flag at all, and no-mistakes runs --help exposes only --limit int and emits no id column - so there is no id-bearing listing with a caller-chosen limit. Concrete reachable sequence: crew on branch B, several runs/day created in that repo, bare axi status answers another crew's run so RUN_SOURCE=coarse, the ledger proves B's run running at head H -> nm_inspect_live_run reads the status answer's runs table (10 rows) then the home view (10 rows), B's run is row 11+ because it was created a day earlier and has been sitting at an unanswered gate the whole time -> no id -> return 1 -> COARSE_STATUS=running -> state: working - source: run-step - validating (background run) and bin/fm-fleet-snapshot.sh counts it as active work again. Not hypothetical scale: run 01M2H391VCGZ6TX3ZKQKBHDBZ4 is currently awaiting_agent: parked 1d2h, exactly the >1-day park the intent calls out, and bin/fm-crew-state.sh:126-131 already justifies 200 ledger rows because a busy multi-crew fleet pushes a branch's own run deep - the same condition that pushes it out of the 10-row id surface. docs/architecture.md:89 states the outcome unconditionally, which this horizon falsifies. Flagged ask-user, and the REMEDY is what needs authorization rather than the defect: the change degrades to the pre-change word rather than to a wrong one, and closing the gap needs an upstream CLI surface that lists run ids with a caller-chosen limit (or reading ~/.no-mistakes/state.sqlite directly) - both well beyond this change - so the alternative is to accept the bound explicitly and scope the architecture claim to it.
  • ⚠️ bin/fm-crew-state.sh:595 - The new axi status --run at line 595 runs under NM_TIMEOUT (10s), and because a successful inspection sets RUN_SOURCE=full (line 606), the full path's ci branch becomes reachable from the coarse entry point for the first time: an inspected run whose steps table carries ci,running,... reaches nm_effective_ci_step_status -> nm_ci_checks_state (bin/fm-crew-state.sh:518) -> a FIFTH bounded call, nm_run axi logs --step ci --run &lt;id&gt;, also at 10s. Serial worst-case bounds are now 10 (axi status) + 10 (runs --limit 200) + 3 (home view) + 10 (axi status --run) + 10 (axi logs) = 43s, against bin/fm-fleet-snapshot.sh:146's FM_SNAPSHOT_CREW_STATE_TIMEOUT of 10s for the ENTIRE per-crew read. Before the change the same path bounded at 20s and could not reach the ci-log call at all. Stated fairly so it can be dismissed quickly: measured warm cost is ~0.2s per invocation, so the normal added cost is ~0.4s, not the ~0.7s the comment at lines 585-591 assumes; and when the outer bound does trip, the crew folds to state: unknown, which bin/fm-fleet-snapshot.sh:1060-1062 surfaces loudly as child current state unavailable and marks the snapshot invalid - a visible degradation, not a recurrence of the silent false-working report the intent targets. The tail only opens when the CLI's own network update check stalls, which is precisely the condition the line 585-591 comment cites as "unbounded by anything local when the network is slow". The author applied that reasoning to the home view (NM_HOME_VIEW_TIMEOUT=3) but classified axi status --run as authoritative; since failing it only keeps the coarse word, it is as much pure enrichment as the home view. ask-user because which call is authoritative is the author's deliberate call, and the alternative remedy (budgeting each inner bound against the caller's remaining budget) would extend the change beyond its intent.

🔧 Fix applied.
✅ Re-checked - no issues remain.

✅ **Test** - passed

✅ No issues found.

  • Live validation: ✅ go - 10 of 10 scenarios driven live against the product
Scenario Result Live Evidence
A crew whose live run is parked at an unanswered gate reports the gate, not working ✅ pass live bin/fm-crew-state.sh parkedlive printed state: parked · source: run-step · parked at review: 2 finding(s) (ask-user: authority decision) on the target tree vs `state: working · source: run-step ·…
The captain's fleet view counts that crew as a wait, not as active work ✅ pass live bin/fm-fleet-snapshot.sh --secondmate-home-summary moved parkedlive from active_children (counts.active_children 1) into holds with reason parked at review... and source child-state; `bin/…
The same gate is reached when another crew's run is what the direct query answered ✅ pass live test_other_branch_answer_still_reports_the_live_gate in tests/fm-crew-state.test.sh passes on target and fails on base with the incident's working line
Adversarial: an inspected run naming another branch never answers for this worktree ✅ pass live Drove bin/fm-crew-state.sh with axi status --run returning branch fm/some-other-crew; output stayed state: working · source: run-step · validating (background run) - adversarial-guards.txt
Adversarial: an inspected run that comes back terminal never manufactures a failure verdict ✅ pass live Drove bin/fm-crew-state.sh with the inspection answering status: failed; the proven live coarse word stood and no state: failed was emitted - adversarial-guards.txt
Adversarial: an id-bearing row whose head is not the attributed run's is never inspected ✅ pass live Drove bin/fm-crew-state.sh with the home-view row carrying head 9999abc instead of the ledger's 0123abc; the gate detail was not adopted and the coarse word stood - adversarial-guards.txt
Accepted limitation: a run past the ten id-bearing rows degrades to the pre-change word, not a wrong one ✅ pass live Drove bin/fm-crew-state.sh with a home view of ten newer foreign runs (count: 10 of 60 total); output was state: working · source: run-step · validating (background run), exit 0 - adversarial-gu…
Adversarial: hung enrichment calls stay inside the fleet snapshot's per-crew budget ✅ pass live With both the home view and axi status --run sleeping 30s, a single bin/fm-crew-state.sh read measured 5.09s/5.18s/5.10s against FM_SNAPSHOT_CREW_STATE_TIMEOUT=10, and the snapshot reported the cr…
A live run the pipeline rebased ahead of a dead one is not reported as that dead run ✅ pass live test_rebased_live_run_outranks_terminal_row_at_worktree_head passes on target; on base bdcacb9 it fails with state: failed · source: run-step · run failed
The narrowing guards around that widening still reject rows they must not bind ✅ pass live tests/fm-crew-state.test.sh cases test_rebased_live_run_without_exact_anchor_binds_nothing, test_diverged_terminal_newest_row_is_never_anchored, test_ancestor_live_newest_row_is_never_anchored…
  • bash tests/fm-crew-state.test.sh - full file passed (all crew-state tests, including the 11 new parked-gate/rebased-run cases)
  • bash tests/fm-fleet-snapshot-view.test.sh - full file passed, including test_parked_live_run_is_a_hold_not_active_work
  • Red-before-green: ran test_parked_live_sibling_reports_its_gate_not_working, test_other_branch_answer_still_reports_the_live_gate, test_status_answer_carrying_the_runs_table_needs_no_home_view, test_rebased_live_run_outranks_terminal_row_at_worktree_head against the base commit's bin/ - all four fail with the incident's own wrong line
  • Manual E2E: isolated FM_HOME + real git repo on fm/feat-parkedlive + stand-in no-mistakes CLI in the v1.72.0 shapes; bin/fm-crew-state.sh parkedlive driven against base bdcacb9 bin/ and target 4c232d7 bin/
  • bin/fm-fleet-snapshot.sh --secondmate-home-summary driven before/after, inspecting .holds, .active_children, .counts.active_children with jq
  • bin/fm-fleet-view.sh rendered before/after (the captain's human surface)
  • Adversarial drives of bin/fm-crew-state.sh: inspected run on another branch, inspected run answering terminal, id-bearing row with a different head, live run past the ten most-recently-created id-bearing rows
  • Timing drive: both enrichment calls (axi home view, axi status --run) hung 30s, single crew-state read measured 3x at ~5.1s against the 10s FM_SNAPSHOT_CREW_STATE_TIMEOUT; normal-path cost measured 1.28s (base) vs 1.91s (target) per read
✅ **Document** - passed

✅ No issues found.

⚠️ **Lint** - 1 warning
  • ⚠️ linter found issues (exit code 1)
✅ **Push** - passed

✅ No issues found.

Merge order, checks, and a known limitation

This work stacks on #4590 and must not be merged before it. The branch carries that pull request's seven commits because they are on the fork's main branch but not yet on this repository's main. Keeping them bundled was a deliberate decision at the rebase gate, not accidental work.

Checks have not run. GitHub left both workflow runs for this pull request at action_required, the standard hold for a pull request opened from a fork: run 35052281258 (CI) and run 35052281321 (Require no-mistakes). Nothing was forced or retried from this side. A maintainer of this repository has to authorize the workflows before any check can report. The validation pipeline that produced this branch ran locally through review, test, document and lint before pushing; the lint step's warning above is the pipeline's environment lacking ShellCheck, not a lint finding - bin/fm-lint.sh with the pinned ShellCheck 0.11.0 reports no findings on every changed shell file at this head.

Known limitation, accepted deliberately. The gate detail is recovered by naming the live run through the CLI's recent-runs table, and that table exposes only the ten most recently created runs for a repository, while the ledger that proves the run reads far more rows. A run parked for longer than those ten runs take to accumulate has no recoverable id, so it falls back to exactly the coarse word it reported before this change, never to a wrong one. Measured against this fleet's own run records, the ten newest runs in its busiest repository span about 72 hours, so the gate detail holds for roughly three days of continuous parking there, and longer in quieter repositories. Closing the gap would need an id-bearing listing that accepts a caller-chosen limit, which the CLI does not offer today, or reading the tool's internal state directly.

bin/fm-crew-state.sh reported "state: failed - source: run-step - run
failed" for a task whose no-mistakes run was alive and parked at its
gate, and the watcher turned that into a terminal-outcome wake for an
outcome that never happened.

The ledger reader ended its scan at any newest row whose head the head
rule could not bind. That is correct for a terminal or unclassifiable
row, but a LIVE row whose head the pipeline replayed onto an advanced
upstream shares no ancestry with the worktree HEAD in either direction,
so it was discarded exactly like a foreign run - leaving the previous
run's terminal record standing as the present.

A rebased head is exactly as unprovable as an unfetched one, so both now
reach the same anchored pipeline-continuation recognition. The anchor
itself is unchanged: the immediately older row for the same branch must
still resolve to EXACTLY the worktree HEAD, so branch-name coincidence
and other tasks' runs still never match.

Observed 2026-09-07 on nutrifam-cerrar-allow-authenticated.
When the direct `axi status` answer cannot be bound to the live run, crew
state fell back to the runs ledger, which records only a status word. A run
sitting at an unanswered gate therefore reported `working - validating
(background run)`: the incident's real state was fix_review, parked 23m57s,
with 2 findings and one of them ask-user. fm-fleet-snapshot.sh counts
`working` as active work and routes only parked/paused/blocked to its waiting
list, so a crew waiting on a decision nobody was told about could sit there
indefinitely.

The ledger has no run id column, so the gate is reachable only by naming that
run: the id now comes from the recent-runs table of the answer already in hand,
or, when that answer carries none, from one short-bounded home-view call, and
only that one run is then inspected with `axi status --run <id>`. The
inspection recovers detail for the run the ledger already attributed and never
widens attribution - the id must name the attributed row, the answer must be
that id, on this crew's branch, and still live - and anything unproven leaves
the ledger's own word exactly as it was, including any terminal answer, which
must never be turned into a failure verdict by a second query.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant