fix(bin): restore beads backend support and dispatch validation - #11
Conversation
…guid#1233) * fix: preserve dispatch harness identity * no-mistakes(review): Fix Grok counterfactual tuple validation * no-mistakes(document): Scope dispatch authentication to selected tuple * fix: restore dispatch instruction budget * no-mistakes(review): Scope dispatch authentication after candidate selection
* Add internal status skill * no-mistakes(document): register /status skill in documentation-audiences inventory * no-mistakes(lint): replace grep|wc -l with grep -c in status skill test * test: silence literal status skill patterns * Refactor bearings default to chat-only --------- Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com>
…#1282) * test: remove source-content assertions * no-mistakes(review): Replace source assertions with runtime behavior coverage * no-mistakes(review): Isolate Kimi task temp runtime coverage * no-mistakes(document): Refresh test cleanup documentation * no-mistakes: apply CI fixes
…henguid#1349) * fix(dispatch): scope candidate authentication to its own surface A locally expired timestamp in one credential store was reported to the captain as a sign-out, including for dispatch candidates that never read that store. A `harness=pi, model=xai/grok-*` candidate authenticates through Pi's own xAI credential, but the only Grok quota reading available was gated on the standalone Grok CLI's separate token, whose expiry clock drifts independently. The always-loaded intake rule then turned that unreadable quota into a mandatory captain escalation. Add `bin/fm-auth-preflight.sh` as the deterministic owner of the parts that must not depend on agent memory: it resolves a tuple's authentication surface from quota-axi's own emitted auth sources rather than from a harness or model name, so another harness's CLI can never gate a candidate that does not use it. A vendor CLI is launched only when the tuple's own harness owns the credential store under test and a non-destructive discovery command is registered for it, which today is `grok models` alone. That probe runs at most once with stdin closed and a hard timeout, reads its verdict from the first stdout line because the command exits 0 either way, treats unrecognized output as indeterminate, and never invokes login, logout, or the interactive TUI. Quota is read at most twice, and unknown headroom never makes a candidate ineligible on its own. Update the dispatch procedure to match: usable authentication with unmeasurable headroom stays eligible at lower preference with the unknown disclosed, and stop-and-report is reserved for unresolved authentication, an unresolved relationship, or malformed configuration. Record that Grok's `credits.remaining` is a prepaid balance rather than window headroom. Gate quota-axi at 0.1.16 in bootstrap, the first build reporting per-credential auth sources. A stale install previously passed the presence check silently, which is why a fix published two days earlier was still not in effect. Replace the orphaned quota-array-dispatch fixtures, which encoded a `provider: "xai"` shape the tool never emits and had no consumer, with fixtures shaped like real 0.1.16 output that the new suite drives the script against. The suite asserts the verdict and, separately, which vendor CLIs were launched, so a Pi/xAI candidate reaching the Grok CLI fails. Map `tests/fixtures/<dir>` to its consuming suite so a fixture change selects the right tests instead of refusing. * refactor(bootstrap): give the quota-axi floor one owner The floor was stated twice - once in bootstrap's gate and once inline in the auth preflight - so bumping it needed two edits that could drift. Move it to bin/fm-quota-axi-lib.sh alongside its rationale, matching the existing tasks-axi library, and derive the comparison from the constant so the number appears exactly once. Bootstrap turns a failing check into the operator diagnostic; the preflight refuses to emit an unscoped verdict. Map the new library to both consuming suites so a bump re-runs them, and record that any usable source means the surface authenticates. * no-mistakes(review): Captain: bound quota checks and removed Python dependency * no-mistakes(review): Captain: enforce conservative headroom and exact preflight retry * no-mistakes(review): Captain: preserve OpenCode eligibility without auth-surface guessing * no-mistakes(review): Captain: reject malformed OpenCode model relationships * no-mistakes(review): Captain: exempt verified unmodeled tuples from intake escalation * no-mistakes(document): Updated dispatch authentication documentation * no-mistakes: apply CI fixes
…chenguid#1327) * feat: add semantic busy-state contract owner and event writer One owner (bin/fm-busy-lib.sh) for the captain-approved semantic busy-state redesign: a per-task gen-bound record written only by bin/fm-busy-event.sh, per-harness trusted-source classification with explicit source attribution, busy/idle/unknown/dead semantics where missing, malformed, stale, or untrusted semantic data is unknown - never idle - and endpoint death is the only process-level override. The Grok-only rendered-tail fallback and the standalone-Kimi verification gate live behind the same classifier. * feat: arm busy-state at spawn and convert Pi to the semantic extension path fm-spawn arms the busy-state contract for converted adapters and seeds busy/fm-spawn (the launch brief is a submitted turn). The Pi/pi-signed per-task extension now reports agent_start -> busy and agent_settled -> idle confirmed by ctx.isIdle(), covering auto-retries, compaction retries, tool loops, and queued continuations, while turn_end stays a wake notification touch. Teardown removes the new record, gen sidecar, and lock. Live-verified on Pi 0.82.0: seed -> agent-start busy -> agent-settled idle with the marker still touched. * feat: convert OpenCode to the semantic session.status plugin path The per-task plugin (renamed .opencode/plugins/fm-busy-state.js) now classifies from OpenCode's semantic session.status events - busy and retry are active, idle is inactive - latched to the worker's own session so a subagent child session can never clear the worker's busy state. The session.idle marker touch stays a wake notification. Teardown removes both the new and the legacy plugin filenames. Live-verified on OpenCode 1.17.18 in a real TUI pane: seed -> session-busy -> session-status-idle. * feat: convert Claude to the full lifecycle hooks path The per-task settings.local.json now wires UserPromptSubmit -> busy and Stop, StopFailure, and SessionEnd -> idle, so API-error and shutdown turn ends can never strand a busy record; Stop keeps the turn-ended notification touch. A refused (stale-gen) event exits 0 and stays silent so Claude's own lifecycle is never broken. Live-verified on Claude Code 2.1.220: UserPromptSubmit fires for the argv launch prompt, Stop closes each turn, a mid-stream Escape interrupt fires no closing hook, and the firstmate-controlled idle/fm-interrupt clear resolves it. * feat: gate Codex busy state behind verified semantic sources The approved contract prefers Codex's app-server turn lifecycle with capability negotiation and sanctions its lifecycle hooks as the intermediate. Live probes on codex-cli 0.145.0 show neither is usable for a pane worker: the app-server daemon is unreachable for a TUI thread and refuses to start outside the managed standalone install, and firstmate-written project hooks never fired (interactive with directory trust granted, and exec, both with --dangerously-bypass-hook-trust) while global hooks fired in the same runs. Codex therefore classifies unknown codex-unverified behind an explicit probe rather than falling back to idle or footer text, and fm-spawn installs no unverified Codex wiring. * feat: gate standalone Kimi busy state on live verification Standalone Kimi has no installed binary here, so per the approved contract its semantic path stays guarded and it classifies unknown kimi-unverified rather than idle - and never from its locale-sensitive moon-phase spinner, which the redesign forbids inventing as a state source. The gate records the preferred source order (Wire prompt request lifetime, which brackets a turn and reports cancellation, then the documented hooks including Interrupt because Stop does not fire on interrupts) and the exact evidence required to open it. Arming without wiring would seed a busy record nothing could clear, so both land together behind the same gate. * feat: route busy consumers through the contract and drop the global OR The watcher, crew-state reader, and away-mode daemon now decide busy state through bin/fm-busy-lib.sh: only an exact busy verdict counts as working, and unknown never becomes working or a silent idle, so a crew whose semantic state is missing, malformed, stale, or unverified surfaces instead of being absorbed. Crew-state reports the producing source in its detail. The watcher's global OR regex default is gone; Grok keeps its isolated fallback inside the contract. The daemon's supervisor-pane reader stays rendered-text - that pane is not a recorded task - but is now scoped to firstmate's own detected harness instead of every vendor signature. Secondmate pending-reply observation is deliberately unchanged and documented as a delivery-confirmation signal, not task state. * docs: point busy-state documentation at the single contract owner Adds a maintainer-architecture section naming bin/fm-busy-lib.sh as the owner of what busy means, with per-adapter sources, the unknown-never-idle rule, the endpoint-death override, and the two rendered-text readers that deliberately stay outside the contract. Replaces the stale regex-first prose in architecture, tmux-backend, herdr-backend, and configuration; converts the harness-adapters per-harness rows from UI signatures to the semantic source each harness uses; and records the live verification evidence, including why Codex and standalone Kimi stay unknown. * fix: arm away-launch signal handlers before acquiring the lifecycle lock fm_afk_launch_main acquired its lock and only then installed the EXIT, INT, and TERM traps. A signal arriving in that window terminated the process by default action and left the lock directory behind, which blocks the next away-mode launch until the stale-owner reclaim path clears it. The release helper only removes a lock this process owns, so the handlers are now armed first. The accompanying test also killed the child whether or not the lock had appeared and sampled cleanup the instant wait returned; it now requires the lock, then allows a bounded settle, so it proves the guarantee instead of racing it. * test: align fleet, Kimi, lifecycle, and detection suites with the contract The fleet snapshot and wake-daemon lifecycle fixtures now prove a working crew through its own semantic busy-state record instead of rendered pane text, which is what those consumers read. The Kimi watcher test asserts the approved contract directly: a standalone Kimi task classifies unknown rather than matching its moon-phase spinner, while Grok's isolated fallback still classifies only Grok. The pi-signed detection cases clear ambient harness markers, fixing a pre-existing failure where the running session's own CLAUDECODE outranked the fixture's marker. * fix: stop teardown from deleting a project's own Codex hooks file An intermediate revision wired Codex through a firstmate-written <worktree>/.codex/hooks.json, and teardown removed it alongside the other generated wiring. The Codex wiring was dropped when its probes came back unverified, so that removal now targets a file firstmate never creates - and a project may legitimately track its own .codex/hooks.json, which teardown would then delete from a pooled worktree. * fix: keep busy-record parsing from disturbing its sourcing caller The record parser split fields with set -- under a temporary noglob, which clobbers a sourcing caller's positional parameters and restores glob expansion even when the caller had disabled it. The watcher, the daemon, and the crew-state reader all source this library, so it now reads fields with read -a, which never globs and never touches caller state. * docs: state exactly which Claude hook paths were reproduced live The busy-state record listed all four wired Claude hooks in the source column, which could read as a claim that every one fired during the pass. UserPromptSubmit and Stop did; StopFailure and SessionEnd are wired from hook names confirmed present in the installed binary, but the abnormal turn ends they cover were not reproduced. * test: let reset_fakes own the crew-state busy-text fixture lifecycle The Grok fallback case set FM_FAKE_BUSY_TEXT and cleared it inline, so the variable's lifetime was owned by one test rather than by the shared reset that every other fake already uses. * no-mistakes(review): Fix semantic busy-state lifecycle races * no-mistakes(review): Make busy-state retirement idempotent * no-mistakes(review): Enforce semantic state boundaries for status and injection * no-mistakes(review): Restore harness-scoped away-mode busy guard * no-mistakes(document): Refresh semantic busy-state documentation * no-mistakes: apply CI fixes
* fix(dispatch): judge candidate provider relations instead of rejecting them Firstmate deterministically dropped supported Pi candidates in the openai-codex family. bin/fm-auth-preflight.sh resolved a harness=pi tuple's credential surface by constructing the source id `pi:<model-prefix>`, so `pi + openai-codex/gpt-5.6-terra` looked for a `pi:openai-codex` source. That source does not exist, because Pi's Codex family authenticates through the Codex store quota-axi already lists as `auth-json`/`cli-rpc`. The tuple returned `eligible=no reason=surface-unresolved` while the Pi catalog listed the model and the Codex provider reported fresh, usable credentials with 64 effective percent remaining on its all-model scope. The prefix construction was only ever valid where Pi holds its own credential (`pi:xai`, `pi:kimi-coding`), which is why every previously configured Pi tuple resolved and the defect stayed hidden until a Codex-family Pi model was configured. Retire dispatch eligibility from deterministic shell. The dispatching first mate now establishes model support and provider family from each harness's authoritative catalog, applies quota at the granularity the vendor supplies, and shows that reasoning. Provider-level and all-model evidence bounds every model established in that family; a named-model window bounds only its own model. Missing model-level quota, a missing auth source, unmeasurable headroom, and unmodeled authentication are disclosed uncertainty. Only concrete contradictory evidence blocks a candidate. Replace the preflight with bin/fm-vendor-auth-probe.sh, which keeps the captain's approved bounded probe envelope without any routing knowledge: it takes no harness, model, or provider, reads no quota, renders no verdict, and holds only a fixed-argv safety allowlist. Its behavior suite proves the absent identity surface, the untouched quota, the uniform exit status, the fixed argv with stdin closed, and a real bound even when the configured bound is zero. Also fixed along the way: a zero FM_*_TIMEOUT silently removed the hard bound, the pinned Grok version had drifted to 0.2.117, and --changed selection refused outright on any deleted bin/ script. AGENTS.md section 4 and quota-array-dispatch own the corrected policy, harness-adapters gets the catalog-responsibility correction, and docs/verification/dispatch-auth.md records the 2026-07-30 evidence on Pi 0.82.0, quota-axi 0.1.16, and grok 0.2.117. * no-mistakes(review): Reject all-zero vendor probe timeouts
…claude pid (#2) * fix(session-lock): resolve Claude bg-spare ancestry to the outermost claude pid fm_harness_ancestry_pid() previously returned the first ancestor process whose command matched a verified harness name. Claude Code's Stop hook fires as a bg-spare worker several levels below the session's actual lock-owning claude process (hook shell -> claude bg-spare -> claude bg-pty-host -> claude -> claude(lock)), so the first match was the bg-spare worker, not the lock owner. fm_session_lock_owned_by_self() then never matched state/.lock, and the Claude Stop auto-arm silently treated its own primary session as an unrelated live owner and never armed the watcher. The walk now keeps going past a claude-named match, looking for a still more ancestral claude-named match, and stops the instant a non-match follows an already-found match (bounding it to a contiguous run rather than the literal ancestry top, so an unrelated claude-named process further up the real process tree is never mistaken for part of this session's own nested chain). Every other harness keeps the original first-match-wins behavior, since e.g. Pi's shared signed-wrapper ancestry actually holds the session at the inner engine pid, not an outer wrapper pid. Hop limit raised from 8 to 16 to cover the deeper bg-spare chain. * no-mistakes(review): Add nested-claude-ancestry regression test; fix nudge doc depth claim
* Wire fm-spawn.sh and fm-teardown.sh to Parlay chat-panel enrollment fm-spawn.sh best-effort enrolls a confirmed launch via `parlay listen --agent <id>`, backgrounded with its pid recorded to state/<id>.parlay-listen-pid. fm-teardown.sh best-effort deregisters on clean teardown: the recorded pid is always killed, and `parlay agent-down <id>` is called only when parlay is on PATH. Neither side ever blocks or fails spawn/teardown when parlay is absent or fails. Adds tests/fm-spawn-parlay.test.sh and two new tests in tests/fm-teardown.test.sh, plus a new fm_path_without test helper in tests/lib.sh for simulating parlay's genuine absence from PATH. * no-mistakes(lint): tests/lib.sh: rename fm_path_without's out array to avoid shellcheck SC2178/SC2128
…ment - Add 'or data/backlog.md' fallback to beads backend pointer for consistency - Quote basename argument in fm-session-lock-lib.sh to handle dash-leading process names (fixes 'basename: missing operand' when ps output starts with '-') - Fixes tests/fm-session-start.test.sh:1148 and tests/fm-secondmate-harness.test.sh
📝 WalkthroughWalkthroughThe change refines quota credential-surface authentication guidance and adds an end-to-end test for the Bearings skill’s four-section chat contract. ChangesCredential Surface Guidance
Bearings Chat Contract
Estimated code review effort: 2 (Simple) | ~10 minutes 🚥 Pre-merge checks | ✅ 5✅ Passed checks (5 passed)
✨ Finishing Touches📝 Generate docstrings
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
There was a problem hiding this comment.
Actionable comments posted: 2
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In `@tests/fm-bearings-snapshot.test.sh`:
- Around line 1904-1907: Update the four empty-state assertions in the snapshot
test to validate the complete exact line, including the final period, rather
than matching partial substrings. Use the existing body-checking utilities or
exact-line assertion pattern for the entries associated with “Captain's Call,”
“Recently Landed,” “Underway,” and “Charted Next,” while preserving their
expected sentences.
- Around line 1908-1909: Update the report_headings extraction in the
detailed-report contract test to capture every indented bold heading, not just
the four recognized names, then compare that complete heading list with expected
so unrecognized sections such as “At Anchor” cause the assertion to fail.
🪄 Autofix (Beta)
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: defaults
Review profile: CHILL
Plan: Pro Plus
Run ID: a05f2bde-fdc6-4eef-ad0c-4e3a0b94640e
📒 Files selected for processing (2)
.agents/skills/quota-array-dispatch/SKILL.mdtests/fm-bearings-snapshot.test.sh
| assert_contains "$body" "Nothing needs your action right now" "Captain's Call empty-state sentence" | ||
| assert_contains "$body" "No recent completions are in the current baseline" "Recently Landed empty-state sentence" | ||
| assert_contains "$body" "Nothing is underway" "Underway empty-state sentence" | ||
| assert_contains "$body" "Nothing is queued" "Charted Next empty-state sentence" |
There was a problem hiding this comment.
🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win
Match the complete empty-state sentences.
The contract defines exact sentences, including the final period. These assert_contains calls check only substrings. A changed sentence can retain the substring and still pass. Use an exact-line check for each Empty-state: "..." entry.
Proposed fix
- assert_contains "$body" "Nothing needs your action right now" "Captain's Call empty-state sentence"
+ printf '%s\n' "$body" | grep -Eq '^[[:space:]]*Empty-state: "Nothing needs your action right now\."$' \
+ || fail "Captain's Call empty-state sentence"
- assert_contains "$body" "No recent completions are in the current baseline" "Recently Landed empty-state sentence"
+ printf '%s\n' "$body" | grep -Eq '^[[:space:]]*Empty-state: "No recent completions are in the current baseline\."$' \
+ || fail "Recently Landed empty-state sentence"
- assert_contains "$body" "Nothing is underway" "Underway empty-state sentence"
+ printf '%s\n' "$body" | grep -Eq '^[[:space:]]*Empty-state: "Nothing is underway\."$' \
+ || fail "Underway empty-state sentence"
- assert_contains "$body" "Nothing is queued" "Charted Next empty-state sentence"
+ printf '%s\n' "$body" | grep -Eq '^[[:space:]]*Empty-state: "Nothing is queued\."$' \
+ || fail "Charted Next empty-state sentence"📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| assert_contains "$body" "Nothing needs your action right now" "Captain's Call empty-state sentence" | |
| assert_contains "$body" "No recent completions are in the current baseline" "Recently Landed empty-state sentence" | |
| assert_contains "$body" "Nothing is underway" "Underway empty-state sentence" | |
| assert_contains "$body" "Nothing is queued" "Charted Next empty-state sentence" | |
| printf '%s\n' "$body" | grep -Eq '^[[:space:]]*Empty-state: "Nothing needs your action right now\."$' \ | |
| || fail "Captain's Call empty-state sentence" | |
| printf '%s\n' "$body" | grep -Eq '^[[:space:]]*Empty-state: "No recent completions are in the current baseline\."$' \ | |
| || fail "Recently Landed empty-state sentence" | |
| printf '%s\n' "$body" | grep -Eq '^[[:space:]]*Empty-state: "Nothing is underway\."$' \ | |
| || fail "Underway empty-state sentence" | |
| printf '%s\n' "$body" | grep -Eq '^[[:space:]]*Empty-state: "Nothing is queued\."$' \ | |
| || fail "Charted Next empty-state sentence" |
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@tests/fm-bearings-snapshot.test.sh` around lines 1904 - 1907, Update the four
empty-state assertions in the snapshot test to validate the complete exact line,
including the final period, rather than matching partial substrings. Use the
existing body-checking utilities or exact-line assertion pattern for the entries
associated with “Captain's Call,” “Recently Landed,” “Underway,” and “Charted
Next,” while preserving their expected sentences.
| report_headings=$(sed -nE 's/^ - \*\*(Captain.s Call|Recently Landed|Underway|Charted Next)\*\*.*/\1/p' "$skill") | ||
| [ "$report_headings" = "$expected" ] || fail "detailed report contract must contain the same four complete sections, got: $report_headings" |
There was a problem hiding this comment.
🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win
Reject unrecognized report headings.
The sed expression selects only the four expected names. It drops every other indented bold heading before the equality check. If .agents/skills/bearings/SKILL.md contains At Anchor, report_headings can still equal expected, even though the contract forbids that section. Capture every heading in the detailed-report block before comparing the complete list.
Proposed fix
- report_headings=$(sed -nE 's/^ - \*\*(Captain.s Call|Recently Landed|Underway|Charted Next)\*\*.*/\1/p' "$skill")
+ report_headings=$(sed -nE 's/^ - \*\*([^*]+)\*\*.*/\1/p' "$skill")📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| report_headings=$(sed -nE 's/^ - \*\*(Captain.s Call|Recently Landed|Underway|Charted Next)\*\*.*/\1/p' "$skill") | |
| [ "$report_headings" = "$expected" ] || fail "detailed report contract must contain the same four complete sections, got: $report_headings" | |
| report_headings=$(sed -nE 's/^ - \*\*([^*]+)\*\*.*/\1/p' "$skill") | |
| [ "$report_headings" = "$expected" ] || fail "detailed report contract must contain the same four complete sections, got: $report_headings" |
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@tests/fm-bearings-snapshot.test.sh` around lines 1908 - 1909, Update the
report_headings extraction in the detailed-report contract test to capture every
indented bold heading, not just the four recognized names, then compare that
complete heading list with expected so unrecognized sections such as “At Anchor”
cause the assertion to fail.
…tness (#65) * fix: hash_pane() md5 fallback chain for restricted PATH on macOS * feat: herdr->firstmate spur bridge prototype fm-herdr-spur.sh: detached daemon that watches herdr agent status via the native pane.agent_status_changed event stream (reusing herdr-eventwait.py) and, on a working->idle/done edge for a tracked EXTERNAL agent, enqueues a check wake into state/.wake-queue so fm-wake-drain surfaces it. Fills the gap where parlay-spawned herdr agents have no firstmate status file or turn-end hook. Poll fallback via herdr agent list when events incapable. Debounced, keyed by agent name, configurable via --agent / config/herdr-spur.agents / all-agents. Read-only against herdr; bash 3.2 safe. * feat: fm-groom proactive ideas-store work generator Grooms ready ideas into dispatched work without the captain writing a brief. Per idea: formulate a concrete runnable brief via PAI Inference (plain prompts to dodge PromptGuard), classify safe (research/design/prototype) vs escalate (merge/deploy/production), then dispatch safe via fm-spawn scout or file unsafe to the review store. Safety rails: OFF by default (dry-run unless FM_GROOM_ENABLED=1), rate-limited (FM_GROOM_MAX_IN_FLIGHT), bounded per run (FM_GROOM_MAX_PER_RUN), idempotent (groom:* state label), fail-safe classify (any error escalates). 7 hermetic tests cover the rails; shellcheck-clean. * feat(review-page): render review-store items to visitable Pulse pages bin/fm-review-page.ts: for each open review item, write a self-contained, phone-readable HTML page under ~/pulse-pages/review/<id>/ served by Pulse, plus an index at the review root. --all sweeps every open item; positional ids render specific ones. Own dependency-free markdown renderer (headings, fenced code, lists, blockquote, links, bold/italic) in the GitHub-dark house style matching existing pulse-pages. Wires each page URL back onto the item (page_url metadata + a Page: note), idempotently. Read-only on item content except that safe append; parameterized by env for scheduling. * test(review-page): hermetic behavior tests for fm-review-page tests/fm-review-page.test.sh: network-free suite stubbing the review CLI on a fakebin PATH (read verbs emit canned JSON, write verbs log calls) and writing pages to a temp FM_REVIEW_PAGE_OUT. Covers single-id render, --all sweep over N items, idempotent re-render (no dup dirs, note skipped on unchanged url), artifact link rendering (url/brain/branch), markdown body rendering, wire-back assertions, empty-queue index, unknown id, and the no-args usage error. shellcheck-clean under the canonical whole-set invocation. * fix(review-page): eliminate code-span placeholder leak in markdown renderer The inline renderer used a text placeholder to protect `code` spans across the bold/italic/link passes; that placeholder corrupted into NUL bytes and leaked bare CODE0/CODE1 tokens into rendered pages (visible on the live fm-groom item). Replace it with a split-based renderer: split on the code-span capture group so even indices are prose (escaped + emphasized) and odd indices are raw code (escaped, wrapped in <code>). No placeholder token can survive to output. Adds a regression test with a code-span-heavy body (adjacent spans, em-dashes, a multi-span list line, parenthesized spans) asserting every span renders and no CODE placeholder leaks. * Add fm-idea-mine.ts: mine chat history for uncaptured ideas, re-evaluate, file harvest * fm-idea-mine: pass generous --timeout to default Inference.ts cmd for large mined tails * Add hermetic tests for fm-idea-mine: generate/evaluate/file, dedup, dry-run, empty, malformed, idempotency, cap * fm-review-decision.sh: route a captain's review decision to firstmate The firstmate-side half of the interactive review loop. Given <id> <verdict> [comment], it durably (a) enqueues a check-kind wake into state/.wake-queue via the sanctioned fm_wake_append helper, keyed review-decision:<id>, so fm-wake-drain surfaces it on the next supervision cycle; (b) annotates the item via 'review note'; (c) appends a JSONL audit record. Fails LOUDLY (non-zero) if the wake or annotation cannot land — never a silent ok (robots-5l8). 8 hermetic tests cover every verdict, the fail-loud path, and --stdin comments. * fm-review-page: interactive decision panel + single top bar Turn read-only review pages into decision surfaces. Every page now carries an Approve/Decline/Comment panel that POSTs same-origin to /api/review/decision with in-page success/error feedback, plus a structured What/Why/Stakes/Recommendation/Artifact breakdown parsed from the body so the decision is answerable in place. An already-recorded 'Captain decision:' note renders as a standing-decision banner. Phone-friendly (48px tap targets, no horizontal scroll, self-contained inline JS/CSS). Also fixes the double-nav-bar bug: portal's injectShell stacked /_pulse/nav.js on top of the page's own .topbar. The page now emits <meta name="pulse-shell" content="off"> to opt out — exactly one top bar, matching how /status and /plans compose. 5 new tests (13 total). * wip: local firstmate customizations snapshot before upstream merge Snapshot of uncommitted local edits (AGENTS.md, .claude/settings.json, .codex/hooks.json) so the working tree is clean for merging origin/main. Reversible; preserves local work per the never-discard rule. * Add fm-isolated-launch.sh for HOME-isolated claude sessions * Document fm-isolated-launch.sh in README * fm-isolated-launch: mirror cwd outside $HOME tree, seed keychain credential HOME override alone doesn't stop Claude Code's ancestor-directory CLAUDE.md walk from re-loading ~/.claude/CLAUDE.md, since firstmate's repo is nested under the real home dir. Launch from a detached worktree mirror under /private/tmp instead (refreshed to HEAD each run), with FM_ROOT_OVERRIDE so bin/ scripts still resolve real state/data/config/projects. Also implements the previously-comment-only keychain credential seeding for first-run auth. Root-caused via a background agent's /context-verified test; confirmed independently by checking the mirror's CLAUDE.md symlink and ancestor chain. * fm-isolated-launch: force bypass-permissions mode, matching crewmate autonomy The isolated session was still stopping for a tool-approval dialog on every command (e.g. bin/fm-session-start.sh) because its fresh $HOME had no bypass-permissions state, unlike ordinary crewmates which fm-spawn.sh already launches with --dangerously-skip-permissions. Fix: seed settings.json (permissions.defaultMode=bypassPermissions + skipDangerousModePermissionPrompt, re-applied every launch) and .claude.json's bypassPermissionsModeAccepted, plus add --dangerously-skip-permissions to the exec line for parity with fm-spawn.sh:319. Verified live: pane now shows 'bypass permissions on' at startup and runs a command with zero approval prompt. * fm-isolated-launch: symlink real data/ for federated bd store access The isolated session overrides HOME, but ~18 federated store wrapper scripts (brain, robots, task, decisions, ...) hardcode BEADS_DIR=$HOME/data/<store>/.beads at runtime, keyed off the actual process HOME rather than a baked-in path. Under isolation that resolved to a nonexistent path instead of the real Dolt-backed stores. Symlink $ISOLATED_HOME/data -> the real ~/data so those wrappers reach the real federated stores. Exposes only store data, not any PAI CLAUDE.md/hooks/skills/agent config - none of that lives under data/. Verified live: HOME inside the isolated session reported as the isolated home, yet 'brain list' and 'robots list' returned real open issues from the actual stores. * feat(bin): beads dispatch lifecycle via spawn/brief hook extension points Add bin/fm-bead-stamp.sh (fail-open: stamps dispatch=sent + assigns a linked bead on spawn) and a --beads <id> flag on fm-brief.sh and fm-spawn.sh. All bead-specific logic lives in two new hook directories rather than being spliced directly into fm-brief.sh/fm-spawn.sh, so those files stay a pure addition target with minimal upstream conflict surface: - bin/fm-brief-hooks.d/beads.sh: sourced before fm-brief.sh writes the Brief section; emits the Bead Receipt (dispatch=claimed on brief read) and Bead Closure (close the bead before done:) sections. - bin/fm-spawn-hooks.d/beads.sh: sourced after a successful spawn; stamps the bead via fm-bead-stamp.sh and registers a watcher check that polls the bead for status=closed and wakes firstmate for teardown. Both hook loops run each hook in its own subshell so a hook's `exit` never terminates the calling script, keeping every hook fail-open by construction. fm-spawn.sh records beads_id= in state/<id>.meta when set. --beads is rejected for --secondmate on both scripts. * no-mistakes(review): Fix bead-close detection false positives, align hook dir resolution * no-mistakes(document): Sync docs/AGENTS.md with beads dispatch hook feature * Add per-account Claude Code launcher and fm-spawn --account flag - bin/claude-account.sh: standalone launcher for per-account Claude Code isolation (CLAUDE_CONFIG_DIR, flock-serialized shared-config symlinks, onboarding/trust-dialog pre-write, settings.json flag pre-write). - bin/claude-1.sh, bin/claude-2.sh: one-line direct launchers. - fm-spawn.sh --account <N>: records account=N in meta, sets CLAUDE_TRUST_DIR to the task worktree, launches through claude-account.sh N. Optional; absent behavior is unchanged. - docs/configuration.md: Multi-account Claude Code section. - tests/claude-account.test.sh, tests/fm-spawn-account.test.sh. * feat(skills): add herdr-navigation internal skill; gitignore CHANGELOG.md Adds .agents/skills/herdr-navigation/SKILL.md so agents know how to navigate herdr panes using existing primitives (herdr pane current, neighbor, split, send-text, list). Agents commonly don't know these exist; the skill surfaces them with working examples. Also adds CHANGELOG.md to .gitignore — it is a generated session activity log, not source content. * fix(session-lock): resolve Claude bg-spare ancestry to the outermost claude pid (#2) * fix(session-lock): resolve Claude bg-spare ancestry to the outermost claude pid fm_harness_ancestry_pid() previously returned the first ancestor process whose command matched a verified harness name. Claude Code's Stop hook fires as a bg-spare worker several levels below the session's actual lock-owning claude process (hook shell -> claude bg-spare -> claude bg-pty-host -> claude -> claude(lock)), so the first match was the bg-spare worker, not the lock owner. fm_session_lock_owned_by_self() then never matched state/.lock, and the Claude Stop auto-arm silently treated its own primary session as an unrelated live owner and never armed the watcher. The walk now keeps going past a claude-named match, looking for a still more ancestral claude-named match, and stops the instant a non-match follows an already-found match (bounding it to a contiguous run rather than the literal ancestry top, so an unrelated claude-named process further up the real process tree is never mistaken for part of this session's own nested chain). Every other harness keeps the original first-match-wins behavior, since e.g. Pi's shared signed-wrapper ancestry actually holds the session at the inner engine pid, not an outer wrapper pid. Hop limit raised from 8 to 16 to cover the deeper bg-spare chain. * no-mistakes(review): Add nested-claude-ancestry regression test; fix nudge doc depth claim * Wire fm-spawn.sh and fm-teardown.sh to Parlay chat-panel enrollment (#4) * Wire fm-spawn.sh and fm-teardown.sh to Parlay chat-panel enrollment fm-spawn.sh best-effort enrolls a confirmed launch via `parlay listen --agent <id>`, backgrounded with its pid recorded to state/<id>.parlay-listen-pid. fm-teardown.sh best-effort deregisters on clean teardown: the recorded pid is always killed, and `parlay agent-down <id>` is called only when parlay is on PATH. Neither side ever blocks or fails spawn/teardown when parlay is absent or fails. Adds tests/fm-spawn-parlay.test.sh and two new tests in tests/fm-teardown.test.sh, plus a new fm_path_without test helper in tests/lib.sh for simulating parlay's genuine absence from PATH. * no-mistakes(lint): tests/lib.sh: rename fm_path_without's out array to avoid shellcheck SC2178/SC2128 * fix(watch): fall back to /sbin/md5 and shasum when md5/md5sum aren't on PATH (#3) hash_pane() only checked 'command -v md5', which fails in any launch context whose PATH omits /sbin (observed here: a background task shell with PATH=~/.local/bin:~/.bun/bin:/opt/homebrew/bin:/usr/local/bin:/usr/bin:/bin, no /sbin). It then fell through to md5sum, which macOS does not ship at all, producing a 'md5sum: command not found' error on every pane hash and likely starving the pane-change detection that watcher cycles rely on to report an actionable reason. * feat(backlog): add beads as third backend option alongside tasks-axi and manual (#7) * fix(bin): allow session-local todo tools in the subagent guard (#1204) * fix(guard): allow session-local todo tools in the primary The delegation-shape guard denied TaskCreate and TaskUpdate because their normalized names contain the `task` stem. Those tools write only the harness's session-local todo list, which has no executor: it spawns no agent, allocates no worktree, registers no schedule, and starts nothing that outlives the session. That is not the unaccounted work the guard exists to stop, so the stem match was a false positive, and the deny text told the primary to run bin/fm-brief.sh and bin/fm-spawn.sh to create a todo entry. Add a separately-reasoned PLAN_ONLY_TOOLS exact-name exclusion rather than widening OBSERVE_ONLY_TOOLS, whose documented contract is tools that only observe or stop existing work. Both lists stay exact-name so neither can widen by substring. Tests cover the two allowed names and six near-miss names that a substring or shortened-stem widening would release; both mutations were watched red. * no-mistakes(review): drop session-local todo tools from recommended deny list * no-mistakes: apply CI fixes * fix(session-lock): resolve Claude bg-spare ancestry to the outermost claude pid (#1206) * fix(session-lock): resolve Claude bg-spare ancestry to the outermost claude pid fm_harness_ancestry_pid() previously returned the first ancestor process whose command matched a verified harness name. Claude Code's Stop hook fires as a bg-spare worker several levels below the session's actual lock-owning claude process (hook shell -> claude bg-spare -> claude bg-pty-host -> claude -> claude(lock)), so the first match was the bg-spare worker, not the lock owner. fm_session_lock_owned_by_self() then never matched state/.lock, and the Claude Stop auto-arm silently treated its own primary session as an unrelated live owner and never armed the watcher. The walk now keeps going past a claude-named match, looking for a still more ancestral claude-named match, and stops the instant a non-match follows an already-found match (bounding it to a contiguous run rather than the literal ancestry top, so an unrelated claude-named process further up the real process tree is never mistaken for part of this session's own nested chain). Every other harness keeps the original first-match-wins behavior, since e.g. Pi's shared signed-wrapper ancestry actually holds the session at the inner engine pid, not an outer wrapper pid. Hop limit raised from 8 to 16 to cover the deeper bg-spare chain. * no-mistakes(review): Add nested-claude-ancestry regression test; fix nudge doc depth claim * no-mistakes: apply CI fixes * fix: conferma l'avvio del watcher su Windows/MSYS (#1212) * fix: confirm watcher startup on MSYS * no-mistakes(review): gate MSYS arm ready timeout, cache uname, harden locale test * no-mistakes(review): validate OpenCode ready timeout, make uname cache internal * fix(spawn): forward CLAUDE_CONFIG_DIR to claude crewmates (#1195) * fix(spawn): forward firstmate's CLAUDE_CONFIG_DIR to claude crewmates Crewmate panes are created by a long-lived tmux/herdr daemon that does not inherit firstmate's current environment. When firstmate runs under a non-default CLAUDE_CONFIG_DIR (for example a work-vs-personal subscription split), a bare `claude` in the crewmate pane fell back to the default ~/.claude store and launched unauthenticated, blocking the crewmate before it could do any work. fm-spawn now prefixes the claude launch with firstmate's own resolved CLAUDE_CONFIG_DIR when set, so the crewmate uses the same credential/config store firstmate is authenticated with. An unset value is the single-store default and adds no prefix; non-claude harnesses are unaffected. Adds three tests in fm-spawn-dispatch-profile.test.sh (forwarded-when-set, omitted-when-unset, non-claude-ignored) and pins CLAUDE_CONFIG_DIR in the test helper so launch assertions no longer depend on the developer's environment. * no-mistakes: apply CI fixes * fix: preserve dispatch identity across authentication checks (#1233) * fix: preserve dispatch harness identity * no-mistakes(review): Fix Grok counterfactual tuple validation * no-mistakes(document): Scope dispatch authentication to selected tuple * fix: restore dispatch instruction budget * no-mistakes(review): Scope dispatch authentication after candidate selection * fix(bin): normalize relative durable paths (#1256) * fix(bin): handle dash-leading harness process names (#2) * fix: handle dash-leading harness process names * no-mistakes(review): Make dash-leading harness regression hermetic * fix: preserve secondmate reply routes across relative homes Resolve relative home, data, and state inputs before durable charter generation, and fail when caller-relative directories cannot be resolved. Use absolute paths at the related spawn, AFK daemon, and X-mode cross-process handoffs so later processes cannot reinterpret them from another working directory. * no-mistakes(review): Preserve absolute overrides and normalize relative durable paths * no-mistakes(review): Normalize relative home before deriving durable paths * no-mistakes(document): Document relative durable-path normalization * no-mistakes(review): Captain: Ignore inherited CDPATH during relative path normalization * no-mistakes(lint): Fix empty CDPATH assignments for ShellCheck * refactor(skills): make Bearings chat-only by default (#1136) * Add internal status skill * no-mistakes(document): register /status skill in documentation-audiences inventory * no-mistakes(lint): replace grep|wc -l with grep -c in status skill test * test: silence literal status skill patterns * Refactor bearings default to chat-only --------- Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com> * Clarify follow-up routing during validation (#1277) * fix: honor concrete approval for project operations (#1272) * docs: add captain-approved project operation exception to hard rule 1 Firstmate stays read-only over projects by default, but when the captain clearly approves a concrete project operation and scope in the moment, firstmate may perform exactly that approved operation with its own tools. The approval is never inferred, broadened, or standing, and it does not relax the existing force, discard, unlanded-work, or merge-authority boundaries. * no-mistakes(review): Clarify captain-approved project operation boundaries * no-mistakes(document): Clarify captain-approved project operation scope * docs: cover directories and preserve the operation-or-scope alternative Widen the captain-approved project operation exception in AGENTS.md to files or directories, and restore the explicit operation-or-scope alternative that a prior pipeline auto-fix had collapsed into "and". Rework project-management SKILL.md's Remove section, which previously told firstmate to refuse project removal until a guarded helper existed; that helper was never built, so the text directly contradicted the new instruction-only exception. It now points at the exception plus the existing removal preflight it still requires unchanged. Update the one instruction-owners test assertion that hard-coded the sentence removed above, so the suite tracks current, not obsolete, text. * docs: add captain-approved project operation exception to hard rule 1 Firstmate stays read-only over projects by default, but when the captain clearly approves a concrete project operation and scope in the moment, firstmate may perform exactly that approved operation with its own tools. The approval is never inferred, broadened, or standing, and it does not relax the existing force, discard, unlanded-work, or merge-authority boundaries. * no-mistakes(review): Clarify captain-approved project operation boundaries * no-mistakes(document): Clarify captain-approved project operation scope * docs: cover directories and preserve the operation-or-scope alternative Widen the captain-approved project operation exception in AGENTS.md to files or directories, and restore the explicit operation-or-scope alternative that a prior pipeline auto-fix had collapsed into "and". Rework project-management SKILL.md's Remove section, which previously told firstmate to refuse project removal until a guarded helper existed; that helper was never built, so the text directly contradicted the new instruction-only exception. It now points at the exception plus the existing removal preflight it still requires unchanged. Update the one instruction-owners test assertion that hard-coded the sentence removed above, so the suite tracks current, not obsolete, text. * no-mistakes(review): Align project removal preflight with approved exception * no-mistakes(document): Align project removal documentation with approved exception * fix: restore removal test byte-for-byte and preserve the default sentence tests/fm-instruction-owners.test.sh had been changed to assert different text; restore it byte-for-byte to origin/main. project-management SKILL.md's Remove section now keeps the exact default "Never issue a raw removal command from Firstmate." sentence that test still asserts, immediately followed by the already-approved captain-operation-or-scope exception, so the default and the exception both stay explicit and consistent. * no-mistakes(document): Align project-write boundary documentation * fix(skills): route new project intake through secondmate scopes (#1275) * Route project intake through secondmate scopes * no-mistakes(test): Guard all main-home project registry mutations * no-mistakes(document): Consolidate secondmate routing documentation * no-mistakes: apply CI fixes * Restore new-project routing scope * no-mistakes(document): Clarify secondmate routing for new-project intake * no-mistakes: apply CI fixes * fix: scope validation corrections by accepted behavior (#1281) * fix: scope validation corrections by accepted behavior * no-mistakes(review): Classify stale delivery evidence as an autonomous correction * test: replace source assertions with behavioral coverage (#1282) * test: remove source-content assertions * no-mistakes(review): Replace source assertions with runtime behavior coverage * no-mistakes(review): Isolate Kimi task temp runtime coverage * no-mistakes(document): Refresh test cleanup documentation * no-mistakes: apply CI fixes * fix(watch): escalate busy workers with no completed turn (#1286) * fix(watch): bound how long a busy pane may run with no completed turn A busy pane (backend busy state or the harness's rendered footer) was unconditional, unbounded proof of liveness in every escalation path, so a hung foreground tool call behind a busy signature could run for hours undetected (2026-07 hibit-agent-focus-nonsteal-r1 incident: a catastrophic- backtracking regex hung one bash call for 25h behind an unchanging "Working..." footer). FM_BUSY_TURN_MAX_SECS (default 3600s) now bounds how long a busy pane may run with no completed turn (state/<id>.turn-ended, or its spawn record before any turn has completed). Past the bound, busy_turn_over_age routes the pane through the existing wedge_timer_check, reusing the identical stale reason, escalation counter, and demand-deep-inspection marker for human inspection only - never an automatic interrupt, signal, or restart of the worker or its tool process. A completed turn resets the age. Reproduced end-to-end against the real installed Pi TUI: a foreground `sleep 999999` bash call with no timeout renders the actual busy footer, and two captures ~15s apart show the elapsed counter changing the pane hash while the same turn stays unfinished. Running the pre-fix watcher against the real captures showed it never starts a wedge timer no matter how long the pane stays busy; the fixed watcher starts and escalates the timer through the same mechanism, while the real hung process remained untouched and alive throughout. * no-mistakes(review): fix: parse enriched AFK stale reasons * no-mistakes(review): fix: preserve enriched wedges during AFK supervision * no-mistakes(review): fix: route all enriched AFK wedges * no-mistakes(document): Clarify busy-turn age supervision documentation * fix(gitignore): ignore config/ as a directory, not by exact filename (#1261) A name-by-name list of config/ entries silently stops ignoring any new or home-local file placed there, which makes the working tree read as dirty and blocks guarded sync paths that refuse to touch a dirty home. AGENTS.md already documents config/ as captain-private and gitignored as a category; this makes .gitignore match that contract. * fix(tests): replace source-content .gitignore assertion with behavioral coverage (#1304) The second assertion in fm-gitignore-config.test.sh (added by #1261) greps .gitignore for a specific spelling of the config/ ignore pattern. It fails on a semantically equivalent pattern like config/** and does not prove Git actually ignores anything, per the completed source-content-test audit. Replace it with a real git check-ignore control test on a generated unrelated path, and strengthen the existing directory-coverage test with generated unpredictable direct and nested config/ paths. * feat: bound and consolidate startup memory during stow (#1303) * Add bounded startup memory curation * no-mistakes(review): Record reproducible stow verification evidence * no-mistakes(review): Validate inherited secondmate stow evidence * no-mistakes(document): Document editable startup-memory budget propagation * fix(herdr): place workers in the launching workspace (#1328) * fix(herdr): place workers in the launching agent's exact workspace Herdr enforces no workspace-label uniqueness, and spawn resolved its container by taking the FIRST workspace whose label matched the home label. With two workspaces both labeled "firstmate", a worker launched from the second one was created in the first, so it appeared in a different space than the Firstmate the captain was watching. Reproduced end to end on Herdr 0.7.5 protocol 17 by running the real bin/fm-spawn.sh inside a launcher pane in the second "firstmate" workspace: the worker landed in w1 while its launcher was in w2, with an unrelated third workspace focused throughout, which also rules out any dependence on the focused workspace. Placement now binds to the launching process's own Herdr identity. Herdr injects HERDR_PANE_ID, HERDR_SESSION, and HERDR_SOCKET_PATH into every process it manages a pane for, and fm_backend_herdr_launcher_identity resolves that pane's current owning tab and workspace live from Herdr, cross-checking the pane against its tab and confirming the workspace exists exactly once in the session. The injected HERDR_TAB_ID and HERDR_WORKSPACE_ID are creation-time snapshots and are deliberately not read as current identity. Labels are no longer placement authority. A claimed parent identity that is unreadable, contradictory, stale, or from another named session or Herdr server stops the spawn before any worker endpoint exists, rather than degrading to a label search. A launcher with no Herdr ancestry has no workspace to inherit and keeps the per-home labeled container, which must now resolve to exactly one workspace; two same-labeled candidates refuse instead of adopting either. A --secondmate launch keeps standing up that home's own workspace by design. With presentation spaces enabled, the projected child is created and bound under that same exact parent and anchors its ordering on it, so a duplicated home label no longer makes the layout ambiguous. Projection, focus restoration, restart binding, and quarantine rules are unchanged, and children are never collapsed into the parent. tmux, Zellij, cmux, Orca, and the away-mode daemon terminal were each inspected and are not affected: none resolves a container by searching mutable labels. tests/fm-backend-herdr-launcher-workspace-e2e.test.sh drives the real spawn and teardown against an isolated Herdr lab, with its headline case running fm-spawn.sh inside a real Herdr pane so the identity comes from Herdr's own injection. The refusal matrix and the ordering anchor are covered deterministically in tests/fm-backend-herdr.test.sh. Eight existing real-Herdr suites inherited the developer terminal's own Herdr pane into their isolated lab sessions, which the new cross-session check correctly refuses. tests/herdr-test-safety.sh now owns herdr_forget_inherited_pane and those suites call it, so what they assert no longer depends on where they were launched from. Two unrelated fixes found along the way. tests/fm-secondmate-harness.test.sh had the same class of environment leak through CLAUDECODE, which outranks PI_CODING_AGENT in bin/fm-harness.sh and made its pi-signed ancestry case resolve "claude" whenever the suite ran inside Claude Code. And fm-spawn.sh's usage() printed a fixed line range that had already been truncating its own help mid-sentence. * no-mistakes(review): Enforce exact Herdr launcher and projection identity * no-mistakes(document): Document exact Herdr launcher workspace placement * fix(calm): refine Calm working boat animation (#1339) * feat(calm): replace Pi's working row with an animated ship while Calm is on While Calm is active and one logical agent run is under way, Calm now hides Pi's built-in working row and renders a small two-row SSHHIP-derived boat in its place. When Calm is off, Pi's stock working row is left untouched. The presentation uses only public Pi extension API: setWorkingVisible(false) plus a temporary setWidget() component whose render(width) owns the responsive geometry and whose timer requests a TUI render. Visibility follows agent_start through agent_settled, so the boat does not flicker between tool calls, automatic continuations, retries, or compaction inside the same run, and settle, abort, and failure all reach the same cleanup. fm-calm.ts stays the sole owner of the presentation choice and the only caller of setWorkingVisible(); the new lib owns the sprite geometry and widget. * no-mistakes(review): Guarded Calm-off lifecycle visibility writes; focused tests pass * no-mistakes(test): Fixed Calm E2E wait to include tmux scrollback * no-mistakes(document): Document Calm working boat behavior * no-mistakes: apply CI fixes * feat(calm): slow the Calm boat, animate blue water, and make the sail directional The boat now moves one column every 880ms while a bounded fixed-cell water phase advances every 220ms, so the water ripples several times between boat steps and the presentation reads as calm. One scheduler drives both clocks and disposing the widget stops them together; ticks rather than wall-clock timestamps drive every state change, so tests seek animation time exactly. Colors are standard ANSI foreground codes instead of theme lookups: blue for every water cell and yellow for the complete boat, each run closed with a default-foreground reset so nothing bleeds into padding or later frames. ANSI bytes never enter geometry, so visible width stays exact. The mainsail is directional and trails aft of the mast: <| travelling right and |> travelling left. Direction reverses the moment the boat lands on an endpoint, so the endpoint frame already shows the new heading and no frame at or after a bounce shows the previous sail. * test(calm): wait for the Ctrl+O expansion redraw this block asserts * docs(calm): record the revised working-presentation verification evidence * no-mistakes(document): Fix Calm feasibility document EOF whitespace * fix(dispatch): preflight candidate auth before quota escalation (#1349) * fix(dispatch): scope candidate authentication to its own surface A locally expired timestamp in one credential store was reported to the captain as a sign-out, including for dispatch candidates that never read that store. A `harness=pi, model=xai/grok-*` candidate authenticates through Pi's own xAI credential, but the only Grok quota reading available was gated on the standalone Grok CLI's separate token, whose expiry clock drifts independently. The always-loaded intake rule then turned that unreadable quota into a mandatory captain escalation. Add `bin/fm-auth-preflight.sh` as the deterministic owner of the parts that must not depend on agent memory: it resolves a tuple's authentication surface from quota-axi's own emitted auth sources rather than from a harness or model name, so another harness's CLI can never gate a candidate that does not use it. A vendor CLI is launched only when the tuple's own harness owns the credential store under test and a non-destructive discovery command is registered for it, which today is `grok models` alone. That probe runs at most once with stdin closed and a hard timeout, reads its verdict from the first stdout line because the command exits 0 either way, treats unrecognized output as indeterminate, and never invokes login, logout, or the interactive TUI. Quota is read at most twice, and unknown headroom never makes a candidate ineligible on its own. Update the dispatch procedure to match: usable authentication with unmeasurable headroom stays eligible at lower preference with the unknown disclosed, and stop-and-report is reserved for unresolved authentication, an unresolved relationship, or malformed configuration. Record that Grok's `credits.remaining` is a prepaid balance rather than window headroom. Gate quota-axi at 0.1.16 in bootstrap, the first build reporting per-credential auth sources. A stale install previously passed the presence check silently, which is why a fix published two days earlier was still not in effect. Replace the orphaned quota-array-dispatch fixtures, which encoded a `provider: "xai"` shape the tool never emits and had no consumer, with fixtures shaped like real 0.1.16 output that the new suite drives the script against. The suite asserts the verdict and, separately, which vendor CLIs were launched, so a Pi/xAI candidate reaching the Grok CLI fails. Map `tests/fixtures/<dir>` to its consuming suite so a fixture change selects the right tests instead of refusing. * refactor(bootstrap): give the quota-axi floor one owner The floor was stated twice - once in bootstrap's gate and once inline in the auth preflight - so bumping it needed two edits that could drift. Move it to bin/fm-quota-axi-lib.sh alongside its rationale, matching the existing tasks-axi library, and derive the comparison from the constant so the number appears exactly once. Bootstrap turns a failing check into the operator diagnostic; the preflight refuses to emit an unscoped verdict. Map the new library to both consuming suites so a bump re-runs them, and record that any usable source means the surface authenticates. * no-mistakes(review): Captain: bound quota checks and removed Python dependency * no-mistakes(review): Captain: enforce conservative headroom and exact preflight retry * no-mistakes(review): Captain: preserve OpenCode eligibility without auth-surface guessing * no-mistakes(review): Captain: reject malformed OpenCode model relationships * no-mistakes(review): Captain: exempt verified unmodeled tuples from intake escalation * no-mistakes(document): Updated dispatch authentication documentation * no-mistakes: apply CI fixes * feat(x-mode): reconcile promised public replies deterministically (#1350) * feat(x-mode): reconcile promised public replies deterministically A promised final reply in an X or Discord thread was only kept while the primary remembered it. Compaction or restart erased that memory, so a typed public-followup obligation could sit at pending-work after its PR merged and the original thread never got its reply. Make the promise durable state instead: - bin/fm-public-followup-emit.sh reports a typed terminal work result (source home, work id, generation, outcome, safe deliverables, bounded public-safe text) into the owning home's private inbox. The event id is derived from that identity tuple, so duplicate reports and restart replay converge with no coordination, and nothing ever parses a free-form done: sentence. - bin/fm-public-followup.sh registers a commitment, reconciles events through tasks-axi public-followup, and runs the idempotent delivery sequence (begin-delivery with the payload hash, post, record the posted receipt or a typed error) against the stored platform and opaque thread binding. A delivery interrupted between post and receipt refuses rather than risk a second public reply. - Session start surfaces unresolved commitments from disk, the existing relay poll surfaces a new terminal-result set once, and teardown refuses while this home still owes a public reply for that exact work. tasks-axi public-followup remains the only owner of the obligation state machine, state/x-context/ the only owner of the private request context, and fm-x-reply.sh the only thing that posts. Its new optional --receipt-file is the one addition there, so a caller can record how many messages were sent. A home that never opted into the myfirstmate relay gates out on a single [ -f "$FM_HOME/.env" ] test: no tasks-axi call, no backlog or context scan, no output, and no artifact. Evidence in docs/verification/public-followup.md. * no-mistakes(review): Hardened public-followup reconciliation and ownership guards * no-mistakes(review): Hardened typed terminal cleanup and receipt reconciliation * no-mistakes(review): Automated typed-delivery cleanup and strict backlog validation * no-mistakes(review): Fail-closed parent resolution and registration-safe delivery * no-mistakes(review): Harden relay gating and validate secondmate bindings * no-mistakes(review): Use owner-aware single-gate teardown protection * no-mistakes(document): Correct public-followup documentation drift * no-mistakes(lint): Quote done literals to fix ShellCheck warnings * no-mistakes: apply CI fixes * feat(bin): replace busy heuristics with semantic lifecycle state (#1327) * feat: add semantic busy-state contract owner and event writer One owner (bin/fm-busy-lib.sh) for the captain-approved semantic busy-state redesign: a per-task gen-bound record written only by bin/fm-busy-event.sh, per-harness trusted-source classification with explicit source attribution, busy/idle/unknown/dead semantics where missing, malformed, stale, or untrusted semantic data is unknown - never idle - and endpoint death is the only process-level override. The Grok-only rendered-tail fallback and the standalone-Kimi verification gate live behind the same classifier. * feat: arm busy-state at spawn and convert Pi to the semantic extension path fm-spawn arms the busy-state contract for converted adapters and seeds busy/fm-spawn (the launch brief is a submitted turn). The Pi/pi-signed per-task extension now reports agent_start -> busy and agent_settled -> idle confirmed by ctx.isIdle(), covering auto-retries, compaction retries, tool loops, and queued continuations, while turn_end stays a wake notification touch. Teardown removes the new record, gen sidecar, and lock. Live-verified on Pi 0.82.0: seed -> agent-start busy -> agent-settled idle with the marker still touched. * feat: convert OpenCode to the semantic session.status plugin path The per-task plugin (renamed .opencode/plugins/fm-busy-state.js) now classifies from OpenCode's semantic session.status events - busy and retry are active, idle is inactive - latched to the worker's own session so a subagent child session can never clear the worker's busy state. The session.idle marker touch stays a wake notification. Teardown removes both the new and the legacy plugin filenames. Live-verified on OpenCode 1.17.18 in a real TUI pane: seed -> session-busy -> session-status-idle. * feat: convert Claude to the full lifecycle hooks path The per-task settings.local.json now wires UserPromptSubmit -> busy and Stop, StopFailure, and SessionEnd -> idle, so API-error and shutdown turn ends can never strand a busy record; Stop keeps the turn-ended notification touch. A refused (stale-gen) event exits 0 and stays silent so Claude's own lifecycle is never broken. Live-verified on Claude Code 2.1.220: UserPromptSubmit fires for the argv launch prompt, Stop closes each turn, a mid-stream Escape interrupt fires no closing hook, and the firstmate-controlled idle/fm-interrupt clear resolves it. * feat: gate Codex busy state behind verified semantic sources The approved contract prefers Codex's app-server turn lifecycle with capability negotiation and sanctions its lifecycle hooks as the intermediate. Live probes on codex-cli 0.145.0 show neither is usable for a pane worker: the app-server daemon is unreachable for a TUI thread and refuses to start outside the managed standalone install, and firstmate-written project hooks never fired (interactive with directory trust granted, and exec, both with --dangerously-bypass-hook-trust) while global hooks fired in the same runs. Codex therefore classifies unknown codex-unverified behind an explicit probe rather than falling back to idle or footer text, and fm-spawn installs no unverified Codex wiring. * feat: gate standalone Kimi busy state on live verification Standalone Kimi has no installed binary here, so per the approved contract its semantic path stays guarded and it classifies unknown kimi-unverified rather than idle - and never from its locale-sensitive moon-phase spinner, which the redesign forbids inventing as a state source. The gate records the preferred source order (Wire prompt request lifetime, which brackets a turn and reports cancellation, then the documented hooks including Interrupt because Stop does not fire on interrupts) and the exact evidence required to open it. Arming without wiring would seed a busy record nothing could clear, so both land together behind the same gate. * feat: route busy consumers through the contract and drop the global OR The watcher, crew-state reader, and away-mode daemon now decide busy state through bin/fm-busy-lib.sh: only an exact busy verdict counts as working, and unknown never becomes working or a silent idle, so a crew whose semantic state is missing, malformed, stale, or unverified surfaces instead of being absorbed. Crew-state reports the producing source in its detail. The watcher's global OR regex default is gone; Grok keeps its isolated fallback inside the contract. The daemon's supervisor-pane reader stays rendered-text - that pane is not a recorded task - but is now scoped to firstmate's own detected harness instead of every vendor signature. Secondmate pending-reply observation is deliberately unchanged and documented as a delivery-confirmation signal, not task state. * docs: point busy-state documentation at the single contract owner Adds a maintainer-architecture section naming bin/fm-busy-lib.sh as the owner of what busy means, with per-adapter sources, the unknown-never-idle rule, the endpoint-death override, and the two rendered-text readers that deliberately stay outside the contract. Replaces the stale regex-first prose in architecture, tmux-backend, herdr-backend, and configuration; converts the harness-adapters per-harness rows from UI signatures to the semantic source each harness uses; and records the live verification evidence, including why Codex and standalone Kimi stay unknown. * fix: arm away-launch signal handlers before acquiring the lifecycle lock fm_afk_launch_main acquired its lock and only then installed the EXIT, INT, and TERM traps. A signal arriving in that window terminated the process by default action and left the lock directory behind, which blocks the next away-mode launch until the stale-owner reclaim path clears it. The release helper only removes a lock this process owns, so the handlers are now armed first. The accompanying test also killed the child whether or not the lock had appeared and sampled cleanup the instant wait returned; it now requires the lock, then allows a bounded settle, so it proves the guarantee instead of racing it. * test: align fleet, Kimi, lifecycle, and detection suites with the contract The fleet snapshot and wake-daemon lifecycle fixtures now prove a working crew through its own semantic busy-state record instead of rendered pane text, which is what those consumers read. The Kimi watcher test asserts the approved contract directly: a standalone Kimi task classifies unknown rather than matching its moon-phase spinner, while Grok's isolated fallback still classifies only Grok. The pi-signed detection cases clear ambient harness markers, fixing a pre-existing failure where the running session's own CLAUDECODE outranked the fixture's marker. * fix: stop teardown from deleting a project's own Codex hooks file An intermediate revision wired Codex through a firstmate-written <worktree>/.codex/hooks.json, and teardown removed it alongside the other generated wiring. The Codex wiring was dropped when its probes came back unverified, so that removal now targets a file firstmate never creates - and a project may legitimately track its own .codex/hooks.json, which teardown would then delete from a pooled worktree. * fix: keep busy-record parsing from disturbing its sourcing caller The record parser split fields with set -- under a temporary noglob, which clobbers a sourcing caller's positional parameters and restores glob expansion even when the caller had disabled it. The watcher, the daemon, and the crew-state reader all source this library, so it now reads fields with read -a, which never globs and never touches caller state. * docs: state exactly which Claude hook paths were reproduced live The busy-state record listed all four wired Claude hooks in the source column, which could read as a claim that every one fired during the pass. UserPromptSubmit and Stop did; StopFailure and SessionEnd are wired from hook names confirmed present in the installed binary, but the abnormal turn ends they cover were not reproduced. * test: let reset_fakes own the crew-state busy-text fixture lifecycle The Grok fallback case set FM_FAKE_BUSY_TEXT and cleared it inline, so the variable's lifetime was owned by one test rather than by the shared reset that every other fake already uses. * no-mistakes(review): Fix semantic busy-state lifecycle races * no-mistakes(review): Make busy-state retirement idempotent * no-mistakes(review): Enforce semantic state boundaries for status and injection * no-mistakes(review): Restore harness-scoped away-mode busy guard * no-mistakes(document): Refresh semantic busy-state documentation * no-mistakes: apply CI fixes * fix: preserve Calm boat continuity across working periods (#1356) * fix(calm): resume working boat from frozen column across runs Keep one extension-owned boat animation for the Pi session so settling freezes column and direction, the next working period resumes there without hidden-time jumps, and only a fresh session resets to the left edge. * no-mistakes(review): Freeze Calm boat from last rendered state * no-mistakes(document): Document Calm boat continuity contract * fix: restore evidence-based dispatch eligibility (#1358) * fix(dispatch): judge candidate provider relations instead of rejecting them Firstmate deterministically dropped supported Pi candidates in the openai-codex family. bin/fm-auth-preflight.sh resolved a harness=pi tuple's credential surface by constructing the source id `pi:<model-prefix>`, so `pi + openai-codex/gpt-5.6-terra` looked for a `pi:openai-codex` source. That source does not exist, because Pi's Codex family authenticates through the Codex store quota-axi already lists as `auth-json`/`cli-rpc`. The tuple returned `eligible=no reason=surface-unresolved` while the Pi catalog listed the model and the Codex provider reported fresh, usable credentials with 64 effective percent remaining on its all-model scope. The prefix construction was only ever valid where Pi holds its own credential (`pi:xai`, `pi:kimi-coding`), which is why every previously configured Pi tuple resolved and the defect stayed hidden until a Codex-family Pi model was configured. Retire dispatch eligibility from deterministic shell. The dispatching first mate now establishes model support and provider family from each harness's authoritative catalog, applies quota at the granularity the vendor supplies, and shows that reasoning. Provider-level and all-model evidence bounds every model established in that family; a named-model window bounds only its own model. Missing model-level quota, a missing auth source, unmeasurable headroom, and unmodeled authentication are disclosed uncertainty. Only concrete contradictory evidence blocks a candidate. Replace the preflight with bin/fm-vendor-auth-probe.sh, which keeps the captain's approved bounded probe envelope without any routing knowledge: it takes no harness, model, or provider, reads no quota, renders no verdict, and holds only a fixed-argv safety allowlist. Its behavior suite proves the absent identity surface, the untouched quota, the uniform exit status, the fixed argv with stdin closed, and a real bound even when the configured bound is zero. Also fixed along the way: a zero FM_*_TIMEOUT silently removed the hard bound, the pinned Grok version had drifted to 0.2.117, and --changed selection refused outright on any deleted bin/ script. AGENTS.md section 4 and quota-array-dispatch own the corrected policy, harness-adapters gets the catalog-responsibility correction, and docs/verification/dispatch-auth.md records the 2026-07-30 evidence on Pi 0.82.0, quota-axi 0.1.16, and grok 0.2.117. * no-mistakes(review): Reject all-zero vendor probe timeouts * fix(session-lock): resolve Claude bg-spare ancestry to the outermost claude pid (#2) * fix(session-lock): resolve Claude bg-spare ancestry to the outermost claude pid fm_harness_ancestry_pid() previously returned the first ancestor process whose command matched a verified harness name. Claude Code's Stop hook fires as a bg-spare worker several levels below the session's actual lock-owning claude process (hook shell -> claude bg-spare -> claude bg-pty-host -> claude -> claude(lock)), so the first match was the bg-spare worker, not the lock owner. fm_session_lock_owned_by_self() then never matched state/.lock, and the Claude Stop auto-arm silently treated its own primary session as an unrelated live owner and never armed the watcher. The walk now keeps going past a claude-named match, looking for a still more ancestral claude-named match, and stops the instant a non-match follows an already-found match (bounding it to a contiguous run rather than the literal ancestry top, so an unrelated claude-named process further up the real process tree is never mistaken for part of this session's own nested chain). Every other harness keeps the original first-match-wins behavior, since e.g. Pi's shared signed-wrapper ancestry actually holds the session at the inner engine pid, not an outer wrapper pid. Hop limit raised from 8 to 16 to cover the deeper bg-spare chain. * no-mistakes(review): Add nested-claude-ancestry regression test; fix nudge doc depth claim * Wire fm-spawn.sh and fm-teardown.sh to Parlay chat-panel enrollment (#4) * Wire fm-spawn.sh and fm-teardown.sh to Parlay chat-panel enrollment fm-spawn.sh best-effort enrolls a confirmed launch via `parlay listen --agent <id>`, backgrounded with its pid recorded to state/<id>.parlay-listen-pid. fm-teardown.sh best-effort deregisters on clean teardown: the recorded pid is always killed, and `parlay agent-down <id>` is called only when parlay is on PATH. Neither side ever blocks or fails spawn/teardown when parlay is absent or fails. Adds tests/fm-spawn-parlay.test.sh and two new tests in tests/fm-teardown.test.sh, plus a new fm_path_without test helper in tests/lib.sh for simulating parlay's genuine absence from PATH. * no-mistakes(lint): tests/lib.sh: rename fm_path_without's out array to avoid shellcheck SC2178/SC2128 * feat: add beads as third backlog backend option When config/backlog-backend=beads is set, firstmate uses the beads federated 'task' store as the queue source instead of data/backlog.md. Session-start's digest lists items with status:ready label from the beads store. - Add fm_beads_backend_available() check in fm-tasks-axi-lib.sh - Add print_backlog_beads_compact() rendering function in fm-session-start.sh - Update print_backlog_compact() to prioritize beads backend when configured - Add beads backend validation to bootstrap (checks task CLI and store reachability) - Update docs/configuration.md to document beads backend option - Update AGENTS.md section 10 to reference beads backend in backlog contract - Beads backend reuses existing task linkage machinery (task set-state, task close) The beads backend is fail-open: if task CLI is missing or store is unreachable, bootstrap reports a MISSING: diagnostic line and the home can still operate. * test: add beads backend integration tests Tests for: - fm_backlog_backend_value() reading beads config - fm_beads_backend_available() checking task CLI and store - fm_tasks_axi_backend_available() returning false when beads is set - whitespace handling in backend config values * fix: add install_cmd support for task (beads CLI) The install_cmd() function now recognizes 'task' and provides an install command for the beads CLI tool. * fix: make print_backlog_pointer() backend-aware Update print_backlog_pointer() to provide backend-specific guidance: - beads backend: suggest 'task show <id>' for beads task store - manual backend: suggest 'inspect data/backlog.md' - default/tasks-axi: original message with tasks-axi and data/backlog.md * no-mistakes(document): Add beads backlog backend documentation about feature limitations for handoff and decision holds. * no-mistakes(review): Fix beads backend binary name mismatch and add unsupported operation checks * no-mistakes(document): Add beads as third supported backlog backend - sync documentation * no-mistakes(document): Add beads as third supported backlog backend - sync documentation * no-mistakes(lint): Remove duplicate deregister_parlay_agent function definition * no-mistakes(document): Document beads as third backlog backend option * Fix: restore backlog pointer for all backends and quote basename argument - Add 'or data/backlog.md' fallback to beads backend pointer for consistency - Quote basename argument in fm-session-lock-lib.sh to handle dash-leading process names (fixes 'basename: missing operand' when ps output starts with '-') - Fixes tests/fm-session-start.test.sh:1148 and tests/fm-secondmate-harness.test.sh * no-mistakes(document): Documented beads backlog backend throughout project - one clarification edit to configuration.md made * Fix: manual backend pointer must include 'or data/backlog.md' fallback * Fix: manual backend pointer wording - include 'or data/backlog.md' with manual guidance * no-mistakes(review): Add beads backlog fallback to manual when query fails * no-mistakes(review): Add beads backend check to backlog_refresh_reminder messaging --------- Co-authored-by: Daniel Kuykendall IV <danielkuykendall23@gmail.com> Co-authored-by: Unknownzed <45267749+Unknownzed@users.noreply.github.com> Co-authored-by: lhalbert <lucashalbert@users.noreply.github.com> Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Co-authored-by: AG <ag@agw3.org> Co-authored-by: deeto15 <92119640+deeto15@users.noreply.github.com> Co-authored-by: Christopher McKay <101884182+karotkriss@users.noreply.github.com> * Add documentation audience classification for no-mistakes-run-liveness.md * feat(beads): Add lifecycle labels to dispatch and claim flow for mg badge rendering (#10) * wire lifecycle labels to bd dispatch and claim flow for mg badge rendering - fm-bead-stamp.sh: add lifecycle=sent label on dispatch (alongside dispatch=sent) - fm-brief-hooks.d/beads.sh: add lifecycle=claimed label to Bead Receipt (alongside dispatch=claimed) This enables the mg board to render lifecycle state as gold badges. The dispatch/lifecycle metadata remains for data tracking; set-state labels are the render channel for mg visibility. * no-mistakes(document): Synced all documentation with beads lifecycle-tracking feature implementation. * feat: add optional --label flag to fm-spawn (#9) * Add optional --label flag to fm-spawn.sh - Add --label <slug> flag to fm-spawn.sh flag parsing - When provided, use it for the tab label with fm- prefix (e.g., fm-<slug>) - Fall back to default fm-<ID> behavior when --label is absent (fully backward compatible) - Record label=<slug> in state/<id>.meta when a non-default label is used - Document the new optional label field in docs/configuration.md * no-mistakes(document): Document --label flag in fm-spawn.sh and related docs. * feat(no-mistakes): Add run liveness detection tools (#8) * Add fm-no-mistakes-liveness.sh: check no-mistakes run liveness and task ownership * Add fm-nm-run-is-live.sh: lightweight run liveness check for supervision * Add documentation for no-mistakes run liveness checks * Improve active_steps parsing for accurate liveness detection - Extract active_steps section before parsing to avoid mixing with steps table - Parse last_activity timestamp from active_steps for real-time progress tracking - Report actual elapsed time since last activity instead of cumulative step duration * no-mistakes(review): Fixed hardcoded path, unused functions, inconsistent timeout config, and liveness detection logic * no-mistakes(review): Fixed error handling, directory contexts, removed dead code, unified duplicate utilities * no-mistakes(document): Add liveness checking scripts to documentation index and environment variables * no-mistakes: apply CI fixes * fix(bin): restore beads backend support and dispatch validation (#11) * Add fork URL mapping fix and helper script for no-mistakes daemon (#13) * Add helper script to fix no-mistakes fork URL mapping for forked repos When a repo is a fork (e.g., trillium/firstmate is a fork of kunchenguid/firstmate), the no-mistakes daemon may use the parent repo URL for PR creation. This script detects and fixes the fork_url in the daemon's state database to ensure PRs are created against the correct fork. The issue occurred with beads-backlog-backend task when it tried to create a PR against the daemon's cached fork URL, which was incorrectly pointing to the parent repository instead of the user's fork. * no-mistakes(review): Add SQL injection protection, error visibility, precise error messages, exact URL matching * no-mistakes(review): Add SQL injection protection, error visibility, precise error messages, exact URL matching - Add proper SQL escaping to prevent SQL injection if URLs contain quotes - Remove error suppression (2>/dev/null) from sqlite3 calls to surface database errors - Distinguish between 'repository not found' and 'database update failed' for clearer errors - Replace loose LIKE pattern matching with exact equality to prevent wrong repo updates * feat(brief): inject fork-first push rule for non-Trillium origins (#12) * fm-brief: add fork-first push rule for non-Trillium origins Push-mode ship briefs (direct-PR, no-mistakes) whose project clone has a git origin outside the trillium/ namespace now tell the worker to push its branch to the trillium/<repo> fork and open the PR from there, instead of stalling on a refused upstream push. Detection reads the clone's real origin remote, not registry prose. Trillium-origin, unreadable, or absent origins and every local-only brief are unchanged. Extends tests/fm-brief.test.sh with the rule-present (HTTPS and SSH origins), Trillium rule-absent, and local-only exemption cases. * no-mistakes(review): Reword fork-first rule to be mode-agnostic with upstream-safety guard * no-mistakes(document): Document fork-first brief rule in AGENTS.md section 11 * no-mistakes: apply CI fixes * docs: add dirty-worktree default pattern [quality:documented] Default for exploratory changes during ship tasks: 1. Commit the work (never discard) 2. Mark with [quality:needs-review] 3. File tracking task for curation Preserves thought, keeps decisions durable, makes exploration visible. * fix(watcher): prevent false-FAILED alarm on benign one-shot cycle end (#15) * Fix false watcher FAILED alarm on benign one-shot cycle end fm-watch-arm.sh reported 'watcher: FAILED - cycle ended without an actionable reason' whenever a cycle ended with no verified healthy successor, ignoring the liveness beacon. A one-shot watcher exiting after surfacing an actionable wake, or an attached/duplicate arm outliving that exit, hit this path constantly: every false-FAILED record in the cycle ledger had a beacon age of 12-44s, well within the 300s grace. The Stop hook and turn-end guard read that FAILED as supervision-down, producing a false-failure -> re-arm -> re-detect loop that spammed the primary and re-emitted stale events for already-parked panes on each restart. Gate the terminal classification on the beacon, the same liveness authority the turn-end guard and fm-guard use. A beacon still fresh within FM_GUARD_GRACE proves a watcher was healthy up to the moment the cycle ended (a clean one-shot exit whose wake fm-watch.sh already enqueued durably, or a benign empty poll), so the arm now prints a non-FAILED 'watcher: idle' line and exits zero, leaving continuity to the adapter re-arm. Only a stale, expired, or absent beacon still emits the typed nonzero FAILED. The FAILED string and nonzero exit are unchanged for genuine failures, so the Stop-hook and turn-end-guard contract is intact. Tests: convert the existing no-successor FAILED cases to expire the beacon (genuine supervision-down) and add a regression test proving a fresh-beacon cycle end with no successor is benign, not a false FAILED. * no-mistakes(lint): Remove unused ROWS_AFFECTED variable in fork-mapping script * feat(watcher): Wake on automated PR reviews (#16) * Add bot-review-comment wake to the PR poll The watcher's PR poll wakes firstmate only on merge. Extend fm-pr-poll.sh so a new automated-reviewer review (CodeRabbit or any Bot-type GitHub account) also wakes firstmate, surfacing review feedback promptly instead of only at merge time. - fm-pr-poll.sh emits a distinct bot-review token when the highest Bot-authored review id exceeds the last one surfaced, deduping through a private state/<id>.pr-review-seen sidecar. merged takes priority and short-circuits; the poll stays armed after a bot-review wake. GitHub only (glab has no reviews API); every failure path stays silent. The seen sidecar is derived from the check basename in standalone mode and passed as the seventh validated argument by the watcher. - fm-watch.sh passes the per-task seen path to the validated poll. - fm-teardown.sh removes the seen sidecar with the other poll artifacts. - Extend the PR-check security suite with a bot-review wake test (silence, dedup, id advance, reviews-API failure, merged priority, and an end-to-end watcher-bounded run that stays armed) and cover seen-sidecar teardown. - Record the new sidecar in the AGENTS.md state inventory. * no-mistakes(document): Document bot-review wake and GitHub/GitLab polling differences * no-mistakes(lint): Remove unused ROWS_AFFECTED variable assignment * feat(watcher): implement idle-task-discovery and auto-dispatch (#17) * feat: implement idle-task-discovery in watcher When the fleet is idle (no active crewmates/secondmates running), the watcher now automatically discovers ready tasks from the task store and dispatches them. Key features: - Detects idle state when recorded_windows is empty - Queries tasks-axi ready for unblocked, unheld queued work - Selects first ready task (highest priority by default) - Dispatches secondmate tasks autonomously - Skips ship/scout tasks pending project-routing implementation - Throttled to run every 60s to avoid constant polling - Logged via triage_log for watcher debug visibility This reduces captain context switches by allowing the watcher to pick up ready work autonomously during idle periods. * no-mistakes(document): Document FM_IDLE_DISCOVERY_INTERVAL env var in configuration.md * no-mistakes(lint): Fix ShellCheck warnings: remove unused variables and fix trap quoting * docs: add gascity integration feasibility analysis (#18) * chore: gascity integrati…
Intent
add beads backend error handling and backend-aware user guidance
What Changed
Risk Assessment
✅ Low: Changes are straightforward documentation clarification and a well-implemented contract validation test; no correctness, security, or performance concerns identified.
Testing
Validated beads backend error handling and user guidance by running unit tests (fm-beads-backend.test.sh passing all 4 test cases) and bearings snapshot test (test_chat_contract_four_sections passing), plus manual verification of all three error scenarios: task CLI available (shows bootstrap info), task CLI missing (shows install instructions), and task store unreachable (shows connectivity error). All backend configuration logic works correctly including default fallback, whitespace handling, and proper disabling of tasks-axi when beads is configured.
Evidence: Beads Backend Validation Report
Evidence: Test Execution Summary
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
🔧 **Rebase** - 4 issues found → auto-fixed ✅
.agents/skills/harness-adapters/SKILL.md- merge conflict rebasing onto origin/main.agents/skills/quota-array-dispatch/SKILL.md- merge conflict rebasing onto origin/maintests/fixtures/quota-array-dispatch/cases.json- merge conflict rebasing onto origin/maintests/fm-quota-array-dispatch.test.sh- merge conflict rebasing onto origin/main🔧 Fix applied.
✅ Re-checked - no issues remain.
🔧 **Review** - 2 issues found → auto-fixed (2) ✅
bin/fm-teardown.sh:315- Duplicate function definition: retire_busy_state() is defined twice (lines 290-297 and 315-322), creating dead code since the second definition shadows the first.bin/fm-teardown.sh:328- Duplicate function definition: deregister_parlay_agent() is defined twice (lines 303-313 and 328-338), creating dead code since the second definition shadows the first.🔧 Fix: Remove duplicate function definitions in fm-teardown.sh
2 issues (1 error, 1 info) still open:
tests/fm-bearings-snapshot.test.sh:1896- Test function test_chat_contract_four_sections() is defined but never invoked. The function is not added to the test invocation list at the end of the file (lines 1920+), so it will never execute.tests/fm-bearings-snapshot.test.sh:1908- Regex pattern 'Captain.s Call' uses '.' as a wildcard and will match 'CaptainAs Call' or similar unintended variations. Should use 'Captain&feat(fm-backend): extend legacy-metadata self-repair to zellij and cmux backends #39;s' or similar for precision, though this won't cause test failures in practice.🔧 Fix: Add missing test invocation to bearings snapshot suite
✅ Re-checked - no issues remain.
✅ **Test** - passed
✅ No issues found.
bash tests/fm-beads-backend.test.shtest_chat_contract_four_sections() from fm-bearings-snapshot.test.shManual validation: beads backend detection with available task CLIManual validation: error handling with missing task CLIManual validation: error handling with unreachable task storeBackend configuration parsing with whitespace handlingTasks-axi backend correctly disabled when beads is configured✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.
Summary by CodeRabbit
Documentation
Tests