feat(fleet): add per-crew usage rows - #36
Merged
Merged
Conversation
Slices 0-2 of the captain-approved fleet usage-bar report (data/fm-harness-usage-bar/report.md): - Slice 0: verified the Codex statusline proxy (~/.codex/statusline.sh -> LIFEOS_StatusLine.sh) renders live and is input-sensitive, not a bare fallback. Evidence in docs/verification/fm-harness-usage-bar.md. - Slice 1: bin/fm-crew-usage-lib.sh adds a harness/model/context_pct/quota row per task, wired into fm-fleet-snapshot.sh's canonical per-task JSON and projected into fm-bearings-snapshot.sh's Underway rows as flat usage_* fields. context_pct is always "n/a" (no harness exposes it externally). quota (quota-axi spendPriority) is opt-in via FM_CREW_USAGE_ENABLE_QUOTA=1 because it is a live per-account network call; on by default it would multiply into several network round trips per snapshot (secondmate homes recurse into their own snapshot) and measurably slowed/flaked the existing test suite during development. No new poller either way. - Slice 2: confirmed opencode's "stats" command has no per-task scoping (cross-session aggregate only); confirmed pi --mode json and cline --json do carry real per-turn token usage, but only from a freshly invoked print turn, not a read of an already-running interactive task's state, so no adapter was added for either. Findings recorded in the verification doc. Tests: tests/fm-crew-usage-lib.test.sh (7 cases). Existing tests/fm-fleet-snapshot-view.test.sh and tests/fm-bearings-snapshot.test.sh still pass with the new usage field present. Co-Authored-By: Claude Sonnet 5 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01B92B7nHnZmk1rzCbtsAU2i
The usage row read each task's statusline on every canonical snapshot. That is not a background cost: bin/fm-watch.sh backgrounds two snapshot consumers on EVERY poll - fm-home-summary-refresh.sh (--secondmate-home-summary) and fm-secondmate-reconcile.sh process-requests (--json) - so the snapshot became a second pane reader racing the watcher's own capture. The watcher proves pane churn by comparing consecutive captures and uses that proof to absorb a bare turn-end. A competing capture destroys the evidence, so the watcher resurfaced a wake it had proof to absorb: tests/fm-watch-triage.test.sh "pane churn resets prior wedge escalation state before the stale-path poll" failed on every run, and passed with bin/fm-fleet-snapshot.sh reverted (isolated by reverting each changed file in turn). A bound does not fix this - a fast extra capture is still an extra capture. Gate the live read behind FM_CREW_USAGE_ENABLE_CONTEXT, matching the existing FM_CREW_USAGE_ENABLE_QUOTA opt-in, and let only bin/fm-bearings-snapshot.sh - the on-demand human reader that renders the bar - turn it on. Every supervision-path caller keeps main's capture behaviour exactly. Also stop presenting a placeholder as a model. bin/fm-spawn.sh records model=default when no model was chosen, so live codex rows read "default". Report the placeholders as unrecorded, and under the same opt-in recover the real model from codex's own native footer, which is the one thing that footer does carry. Verified read-only on three running codex panes: all recorded model=default while their footers read "gpt-5.6-terra high"/"xhigh"; the default path captures nothing, the opted-in path returns gpt-5.6-terra. Codex context percentage stays "n/a" and is not synthesized: fm-statusline-quota.sh returns status=unknown source=none for those panes and the footer carries no percentage, because codex's TUI does not render the configured external proxy. That boundary is recorded rather than papered over. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HRyaSpa65jHGqWXgT6sqX1
tests/federation/test_spawn_account.sh "wrapper preserves quoted account launch
harness" failed before this change: an account spawn passing neither --model nor
--effort left ACCOUNT_PROFILE_ARGS empty, and stock macOS bash 3.2 treats
"${arr[@]}" on an EMPTY array as an unbound variable under set -u, so the
wrapper died before reaching fm-spawn. PASS already carried the +-guard; this
gives ACCOUNT_PROFILE_ARGS the same one.
Also correct a stale fixtures comment: fm_test_make_spawn_fakebin no longer
stubs treehouse, so the comment now describes what it actually creates.
Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01HRyaSpa65jHGqWXgT6sqX1
adibirzu
force-pushed
the
fm/fm-harness-usage-bar
branch
from
September 7, 2026 19:51
55d17b8 to
2fa8d37
Compare
… guard alarm Two blocking defects from independent review of the previous repair. B1 - an invalid usage timeout silently blinded the canonical fleet snapshot. bin/fm-crew-usage-lib.sh refuses to load on a non-positive bound, because fm_run_timed's contract DISABLES the deadline for one. bin/fm-fleet-snapshot.sh runs under set -u but not set -e, so it continued without the library: every task's `jq --argjson usage ""` failed, whole task rows were dropped, and it still exited 0. Every supervision consumer - the two per-poll watcher backgrounds, secondmate reconcile, bearings, session start - would have read an empty fleet and reported no work in flight, with the only signal on a discarded stderr. Now the snapshot fails closed and says why. The refusal also named FM_CREW_USAGE_QUOTA_TIMEOUT even when FM_CREW_USAGE_CONTEXT_TIMEOUT was the invalid knob, because both went through one hard-coded validator. It now names the knob the operator actually set. B2 - the opted-in usage read silently spent the watcher-down alarm. fm-statusline-quota.sh and fm-peek.sh both run fm-guard.sh, which prints its WATCHER DOWN banner once per down-episode and CLAIMS that episode with a marker. Both usage reads discard stderr, so a /bearings run during a supervision lapse - exactly when the captain is catching up - swallowed the banner and left the next guarded command reporting that the full banner had "already been printed this episode", about an alarm nobody ever saw. Both diagnostics now run under FM_GUARD_READ_ONLY=1, the supported mode that claims nothing. Verified end to end: after an opted-in read the marker is unclaimed and the next guard run still prints the full banner. Also drop the dead RELAUNCH=0 in bin/fm-spawn.sh (nothing in the repo reads it; main uses REUSE_WORKTREE and RELAUNCH_STRICT), and correct the verification doc's overstated "failed on every run": the regression is a race, frequent but not deterministic, so the measured 3/3 and independently replicated 3/4 are recorded as such. The doc also now records the guard read-only requirement and the accepted residual that /bearings can still read a pane while the watcher is armed. Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01HRyaSpa65jHGqWXgT6sqX1
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
Add a fleet-wide per-crew usage row (harness, model, context percentage, provider quota) to firstmate's canonical fleet snapshot and the bearings Underway view, and verify the Codex statusline proxy renders live, per the captain-approved fm-harness-usage-bar report Slices 0-2: (1) verify the Codex statusline proxy already wired to LIFEOS_StatusLine.sh actually renders and is input-sensitive; (2) extend fm-fleet-snapshot.sh and bin/fm-crew-usage-lib.sh so every live pane's row carries harness, model, context percentage where available, and provider quota from quota-axi spendPriority, with no new poller and quota gated opt-in since it is a live network call; (3) add light adapters for opencode/pi/cline usage data only where their output was confirmed to carry real per-task token/context data usable for a live row, and record findings where it does not. Slices 3 and 4 (upstream feature requests, Cursor editor extension) are explicitly out of scope. Captain-decided review-gate fixes to apply: R2 wire context_pct through fm-statusline-quota.sh where supported; R3 propagate secondmate child usage through secondmate_home_summary_json into the parent snapshot and Bearings view; R6 capture the rendered bar from a live Codex pane read-only (do not launch a new Codex session) and record it in the verification doc; plus auto-fixes R1 (quota caching survives the production call path, not just subshells), R4 (harness mapping covers cursor and agy, not only cursor-agent), R5 (quota reads/caching are bound to the worker's own account identity, reporting unavailable when that identity cannot be established), and R7 (FM_CREW_USAGE_QUOTA_TIMEOUT validated positive before use).
What Changed
Risk Assessment
✅ Low: The change is bounded to on-demand snapshot telemetry, preserves account isolation before quota attribution, and handles unsupported or stale readings as unavailable.
Testing
Focused behavior tests passed for context validation/timeouts, opt-in account-bound quota caching, and account-launch isolation. A controlled end-to-end snapshot-to-Bearings run produced the reviewer-visible Codex usage row with model, live context percentage, and provider quota; the worktree remains clean.
Evidence: End-to-end Bearings usage row
Source: End-to-end Bearings usage row
{ "schema": "fm-bearings.v1", "in_flight": [ { "id": "usage-demo", "usage_harness": "codex", "usage_model": "gpt-5.6-terra", "usage_context_pct": "40", "usage_quota": "-3.25" } ] }Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
🔧 **Rebase** - 3 issues found → auto-fixed ✅
bin/fm-bearings-snapshot.sh- merge conflict rebasing onto origin/mainbin/fm-fleet-snapshot.sh- merge conflict rebasing onto origin/maindocs/documentation-audiences.json- merge conflict rebasing onto origin/main🔧 Fix applied.
✅ Re-checked - no issues remain.
🔧 **Review** - 6 issues found → auto-fixed (10) ✅
bin/fm-fleet-snapshot.sh:758- Required R1 says quota-axi must be fetched once in the task-loop shell so every mapped worker shares one call. This hunk invokesfm_crew_usage_jsonvia command substitution; it then command-substitutes the quota helper, so its cache mutations are discarded for every row. With two mapped tasks and quota enabled, quota-axi runs twice instead of once.bin/fm-crew-usage-lib.sh:84- Required R2 says to wirecontext_pctthrough fm-statusline-quota where supported. The new function always returnsn/a; a live Codex pane withContext 40% leftstill producesusage.context_pct="n/a".bin/fm-crew-usage-lib.sh:68- Required R4 covers normalcursorplusagy. The new lookup delegates to the existing mapping, which has neither mapping, so enabled quota rows for either harness silently rendern/adespite available provider telemetry.bin/fm-crew-usage-lib.sh:58- Required R5 requires worker-account-bound quota or unavailable. This path runs barequota-axiunder the snapshot process credentials and accepts no worker account/isolation identity, so two Codex workers on different accounts both receive the snapshot account's spendPriority.bin/fm-crew-usage-lib.sh:48- Required R7 requires a positive timeout validation.FM_CREW_USAGE_QUOTA_TIMEOUTis only defaulted; setting it to0reachesfm_run_timed, whose documented contract disables the deadline for non-positive bounds.docs/verification/fm-harness-usage-bar.md:32- Required R6 requires evidence captured read-only from an already-running Codex pane. The added document explicitly says its evidence is synthetic stdin only and does not prove the interactive TUI path, so the required live-pane capture remains absent.🔧 Fix: Fix account-bound cached crew usage telemetry
3 errors still open:
docs/verification/fm-harness-usage-bar.md:32- Required R6 says to “capture the rendered bar from a live Codex pane read-only … and record it.” This document still records synthetic stdin invocation and explicitly says it does not prove the interactive TUI path, so the required live-pane evidence remains absent..agents/skills/bearings/SKILL.md:145- The required usage row includes context percentage in the Bearings Underway view, but this new renderer instruction saysusage_context_pctis alwaysn/aand to never render it. A Codex row withcontext_pct:"40"is therefore silently rendered without its available context value.bin/fm-crew-usage-lib.sh:118- R2 requires context_pct through the statusline path where supported, including Claude panes reportingContext N% left. The statusline parser labels that shapesource=codex; this equality check rejects it for a Claude task, returningn/adespite a valid displayed context percentage.🔧 Fix: Render available crew context percentages
2 issues (1 error, 1 warning) still open:
docs/verification/fm-harness-usage-bar.md:32- The required criterion says to “capture the rendered bar from a live Codex pane read-only … and record it in the verification doc.” The added verification record instead says its only evidence is synthetic stdin and explicitly admits it does not prove the interactive TUI path, so the required live-pane evidence is still absent.bin/fm-crew-usage-lib.sh:118-context_pctaccepts any digit string from the statusline diagnostic. For a captured footer such asContext 999% left, it silently emits"999"as a percentage rather than marking the malformed reading unavailable. Restrict accepted values to 0–100 before placing them in the fleet row.🔧 Fix: Validate context percentage bounds
1 error still open:
bin/fm-spawn-acct.sh:43- The wrapper now deliberately forwards--account, but the existing executable wrapper-contract test still asserts exactly three forwarded arguments. Its normal invocation will receive four (T-1,/proj, launch command,--account,claude-altactually five—note the test's three-argument premise), incrementfails, and exit 1. Update that test to assert the account flag/value in addition to the composed launch.🔧 Fix: Update account spawn wrapper contract test
1 warning still open:
bin/fm-fleet-snapshot.sh:70- The newly documented snapshot contract sayscontext_pctis always"n/a", but the same change now emits a validated live Codex/Claude statusline value (for example40) when available. Update this contract text (and the matching stale library header) to describe"n/a"only when unavailable, so consumers are not instructed to ignore a required output field.🔧 Fix: Clarify live context output contract
1 error still open:
bin/fm-spawn.sh:1460- Required R5 says quota reads must be “bound to the worker's own account identity” and unavailable when it cannot be established. This hunk preservesaccount=codex-altduringfm-spawn.sh --relaunch, but relaunch deliberately clears the original raw account-isolated launch and starts the normal harness command. Concrete path: a task originally launched throughfm-spawn-acctwithCODEX_HOME=...codex-altis relaunched byfm-control.sh; the replacement uses the default Codex account while its metadata still sayscodex-alt, so the snapshot invokes quota-axi undercodex-altand silently attributes that quota to the default-account worker. Bind account metadata at the spawn boundary to a launch isolation that is actually reapplied, or drop it on generic relaunches.🔧 Fix: Drop unproven account inheritance on relaunch
3 issues (1 error, 2 warnings) still open:
bin/fm-spawn.sh:4156---accountis accepted and persisted without proving that the launch applied that account’s isolation. For example,fm-spawn.sh T /project codex --account codex-altlaunches normal/default Codex but writesaccount=codex-alt; an opt-in snapshot then runs quota-axi ascodex-altand attributes that quota to the default-account pane. This leaves required R5 (“bound to the worker's own account identity … unavailable when that identity cannot be established”) reachable. Persist an account only from a spawn path that applies/verifies its isolation, or omit/refuse direct unverified--accountattribution.bin/fm-crew-usage-lib.sh:114- The new live context read invokesfm-statusline-quota.shsynchronously for every Codex/Claude row, but has no timeout. Its backend capture can block on an unresponsive pane/backend CLI, serially stalling the canonical snapshot even with quota opt-in disabled. Run this diagnostic through a validated bounded call and emitn/aon expiry.bin/fm-fleet-snapshot.sh:760- Usage context is read from the current task-id selector after the captured-generation check. If a task relaunches between prefetch and this call, the selector reads the replacement pane while the row retains old captured harness/model metadata, silently misattributing its context percentage. Recheck generation after this read and discard usage on mismatch, or collect it inside the existing generation-guarded observation boundary.🔧 Fix: Bound live usage reads and account attribution
1 error still open:
bin/fm-spawn.sh:4156- Required criterion: “quota reads/caching are bound to the worker's own account identity, reporting unavailable when that identity cannot be established.” The new metadata guard trusts the caller-controlledFM_SPAWN_ACCT_ISOLATED=1as that proof. For example,fm-spawn-acct.sh T /proj --account codex-alt --harness codexforwards--harnessafter its isolated raw command;fm-spawn.shreplaces that raw command with normal/default Codex, but this line still recordsaccount=codex-alt. (A directFM_SPAWN_ACCT_ISOLATED=1 fm-spawn.sh ... codex --account codex-althas the same result.) With quota enabled, the snapshot reads codex-alt quota for a default-account pane. Verify/own the isolated launch contract at the spawn boundary before persistingaccount, otherwise emit unavailable.🔧 Fix: Verify isolated account launch metadata
1 error still open:
bin/fm-spawn-acct.sh:45- The required criterion says “every live pane's row carries harness, model, context percentage where available, and provider quota.” A valid account config directory containing spaces is shell-quoted byfm_account_compose_launch, then this wrapper sends it as a raw launch.fm-spawn.shsplits that quoted string on whitespace and recordsspaces'rather thanclaude/codexas the harness. The pane launches correctly, but its usage row loses the supported context/quota mapping. Preserve the verified account harness at the spawn boundary (or parse the composed command without shell word-splitting).🔧 Fix: Preserve verified account harness metadata
1 error still open:
docs/verification/fm-harness-usage-bar.md:41- Required intent says to “verify the Codex statusline proxy already wired to LIFEOS_StatusLine.sh actually renders” and capture that bar from a live pane. This capture shows only a native-lookinggpt-5.6-terra high · <cwd>footer, not any LifeOS bar fields described at lines 26–29, so it does not establish that the configured proxy rendered live.🔧 Fix: Correct live proxy verification evidence
✅ Re-checked - no issues remain.
✅ **Test** - passed
✅ No issues found.
bash tests/fm-crew-usage-lib.test.shbash tests/federation/test_spawn_account.shbash /Users/adrianb/.no-mistakes/evidence/01M1XCRT7JYY1Y6VQWHYYYXDEN/fm-usage-e2e.shwith controlled Codex statusline, account-isolated quota-axi, canonical snapshot, and Bearings projection✅ **Document** - passed
✅ No issues found.
🔧 **Lint** - 1 issue found → auto-fixed ✅
🔧 Fix: Suppress intentional inner-shell expansion lint warnings
✅ Re-checked - no issues remain.
✅ **Push** - passed
✅ No issues found.