Conversation
…#1778) Discord mentions already ride the same pairing-token opt-in, relay poll, and platform-aware reply path as X mentions, but the docs still read as X-only, so a stranger could not self-serve the Discord path. Add the numbered turn-on steps to the X mode configuration reference, pointing at the myfirstmate dashboard for account creation, bot install, and token issuance rather than duplicating operator setup here, and drop the X-only framing from the README bullet, the documentation index, and the architecture overview.
…#1781) * feat(bin): run session start deterministically on hook-capable harnesses Session start relied on a native nudge that only asked the agent to run bin/fm-session-start.sh, and an agent can defer that. Observed 2026-08-01: an /ahoy-first session followed the recap path and did not take the helm until a later request forced it. Claude, Codex, and Pi now RUN the digest in their session-open hook through the new bin/fm-sessionstart-run.sh, so the full ordered digest is in model context before the first turn. That wrapper is the single owner of what a session-open source means: startup and Pi's "new" take the helm, clear and compact re-emit, resume/reload/fork delegate to the nudge, and an unreadable source takes the helm because doing that redundantly is idempotent while skipping it is the bug. Grok and OpenCode keep the nudge as the floor, since neither can carry hook stdout into a model turn. Because the hook now blocks session initialization, fm-session-start.sh bounds itself first. Its steps are not all individually bounded - bootstrap's gh auth probe, tool version probes, the backlog listing and per-task endpoint reads are unbounded - so the whole digest runs as one bounded child (default 120s). Whatever it emitted before the bound survives, and the parent adds a loud STARTUP TRUNCATED banner naming the stage that stalled and every stage that never ran, still exiting 0. --reemit skips only the sweeps startup already reconciled. It still re-verifies lock ownership and still drains queued wakes, which arrived after startup and are the turn's work. fm-bootstrap.sh gains FM_BOOTSTRAP_LOCKED so a re-emit keeps repair ownership instead of deferring to a lock holder that is itself. Also adds bin/fm-timeout-lib.sh as the single owner of bounded execution, replacing three near-identical copies, and gives the ahoy skill a helm check so a nudge-tier harness cannot recap before taking the helm. Verified live on 2026-08-05 against Claude 2.1.222, Codex 0.146.0, and Pi 0.82.0; docs/verification/supervision.md records the per-harness source vocabulary, the two named gaps, and the refresh command. * no-mistakes(review): Harden session-start completion, timeout, and Pi delivery * no-mistakes(review): Harden completion ownership and portable timeout escalation * no-mistakes(review): Normalize watchdog KILL exits without masking command status * no-mistakes(review): Guarantee startup bounds and align harness delivery tiers * no-mistakes(test): Fix Pi session-start live verification fixture * no-mistakes(document): Align session-start documentation with deterministic hooks * no-mistakes(lint): Silence intentional child-shell expansion lint warning * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes
* docs: rename the user-facing product name to Relay The public-mention integration gated by the `.env` pairing token is now called Relay across user-facing prose, covering X and Discord alike instead of implying a single network. Renames the product-name strings only: README, docs, the captain-facing skill descriptions, and the AGENTS.md operating prose, including the `X mode (.env)` and `Optional X mode` headings and every link anchor that pointed at them. AGENTS.md section 14 carries a one-line bridge note so the older name and the unchanged identifier spellings stay discoverable. Internal identifiers are untouched: `FMX_*`, `config/x-mode.env`, `state/x-*`, `bin/fm-x-*`, the `fmx-respond` skill path, `__FM_X_MODE_ENV__`, and `x-mode-error`. Platform references to X and Discord as networks stay as they are, and the bootstrap-diagnostics entry still quotes bootstrap's emitted `FMX: X mode on/off` line verbatim because `bin/` output is out of scope for this pass. * no-mistakes(review): Complete Relay prose rename in maintained docs * no-mistakes: apply CI fixes
* feat(harness): add a verified muse crewmate adapter Muse Code joins the fleet as a crewmate/scout adapter, verified live against Muse Code 0.1.0-R708.1 in an isolated lab. Detection matches the anchored prefix muse-bin*, because the installed launcher execs a version-suffixed binary whose name changes on every auto-update and whose install path carries no muse component to fall back on. The same identity is taught to the tmux liveness classifier, without which a healthy muse pane would have read as a dead endpoint. Busy state folds muse's own durable session event log, bound per task by a sessions-root/worktree sidecar. It is a pull source with no writer, so nothing is armed and no record is ever seeded. The fold is anchored on the full run lifecycle prefix so muse's nested cleanup "terminal" payloads cannot settle an in-flight run, and it is depth-bounded so muse's native sub-agent logs cannot be mistaken for the parent's. The idle half stays gated: an open run proves busy, but a settled log reads unknown until a credentialed multi-step run proves one turn stays inside one run. Two findings corrected the scout report. The exec-only --no-foreign-personal-context flag is rejected by the interactive TUI, so the privacy control that actually reaches a pane worker is MUSE_EXPERIMENTAL_FOREIGN_PERSONAL_CONTEXT_KILL, verified to drop the operator's foreign personal rules while keeping the project's own AGENTS.md. And an unauthenticated muse pane never exits, it waits on a device-code prompt, so credentials are a spawn preflight rather than a screen check. muse is refused for secondmates: it has no primary supervision protocol and its hook dialect rejects the reawakening handlers that protocol needs. Per the captain's decision, auto-update is not pinned, and the credentialed multi-step smoke is deferred with an explicit checklist in docs/verification/muse.md. * no-mistakes(review): Accept Muse dispatch profiles and shared efforts * no-mistakes(review): Bind Muse busy state to current session * no-mistakes(review): Compare Muse workspace bindings literally * no-mistakes(review): Harden Muse worker credentials and live signal verification * no-mistakes(review): Cache Muse session bindings and clarify worker credentials * no-mistakes(review): Clear Muse marker inheritance and normalize interrupt aliases * no-mistakes(review): Verify Muse glyph effective foreground color * no-mistakes(review): Harden Muse XDG paths, session cache, and glyph parsing * no-mistakes(document): Document Muse adapter boundaries
…d#1787) * fix(herdr): floor default-on presentation spaces at Herdr 0.8.0 Default-on presentation projection turns every crewmate teardown into a workspace-emptying removal. The focus-safe removal plan avoids Herdr's focus-stealing explicit close only while the doomed pane's shell can be proved lone, childless, and idle; a persistent child of that shell (gitstatusd, a zsh-async worker, direnv) fails that proof permanently and forces the plain close, which on every release before Herdr 0.8.0 moves the captain's active workspace for ~140ms on each teardown. Gate the unconfigured default behind a Herdr 0.8.0 floor. At or above it, project as before; below it, fall back to the flat per-home layout with one warning per home per detected release naming the version and the upgrade. An explicit "on" - including the historical empty opt-in file - is still honored below the floor, so a deliberate opt-in is never silently downgraded. The floor reads two independent signals from the client's own status, either of which can establish a supported release: the protocol number and the release core of the version string. Measured against the real release binaries, no build lacking both upstream focus fixes reaches protocol 19 and every pre-fix build tops out at 17, so protocol 19 is a safe structural expression of the floor. A release that reports neither signal readably is treated as unsupported rather than guessed at. Also: - Correct the adapter comment claiming the mitigation "stays safe without any version gate". That holds for the pane-death route only; the plain-close fallback is reachable precisely on the releases where it is unsafe. - Stop discarding the projected-close helper's stderr at teardown, so a refused or failed focus restore is visible instead of silent. The close stays non-fatal; the presence gate still decides record removal. - Add Part C to the focus-flash regression: a doomed pane whose shell holds a persistent child, in the geometry where the closing workspace's right neighbour is not the anchor. That is the fallback branch the suite could not structurally reach. On 0.7.5 it observes a bounded four-sample wrong-focus window restored exactly; on 0.8.0 it observes none. It also cross-checks its own measurement against the floor classifier, so a drifted protocol mapping fails loudly. - Make the projection suite's unconfigured-home case release-aware, so the whole real-Herdr lane passes on both the CI-pinned 0.7.4 and 0.8.0. - Add an opt-in live guard that re-measures the release-to-protocol mapping against the pinned upstream binaries. The immediate no-code mitigation for a home that cannot upgrade remains writing "off" into config/herdr-presentation-spaces. * no-mistakes(review): Pin Herdr live-guard digests across supported platforms * no-mistakes(review): Document authorized Herdr cleanup containment * no-mistakes(review): Harden Herdr warning marker publication * no-mistakes(review): Honor running Herdr server presentation floor * no-mistakes(review): Recheck Herdr floor after server ensure * no-mistakes(review): Refresh 0.7.5 and 0.8.0 focus transcripts * no-mistakes(review): Route Herdr floor probe through lab session * no-mistakes(document): Align Herdr floor documentation and comments * no-mistakes(lint): Document Herdr presentation out-parameter consumer
* fix(muse): trust the settled session log as idle The credentialed multi-step smoke on Muse Code 0.1.0-R708.1 answered the one question the idle half was held back for: one real 75-second tool-loop turn with 23 tool batches stays inside exactly one run started/terminal pair, and an Escape mid tool loop closes that run as cancelled rather than leaving the turn to continue in another run. A settled log is therefore a finished turn, not a pause between the runs of one turn. Remove fm_busy_muse_idle_verified and FM_BUSY_MUSE_IDLE_VERIFIED_VERSIONS outright rather than pinning them to a version: the session log's own metadata carries only semver 0.1.0 and a build sha, so a version allowlist could not actually match the running build and would be false precision. A settled log now classifies idle, an open run still classifies busy, and only a resolution failure - no binding, no matching log, an unreadable or run-free log - stays unknown. Record the evidence in docs/verification/muse.md, including the run-scoped grep the counts must use, and keep the post-upgrade re-check guidance. * no-mistakes(review): Document Muse idle trust and remove stale gate reference * no-mistakes(document): Clarify Muse idle verification ownership
📝 WalkthroughWalkthroughThe change adds routed and bounded session-start execution, verified Muse harness support, shared timeout handling, and Herdr presentation capability gating. It also updates Relay terminology, operational documentation, verification records, and regression coverage. ChangesRuntime behavior
Documentation and contracts
Estimated code review effort: 5 (Critical) | ~120 minutes Possibly related PRs
Suggested reviewers: 🚥 Pre-merge checks | ✅ 4 | ❌ 1❌ Failed checks (1 warning)
✅ Passed checks (4 passed)
✨ Finishing Touches 💡 1📝 Generate docstrings 💡
🧪 Generate unit tests (beta)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
…ion inventory duplicates
There was a problem hiding this comment.
Actionable comments posted: 15
🧹 Nitpick comments (7)
docs/verification/runtime-backends.md (1)
346-346: 🗄️ Data Integrity & Integration | 🔵 Trivial | 💤 Low valueClarify the fixture that produced the absent-file default result.
a home that configured nothing is projected by defaultfollows the empty-file opt-in claim, but line 357 should identify which fixture run exercised noconfig/herdr-presentation-spaces. The release-floor section depends on that no-file behavior.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@docs/verification/runtime-backends.md` at line 346, Update the projection suite documentation to identify the specific fixture run that had no config/herdr-presentation-spaces file and produced the absent-file default result. Distinguish this no-file fixture from the empty-file opt-in case, while preserving the documented release-floor behavior under “Presentation version floor.”bin/fm-sessionstart-run.sh (1)
101-116: 🩺 Stability & Availability | 🔵 Trivial | ⚡ Quick win
execbreaks the documented exit-0 invariant on the resume path.The header states that every path exits 0.
exectransfers the exit status offm-sessionstart-nudge.shto the harness, and an exec failure returns 126 or 127. Call the wrapper like the other branches so the finalexit 0still applies.♻️ Proposed change
resume|reload|fork) - exec "$SCRIPT_DIR/fm-sessionstart-nudge.sh" + "$SCRIPT_DIR/fm-sessionstart-nudge.sh" || true ;;🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@bin/fm-sessionstart-run.sh` around lines 101 - 116, Replace the exec invocation in the resume|reload|fork branch of the SOURCE case with a normal wrapper call, matching the other branches, so control reaches the final exit 0 even if fm-sessionstart-nudge.sh fails or cannot be executed.bin/fm-spawn.sh (1)
2144-2154: 🗄️ Data Integrity & Integration | 🔵 Trivial | ⚡ Quick winWrite the Muse binding sidecar atomically.
The block streams directly into
$STATE/$ID.muse-session. A classifier that reads the file during that write can observesessions_root,workspace_root, andbinding_idwithout the trailingprior_log=lines.fm_busy_muse_session_logthen treats a pre-existing log as the one new log, andfm_busy_muse_cache_session_logpersists that wrong selection until the complete sidecar invalidates it.fm_busy_muse_cache_session_logalready uses a temp file plusmv -f; use the same shape here.♻️ Proposed atomic write
MUSE_SESSIONS_ROOT="${MUSE_DATA_HOME:-${XDG_DATA_HOME:-$HOME/.local/share}}/muse/sessions" MUSE_BINDING_ID="$$.$RANDOM.$(date +%s)" rm -f "$STATE/$ID.muse-session-current" { printf 'sessions_root=%s\n' "$MUSE_SESSIONS_ROOT" printf 'workspace_root=%s\n' "$WT" printf 'binding_id=%s\n' "$MUSE_BINDING_ID" while IFS= read -r MUSE_PRIOR_LOG; do [ -n "$MUSE_PRIOR_LOG" ] && printf 'prior_log=%s\n' "$MUSE_PRIOR_LOG" done <<EOF $(fm_busy_muse_matching_logs "$MUSE_SESSIONS_ROOT" "$WT" || true) EOF - } > "$STATE/$ID.muse-session" + } > "$STATE/$ID.muse-session.tmp.$$" + mv -f -- "$STATE/$ID.muse-session.tmp.$$" "$STATE/$ID.muse-session" ;;🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@bin/fm-spawn.sh` around lines 2144 - 2154, Write the Muse binding sidecar through a temporary file in the same state directory, preserving the existing sessions_root, workspace_root, binding_id, and prior_log output, then atomically replace $STATE/$ID.muse-session with mv -f after the complete block finishes. Follow the temp-file and replacement pattern used by fm_busy_muse_cache_session_log, and ensure the temporary file is cleaned up if writing fails.tests/fm-session-start.test.sh (2)
1245-1273: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick winUnguarded perl dependency in the new timeout tests. Both new cases execute perl programs without checking that perl exists. A missing interpreter yields exit 126 or 127, and the
expect_code 124andexpect_code 137assertions then report a timeout-contract regression that did not occur.
tests/fm-session-start.test.sh#L1245-L1273: guardmake_term_escalating_timeoutwithcommand -v perl, or skip the case when perl is absent.tests/fm-session-start.test.sh#L1332-L1341: apply the same precondition check before the inline perl watchdog and the perl TERM-resistant victim.🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@tests/fm-session-start.test.sh` around lines 1245 - 1273, Guard the Perl-dependent timeout tests with a command -v perl precondition so they are skipped when Perl is unavailable rather than producing misleading timeout-contract failures. Apply this to make_term_escalating_timeout at tests/fm-session-start.test.sh lines 1245-1273 and to the inline Perl watchdog and TERM-resistant victim at lines 1332-1341; both sites require the same absence-of-Perl handling.
1305-1309: 📐 Maintainability & Code Quality | 🔵 Trivial | ⚡ Quick winDetect the leaked leaf process, not only the wrapper.
pgrep -f "$fakebin/git"matches the fixture script process. The fixture's real leaf issleep 600, whose command line does not contain the fixture path. If the bound terminates the script but not the process group, the orphanedsleepsurvives and this assertion still passes, so the "kills its hung grandchild" claim is unproven.Record the process group and assert it is empty, or make the leaf identifiable.
♻️ One way to make the leaf identifiable
make_hanging_tool() { local fakebin=$1 name=$2 cat > "$fakebin/$name" <<'SH' #!/usr/bin/env bash trap '' TERM -sleep 600 +# A distinct argv keeps the leaf greppable after the wrapper dies. +exec sleep 600.5 SH chmod +x "$fakebin/$name" }- stray=$(pgrep -f "$fakebin/git" 2>/dev/null | wc -l | tr -d ' ') + stray=$(pgrep -f -e "$fakebin/git" -e 'sleep 600\.5' 2>/dev/null | wc -l | tr -d ' ')🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@tests/fm-session-start.test.sh` around lines 1305 - 1309, Update the leaked-process assertion in the process-group timeout test to detect the actual hung leaf, not just the fixture wrapper matched by "$fakebin/git". Make the leaf identifiable or capture its process group and assert that group is empty, while preserving the existing failure when any descendant remains.tests/fm-remote-reply.test.sh (1)
21-28: 🩺 Stability & Availability | 🔵 Trivial | ⚡ Quick winEscalate to SIGKILL when the worker outlives the wait loop.
The loop gives the worker 5 seconds. If the worker ignores TERM or blocks, cleanup deletes
TMP_ROOTwhile the worker still runs. The surviving worker can then recreate state or keep a procevent claim, which leaks into later tests.♻️ Proposed escalation
while kill -0 "$worker_pid" 2>/dev/null && [ "$wait_attempt" -lt 100 ]; do wait_attempt=$((wait_attempt + 1)) sleep 0.05 done + if kill -0 "$worker_pid" 2>/dev/null; then + kill -KILL "$worker_pid" 2>/dev/null || true + wait "$worker_pid" 2>/dev/null || true + fi fi🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@tests/fm-remote-reply.test.sh` around lines 21 - 28, Update the worker cleanup block around worker_pid and the wait_attempt loop to send SIGKILL when the worker is still alive after all 100 wait attempts. Preserve the existing graceful termination first, and only force-kill the surviving worker before cleanup continues.tests/fm-backend-herdr.test.sh (1)
1161-1178: 📐 Maintainability & Code Quality | 🔵 Trivial | 💤 Low valueAdd a case-insensitive value to the preference test.
test_presentation_explicit_off_opts_outpinsOFFandOffat the gate.test_presentation_preference_reports_three_distinct_statesonly exercises lowercase values, so a regression in case folding foronwould not be caught at the parser level.♻️ Proposed additional case
printf 'on\n' > "$config/herdr-presentation-spaces" got=$(preference "$config") [ "$got" = on ] || fail "an explicit on must report on, got '$got'" + printf ' ON \n' > "$config/herdr-presentation-spaces" + got=$(preference "$config") + [ "$got" = on ] || fail "an uppercase padded on must report on, got '$got'" printf 'off\n' > "$config/herdr-presentation-spaces"🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the rest with a brief reason, keep changes minimal, and validate. In `@tests/fm-backend-herdr.test.sh` around lines 1161 - 1178, Extend test_presentation_preference_reports_three_distinct_states to write an uppercase or mixed-case ON value and assert that preference returns on, covering case-insensitive parsing alongside the existing lowercase on, off, and default cases.
🤖 Prompt for all review comments with AI agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
Inline comments:
In @.agents/skills/harness-adapters/SKILL.md:
- Line 415: Update the environment-marker table row to explicitly document both
matched process-name forms, bare `muse` and the anchored `muse-bin-*` prefix,
while preserving the existing process-ancestry detection and marker-clearing
details.
In @.pi/extensions/fm-primary-turnend-guard.ts:
- Around line 170-176: Update the source mapping in the session_start handler to
recognize the reload reason as the "reload" source, so reload events proceed to
injectSessionstart(pi, source) instead of returning early. Preserve the existing
startup, clear, resume, and fork mappings.
In `@bin/fm-busy-lib.sh`:
- Around line 404-423: Update fm_busy_muse_main_log_path_valid to validate that
each rel split consumes an actual slash before assigning year, month, day,
session, and leaf. Reject paths with fewer than four directory components, such
as root/x/session.jsonl, while preserving acceptance of the required
year/month/day/session/session.jsonl layout.
- Around line 323-327: Document Node as a prerequisite for
fm_busy_muse_matching_logs, either in the existing Fleet tools/prerequisites
documentation or in the source contract comment immediately above the function.
Clearly state that Muse busy classification requires the node executable,
without changing the function’s behavior.
In `@bin/fm-session-start.sh`:
- Around line 177-182: Track whether mktemp successfully created a real stage
file before assigning the /dev/null fallback, using a distinct cleanup indicator
near SESSION_START_STAGE_FILE initialization. Update the cleanup at line 207 to
remove the stage path only when that indicator confirms a temporary file was
created, never attempting to remove /dev/null.
In `@bin/fm-spawn.sh`:
- Around line 1142-1143: Update the XDG initialization around
resolve_directory_input for MUSE_CONFIG_HOME and MUSE_DATA_HOME so missing
default roots are created or resolved without requiring pre-existing
directories. Preserve validation for explicitly invalid paths while allowing
absent ~/.config and ~/.local/share defaults to proceed and letting the spawn
launch.
In `@bin/fm-test-run.sh`:
- Line 141: Update the changed-path family mappings consumed by select_changed
so the pure-contract-unit tests run for their corresponding sources: map
bin/fm-teardown.sh to fm-muse-harness.test.sh and map
.pi/extensions/fm-primary-turnend-guard.ts to fm-calm-pi-extension.test.sh,
either by adding pure-contract-unit to the exact source cases or by adding
explicit __script__ entries. Preserve the existing pr-forge, session-bootstrap,
and live-harness-optin mappings.
In `@bin/fm-timeout-lib.sh`:
- Line 47: Update both temporary-file creation sites in bin/fm-timeout-lib.sh:
the command-status file path at lines 47-47 and the external-runner status file
path at lines 92-92 must return the same non-124 error when mktemp fails.
Preserve 124 exclusively for actual deadline expiration so callers do not
classify temporary-file failures as timeouts.
- Around line 118-121: Update fm_run_timed to validate seconds before calling
fm_timeout_mechanism or dispatching to any timeout implementation. Reject zero
and negative timeout values, returning a non-timeout exit status; preserve the
existing mechanism behavior for positive values.
In `@docs/verification/muse.md`:
- Line 109: Replace the concrete authorization code in the authentication URL in
muse.md with a clearly non-functional placeholder, and invalidate the device
authorization flow if the exposed code was real. Remove the credential from
repository history before publication.
- Line 17: Update every fenced code block in muse.md, including the listed
locations, with an explicit language identifier: use sh for command-only blocks
and text for captured output or mixed command/output blocks, without changing
their contents.
In `@docs/verification/supervision.md`:
- Line 123: Clarify the semantic-busy probe sentence by naming the specific
Codex lifecycle event names tested for the Firstmate-written hooks under
<worktree>/.codex/hooks.json. Reconcile the statement about codex exec with the
documented SessionStart behavior, distinguishing which event did or did not fire
while preserving the existing trust and global-probe details.
In `@tests/fm-backend-herdr-presentation-e2e.test.sh`:
- Around line 508-524: Both live tests must predict the presentation gate using
the composed client/server classifier. In
tests/fm-backend-herdr-presentation-e2e.test.sh:508-524, replace the
server-or-client release selection and fm_backend_herdr_release_floor_verdict
call with fm_backend_herdr_presentation_release_supported "$HERDR_LAB_SESSION",
and grep the warning for FM_BACKEND_HERDR_PRESENTATION_RELEASE. In
tests/fm-backend-herdr-focus-flash-e2e.test.sh:341-395, use the same composed
classifier and its reported release for the naming assertion, while retaining
the client-only classifier at lines 354-364 for Part A focus-behaviour
correlation.
In `@tests/fm-muse-harness.test.sh`:
- Around line 164-190: Unset FM_PI_HARNESS in the environment for both mock bash
probes within test_detects_versioned_process_ancestor and
test_detection_is_anchored, alongside the existing agent-marker unsets, so muse
ancestry detection is tested without inherited Pi markers.
In `@tests/fm-sessionstart-nudge.test.sh`:
- Around line 282-284: Update the cp commands in
tests/fm-sessionstart-nudge.test.sh:282-284 and
tests/fm-calm-pi-extension.test.sh:2864-2870 to include
"$ROOT/bin/fm-session-lock-lib.sh" alongside the existing session-start scripts,
ensuring both run-wrapper fixtures copy the library sourced by
fm-sessionstart-run.sh.
---
Nitpick comments:
In `@bin/fm-sessionstart-run.sh`:
- Around line 101-116: Replace the exec invocation in the resume|reload|fork
branch of the SOURCE case with a normal wrapper call, matching the other
branches, so control reaches the final exit 0 even if fm-sessionstart-nudge.sh
fails or cannot be executed.
In `@bin/fm-spawn.sh`:
- Around line 2144-2154: Write the Muse binding sidecar through a temporary file
in the same state directory, preserving the existing sessions_root,
workspace_root, binding_id, and prior_log output, then atomically replace
$STATE/$ID.muse-session with mv -f after the complete block finishes. Follow the
temp-file and replacement pattern used by fm_busy_muse_cache_session_log, and
ensure the temporary file is cleaned up if writing fails.
In `@docs/verification/runtime-backends.md`:
- Line 346: Update the projection suite documentation to identify the specific
fixture run that had no config/herdr-presentation-spaces file and produced the
absent-file default result. Distinguish this no-file fixture from the empty-file
opt-in case, while preserving the documented release-floor behavior under
“Presentation version floor.”
In `@tests/fm-backend-herdr.test.sh`:
- Around line 1161-1178: Extend
test_presentation_preference_reports_three_distinct_states to write an uppercase
or mixed-case ON value and assert that preference returns on, covering
case-insensitive parsing alongside the existing lowercase on, off, and default
cases.
In `@tests/fm-remote-reply.test.sh`:
- Around line 21-28: Update the worker cleanup block around worker_pid and the
wait_attempt loop to send SIGKILL when the worker is still alive after all 100
wait attempts. Preserve the existing graceful termination first, and only
force-kill the surviving worker before cleanup continues.
In `@tests/fm-session-start.test.sh`:
- Around line 1245-1273: Guard the Perl-dependent timeout tests with a command
-v perl precondition so they are skipped when Perl is unavailable rather than
producing misleading timeout-contract failures. Apply this to
make_term_escalating_timeout at tests/fm-session-start.test.sh lines 1245-1273
and to the inline Perl watchdog and TERM-resistant victim at lines 1332-1341;
both sites require the same absence-of-Perl handling.
- Around line 1305-1309: Update the leaked-process assertion in the
process-group timeout test to detect the actual hung leaf, not just the fixture
wrapper matched by "$fakebin/git". Make the leaf identifiable or capture its
process group and assert that group is empty, while preserving the existing
failure when any descendant remains.
🪄 Autofix
Fix all unresolved CodeRabbit comments on this PR:
- Push a commit to this branch (recommended)
- Create a new PR with the fixes
ℹ️ Review info
⚙️ Run configuration
Configuration used: defaults
Review profile: CHILL
Plan: Pro Plus
Run ID: 20a64d92-7397-4b1a-86f0-7e6ee4ace40c
📒 Files selected for processing (64)
.agents/skills/ahoy/SKILL.md.agents/skills/bootstrap-diagnostics/SKILL.md.agents/skills/fmx-respond/SKILL.md.agents/skills/harness-adapters/SKILL.md.claude/settings.json.codex/hooks.json.pi/extensions/fm-primary-turnend-guard.tsAGENTS.mdREADME.mdbin/backends/cmux.shbin/backends/herdr.shbin/backends/tmux.shbin/backends/zellij.shbin/fm-bearings-snapshot.shbin/fm-bootstrap.shbin/fm-busy-lib.shbin/fm-composer-lib.shbin/fm-config-inherit-lib.shbin/fm-fleet-snapshot.shbin/fm-harness.shbin/fm-herdr-session-cleanup.shbin/fm-send.shbin/fm-session-start.shbin/fm-sessionstart-run.shbin/fm-spawn.shbin/fm-teardown.shbin/fm-test-run.shbin/fm-timeout-lib.shbin/fm-vendor-auth-probe.shdocs/architecture.mddocs/configuration.mddocs/documentation-audiences.jsondocs/fork-features.mddocs/herdr-backend.mddocs/scripts.mddocs/sessionstart-nudge.mddocs/subagent-guard.mddocs/supervision-protocols/codex.mddocs/supervision-protocols/grok.mddocs/tmux-backend.mddocs/trace-context.mddocs/turnend-guard.mddocs/verification/muse.mddocs/verification/public-followup.mddocs/verification/runtime-backends.mddocs/verification/supervision.mdtests/fm-backend-herdr-focus-flash-e2e.test.shtests/fm-backend-herdr-presentation-e2e.test.shtests/fm-backend-herdr.test.shtests/fm-bootstrap.test.shtests/fm-calm-pi-extension.test.shtests/fm-composer-ghost.test.shtests/fm-composer-lib.test.shtests/fm-harness-liveness-drift-live-e2e.test.shtests/fm-herdr-version-floor-live-e2e.test.shtests/fm-muse-harness.test.shtests/fm-muse-signals-live-e2e.test.shtests/fm-remote-reply.test.shtests/fm-secondmate-harness.test.shtests/fm-session-start.test.shtests/fm-sessionstart-hook-live-e2e.test.shtests/fm-sessionstart-nudge.test.shtests/fm-teardown.test.shtests/fm-tmux-agent-liveness.test.sh
| | Skill invocation | `/<skill>`, the claude/grok form. | | ||
| | Autonomy | `--yolo`, which disables approval, disables the sandbox, and trusts the workspace for the run. | | ||
| | Trust dialog | `Do you trust this workspace?` with `1 Trust and continue` preselected, accepted by Enter. `--yolo` suppresses it entirely, which is what firstmate relies on because every task gets a fresh worktree path. | | ||
| | Environment marker | None. Detection is process ancestry on the anchored prefix `muse-bin-*`. The launch clears foreign primary markers before Muse starts so their higher detection precedence cannot override that ancestry. `MUSE_CURRENT_SESSION_LOG` is a session-log PATH rather than an identity, and its export to tool subprocesses is unverified. | |
There was a problem hiding this comment.
📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win
Record both matched process names.
The row states detection is process ancestry on the anchored prefix muse-bin-*. bin/fm-harness.sh matches muse|muse-bin-*, and bin/backends/tmux.sh classifies the same two forms. Name the bare launcher too, so a later change cannot drop the muse case as undocumented.
📝 Proposed wording fix
-| Environment marker | None. Detection is process ancestry on the anchored prefix `muse-bin-*`. The launch clears foreign primary markers before Muse starts so their higher detection precedence cannot override that ancestry. `MUSE_CURRENT_SESSION_LOG` is a session-log PATH rather than an identity, and its export to tool subprocesses is unverified. |
+| Environment marker | None. Detection is anchored process ancestry on the launcher name `muse` or the versioned prefix `muse-bin-*`. The launch clears foreign primary markers before Muse starts so their higher detection precedence cannot override that ancestry. `MUSE_CURRENT_SESSION_LOG` is a session-log PATH rather than an identity, and its export to tool subprocesses is unverified. |📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| | Environment marker | None. Detection is process ancestry on the anchored prefix `muse-bin-*`. The launch clears foreign primary markers before Muse starts so their higher detection precedence cannot override that ancestry. `MUSE_CURRENT_SESSION_LOG` is a session-log PATH rather than an identity, and its export to tool subprocesses is unverified. | | |
| | Environment marker | None. Detection is anchored process ancestry on the launcher name `muse` or the versioned prefix `muse-bin-*`. The launch clears foreign primary markers before Muse starts so their higher detection precedence cannot override that ancestry. `MUSE_CURRENT_SESSION_LOG` is a session-log PATH rather than an identity, and its export to tool subprocesses is unverified. | |
🧰 Tools
🪛 SkillSpector (2.5.1)
[error] 340: [AS1] Agent Config Directory Access: Skill reads from agent configuration directories (.claude/, .codex/, .gemini/). These directories may contain API keys, personal settings, and other credentials that the skill has no legitimate need to access.
Remediation: Remove all code or instructions that access agent configuration directories (.claude/, .codex/, .gemini/). If configuration values are needed, pass them explicitly as parameters or environment variables — never read the agent's own config files.
(Agent Snooping (AS1))
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In @.agents/skills/harness-adapters/SKILL.md at line 415, Update the
environment-marker table row to explicitly document both matched process-name
forms, bare `muse` and the anchored `muse-bin-*` prefix, while preserving the
existing process-ancestry detection and marker-clearing details.
| pi.on?.("session_start", async (event) => { | ||
| const reason = String((event as { reason?: unknown }).reason ?? ""); | ||
| const nudge = ["startup", "new", "resume"].includes(reason) ? runSessionstartNudge() : ""; | ||
| const source = { startup: "startup", new: "clear", resume: "resume", fork: "fork" }[reason]; | ||
| markLoaded(); | ||
| if (!nudge) return; | ||
| try { | ||
| pi.sendMessage({ | ||
| customType: "firstmate-sessionstart-nudge", | ||
| content: nudge, | ||
| display: false, | ||
| details: { kind: "session-start" }, | ||
| }); | ||
| } catch { | ||
| } | ||
| if (!source) return; | ||
| await injectSessionstart(pi, source); | ||
| }); |
There was a problem hiding this comment.
🎯 Functional Correctness | 🟠 Major | ⚡ Quick win
🧩 Analysis chain
🏁 Script executed:
#!/bin/bash
# Confirm every documented Pi session_start reason has an adapter mapping.
rg -n -C4 'session_start|session_compact' .pi/extensions/ docs/sessionstart-nudge.md
rg -n -C3 'resume\|reload\|fork|reload' bin/fm-sessionstart-run.sh docs/sessionstart-nudge.mdRepository: trillium/firstmate
Length of output: 16217
Treat Pi reload as a session-start source.
The extension docs and .pi/examples/fm-calm.ts evidence Pi emits reload, and bin/fm-sessionstart-run.sh routes reload to the nudge wrapper. Since source is empty for reload in the current map, Pi reloads return before injectSessionstart(pi, "reload"). Add reload as a recognized Pi source.
🧰 Tools
🪛 ast-grep (0.45.0)
[warning] Importing child_process exposes a command-execution surface; ensure any command/argument built from input is validated, and prefer execFile/spawn with an argument array over exec.
Context: import { spawn, spawnSync } from "node:child_process";
Note: [CWE-78] Improper Neutralization of Special Elements used in an OS Command ('OS Command Injection').
(detect-child-process-typescript)
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In @.pi/extensions/fm-primary-turnend-guard.ts around lines 170 - 176, Update
the source mapping in the session_start handler to recognize the reload reason
as the "reload" source, so reload events proceed to injectSessionstart(pi,
source) instead of returning early. Preserve the existing startup, clear,
resume, and fork mappings.
| fm_busy_muse_matching_logs() { # <sessions-root> <workspace-root> | ||
| local root=$1 ws=$2 | ||
| [ -d "$root" ] || return 1 | ||
| command -v node >/dev/null 2>&1 || return 1 | ||
| node - "$root" "$ws" <<'NODE' |
There was a problem hiding this comment.
📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win
🧩 Analysis chain
🏁 Script executed:
#!/bin/bash
# Check whether node is already a declared prerequisite for the fleet scripts.
rg -n 'command -v node|require\("fs"\)|node -' bin | head -50
fd -t f 'README.md|AGENTS.md' -d 2 | xargs rg -n -i 'node|prerequisit|dependenc' | head -40Repository: trillium/firstmate
Length of output: 2037
🏁 Script executed:
#!/bin/bash
set -euo pipefail
printf '--- relevant pretool checks ---\n'
sed -n '120,165p' bin/fm-cd-pretool-check.sh
printf '\n--- fm-arm-pretool-check node context ---\n'
sed -n '130,158p' bin/fm-arm-pretool-check.sh
printf '\n--- fm-busy-lib busy source context ---\n'
sed -n '300,340p' bin/fm-busy-lib.sh
printf '\n--- node usage summary in bin ---\n'
rg -n 'command -v node|node ' bin | wc -l
rg -n 'command -v node|node ' bin
printf '\n--- README dependency section ---\n'
sed -n '45,75p' README.mdRepository: trillium/firstmate
Length of output: 247
🏁 Script executed:
#!/bin/bash
set -euo pipefail
printf '%s\n' '--- relevant pretool checks ---'
sed -n '120,165p' bin/fm-cd-pretool-check.sh
printf '%s\n' ''
printf '%s\n' '--- fm-arm-pretool-check node context ---'
sed -n '130,158p' bin/fm-arm-pretool-check.sh
printf '%s\n' ''
printf '%s\n' '--- fm-busy-lib busy source context ---'
sed -n '300,340p' bin/fm-busy-lib.sh
printf '%s\n' ''
printf '%s\n' '--- node usage summary in bin ---'
rg -n 'command -v node|node ' bin 2>/dev/null || true
printf '%s\n' ''
printf '%s\n' '--- README dependency section ---'
sed -n '45,75p' README.mdRepository: trillium/firstmate
Length of output: 9793
Document Node as a prerequisite for busy classification.
fm_busy_muse_matching_logs depends on node, so Muse busy checks return unknown muse-session-log when Node is missing. Add Node to the documented Fleet tools/prerequisites, or spell out that dependency in the source contract comment above.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@bin/fm-busy-lib.sh` around lines 323 - 327, Document Node as a prerequisite
for fm_busy_muse_matching_logs, either in the existing Fleet tools/prerequisites
documentation or in the source contract comment immediately above the function.
Clearly state that Muse busy classification requires the node executable,
without changing the function’s behavior.
| fm_busy_muse_main_log_path_valid() { # <sessions-root> <session-log> | ||
| local root=${1%/} log=$2 rel year month day session leaf | ||
| while :; do | ||
| case "$root" in | ||
| *'//'*) root=${root//\/\//\/} ;; | ||
| *) break ;; | ||
| esac | ||
| done | ||
| [ -n "$root" ] && [ -f "$log" ] && [ ! -L "$log" ] || return 1 | ||
| case "$log" in | ||
| "$root"/*) rel=${log#"$root"/} ;; | ||
| *) return 1 ;; | ||
| esac | ||
| year=${rel%%/*}; rel=${rel#*/} | ||
| month=${rel%%/*}; rel=${rel#*/} | ||
| day=${rel%%/*}; rel=${rel#*/} | ||
| session=${rel%%/*}; leaf=${rel#*/} | ||
| [ -n "$year" ] && [ -n "$month" ] && [ -n "$day" ] && [ -n "$session" ] \ | ||
| && [ "$leaf" = session.jsonl ] | ||
| } |
There was a problem hiding this comment.
🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win
The path-shape check accepts fewer than four components.
${rel#*/} returns the string unchanged when it holds no /. For log="$root/x/session.jsonl" the splits yield year=x, then month, day, session, and leaf all equal session.jsonl, so every emptiness test passes and leaf = session.jsonl matches. The function reports a valid main-log layout for a two-component path. Test the remainder at each step.
🐛 Proposed fix
case "$log" in
"$root"/*) rel=${log#"$root"/} ;;
*) return 1 ;;
esac
- year=${rel%%/*}; rel=${rel#*/}
- month=${rel%%/*}; rel=${rel#*/}
- day=${rel%%/*}; rel=${rel#*/}
- session=${rel%%/*}; leaf=${rel#*/}
+ case "$rel" in */*/*/*/*) return 1 ;; esac
+ year=${rel%%/*}; case "$rel" in */*) rel=${rel#*/} ;; *) return 1 ;; esac
+ month=${rel%%/*}; case "$rel" in */*) rel=${rel#*/} ;; *) return 1 ;; esac
+ day=${rel%%/*}; case "$rel" in */*) rel=${rel#*/} ;; *) return 1 ;; esac
+ session=${rel%%/*}; case "$rel" in */*) leaf=${rel#*/} ;; *) return 1 ;; esac
[ -n "$year" ] && [ -n "$month" ] && [ -n "$day" ] && [ -n "$session" ] \
&& [ "$leaf" = session.jsonl ]📝 Committable suggestion
‼️ IMPORTANT
Carefully review the code before committing. Ensure that it accurately replaces the highlighted code, contains no missing lines, and has no issues with indentation. Thoroughly test & benchmark the code to ensure it meets the requirements.
| fm_busy_muse_main_log_path_valid() { # <sessions-root> <session-log> | |
| local root=${1%/} log=$2 rel year month day session leaf | |
| while :; do | |
| case "$root" in | |
| *'//'*) root=${root//\/\//\/} ;; | |
| *) break ;; | |
| esac | |
| done | |
| [ -n "$root" ] && [ -f "$log" ] && [ ! -L "$log" ] || return 1 | |
| case "$log" in | |
| "$root"/*) rel=${log#"$root"/} ;; | |
| *) return 1 ;; | |
| esac | |
| year=${rel%%/*}; rel=${rel#*/} | |
| month=${rel%%/*}; rel=${rel#*/} | |
| day=${rel%%/*}; rel=${rel#*/} | |
| session=${rel%%/*}; leaf=${rel#*/} | |
| [ -n "$year" ] && [ -n "$month" ] && [ -n "$day" ] && [ -n "$session" ] \ | |
| && [ "$leaf" = session.jsonl ] | |
| } | |
| fm_busy_muse_main_log_path_valid() { # <sessions-root> <session-log> | |
| local root=${1%/} log=$2 rel year month day session leaf | |
| while :; do | |
| case "$root" in | |
| *'//'*) root=${root//\/\//\/} ;; | |
| *) break ;; | |
| esac | |
| done | |
| [ -n "$root" ] && [ -f "$log" ] && [ ! -L "$log" ] || return 1 | |
| case "$log" in | |
| "$root"/*) rel=${log#"$root"/} ;; | |
| *) return 1 ;; | |
| esac | |
| case "$rel" in */*/*/*/*) return 1 ;; esac | |
| year=${rel%%/*}; case "$rel" in */*) rel=${rel#*/} ;; *) return 1 ;; esac | |
| month=${rel%%/*}; case "$rel" in */*) rel=${rel#*/} ;; *) return 1 ;; esac | |
| day=${rel%%/*}; case "$rel" in */*) rel=${rel#*/} ;; *) return 1 ;; esac | |
| session=${rel%%/*}; case "$rel" in */*) leaf=${rel#*/} ;; *) return 1 ;; esac | |
| [ -n "$year" ] && [ -n "$month" ] && [ -n "$day" ] && [ -n "$session" ] \ | |
| && [ "$leaf" = session.jsonl ] | |
| } |
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@bin/fm-busy-lib.sh` around lines 404 - 423, Update
fm_busy_muse_main_log_path_valid to validate that each rel split consumes an
actual slash before assigning year, month, day, session, and leaf. Reject paths
with fewer than four directory components, such as root/x/session.jsonl, while
preserving acceptance of the required year/month/day/session/session.jsonl
layout.
| SESSION_START_STAGE_FILE=$(mktemp "${TMPDIR:-/tmp}/fm-session-start-stage.XXXXXX" 2>/dev/null) || SESSION_START_STAGE_FILE= | ||
| if [ -z "$SESSION_START_STAGE_FILE" ]; then | ||
| # Without a breadcrumb the bound still holds; only the banner's precision | ||
| # is lost, so the child still runs bounded. | ||
| SESSION_START_STAGE_FILE=/dev/null | ||
| fi |
There was a problem hiding this comment.
🩺 Stability & Availability | 🟠 Major | ⚡ Quick win
Do not rm -f the /dev/null fallback path.
When mktemp fails, SESSION_START_STAGE_FILE becomes /dev/null. Line 207 then runs rm -f /dev/null. A non-root run fails silently, but a root run (container or CI image) deletes the device node and breaks redirection for every later process on that host. Track whether a real temp file was created and remove only that file.
🛡️ Proposed fix
SESSION_START_STAGE_FILE=$(mktemp "${TMPDIR:-/tmp}/fm-session-start-stage.XXXXXX" 2>/dev/null) || SESSION_START_STAGE_FILE=
+ SESSION_START_STAGE_TMP=1
if [ -z "$SESSION_START_STAGE_FILE" ]; then
# Without a breadcrumb the bound still holds; only the banner's precision
# is lost, so the child still runs bounded.
SESSION_START_STAGE_FILE=/dev/null
+ SESSION_START_STAGE_TMP=0
fi- rm -f "$SESSION_START_STAGE_FILE" 2>/dev/null || true
+ [ "$SESSION_START_STAGE_TMP" -eq 1 ] && rm -f "$SESSION_START_STAGE_FILE" 2>/dev/null
+ trueAlso applies to: 207-207
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@bin/fm-session-start.sh` around lines 177 - 182, Track whether mktemp
successfully created a real stage file before assigning the /dev/null fallback,
using a distinct cleanup indicator near SESSION_START_STAGE_FILE initialization.
Update the cleanup at line 207 to remove the stage path only when that indicator
confirms a temporary file was created, never attempting to remove /dev/null.
|
|
||
| ``` | ||
| Sign in at this page: | ||
| https://auth.meta.com/oauth/device/?code=DGXZ-NRPR |
There was a problem hiding this comment.
🔒 Security & Privacy | 🟠 Major | ⚡ Quick win
Redact the device authorization code.
Line 109 contains a concrete code in an authentication URL. Treat it as credential material. Replace it with a placeholder and invalidate the device flow if it was real. Scrub it from repository history before publication.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@docs/verification/muse.md` at line 109, Replace the concrete authorization
code in the authentication URL in muse.md with a clearly non-functional
placeholder, and invalidate the device authorization flow if the exposed code
was real. Remove the credential from repository history before publication.
|
|
||
| The daemon refused with `managed standalone Codex install not found`, and an interactive TUI worker neither starts nor attaches to the app-server control socket, so no client can observe its turns. | ||
| Firstmate-written project hooks under `<worktree>/.codex/hooks.json` fired for neither an interactive pane whose directory trust was granted nor `codex exec`, in both cases with `--dangerously-bypass-hook-trust`, while global `~/.codex/hooks.json` `SessionStart` hooks fired in the same runs. | ||
| In this 2026-07-28 Codex 0.145.0 semantic-busy probe, Firstmate-written lifecycle project hooks under `<worktree>/.codex/hooks.json` fired for neither an interactive pane whose directory trust was granted nor `codex exec`, in both cases with `--dangerously-bypass-hook-trust`, while an untracked global probe fired in the same runs; Firstmate does not ship, install, recommend, or depend on that global path. |
There was a problem hiding this comment.
🗄️ Data Integrity & Integration | 🟡 Minor | ⚡ Quick win
🧩 Analysis chain
🏁 Script executed:
#!/bin/bash
set -euo pipefail
echo "== Locate files =="
git ls-files | rg '(^|/)docs/(verification/supervision|sessionstart-nudge)\.md$|hooks\.json|codex' || true
echo
echo "== supervision.md outline/sections around relevant lines =="
if [ -f docs/verification/supervision.md ]; then
wc -l docs/verification/supervision.md
sed -n '1,180p' docs/verification/supervision.md | cat -n
fi
echo
echo "== sessionstart-nudge.md relevant mentions =="
if [ -f docs/sessionstart-nudge.md ]; then
rg -n -C 4 'codex exec|hook|SessionStart|hooks\.json' docs/sessionstart-nudge.md || true
fi
echo
echo "== Code/docs references to SessionStart hook and codex exec =="
rg -n -C 3 'SessionStart|Codex exec|codex exec|hooks\.json|hook trust|dangerously-bypass-hook-trust|semantic-busy' docs . --glob '!**/.git/**' --glob '!**/node_modules/**' --glob '!**/dist/**' --glob '!**/build/**' || trueRepository: trillium/firstmate
Length of output: 50376
Clarify the Codex hook scope on line 123.
The semantic-busy section should say which Codex lifecycle event names were probed, because other evidence documents that codex exec fires the tracked project SessionStart hook while this sentence says hooks did not fire for codex exec.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@docs/verification/supervision.md` at line 123, Clarify the semantic-busy
probe sentence by naming the specific Codex lifecycle event names tested for the
Firstmate-written hooks under <worktree>/.codex/hooks.json. Reconcile the
statement about codex exec with the documented SessionStart behavior,
distinguishing which event did or did not fire while preserving the existing
trust and global-probe details.
| # A home that configured nothing at all follows the version floor: it is | ||
| # projected on a release at or above it, and takes the ordinary flat layout with | ||
| # one naming warning below it. The only difference from the opted-out spawn | ||
| # above is the removed file, so this case is the floor's live end-user proof on | ||
| # whichever Herdr this lab is running. | ||
| rm -f "$HOME_DIR/config/herdr-presentation-spaces" | ||
| FLOOR_STATUS=$(lab status --json) || fail 'could not read the lab release for the presentation floor' | ||
| FLOOR_VERSION=$(printf '%s' "$FLOOR_STATUS" | jq -r 'if .server.running then .server.version else .client.version end') | ||
| FLOOR_PROTOCOL=$(printf '%s' "$FLOOR_STATUS" | jq -r 'if .server.running then .server.protocol else .client.protocol end') | ||
| FLOOR_VERDICT=$(bash -c ' | ||
| . "$0/bin/backends/herdr.sh" | ||
| status=0 | ||
| fm_backend_herdr_release_floor_verdict "$1" "$2" || status=$? | ||
| printf "%s\n" "$status" | ||
| ' "$ROOT" "$FLOOR_PROTOCOL" "$FLOOR_VERSION") | ||
| [ "$FLOOR_VERDICT" = 0 ] || [ "$FLOOR_VERDICT" = 1 ] \ | ||
| || fail "herdr $FLOOR_VERSION protocol $FLOOR_PROTOCOL could not be classified against the presentation floor" |
There was a problem hiding this comment.
🎯 Functional Correctness | 🟠 Major | ⚡ Quick win
Both live tests predict the gate with the wrong classifier. Each test derives its expected outcome from fm_backend_herdr_release_floor_verdict on a single release pair, but the gate they exercise resolves through fm_backend_herdr_presentation_release_supported, which classifies the client and the selected running server and returns the conservative composition (bin/backends/herdr.sh:231-277). On a lab whose client and server releases straddle the floor, the predicted verdict and the observed gate disagree, and the version named in the warning can differ from the version the test greps for.
tests/fm-backend-herdr-presentation-e2e.test.sh#L508-L524: replace the server-or-client selection at lines 515-516 and the classifier call at 517-522 withfm_backend_herdr_presentation_release_supported "$HERDR_LAB_SESSION", and grep the warning forFM_BACKEND_HERDR_PRESENTATION_RELEASEat line 546.tests/fm-backend-herdr-focus-flash-e2e.test.sh#L341-L395: replace the client-only classification at lines 346-353 with the same composed call for$HERDR_LAB_SESSION, and use its reported release for the naming assertion at line 386. Keep the client-only classifier for the Part A focus-behaviour correlation at lines 354-364, where a client-only verdict is the intended signal.
📍 Affects 2 files
tests/fm-backend-herdr-presentation-e2e.test.sh#L508-L524(this comment)tests/fm-backend-herdr-focus-flash-e2e.test.sh#L341-L395
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@tests/fm-backend-herdr-presentation-e2e.test.sh` around lines 508 - 524, Both
live tests must predict the presentation gate using the composed client/server
classifier. In tests/fm-backend-herdr-presentation-e2e.test.sh:508-524, replace
the server-or-client release selection and
fm_backend_herdr_release_floor_verdict call with
fm_backend_herdr_presentation_release_supported "$HERDR_LAB_SESSION", and grep
the warning for FM_BACKEND_HERDR_PRESENTATION_RELEASE. In
tests/fm-backend-herdr-focus-flash-e2e.test.sh:341-395, use the same composed
classifier and its reported release for the naming assertion, while retaining
the client-only classifier at lines 354-364 for Part A focus-behaviour
correlation.
| test_detects_versioned_process_ancestor() { | ||
| local dir bin out | ||
| dir="$TMP_ROOT/detect" | ||
| mkdir -p "$dir" | ||
| for bin in muse-bin-0.1.0-R708.1 muse-bin-9.9.9-RZZZ.9 muse; do | ||
| cp "$(command -v bash)" "$dir/$bin" | ||
| out=$(env -u CLAUDECODE -u PI_CODING_AGENT -u GROK_AGENT \ | ||
| "$dir/$bin" -c "r=\$(\"$HARNESS\"); printf '%s' \"\$r\"") | ||
| [ "$out" = muse ] || fail "fm-harness.sh under process '$bin' reported '$out', expected muse" | ||
| done | ||
| pass "muse is detected through any versioned muse-bin ancestor" | ||
| } | ||
|
|
||
| # The match must be anchored: an unrelated command whose name merely CONTAINS | ||
| # muse is a different program and must not be claimed by this adapter. | ||
| test_detection_is_anchored() { | ||
| local dir bin out | ||
| dir="$TMP_ROOT/detect-neg" | ||
| mkdir -p "$dir" | ||
| for bin in musescore amuse notmuse-bin muse-binary muse-bind; do | ||
| cp "$(command -v bash)" "$dir/$bin" | ||
| out=$(env -u CLAUDECODE -u PI_CODING_AGENT -u GROK_AGENT \ | ||
| "$dir/$bin" -c "r=\$(\"$HARNESS\"); printf '%s' \"\$r\"") | ||
| [ "$out" != muse ] || fail "fm-harness.sh misdetected unrelated process '$bin' as muse" | ||
| done | ||
| pass "muse detection does not claim unrelated muse-containing commands" | ||
| } |
There was a problem hiding this comment.
🎯 Functional Correctness | 🟡 Minor | ⚡ Quick win
🧩 Analysis chain
🏁 Script executed:
#!/bin/bash
# Description: Check how fm-harness.sh consumes FM_PI_HARNESS relative to PI_CODING_AGENT.
rg -n 'FM_PI_HARNESS|PI_CODING_AGENT' bin/fm-harness.sh -C4Repository: trillium/firstmate
Length of output: 790
🏁 Script executed:
#!/bin/bash
set -euo pipefail
echo "== relevant harness section =="
sed -n '1,110p' bin/fm-harness.sh
echo "== relevant test section =="
sed -n '140,200p' tests/fm-muse-harness.test.sh
echo "== occurrences of FM_PI_HARNESS unsetting =="
rg -n 'env -u .*FM_PI_HARNESS|FM_PI_HARNESS|PI_CODING_AGENT|GROK_AGENT|CLAUDECODE' tests -C2Repository: trillium/firstmate
Length of output: 34835
Unset FM_PI_HARNESS in the muse ancestry detection loops.
FM_PI_HARNESS=pi-signed and FM_PI_HARNESS=pi both take precedence over the muse ancestry cases, so these tests can return Pi instead of muse when run with an inherited Pi marker. Add -u FM_PI_HARNESS to both mock bash probes.
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@tests/fm-muse-harness.test.sh` around lines 164 - 190, Unset FM_PI_HARNESS in
the environment for both mock bash probes within
test_detects_versioned_process_ancestor and test_detection_is_anchored,
alongside the existing agent-marker unsets, so muse ancestry detection is tested
without inherited Pi markers.
| cp "$ROOT/bin/fm-sessionstart-run.sh" "$ROOT/bin/fm-sessionstart-nudge.sh" \ | ||
| "$ROOT/bin/fm-primary-scope-lib.sh" "$ROOT/bin/fm-gate-refuse-lib.sh" \ | ||
| "$ROOT/bin/fm-operational-input.sh" "$fixture/bin/" |
There was a problem hiding this comment.
📐 Maintainability & Code Quality | 🟡 Minor | ⚡ Quick win
🧩 Analysis chain
🏁 Script executed:
#!/bin/bash
# Resolve the full source closure of bin/fm-sessionstart-run.sh.
rg -n '^\s*\.\s+"\$SCRIPT_DIR/' bin/fm-sessionstart-run.sh bin/fm-sessionstart-nudge.sh \
bin/fm-primary-scope-lib.sh bin/fm-gate-refuse-lib.sh bin/fm-session-lock-lib.sh
# Show what each fixture copies.
rg -n -A8 'fm-sessionstart-run\.sh' tests/fm-sessionstart-nudge.test.sh tests/fm-calm-pi-extension.test.shRepository: trillium/firstmate
Length of output: 2893
🏁 Script executed:
#!/bin/bash
set -u
printf 'bin/fm-sessionstart-run.sh outline/relevant:\n'
wc -l bin/fm-sessionstart-run.sh
sed -n '1,90p' bin/fm-sessionstart-run.sh
printf '\nbin/fm-session-lock-lib.sh existence and functions:\n'
wc -l bin/fm-session-lock-lib.sh
sed -n '1,220p' bin/fm-session-lock-lib.sh
printf '\nReferences to fm_session_lock_owned_by_self:\n'
rg -n 'fm_session_lock_owned_by_self' .
printf '\nRelevant fixture setup excerpts:\n'
sed -n '276,295p' tests/fm-sessionstart-nudge.test.sh
sed -n '2858,2878p' tests/fm-calm-pi-extension.test.sh
printf '\nBehavioral probe: source closure from run-wrapper (static, based on . "$SCRIPT_DIR/... patterns only).\n'
python3 - <<'PY'
from pathlib import Path
import re
base = Path('bin')
wrap = base / 'fm-sessionstart-run.sh'
seen = {wrap}
q = [wrap]
while q:
p = q.pop()
for line in p.read_text(errors='replace').splitlines():
m = re.match(r'\s*\.\ "\$"SCRIPT_DIR/([^"]+)"', line)
if not m:
m = re.match(r'\. "\$"SCRIPT_DIR/([^"]+)"', line)
if m:
target = base / m.group(1)
print(f'{p}:{p.read_text(errors="replace").splitlines().index(line)+1}:{target}')
seen.add(target)
if target not in seen:
q.append(target)
PYRepository: trillium/firstmate
Length of output: 15131
Copy the session lock library into both run-wrapper fixtures. bin/fm-sessionstart-run.sh sources fm-session-lock-lib.sh, and both fixtures also copy executable script paths while missing that library. Include "$ROOT/bin/fm-session-lock-lib.sh" in both cp commands.
📍 Affects 2 files
tests/fm-sessionstart-nudge.test.sh#L282-L284(this comment)tests/fm-calm-pi-extension.test.sh#L2864-L2870
🤖 Prompt for AI Agents
Verify each finding against current code. Fix only still-valid issues, skip the
rest with a brief reason, keep changes minimal, and validate.
In `@tests/fm-sessionstart-nudge.test.sh` around lines 282 - 284, Update the cp
commands in tests/fm-sessionstart-nudge.test.sh:282-284 and
tests/fm-calm-pi-extension.test.sh:2864-2870 to include
"$ROOT/bin/fm-session-lock-lib.sh" alongside the existing session-start scripts,
ensuring both run-wrapper fixtures copy the library sourced by
fm-sessionstart-run.sh.
|
Closing in favor of pipeline-opened replacement. See #XXX for the re-raised PR via no-mistakes pipeline. Branch fm/upstream-merge-into-fork preserved with all commits (head 00ccc5a). |
Syncing the fork with upstream/main to keep up with the latest changes.
Summary by CodeRabbit
New Features
Bug Fixes