feat(codex): propagate assigned display titles - #14
Conversation
🤖 CodeAnt AI — Review Status
|
Thanks for using CodeAnt! 🎉We're free for open-source projects. if you're enjoying it, help us grow by sharing. Share on X · |
🏁 CodeAnt Quality Gate ResultsCommit: ✅ Overall Status: PASSEDQuality Gate Details
|
| while [ "$i" -lt 30 ]; do | ||
| agent_json=$(lab agent get "$PANE" 2>/dev/null || true) | ||
| after_uuid=$(printf '%s' "$agent_json" | jq -r '.result.agent.agent_session.value // empty' 2>/dev/null || true) | ||
| pane_json=$(lab pane get "$PANE" 2>/dev/null || true) | ||
| terminal_title=$(printf '%s' "$pane_json" | jq -r '.result.pane.terminal_title // empty' 2>/dev/null || true) | ||
| header=$(printf '%s' "$pane_json" | jq -r '.result.pane.title // empty' 2>/dev/null || true) | ||
| [ "$after_uuid" = "$SESSION_ID" ] \ | ||
| || fail "conversation UUID changed during rename (expected $SESSION_ID, got ${after_uuid:-none})" | ||
| case "$header" in *"$EXPECTED_HEADER_SUFFIX") header_matches=true ;; *) header_matches=false ;; esac | ||
| [ "$terminal_title" = "$EXPECTED_TERMINAL_TITLE" ] && [ "$header_matches" = true ] && break | ||
| i=$((i + 1)) | ||
| sleep 0.25 | ||
| done |
There was a problem hiding this comment.
Suggestion: Extract the repeated UUID, pane-title, and terminal-title polling into a single helper that accepts the timeout and expected values, then reuse it for the post-rename and post-resume checks to keep the asynchronous verification logic consistent. [custom_rule]
Severity Level: Major
Why it matters? ⭐
The post-rename polling loop duplicates the same UUID, pane-title, terminal-title, and header polling logic used again during exact resume at lines 202–217. This duplicated asynchronous verification is a meaningful maintainability concern because changes to polling behavior or validation can diverge between the two paths. Extracting the shared logic into a helper would improve consistency without changing behavior.
Rule source 📖
CodeAnt dashboard (rule "agent-persona-maintainability")
Prompt for AI Agent 🤖
This is a comment left during a code review.
**Path:** tests/fm-codex-title-live-e2e.test.sh
**Line:** 154:166
**Comment:**
*Custom Rule: Extract the repeated UUID, pane-title, and terminal-title polling into a single helper that accepts the timeout and expected values, then reuse it for the post-rename and post-resume checks to keep the asynchronous verification logic consistent.
Validate the correctness of the flagged issue. If correct, How can I resolve this? If you propose a fix, implement it and please make it concise.
Once fix is implemented, also check other comments on the same PR, and ask user if the user wants to fix the rest of the comments as well. if said yes, then fetch all the comments validate the correctness and implement a minimal fix| Legacy registry lines without a code retain their delivery behavior, while `bin/fm-project-mode.sh --code` refuses to guess a presentation acronym. | ||
|
|
||
| Worker presentation identity keeps five independent fields: the persistent callsign, registered project code, explicit one- or two-word task label, internal task id, and harness session UUID. | ||
| `bin/fm-display-title-lib.sh` owns the harness-neutral `Callsign · ProjectCode · TaskLabel` rendering and ASCII fallback; `state/<id>.meta` carries only `project_code=` and `task_label=`, while the callsign and Codex UUID retain their existing separate owners. |
There was a problem hiding this comment.
Suggestion: Clarify that state/<id>.meta carries these two presentation fields in addition to the task's other runtime metadata, rather than saying it carries only them. [custom_rule]
Severity Level: Major
Why it matters? ⭐
The statement that state/<id>.meta carries only these two fields is misleading because the architecture also documents other runtime metadata stored there, including pr= and pr_head=. Clarifying that these are the two presentation fields, rather than the only metadata fields, is a meaningful maintainability improvement that prevents readers from misunderstanding the metadata contract.
Rule source 📖
CodeAnt dashboard (rule "agent-persona-maintainability")
Prompt for AI Agent 🤖
This is a comment left during a code review.
**Path:** docs/architecture.md
**Line:** 255:255
**Comment:**
*Custom Rule: Clarify that `state/<id>.meta` carries these two presentation fields in addition to the task's other runtime metadata, rather than saying it carries only them.
Validate the correctness of the flagged issue. If correct, How can I resolve this? If you propose a fix, implement it and please make it concise.
Once fix is implemented, also check other comments on the same PR, and ask user if the user wants to fix the rest of the comments as well. if said yes, then fetch all the comments validate the correctness and implement a minimal fix| local backend=$1 target=$2 expected=${3:-} attempt=0 | ||
| local max=${FM_CODEX_TITLE_READY_POLLS:-60} interval=${FM_CODEX_TITLE_POLL_INTERVAL:-0.5} | ||
| while [ "$attempt" -lt "$max" ]; do | ||
| if [ "$(fm_backend_composer_state "$backend" "$target" "$expected" 2>/dev/null)" = empty ]; then |
There was a problem hiding this comment.
Suggestion: Replace the fixed-interval full composer polling with a backend-native readiness signal or an adaptive backoff so launch does not perform up to 60 expensive screen or network inspections at 0.5-second intervals. [custom_rule]
Severity Level: Major
Why it matters? ⭐
The readiness loop can invoke the backend composer-state inspection up to 60 times, sleeping for 0.5 seconds between checks. Because this runs during every Codex title delivery and may perform screen or external I/O, a backend-native readiness signal or adaptive polling could avoid repeated inspections and reduce startup overhead. This is a real repeated-I/O performance concern at the stated medium threshold.
Rule source 📖
CodeAnt dashboard (rule "agent-persona-performance")
Prompt for AI Agent 🤖
This is a comment left during a code review.
**Path:** bin/fm-codex-title-lib.sh
**Line:** 25:25
**Comment:**
*Custom Rule: Replace the fixed-interval full composer polling with a backend-native readiness signal or an adaptive backoff so launch does not perform up to 60 expensive screen or network inspections at 0.5-second intervals.
Validate the correctness of the flagged issue. If correct, How can I resolve this? If you propose a fix, implement it and please make it concise.
Once fix is implemented, also check other comments on the same PR, and ask user if the user wants to fix the rest of the comments as well. if said yes, then fetch all the comments validate the correctness and implement a minimal fix| return 1 | ||
| } | ||
| opinput="${FM_ROOT:-$(cd "$FM_CODEX_TITLE_LIB_DIR/.." && pwd)}/bin/fm-operational-input.sh" | ||
| input=$("$opinput" encode "$input_kind" < "$input_file") || return 1 |
There was a problem hiding this comment.
Suggestion: Use a file-aware operational-input delivery path, or enforce a bounded input size before encoding, instead of reading the entire input file into a shell variable and then passing another copy through the backend argument. [custom_rule]
Severity Level: Major
Why it matters? ⭐
The code reads the entire operational-input file into a shell variable, and the encoded result is then passed as an argument to the backend submission function. This creates full-payload in-memory copies and provides no size bound; for a large operational input, shell and argument memory use can become noticeable and the eventual argument-based delivery may hit system limits. A streaming or file-aware delivery path, or an explicit bounded-size check, would be a concrete leaner alternative.
Rule source 📖
CodeAnt dashboard (rule "agent-persona-performance")
Prompt for AI Agent 🤖
This is a comment left during a code review.
**Path:** bin/fm-codex-title-lib.sh
**Line:** 61:61
**Comment:**
*Custom Rule: Use a file-aware operational-input delivery path, or enforce a bounded input size before encoding, instead of reading the entire input file into a shell variable and then passing another copy through the backend argument.
Validate the correctness of the flagged issue. If correct, How can I resolve this? If you propose a fix, implement it and please make it concise.
Once fix is implemented, also check other comments on the same PR, and ask user if the user wants to fix the rest of the comments as well. if said yes, then fetch all the comments validate the correctness and implement a minimal fix| FM_DISPLAY_METADATA_STATE=absent | ||
| FM_DISPLAY_PROJECT_CODE= | ||
| FM_DISPLAY_TASK_LABEL= |
There was a problem hiding this comment.
Suggestion: Extract the shared metadata-state initialization, paired-field validation, value validation, and global assignment into a common helper so the brief and task readers only supply their format-specific parsing logic. [custom_rule]
Severity Level: Major
Why it matters? ⭐
The same metadata-state initialization, validation flow, and global assignments are duplicated in both reader functions, with the corresponding initialization also present at lines 129-131. This creates two implementations that must be kept synchronized; a shared helper would meaningfully reduce maintenance risk without changing behavior.
Rule source 📖
CodeAnt dashboard (rule "agent-persona-maintainability")
Prompt for AI Agent 🤖
This is a comment left during a code review.
**Path:** bin/fm-display-title-lib.sh
**Line:** 93:95
**Comment:**
*Custom Rule: Extract the shared metadata-state initialization, paired-field validation, value validation, and global assignment into a common helper so the brief and task readers only supply their format-specific parsing logic.
Validate the correctness of the flagged issue. If correct, How can I resolve this? If you propose a fix, implement it and please make it concise.
Once fix is implemented, also check other comments on the same PR, and ask user if the user wants to fix the rest of the comments as well. if said yes, then fetch all the comments validate the correctness and implement a minimal fix| project_count=$(awk '/^# Task$/ { exit } /^Firstmate project code: / { n++ } END { print n+0 }' "$brief") || return 1 | ||
| label_count=$(awk '/^# Task$/ { exit } /^Firstmate task label: / { n++ } END { print n+0 }' "$brief") || return 1 | ||
| if [ "$project_count" -eq 0 ] && [ "$label_count" -eq 0 ]; then | ||
| return 0 | ||
| fi | ||
| if [ "$project_count" -ne 1 ] || [ "$label_count" -ne 1 ]; then | ||
| printf 'error: brief display metadata requires exactly one project code and one task label before # Task\n' >&2 | ||
| return 1 | ||
| fi | ||
| project_code=$(awk '/^# Task$/ { exit } /^Firstmate project code: / { sub(/^Firstmate project code: /, ""); print; exit }' "$brief") || return 1 | ||
| task_label=$(awk '/^# Task$/ { exit } /^Firstmate task label: / { sub(/^Firstmate task label: /, ""); print; exit }' "$brief") || return 1 |
There was a problem hiding this comment.
Suggestion: Parse the brief once with a single awk invocation that counts and extracts both metadata fields while preserving the pre-Task boundary, instead of rescanning the entire file four times. [custom_rule]
Severity Level: Major
Why it matters? ⭐
The brief reader launches four separate awk processes and rescans the same file to count and then extract two fields. For commonly invoked metadata reads, a single awk pass can perform both operations, avoiding repeated file I/O and process startup overhead while preserving the pre-# Task boundary.
Rule source 📖
CodeAnt dashboard (rule "agent-persona-performance")
Prompt for AI Agent 🤖
This is a comment left during a code review.
**Path:** bin/fm-display-title-lib.sh
**Line:** 100:110
**Comment:**
*Custom Rule: Parse the brief once with a single awk invocation that counts and extracts both metadata fields while preserving the pre-Task boundary, instead of rescanning the entire file four times.
Validate the correctness of the flagged issue. If correct, How can I resolve this? If you propose a fix, implement it and please make it concise.
Once fix is implemented, also check other comments on the same PR, and ask user if the user wants to fix the rest of the comments as well. if said yes, then fetch all the comments validate the correctness and implement a minimal fix| project_count=$(awk -F= '$1 == "project_code" { n++ } END { print n+0 }' "$meta") || return 1 | ||
| label_count=$(awk -F= '$1 == "task_label" { n++ } END { print n+0 }' "$meta") || return 1 | ||
| if [ "$project_count" -eq 0 ] && [ "$label_count" -eq 0 ]; then | ||
| return 0 | ||
| fi | ||
| if [ "$project_count" -ne 1 ] || [ "$label_count" -ne 1 ]; then | ||
| printf 'error: task display metadata requires exactly one project_code= and one task_label= field\n' >&2 | ||
| return 1 | ||
| fi | ||
| project_code=$(awk -F= '$1 == "project_code" { sub(/^[^=]*=/, ""); print; exit }' "$meta") || return 1 | ||
| task_label=$(awk -F= '$1 == "task_label" { sub(/^[^=]*=/, ""); print; exit }' "$meta") || return 1 |
There was a problem hiding this comment.
Suggestion: Parse the task metadata once with a single awk invocation that counts and extracts both fields, rather than performing four separate scans of the metadata file. [custom_rule]
Severity Level: Major
Why it matters? ⭐
The task metadata reader performs four separate awk invocations over the same file, causing repeated file reads and process startup costs on each invocation. Combining counting and extraction into one awk pass is a concrete performance improvement with no required behavior change.
Rule source 📖
CodeAnt dashboard (rule "agent-persona-performance")
Prompt for AI Agent 🤖
This is a comment left during a code review.
**Path:** bin/fm-display-title-lib.sh
**Line:** 136:146
**Comment:**
*Custom Rule: Parse the task metadata once with a single awk invocation that counts and extracts both fields, rather than performing four separate scans of the metadata file.
Validate the correctness of the flagged issue. If correct, How can I resolve this? If you propose a fix, implement it and please make it concise.
Once fix is implemented, also check other comments on the same PR, and ask user if the user wants to fix the rest of the comments as well. if said yes, then fetch all the comments validate the correctness and implement a minimal fix| if [ "$CODEX_DEFERRED_INPUT" -eq 1 ]; then | ||
| if ! fm_codex_title_deliver \ | ||
| "$BACKEND" "$T" "$W" "$CALLSIGN" "$PROJECT_CODE" "$TASK_LABEL" \ | ||
| "$CODEX_INPUT_KIND" "$CODEX_INPUT_FILE"; then | ||
| printf 'failed: Codex conversation naming or initial task delivery was not confirmed\n' >> "$STATE/$ID.status" | ||
| echo "error: Codex conversation naming or initial task delivery was not confirmed; inspect window $T" >&2 | ||
| exit 1 | ||
| fi | ||
| fi |
There was a problem hiding this comment.
Suggestion: When fm_codex_title_deliver fails after the launch command has already been submitted, this path records a failed status and exits but leaves the Codex process and endpoint alive. The task metadata can therefore report a failed spawn while the agent continues consuming the task endpoint, and later relaunch or recovery can encounter an unexpected live process. Quarantine or terminate the exact backend endpoint before returning failure, while preserving fail-closed metadata if cleanup cannot be confirmed. [incomplete implementation]
Severity Level: Major ⚠️
- ❌ Failed spawns leave live Codex endpoints behind.
- ⚠️ Recovery can encounter stale live task processes.
- ⚠️ Failed metadata disagrees with backend endpoint liveness.Prompt for AI Agent 🤖
This is a comment left during a code review.
**Path:** bin/fm-spawn.sh
**Line:** 3124:3132
**Comment:**
*Incomplete Implementation: When `fm_codex_title_deliver` fails after the launch command has already been submitted, this path records a failed status and exits but leaves the Codex process and endpoint alive. The task metadata can therefore report a failed spawn while the agent continues consuming the task endpoint, and later relaunch or recovery can encounter an unexpected live process. Quarantine or terminate the exact backend endpoint before returning failure, while preserving fail-closed metadata if cleanup cannot be confirmed.
Validate the correctness of the flagged issue. If correct, How can I resolve this? If you propose a fix, implement it and please make it concise.
Once fix is implemented, also check other comments on the same PR, and ask user if the user wants to fix the rest of the comments as well. if said yes, then fetch all the comments validate the correctness and implement a minimal fix| fm_backend_send_text_submit() { | ||
| local backend=$1 target=$2 payload=$3 retries=$4 sleep_s=$5 settle=$6 expected=${7:-} | ||
| printf '%s\t%s\t%s\t%s\t%s\t%s\t%s\n' \ | ||
| "$backend" "$target" "$retries" "$sleep_s" "$settle" "$expected" "$payload" >> "$LOG" | ||
| printf '%s' "${FM_FAKE_SUBMIT_VERDICT:-empty}" | ||
| } |
There was a problem hiding this comment.
Suggestion: The backend stub returns empty for every submission and records no backend-side state, so the test passes even if the /rename command is ignored, the readiness check is ineffective, or the operational input is accepted before the rename takes effect. Model the rename and task-delivery verdicts separately and make the composer state depend on the submitted command so this test verifies the production ordering contract rather than only the two log entries. [incomplete implementation]
Severity Level: Major ⚠️
- ❌ Unit tests can miss ignored Codex rename commands.
- ⚠️ Adapter readiness behavior is only synthetically validated.
- ⚠️ Spawn failures may escape regression coverage.Prompt for AI Agent 🤖
This is a comment left during a code review.
**Path:** tests/fm-codex-title.test.sh
**Line:** 26:31
**Comment:**
*Incomplete Implementation: The backend stub returns `empty` for every submission and records no backend-side state, so the test passes even if the `/rename` command is ignored, the readiness check is ineffective, or the operational input is accepted before the rename takes effect. Model the rename and task-delivery verdicts separately and make the composer state depend on the submitted command so this test verifies the production ordering contract rather than only the two log entries.
Validate the correctness of the flagged issue. If correct, How can I resolve this? If you propose a fix, implement it and please make it concise.
Once fix is implemented, also check other comments on the same PR, and ask user if the user wants to fix the rest of the comments as well. if said yes, then fetch all the comments validate the correctness and implement a minimal fixa2d6dec to
94291fe
Compare
Thanks for using CodeAnt! 🎉We're free for open-source projects. if you're enjoying it, help us grow by sharing. Share on X · |
* feat(codex): propagate assigned display titles * test(codex): keep title resume guard isolated --------- Co-authored-by: wenkxu <v-wenkxu@expediagroup.com>
* fix(composer): stop a blocked pi pane from proving an empty composer (kunchenguid#2811) A pi worker parked on an interactive prompt - a permission dialog, a question menu, a trust dialog - reports agent_status=blocked, because it is waiting on a human keystroke. Pi draws that menu above its separator pair, so the composer region between the rules is blank and structure alone looks like a free composer. _fm_composer_pi_verdict admitted blocked alongside idle and done, so the shared classifier reported an affirmatively empty composer for exactly the pane where typing is unsafe. Every "is it safe to type here?" consumer reads that verdict and proceeds only on an affirmative empty, so both are told yes on a parked prompt: the away-mode injection guard in bin/fm-supervise-daemon.sh, and fm-send's pre-type refusal. The keys then answer the menu instead of composing a message - the highlighted default is selected, the text is discarded, and the record attributes a decision to a human who never made it. blocked now defers to unknown, which every consumer already treats as fail-closed. idle and done still prove an empty composer, so ordinary steering is unchanged, and Cursor is unaffected because its always-blocked panes never reach this pi-only branch. Regression coverage lands first at both levels: the verdict owner (a blocked pi defers) and the herdr adapter (a parked pi prompt is not an empty composer). * fix(bin): require project clone roots during fleet sync (kunchenguid#2849) * fix(bin): require a clone root before fleet-sync touches a project Git repository discovery walks upward, so `git -C projects/<dir>` on a plain directory nested under projects/ resolves to the enclosing repository - in a firstmate home, the firstmate checkout itself. fm-fleet-sync.sh guarded its candidates with `rev-parse --is-inside-work-tree`, which such a directory passes, so every later git call read, pruned and fast-forwarded firstmate's own default branch and reported it under the project directory's label. A running session's AGENTS.md changed underneath it, and the report named a project that had nothing to do with the change. Require each candidate to be the root of its own work tree before any other git command: compare `rev-parse --show-toplevel` against the directory's own physical path. Both sides are physical, so a symlinked clone still compares equal. Anything else is skipped by name, naming the repository that would have been touched, and bootstrap relays that as a FLEET_SYNC line. Regression coverage reproduces the wrong-repo fast-forward against a home nested inside another repository, in both the whole-fleet and single-project forms, and pins that a symlinked clone dir still syncs. * no-mistakes(review): Keep enclosing fixture clean during clone-root regression * fix(bin): retry transient Lavish poll interruptions (kunchenguid#2846) * fix(procevent): retry a transient Lavish poll interruption quietly A live Lavish listener can be cut short by the server with exactly error: Lavish Editor poll response was interrupted code: SERVER_ERROR while the session's marks remain available. Firstmate registered raw `lavish-axi poll` output, so the generic process-event runner captured that transient response as a result and woke the whole fleet over what is really an internal retry. The Lavish adapter now registers its own listener command, which reruns the published blocking poll up to 12 times at 5 second intervals for that one exact two-line response. The match is deliberately narrow: real feedback, ended and missing sessions, any other SERVER_ERROR, and the same interruption still standing once the bound is spent all pass straight through and are captured and announced as before. The retry is a Lavish fact, so the generic runner stays adapter-agnostic. `FM_LAVISH_POLL_RETRY_DELAY` is a bounded 0 to 60 second override for the interval only, refused rather than rounded when malformed, so a test can exercise the real bound without waiting it out. * no-mistakes(review): Harden Lavish retry matching, validation, and cleanup * no-mistakes(review): Bound Lavish retry staging and stabilize regression * no-mistakes(document): docs: explain Lavish retry adoption * no-mistakes(lint): Restore Lavish trap ShellCheck suppression * fix(brief): stop the documented {TASK} fill from corrupting the Herdr gate (kunchenguid#2838) The unguarded Herdr declaration quoted `{TASK}` in its own prose while the scaffold instructs firstmate to replace every `{TASK}` placeholder. The documented global replace therefore spliced the whole task body into the middle of the safety gate's sentence, silently destroying the one contract that exists precisely because the scaffold cannot inspect the task text. Reword the gate to refer to the task text filled in above, leaving the placeholder only at its genuine fill site. Rewording rather than renaming the token keeps the unfilled-charter guards in fm-home-seed.sh and fm-remote-home-seed.sh working unchanged. Add a regression test that performs the documented global fill on ship and scout scaffolds and asserts the body lands once and the gate survives. * fix(bin): resolve the busy-state lock mtime with the platform's own stat form (kunchenguid#2837) The writer lock's stale-lock branch read the lock's mtime with `stat -f %m ... || stat -c %Y ...`. On GNU coreutils `-f` is filesystem stat, so it consumed the format string as a path, complained on stderr, printed a partial filesystem dump (" File: ...") on stdout, and still exited 0. The GNU form in the fallback therefore never ran, and the following arithmetic evaluated the word `File`, aborting the writer under `set -u` with "File: unbound variable". fm-teardown.sh died there after returning the worktree, leaving state/<id>.meta, .status, .busy-gen, .busy-state, .busy-state.lock/ and .turn-ended behind. The surviving metadata kept the watcher monitoring an endpoint whose agent was gone, so a finished task produced stale wakes forever, and every re-run died identically because the abandoned lock was never broken. Detect the platform once and pick the right stat form, the pattern bin/fm-watch.sh already documents, and treat any non-numeric result as "just created" so a future portability surprise degrades to a lock-timeout refusal rather than killing teardown mid-way. * fix(stow): add opt-in pass horizon for memory decay (kunchenguid#2850) * fix(stow): give memory decay a per-pass horizon so the clock fires The tiered decay clocks were wall-clock only, while admission is per-pass: each /stow admits the findings that pass produced. In a home that stows daily those two rates diverge by the stow cadence, an entry the fleet keeps exercising never reaches 30 days unreinforced, and memory only grows while the pass reports decay evaluated. Give each dated marker an optional unreinforced-pass counter and make both tiers stale at whichever horizon comes first: 10 passes or 30 days for aging, 3 passes or 7 days for perishable. Reinforcement clears the counter and nothing else does, so the existing evidence-based restamp rule stays the only way an entry renews its lease. An absent /N means zero, so entries that stay exercised carry no extra marker bytes, and a rarely stowed home keeps its current behaviour through the unchanged date horizon. * no-mistakes(document): Align stow workflow with dual decay clocks * fix(stow): make the per-pass decay horizon opt-in The unreinforced-pass horizon shipped as a new default archival cadence, which is a product default rather than a restoration of the existing wall-clock contract. Keep the 30-day and 7-day horizons as the only default clock, and put the 10-pass and 3-pass horizons behind an explicit opt-in: config/stow-pass-horizon for the firstmate home, and the file's own header pointer for the public skill. With the opt-in absent no counter is written and no counter is read, so a home that does not ask for it decays exactly as it does today. * no-mistakes(review): Preserve frozen counters and correct archive provenance * test(watcher): stop fixture confirmation budgets racing real child startup (kunchenguid#2876) tests/fm-watcher-lock.test.sh passed in isolation but failed intermittently under full-suite and ambient concurrent load. bin/fm-watch-arm.sh computes its confirmation deadline immediately after forking the real child watcher, so the child's entire fork, exec, lock acquisition and beacon publication has to land inside that wall clock. Two cases shrank that budget to one second, leaving a two-second window for work measured at 3.1-4.9s under CPU oversubscription, so the arm honestly reported "FAILED - no live watcher with a fresh beacon" and their premises collapsed. A third case ran on the production budget, but its child must also execute a registered check before exiting: measured at 1.9-2.3s idle and 9.1-13.1s under load, against an 11s budget. The two cases that must confirm a real child now hold the arm to production's own budget instead of a shrunken fixture one, the immediate-wake case gets an explicit budget with headroom over its measured loaded cost, and the two waits for the arm's typed failure are sized off the largest production default rather than a fixed eight seconds. No bin/ change and no default behavior change: the lock's fail-closed semantics, SIGSTOP handling, stale-heartbeat detection and the arm's typed failures are untouched. Verified 4/4 green at 3x CPU oversubscription (loadavg 75-80) after 3/3 red before the change, and CONTRIBUTING.md records the convention. * fix(bin): deterministically order remote tool paths (kunchenguid#2870) * fix(bin): order discovered tool installs by the shell's own expansion fm_remote_job_compose_operator_path built the asdf and mise install directories with `compgen -G`, which does not sort. Bash sorts glob matches in pathexp.c, on the shell's own pathname-expansion path only; `compgen -G` reaches the same glob_filename through pcomplete.c, which sorts nothing. On bash 3.2 (macOS /bin/bash) and every bash before 5.3 that handed the composition raw readdir order, so which install of a multi-version tool a remote job resolved was decided by directory order on disk rather than by this composition. Expand the globs at the call sites and let the function take the matches, so the composition and the documented portable-PATH contract are the same operation. Quoting the account home at the call site also stops a home whose name contains glob metacharacters from being reinterpreted. The colocated regression pins both the order and the mechanism: bash 5.3 moved sorting into the glob library, so an order-only assertion cannot see the defect there. * no-mistakes(review): Remove source-reading PATH regression guard * fix(bin): prevent routed secondmate work from stranding (kunchenguid#2848) * fix: surface stalled secondmate queues and wake handoffs * no-mistakes(review): Make handoff wakes retryable and stall alerts crash-safe * no-mistakes(review): Prevent duplicate handoff wakes and cover remote delivery * no-mistakes(review): Serialize local handoffs and preserve pre-move wake intent * no-mistakes(review): Serialize teardown with handoffs and retain remote wake confirmation * no-mistakes(review): Reconcile correlated handoff wake delivery after crashes * no-mistakes(review): Keep failed wakes retryable and isolate stall receipts * no-mistakes(review): Reset known-undelivered wake attempts for durable retries * no-mistakes(review): Refuse duplicate sends for unresolved delivery attempts * no-mistakes(review): Atomically restore retryability after reconciled send failures * no-mistakes(review): Serialize delivery confirmation with reconciliation * no-mistakes(document): Document routed wake and stall supervision * no-mistakes(lint): Fix ShellCheck expansion and subshell warnings * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes(review): Retire stale wake state and defer pre-move wakes * no-mistakes(review): Secure markers, bind batches, and preserve teardown routes * no-mistakes(review): Preserve unresolved prepared wakes across unrelated handoffs * no-mistakes(review): Preserve prepared wakes before unrelated moving handoffs * no-mistakes(document): Document prepared wake batch ownership * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes(review): Make local wake retirement recoverable * no-mistakes(document): Clarify handoff recovery and teardown documentation * fix: make macOS inbox test path portable (kunchenguid#2857) * feat(bin): deliver local steers through durable task inboxes (kunchenguid#2856) * feat(bin): steer local tasks by durable inbox record plus constant doorbell Stage 1 (local steers) of the captain-adopted reframe in data/fm-send-reliability-reframe-s1/report.md: an ordinary fm-send text steer to a task recorded in this home is appended as a sequenced durable record under state/<id>.inbox/ and the terminal receives only one constant self-describing doorbell line, best-effort. The worker acknowledges by moving the record into handled/; the watcher re-rings an unacknowledged message on an idle pane and escalates once as an ordinary stale wake. --resolve-key closes decisions at enqueue time, because the durable enqueue IS delivery to the task's record. bin/fm-task-inbox-lib.sh owns the record format, doorbell line, and re-ring ladder. The typed plane remains for what must reach the terminal itself: lifecycle keys, harness-native slash and codex $-skill invocations, explicit backend targets, and the remote secondmate leg (unchanged until the remote inbox leg ships separately). The composer classifier is demoted from delivery proof to an advisory ring guard that skips only on a proven pending verdict. Verified live against claude, codex, opencode, pi, grok, and muse: each real worker read its record, acted, and acked with the mv (docs/verification/runtime-backends.md "Steering-inbox doorbell"). * docs(verification): flag the grok 1.0.5 composer-matrix staleness observed by the doorbell run * test(captain-hold): read the chat-channel answer from the durable inbox record * test: migrate fm-control's marker contrast to the inbox record and fix macOS wc padding in the tool-update suite * no-mistakes(review): Harden inbox locking, teardown races, and acknowledgements * no-mistakes(review): Serialize watcher actions with inbox acknowledgements * no-mistakes(review): Bound metadata locking and tighten acknowledgement rechecks * no-mistakes(review): Preserve exact inbox bytes and harden delivery recovery * no-mistakes(review): Harden watcher bookkeeping against concurrent inbox teardown * no-mistakes(document): Update inbox and typed-plane documentation * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * revert(pipeline): keep parser-native secondmate marking and the both-failed exit out of stage 1 The CI monitor's fix changed the secondmate marking contract for parser-native invocations (appending the marker after the text) and softened the both-commit-and-marker-failed branch to exit 0. The merge authority ruled the marking question out of scope for this stage-1 transport PR (follow-up: fm-send-secondmate-harness-invocation-r1) and ruled the both-failed case a loud nonzero local failure. Restore both, keeping the monitor's legitimate migrations and hardening. * no-mistakes(document): Document inbox and typed-plane boundaries * no-mistakes(document): Scope backend transport docs to typed plane * no-mistakes(document): Clarify inbox attempt-budget documentation * no-mistakes: apply CI fixes * fix(send): the durable record alone governs the inbox exit status Captain-refined ruling on the F2/Greptile finding: the durable inbox record is what delivers the steer, so pending-reply bookkeeping trouble after a successful enqueue never exits nonzero - a resend-inviting status would make automated callers enqueue the delivered instruction again under a new sequence. With the recovery marker stored the watcher reconciles silently; with the commit and marker both lost the send surfaces a distinct reply-tracking-degraded do-not-resend warning and still exits 0. Nonzero remains only where nothing was delivered (or a decision close needs its manual command). Regression: record durable + both bookkeeping writes lost -> exit 0, one record, no duplicate. * no-mistakes(review): Preserve inbox ordering with drain-all doorbells * no-mistakes(review): Surface unwritable inbox ladder bookkeeping * no-mistakes(review): Silence ladder failures after inbox acknowledgement * no-mistakes(document): Update steering inbox documentation * no-mistakes: apply CI fixes * feat(bin): add fast local lint mode (kunchenguid#2891) * feat: add fast local lint mode * fix: preserve complete fm-lint help * fix: isolate fast lint mode * no-mistakes(document): Clarify lint mode documentation ownership * no-mistakes: apply CI fixes * feat(bin): deliver remote steers through durable inboxes (kunchenguid#2901) * feat(bin): deliver remote secondmate steers through durable task inboxes Stage 2 of the inbox+doorbell steer channel (stage 1: kunchenguid#2856). A remote secondmate steer now crosses fm-on.sh as a durable record written idempotently into the remote home's steering inbox plus a best-effort remote doorbell, and the last typed-payload steer transport is deleted: - fm-remote-secondmate-control.sh cmd_send writes the record via the new fm_task_inbox_write_idempotent and rings the doorbell; it no longer types the payload through an inner fm-send at an explicit pane target. - fm-send.sh routes every remote text steer (harness-native included, which marking already reduced to chat) onto the remote inbox leg, retries the identical leg once on ssh 255, closes --resolve-key decisions at enqueue for remote too, and preserves a marked request's reply expectation when completion stays unknown. The exit-3-as- delivered remap, the 255 do-not-resend trap, and the remote typed submit block are removed. - fm-task-inbox-lib.sh owns the idempotent enqueue: an exact-body re-run lands on the existing record, handled or not, so an ambiguous transport can always be safely re-run. - Tests pin the new contract end to end (record + doorbell + no typed payload across ssh, one-record idempotence under an ambiguous transport, enqueue-time decision close, loud real failures, and the deleted typed-payload behaviors gone), and AGENTS.md plus docs/remote-secondmates.md describe the remote leg's new semantics. * no-mistakes(review): Harden remote inbox delivery against lifecycle races * no-mistakes(review): Enable correlation-preserving remote steer resends * no-mistakes(review): Fail closed on stale correlation resends * no-mistakes(review): Include home context in remote resend commands * no-mistakes(review): Lock and revalidate remote parent routes * no-mistakes(document): Clarify remote steer retry documentation * no-mistakes: apply CI fixes * feat: add persistent Pi supervision branch (kunchenguid#2858) * wip: forked supervision on Pi (checkpoint before docs) * fix(pi-branch): harden mirror delivery, fallback encoding, and session replacement Peek-then-shift mirror flush so a failed append retries instead of dropping; durable mirror cursor commits only after delivery into the branch; the main fallback wake is operational-encoded like every watcher injection; session_shutdown quiesces the generation and session_start re-arms, so /new and /resume no longer kill the branch permanently. Registers the extension in the strict typecheck, adds the dispatch handshake test, the branch extension suite, the bash-level regression suite, the session-start replay test, and the opt-in real-SDK live guard. * test(fixtures): carry the branch-dispatch lib and lease lib into isolated fixtures The watcher extension now imports lib/fm-branch-dispatch.ts and fm-teardown sources fm-lease-lib.sh, so every fixture that copies or symlinks those files in isolation gains the new sibling. * no-mistakes(review): Prevent shutdown wake loss and serialize lease claims * no-mistakes(review): Durably hand off wakes and retain portable leases * no-mistakes(review): Require durable reports and clear disposed branch leases * no-mistakes(review): Enforce per-wake outcomes and quiescent lease cleanup * no-mistakes(review): Require wake acknowledgements and tighten branch lifecycle boundaries * no-mistakes(review): Require complete acknowledgements and replay cleanup failures * no-mistakes(review): Bind supervision to lock ownership and durable delivery * no-mistakes(review): Activate branch lazily after session lock acquisition * no-mistakes(review): Preserve undelivered mirror context across extension rebinds * no-mistakes(review): Acknowledge startup replay only after main delivery * no-mistakes(review): Isolate replay metadata from untrusted digest content * no-mistakes(review): Reject duplicate reports for active wake sequences * no-mistakes(review): Retain failed fallbacks and deduplicate outcome replay * no-mistakes(review): Deduplicate durable outcomes and cache delivery receipts * no-mistakes(review): Anchor wake sequence matching to outcome fields * no-mistakes(document): Clarify Pi supervision durability contracts * no-mistakes(lint): Fix ShellCheck issues in branch supervision scripts * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * refactor(pi-branch): collapse to confused-agent-grade guards per captain decision Captain decision A: the lease/actor guards target the CONFUSED-AGENT threat model bin/fm-gate-refuse-lib.sh already documents; adversarial-grade separation is impossible in the shared-process design and is filed as separate follow-up work. Rip out the machinery that chased it: the generation fence and shell-provenance markers, the wrapper-tagged ancestry walks, guard auto-claim with per-script release traps, the pending-wake files and ack-receipt correlation (the durable wake queue already re-presents anything unacknowledged), the delivery-receipt store with contiguous cursor advancement, the session-start replay-metadata channel, and the branch tool quiescence counters. Keep the behaviors the board requires, each on its simplest implementation: lazy per-action session-lock ownership (cold start activates after the lock lands; a secondary session stays inert), mirror durability across extension rebinds via the durable cursor, replay-exactly-once from the one read cursor, the awaited operational-encoded fallback, per-generation stray-lease cleanup, session-lock-bound lease liveness (a recycled pid or a non-Pi home never honors a leftover lease), the loud accidental-override guards (readonly actor prelude, cross-actor claim refusal), and the role-partition refinements (no forced teardown, no direct relaunch for the branch). Default-on-for-Pi is unchanged. * no-mistakes(review): Enforce lock ownership and serialize lease mutations * no-mistakes(review): Synchronize guard cleanup and bind leases to lock owner * no-mistakes(review): Report outcomes before acknowledging durable wakes * no-mistakes(review): Restrict leases to Pi and instruct main claims * no-mistakes(review): Reject malformed lease locks and torn outcome tails * no-mistakes(review): Validate complete outcome tails before appending * no-mistakes(review): Guard branch side effects across session replacements * no-mistakes(document): Update Pi supervision durability and lease documentation * no-mistakes(lint): Suppress intentional nested-shell expansion warning * no-mistakes: apply CI fixes * fix(pi-branch): authorize lease releases by caller * fix(lint): break redundant source-analysis path in fm-lease-lib.sh fm-lease-lib.sh's lazy fallback source of fm-wake-lib.sh gave ShellCheck's --external-sources traversal a second path into an already 1540-line file that fm-send.sh and fm-teardown.sh also source directly, blowing up the recursive analysis past CI's lint timeout. Mark it a source=/dev/null analysis boundary, matching the existing fm-task-inbox-lib.sh convention. Also restores bin/fm-lint.sh and tests/fm-lint.test.sh to the shared serial-lint definition (dropping an unrelated parallel-sharding change that was itself hanging and masked this root cause). * no-mistakes(document): Correct lease caller-authorization documentation * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * fix(bin): parallelize startup network sweeps (kunchenguid#2927) * feat(bin): parallelize session-start remote secondmate network sweeps Run per-secondmate liveness and convergence probes concurrently and overlap clone refresh, while replaying each mate's fail-closed diagnostic in original order. Ignore scratchpad* so untracked scratch no longer blocks remote sync. Co-authored-by: Cursor <cursoragent@cursor.com> * no-mistakes(document): Document parallel startup network sweeps * no-mistakes(lint): Fix empty environment assignment lint warning * no-mistakes: apply CI fixes --------- Co-authored-by: Cursor <cursoragent@cursor.com> * test: handle absent watcher wake queues (kunchenguid#2845) * fix(tests): count declared-pause wakes without crashing on an absent queue The exited-declared-pause case counts queued stale wakes by handing state/.wake-queue straight to awk. A watcher that queues nothing never creates that file, and awk aborts on a missing path before its END rule runs, so the count collapses to the empty string. The next comparison then fails as an integer-expression error and surfaces as a wake flood with no number, hiding the real contract breach the following grep names. Read the queue the way the drain-count assertion at the end of this file already does: silence awk's open error and default an absent queue to zero. Applied to all four counts in this case, including the live external-decision gate pair whose queue an acknowledged drain can also leave behind. An absent queue now reports "did not use the bounded paused recheck", while a genuine flood still fails with its real count. Fixes kunchenguid#2628 * no-mistakes: apply CI fixes * no-mistakes: apply CI fixes * style(pi): distinguish routine and captain supervision merge notes by icon (kunchenguid#2934) * style(pi): restyle supervision merge notes with a sailboat and matching pad Secondary-session notes were flush against the TUI edge and fully tinted. Use the sailboat prefix, Pi's default outputPad, boat-only color, and dim remainder so they sit like real messages. * style(pi): distinguish routine and captain merge notes by icon only Visible notes now lead with a sailboat or anchor, then only the dim outcome. Drop the branch-merged wording and verdict brackets so the icon is the only kind signal. * docs(pi): add the approved multi-brain architecture poster (kunchenguid#2938) The markdown contract stays the owner; the still is only the visual of the idea. * feat(pi): default branch supervision and route heartbeats (kunchenguid#2939) * fix(bin): bound remote job worker supervisor restarts (kunchenguid#2942) * fix(bin): bound remote worker supervisors * no-mistakes(review): release incumbent supervisor before starting its replacement * no-mistakes(review): wait out a healthy same-root supervisor instead of replacing it * no-mistakes(review): narrow remote worker change to restart accounting only * no-mistakes(document): clarify supervisor restart guard is a lifetime total * fix: safely split supervision wake handling by actor (kunchenguid#2953) * feat(bin,pi): per-actor wake consume, silent success gating, merge-poll dedup Three related fixes to the shared wake-drain and Pi supervision-branch dispatch machinery so a routine success is never main-blocking and a mixed queue can safely split between actors. 1. Successful routine results no longer create main-blocking wake rows. fm-startup-network.sh only enqueues a check: startup-network wake when the deferred result is actionable (state is not "done", or the report carries a bootstrap-diagnostics actionable prefix); a clean success stays durable in the report file without ever waking the agent. 2. Per-actor wake-drain consume contract. bin/fm-wake-drain.sh now scopes presentation and --ack-through to the current actor (bin/fm-lease-lib.sh's fm_lease_actor): main keeps the original whole-queue cutoff behavior, unaffected. A branch actor (FM_SUPERVISION_ACTOR=branch, set only inside the Pi supervision branch's own bash tool calls) is scoped to an explicit eligible-row snapshot instead of a cutoff comparison, so it can never remove a row it was not granted - the fix for the swallow risk that used to force an all-or-nothing whole-queue fallback to main. .pi/extensions/lib/fm-branch-dispatch.ts's scopeForUnreadWake is the single owner of eligibility: a check-kind row (merge-confirmation polls, Relay mentions, credential/auth failures) is now excluded rather than vetoing the whole scan for a non-heartbeat wake, while a heartbeat review keeps its original all-or-nothing rule unchanged. writeEligibleRowsSnapshot publishes the exact eligible sequence numbers before every branch prompt; fm-primary-pi-watch.ts's offer still refuses a check-kind trigger outright so a main-only close is never itself routed to the branch. 3. A repeat identical merged-PR-poll result for an already-notified task is absorbed instead of enqueued again. A poll's own retirement state is scoped to one registration and cannot see a prior registration's outcome, so a task re-registered after its merge was already surfaced would otherwise wake main a second time for the same event. bin/fm-pr-lib.sh's new per-task pr-poll-merge-notified marker survives across re-registrations to catch that case; the first notification for a task still reaches main unchanged. Regression tests colocated in tests/fm-startup-network.test.sh, tests/fm-wake-queue.test.sh (including the mixed-queue no-swallow property), tests/fm-pi-branch-extension.test.sh, and tests/fm-pr-check-security.test.sh. docs/watcher-continuity.md and docs/pi-supervision-branch.md updated for the new contracts. * no-mistakes(review): Bind merge deduplication to canonical PR identity * no-mistakes(review): Serialize wake row ownership across main and branch * no-mistakes(review): Bind branch grants and deduplicate within actor claims * no-mistakes(review): Fallback main-owned wake claims to main delivery * no-mistakes(review): Clarify silent startup success guidance * no-mistakes(review): Release residual branch grants after settled prompts * no-mistakes(review): Reject truncated wake rows as corrupted * no-mistakes(document): Document per-actor routing and silent startup success * no-mistakes(lint): Fix ShellCheck findings in wake grant and startup test * no-mistakes: apply CI fixes * fix(pi): hide branch outcomes tool rows in Calm (kunchenguid#3024) * Hide branch outcome tool in Pi Calm * no-mistakes(review): Preserve stock outcomes rendering and document tool audit * no-mistakes(review): Document branch read tool audit disposition * no-mistakes(review): Match stock outcomes output sanitization * no-mistakes(document): Document Calm custom-tool visibility * no-mistakes(review): Secure Pi wake batches and complete orchestrator documentation * no-mistakes(document): Correct current orchestrator comparison facts * Add opt-in resumable Codex crewmates (#4) * Add opt-in resumable Codex crewmates * fix(bin): restore parked Codex binding after crash --------- Co-authored-by: wenkxu <v-wenkxu@expediagroup.com> * Add persistent human names and names-first routing (#6) * Add persistent human callsigns * Close persistent callsign safety gaps --------- Co-authored-by: wenkxu <v-wenkxu@expediagroup.com> * Bind tmux task routing to stable pane identity (#7) * Bind tmux task routes to stable pane identity * Handle failed tmux identity binding --------- Co-authored-by: wenkxu <v-wenkxu@expediagroup.com> * fix: gate Codex checkpoints on genuine idle (#12) Co-authored-by: wenkxu <v-wenkxu@expediagroup.com> * feat: register external workspace routes (#13) Co-authored-by: wenkxu <v-wenkxu@expediagroup.com> * feat(codex): propagate assigned display titles (#14) * feat(codex): propagate assigned display titles * test(codex): keep title resume guard isolated --------- Co-authored-by: wenkxu <v-wenkxu@expediagroup.com> * no-mistakes(review): Secure Pi wake batches and complete orchestrator documentation * no-mistakes(document): Correct current orchestrator comparison facts --------- Co-authored-by: Christopher McKay <101884182+karotkriss@users.noreply.github.com> Co-authored-by: Kun Chen <3233006+kunchenguid@users.noreply.github.com> Co-authored-by: Cursor <cursoragent@cursor.com> Co-authored-by: stanzhang <thinking.chang@gmail.com> Co-authored-by: wenkxu <v-wenkxu@expediagroup.com>
User description
Summary
Validation
The no-argument bin/fm-lint.sh was also run with pinned tools. Its changed-path selection follows the shared upstream origin/main remote instead of this authorized fork base, so it reports pre-existing fork warnings in tmux/identity/control and older tests; this change introduces no remaining warning in its owned lines.
CodeAnt-AI Description
Assign stable, user-visible titles to Codex worker conversations
What Changed
Callsign · ProjectCode · TaskLabelbefore the initial task or resumed task is delivered.Impact
✅ Stable worker titles across launch and resume✅ Clearer Codex and Herdr conversation identification✅ Fewer launches with ambiguous or malformed task identity💡 Usage Guide
Checking Your Pull Request
Every time you make a pull request, our system automatically looks through it. We check for security issues, mistakes in how you're setting up your infrastructure, and common code problems. We do this to make sure your changes are solid and won't cause any trouble later.
Talking to CodeAnt AI
Got a question or need a hand with something in your pull request? You can easily get in touch with CodeAnt AI right here. Just type the following in a comment on your pull request, and replace "Your question here" with whatever you want to ask:
This lets you have a chat with CodeAnt AI about your pull request, making it easier to understand and improve your code.
Example
Preserve Org Learnings with CodeAnt
You can record team preferences so CodeAnt AI applies them in future reviews. Reply directly to the specific CodeAnt AI suggestion (in the same thread) and replace "Your feedback here" with your input:
This helps CodeAnt AI learn and adapt to your team's coding style and standards.
Example
Retrigger review
Ask CodeAnt AI to review the PR again, by typing:
Check Your Repository Health
To analyze the health of your code repository, visit our dashboard at https://app.codeant.ai. This tool helps you identify potential issues and areas for improvement in your codebase, ensuring your repository maintains high standards of code health.