Repository navigation
fix: prevent false-positive session-start endpoint liveness - #67
Merged
Merged
Conversation
tmux display-message -t answers success for an absent window or session while any server runs, so session start printed 'endpoint: alive' for a closed window and fm-crew-state read it as a readable pane. fm_backend_target_exists now matches the recorded session:window name against the exact session's window inventory, and fm-crew-state uses the same check.
…rted checks to fail: exact-session selectors now resolve correctly, window inventory preserves both bare and qualified representations, and supervisor fixtures expose numeric pane targets. Seeded the explicit named-supervisor fixture and updated fixture documentation; production liveness behavior and assertions remain unchanged. Verification: secondmate lifecycle, daemon, wake-daemon lifecycle, and backlog-handoff suites passed completely; all 18 focused reconcile delivery cases passed, including the originally failing cooldown case. Crew-state, tmux liveness, and scoped full ShellCheck passed. Full reconcile and session-start runs encountered separate local timing failures: an 8-second collector operation exceeded a 5-second assertion, and a live endpoint read exceeded its 2-second bound. Temporary verification files were removed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Intent
I would like all bugs to be fixed so tomorrow can be focused entirely to Vernant and not problems preventning Vernant from getting built.
A recommended architecture should come with proof.
Context (the backlog record for this bug, as filed):
Correct session-start alive reporting when recorded endpoint is missing (repo: firstmate) (kind: ship) (priority: 0) (since 2026-10-02)
Owner: Main Firstmate commissioning; bounded implementation owner to be assigned through normal isolated dispatch. Confirmed current context reemit saysAPIguard endpointalive while actual fm-crew-state returnsunknown/missing. OriginalAPIguardtaskwas intentionallystopped and sourceholds958c95ed/341b4d0c for dependentpublication; no relaunch authorized by mismatchalone. Next action: compare actual startup presence verdict with authoritativebackend/harnessstate at sameidentity, reproduce in isolatedfixture and repair false-positive report without claimingunknown isdead or sweepingotherhomes. Preserve retainedsource and records. Distinct reportinggap, not orphanedPR or secondAPIguardimplementation. Review due2026-10-03T00:00:00Z.
What Changed
Risk Assessment
✅ Low: The changes are bounded to endpoint-presence verification and reporting, preserve unknown versus absent states and retained records, and maintain fail-closed action gates without introducing new lifecycle machinery.
Testing
Reproduced the baseline reporting defects, then exercised missing endpoints, exact-session identity, pane selectors, real transport unreadability, and remote-record handling through the running CLI product. All scenarios passed with digest transcripts, reconciliation output, backend-call audits, and retained-state hashes captured. Validation used passive reemit interfaces without launching a primary or touching fleet lifecycle; all disposable fixtures were removed.
Evidence: Baseline startup digest reproducing false-alive reporting
Source: Baseline startup digest reproducing false-alive reporting
Evidence: Target startup digest with correct endpoint presence
Source: Target startup digest with correct endpoint presence
Evidence: Closed endpoint crew-state reconciliation
Source: Closed endpoint crew-state reconciliation
Evidence: Neighboring-session reconciliation before and after
Source: Neighboring-session reconciliation before and after
Evidence: Real tmux selector presence results
Source: Real tmux selector presence results
Evidence: Real transport error and unknown reconciliation
Source: Real transport error and unknown reconciliation
Evidence: Remote records remain unknown in startup
Source: Remote records remain unknown in startup
Pipeline
Updates from git push no-mistakes
✅ **intent** - passed
✅ No issues found.
✅ **Rebase** - passed
✅ No issues found.
🔧 **Review** - 3 issues found → auto-fixed (2) ✅
bin/backends/tmux.sh:184- The intent requires repairing reporting "without claiming unknown is dead", but the addedwindows=$(fm_backend_tmux_window_inventory "=$session") || return 1collapses an unreadable inventory into absence. For a retained live task whose inventory query fails non-definitively, window_inventory returns 1, this helper returns 1, and fm-session-start.sh:908-918 printsendpoint: dead; fm-crew-state instead preserves unknown/unreachable. Related changed sites are bin/backends/tmux.sh:168-170 (explicitly endorses this collapse), bin/fm-backend.sh:939-952 (shared presence contract and dispatch), and bin/fm-crew-state.sh:321 (consumer that subsequently distinguishes unreadability). tests/fm-session-start.test.sh:1393 also retains an expectation of dead for unreadability. Preserve the inventory's unreadable-versus-absent distinction through the shared presence boundary and startup rendering, and correct that expectation.bin/backends/tmux.sh:183- The nonnumeric branch mistakes valid pane selectors for literal window names. With sessionfirstmate, windowmainat index 0, and pane 1,firstmate:0.1queries window names and searches for the literal0.1, returning absent although the pane exists.firstmate:main.1fails identically; exact-qualified selectors are also misinterpreted. These forms remain reachable through fm-send.sh:418-426's explicit backend-target interface and the documented FM_SUPERVISOR_TARGET override, causing sends or daemon startup to refuse valid endpoints. Related changed sites are bin/fm-backend.sh:951-952 (dispatch applies this parsing to every caller), bin/backends/tmux.sh:177-191 (selector classification), and docs/tmux-backend.md:49-50 (presence description). Distinguish canonical recorded window names from supported pane/qualified selectors and verify each against its corresponding inventory.bin/fm-crew-state.sh:321- The crew-state fix still permits a missing recorded session to be replaced by a neighboring session. Suppose retained metadata recordswork:fm-mate, sessionworkis gone, and another home'swork-othersession contains a livefm-matewindow. The new exact-session presence check fails, but fm-crew-state.sh:1265 then calls fm_backend_tmux_agent_state, whose inventory query at bin/backends/tmux.sh:329 usesworkwithout=. Tmux's session-prefix resolution can selectwork-other; its matching window and harness producealive, so the fallback continues and a secondmate's stale working log is reported as current at fm-crew-state.sh:1310-1316. Startup correctly reports the recorded endpoint absent, leaving the same-identity reconciliation gap reachable. Related changed sites are bin/backends/tmux.sh:184 (exact presence boundary) and docs/tmux-backend.md:50-51 (identity guarantee). Require exact recorded-session targeting in the shared recovery classifier as well, rather than patching another consumer.🔧 Fix applied.
3 errors still open:
bin/backends/tmux.sh:183- The nonnumeric branch mistakes valid pane selectors for literal window names. With sessionfirstmate, windowmainat index 0, and pane 1,firstmate:0.1queries window names and searches for the literal0.1, returning absent although the pane exists.firstmate:main.1fails identically; exact-qualified selectors are also misinterpreted. These forms remain reachable through fm-send.sh:418-426's explicit backend-target interface and the documented FM_SUPERVISOR_TARGET override, causing sends or daemon startup to refuse valid endpoints. Related changed sites are bin/fm-backend.sh:951-952 (dispatch applies this parsing to every caller), bin/backends/tmux.sh:177-191 (selector classification), and docs/tmux-backend.md:49-50 (presence description). Distinguish canonical recorded window names from supported pane/qualified selectors and verify each against its corresponding inventory.bin/fm-crew-state.sh:321- The crew-state fix still permits a missing recorded session to be replaced by a neighboring session. Suppose retained metadata recordswork:fm-mate, sessionworkis gone, and another home'swork-othersession contains a livefm-matewindow. The new exact-session presence check fails, but fm-crew-state.sh:1265 then calls fm_backend_tmux_agent_state, whose inventory query at bin/backends/tmux.sh:329 usesworkwithout=. Tmux's session-prefix resolution can selectwork-other; its matching window and harness producealive, so the fallback continues and a secondmate's stale working log is reported as current at fm-crew-state.sh:1310-1316. Startup correctly reports the recorded endpoint absent, leaving the same-identity reconciliation gap reachable. Related changed sites are bin/backends/tmux.sh:184 (exact presence boundary) and docs/tmux-backend.md:50-51 (identity guarantee). Require exact recorded-session targeting in the shared recovery classifier as well, rather than patching another consumer.bin/fm-session-start.sh:911- The new inventory-based check falsely reports healthy remote secondmates as dead. A supported remote spawn writeswindow=remote:mate,remote_host=<host>, andremote_target=<actual endpoint>, but nobackend=(bin/fm-spawn.sh:1360-1377). Startup therefore probesremote:mateon the local tmux server; when that server has no session namedremote, the new inventory path returns absence and the digest printsendpoint: dead. The remote endpoint was never queried. This contradicts the required comparison "at sameidentity" and "without claimingunknown isdead". Round 1's unknown-preservation fix left this sibling classification path unhandled. Related changed sites: bin/backends/tmux.sh:181-185 converts local inventory absence into status 1; bin/fm-backend.sh:945-946 dispatches that local probe; bin/fm-session-start.sh:908,917 applies its verdict to every metadata record. Recognizeremote_hostat the fleet-digest metadata boundary before local probing, and report remote liveness as unknown unless the recorded remote route supplies authoritative evidence; bin/fm-crew-state.sh:271-305 already demonstrates this distinction.🔧 Fix applied.
✅ Re-checked - no issues remain.
✅ **Test** - passed
✅ No issues found.
python3 .live-validation-driver.py— disposable driver exercising real Firstmate scripts against a private real tmux server; an initial driver-only assertion was corrected before successful runs. Driver and fixtures were removed afterward.Replayed base-commitfm-session-start.sh --reemitandfm-crew-state.shagainst the same isolated records, reproducing false-alive reporting and stale neighboring-session activity.bin/fm-session-start.sh --reemitandbin/fm-crew-state.sh APIguardafter closing the recorded window; compared retained metadata, status, and source SHA-256 hashes before and after.bin/fm-crew-state.sh missing-sessionandfm_backend_agent_state tmux work:fm-matewith only the neighboringwork-other:fm-mateendpoint present.Executedfm_backend_target_exists tmux <target>for named, indexed, ID, bare-ID, and exact-qualified selectors; exercised pane selectors through the actual startup digest.Ran startup, crew-state, andfm_busy_classify_livethrough a real tmux connection to a non-socket file, then confirmed the original endpoint still existed on the private server.Ran remote-only startup reemit for host-marked, remote-prefix, and legacy no-backend records; audited backend calls and confirmed retained records remained unchanged.Stopped the private tmux server and confirmed all disposable worktree fixtures, sockets, and driver files were removed.✅ **Document** - passed
✅ No issues found.
✅ **Lint** - passed
✅ No issues found.
✅ **Push** - passed
✅ No issues found.