fix(bin): identify the session lock holder from executable identity - #3
Conversation
Two defects with one root cause, plus the fail-open they exposed and a third identity gap the new live guard found. Ownership was decided purely by walking the current shell's ancestry, so a session whose harness rehosted its work under its own pty - a tree that reparents to init and never reaches the pid holding the lock - reported itself as a competing session and degraded session start to read-only. Ownership now also accepts a session-cohort proof: the lock names a live harness that provably STARTED this session, from the launch marker the harness writes into every process it spawns, AND is co-located with it in the same multiplexer pane or on the same controlling terminal. The two signals are AND-ed, not OR-ed, so a genuinely separate concurrent session sharing one pane is still refused, and the absence of either signal falls back to the ancestry verdict rather than widening ownership. A suspended holder never released the lock, because kill -0 succeeds on a stopped process. A durably stopped harness - confirmed over several samples so a momentary stop is not misread - no longer counts as holding: it mutates nothing while stopped and nothing guarantees it resumes, so treating it as a holder converted a recoverable state into a permanent lockout of the whole home. fm-lock.sh status reports such a holder as reclaimable and acquisition names any takeover from a live holder. The turn-end guard's bounded attended fail-open required an auto-arm that had recorded an exhausted failure, so an auto-arm that never ran at all left no failure to advance and the guard re-blocked until Claude's hard 8-block override. It now opens on a second, independent condition - a verified absent auto-arm - under the same block budget and one-shot alarm, so supervision being unstartable for a reason the guard cannot repair always reaches the valve however it arose. The new live guard also found that npm-installed codex reports comm MainThread with no interpreter name, and was not identified as a harness at all. Identity for that shape now comes from the interpreter's script path in argv, by whole path component only. Covered by tests/fm-session-lock-ancestry.test.sh with real processes and no harness, including the divergence each case rests on, and by the opt-in tests/fm-session-lock-identity-live-e2e.test.sh against every installed harness.
FM_SESSION_LAUNCH_MARKERS holds exactly one entry, so the same-session cohort proof can only ever fire for Claude. codex, opencode, pi, pi-signed, grok, kimi, and cursor are decided by process ancestry alone, which means a session those harnesses rehost outside its own process tree still refuses its own home. That is the unfixed half of the defect for those adapters, and nothing in the mechanism said so. Records it at the mechanism's own owner, with the reason it is safe rather than merely incomplete: a missing marker removes an accept path and never a refusal, so a marker-less harness lands exactly on its previous behavior, while a guessed variable name either buys nothing or gets believed as half the proof that keeps a separate concurrent session out. Extending the table is therefore a verification task, not an editing one. The recovery-boundary contract and the per-harness evidence record carry a pointer to it, so the third non-competing holder cannot be read as cross-harness behavior. No mechanism change: the library diff is comment-only.
…pling fail closed
…v path components
The cohort's acceptance typing compared the reported command name for exact equality with a verified harness name, and the Claude Code install actually in use never satisfies it. Its executable is a single-file native build named claude.exe, so a process that execs it reports that name verbatim; and it renames its own background pty worker task to `claude bg-pty-host`, which the 15-character kernel task-name limit cuts to `claude bg-pty-h`. Both shapes were identified as a harness and typed as nothing, so the cohort refused and the session start degraded to read-only against its own home while it was the only session alive - the primary defect, for the shape actually installed. Normalize the reported command name before the SAME exact-equality test: take its first space-delimited word, undoing the truncation, then strip the platform executable suffix. The comparison itself is unchanged, so this cannot become a prefix rule - claudette, claude-code, node and python3 still name nothing - and path components are untouched, so pi keeps the basename anchoring that stops an interior /home/pi component from naming a harness. Neither step reaches the script-path caller, which sees neither artifact. fm_harness_exec_kind opened with a fourth hand-written copy of the loop that fm_harness_basename_name already is; it now delegates, leaving one owner for the exact-equality test. Every fixture in the suite and in the live guard previously symlinked an exact harness name, so neither could observe a suffixed or truncated command name at all, which is why this went unnoticed. Both now carry claude.exe and the truncated worker as acceptance cases with claudette, claude-code, node and python3 as negative controls. The portable cases run real processes, so the kernel and ps produce the name rather than a stub, and each asserts the divergence before the verdict so a platform that stopped producing an artifact fails instead of passing vacuously. The live guard additionally launches each harness as its own INSTALLED executable, which is the only way to see what a vendor actually calls its binary, and fails naming the harness and version; an interpreter script is reported rather than failed, because its own name never reaches a reported command name. The verification record is refreshed from a real run on this machine, and the npm-layout recognition limit is narrowed to what is still true: the observed Claude Code npm install is identified by its own executable name, and only a node-hosted launch out of that layout stays unnamed.
…marker verification pointer
…-96a # Conflicts: # .agents/skills/harness-adapters/SKILL.md # bin/fm-test-run.sh # docs/turnend-guard.md # docs/verification/runtime-backends.md # docs/watcher-continuity.md
…live_gate The live-e2e test predates fork/main's shared live-gate contract (fm_live_gate in tests/lib.sh, enforced fleet-wide by tests/fm-live-gate.test.sh). Switch it to the shared default-on gate so it reports the standard skip line and participates in the fleet-wide FM_LIVE=0 sweep instead of its own private env check.
|
Merge resolution and check-matrix result for this branch update. Merge vs rebase: used a merge, not a rebase. Nearly all 18 commits on this branch touch the same Conflict resolution, one file at a time:
A real integration gap found and fixed (not a git conflict): Check matrix: 12 of 14 green, plus:
GitHub reports the PR mergeable. Merge authority is the captain's; not merged by me. |
Lands the reviewed session-lock identity work on the fork so the full check matrix runs against it.
This branch is the line carried by upstream PR kunchenguid#2929 at head
96a879c5. It was chosen overfm/fm-session-lock-identity-reviewed-2026-09-01(8c3f480), which holds equivalent work on a materially older base and does not registertests/session-signals.shin the changed-test map.Opened to obtain test evidence before merging: upstream runs a single review check, while this fork runs the full matrix.
🤖 Generated with Claude Code