Skip to content

feat(herdr): group project tasks into a shared durable herdr workspace - #3

Merged
friesentius merged 6 commits into
mainfrom
fm/fm-herdr-project-workspaces
Sep 11, 2026
Merged

friesentius merged 6 commits into
mainfrom
fm/fm-herdr-project-workspaces

Conversation

@friesentius

@friesentius friesentius commented Sep 11, 2026 •

Copy link
Copy Markdown
Owner

Intent

In the future, can you create a herdr workspace for each project that you get a request for, and then create a tab in that workspace for each worker/task. Then obviously close the tabs/workspaces when things are complete.

What Changed

  • Add a project value for config/herdr-presentation-spaces that groups every task for one project into that project's own durable herdr workspace, instead of the default per-task disposable projection or the ordinary per-home flat layout.
  • Derive a collision-resistant project key (sanitized directory basename + short hash of the absolute path) and persist the workspace binding per home in state/proj-<project-key>.herdr-workspace; the record is verified live (matching home, session, and workspace presence) before reuse and self-heals with a fresh workspace/record when stale or missing. Unsafe basenames or a herdr release below the version floor fall back to the ordinary flat layout with a warning instead of aborting the spawn.
  • Wire the new project-grouped routing into bin/fm-spawn.sh's herdr case arm (reusing the existing presentation order lock and fm_backend_herdr_create_task path), and add e2e (tests/fm-backend-herdr-project-workspace-e2e.test.sh) and portable unit test coverage (tests/fm-backend-herdr.test.sh).
  • Document the new topology and its contract in docs/herdr-backend.md, docs/verification/runtime-backends.md, and AGENTS.md.

🤖 Generated with Claude Code

Risk Assessment

✅ Low: Both previously-requested fixes (basename-collision disambiguation and graceful fallback for unsafe project-key characters, plus the follow-up hash-out-of-visible-label fix) are correctly and consistently implemented across bin/backends/herdr.sh, bin/fm-spawn.sh, docs, and both test suites, with locking/self-heal/teardown semantics verified by real e2e tests; the only remaining issue found is a cosmetic stale doc-heading reference in test comments.

Testing

Both the herdr backend's unit suite and its real-binary end-to-end test pass cleanly, and a manual real-herdr CLI transcript directly confirms the target behavior: a clean project-labeled workspace (no hash suffix) shared across tasks with one tab each, a persisted record whose project= field carries the collision-resistant hash while its label= field and the visible herdr label stay hash-free, and workspace teardown exactly when the project's last task completes — matching the user's requested herdr-workspace-per-project/tab-per-task/close-on-completion intent and the label-hygiene fix from the last commit.

Evidence: Real herdr CLI transcript: project workspace + tabs + teardown

Source: Real herdr CLI transcript: project workspace + tabs + teardown

herdr workspace list shows one clean-labeled 'proj-frontend' workspace shared by two tasks' tabs (w1:t2, w1:t3); the persisted record's project= field carries 'frontend-4002f115' (hashed) while its label= field stays 'proj-frontend' (hash-free); after tearing down both tasks the workspace list for that id returns empty, confirming teardown removes the drained project workspace.

Manual end-to-end CLI transcript demonstrating the herdr project-workspace
feature (config/herdr-presentation-spaces=project): one herdr workspace per
project, one tab per task/worker within it, and workspace teardown once the
project drains to zero tasks.

Driven with the REAL herdr, jq, and treehouse binaries and the REAL
bin/fm-spawn.sh / bin/fm-teardown.sh, in an isolated throwaway lab session
(never the captain's default session), mirroring the safety conventions of
tests/fm-backend-herdr-project-workspace-e2e.test.sh. All lab artifacts
(session, worktrees, scratch project, scratch home) were cleaned up
automatically at the end of the run.

This transcript also demonstrates the specific fix under test in commit
a6ae310 ("keep the collision-resistant path hash out of the visible
label"): the persisted record's project= field carries the collision-
disambiguating hash ("frontend-4002f115"), while the actual herdr workspace
label and the record's own label= field stay a clean "proj-frontend" with
no hash suffix.

=== Project directory: /tmp/fm-herdr-proj-demo.EInW4D/frontend ===

>>> Spawning task cm1 for project 'frontend' (config/herdr-presentation-spaces=project) <<<
warning: /tmp/fm-herdr-proj-demo.EInW4D/primary-home/data/cm1/launch-brief.md records no delivery contract line (scaffolded before ship briefs recorded one); launching on the explicit --mode no-mistakes - confirm its definition of done matches
spawned cm1 harness=sh kind=ship mode=no-mistakes yolo=off window=fm-lab-herdr-proj-demo-1814532:w1:p2 worktree=/home/friesent/.treehouse/frontend-941995/1/frontend

>>> herdr workspace list after cm1 (real herdr CLI) <<<
{
  "workspace_id": "w1",
  "label": "proj-frontend"
}

>>> Persisted project-workspace record (state/proj-<hashed-key>.herdr-workspace) <<<
--- /tmp/fm-herdr-proj-demo.EInW4D/primary-home/state/proj-frontend-4002f115.herdr-workspace ---
project=frontend-4002f115
home=/tmp/fm-herdr-proj-demo.EInW4D/primary-home
session=fm-lab-herdr-proj-demo-1814532
workspace_id=w1
label=proj-frontend

>>> Spawning a SECOND task, cm2, for the SAME project 'frontend' <<<
warning: /tmp/fm-herdr-proj-demo.EInW4D/primary-home/data/cm2/launch-brief.md records no delivery contract line (scaffolded before ship briefs recorded one); launching on the explicit --mode no-mistakes - confirm its definition of done matches
spawned cm2 harness=sh kind=ship mode=no-mistakes yolo=off window=fm-lab-herdr-proj-demo-1814532:w1:p3 worktree=/home/friesent/.treehouse/frontend-941995/2/frontend

>>> herdr workspace list after cm2 - still ONE workspace for this project <<<
{
  "workspace_id": "w1",
  "label": "proj-frontend"
}

>>> herdr tab list for that shared workspace - one tab per task <<<
{
  "tab_id": "w1:t2",
  "title": null
}
{
  "tab_id": "w1:t3",
  "title": null
}

CONFIRMED: cm1 and cm2 share workspace w1 (one workspace per project, one tab per task).

>>> Tearing down cm1 (not the project's last task) <<<
teardown cm1 complete (window fm-lab-herdr-proj-demo-1814532:w1:p2, worktree /home/friesent/.treehouse/frontend-941995/1/frontend)
>>> herdr workspace list after tearing down cm1 - workspace survives (cm2's tab still open) <<<
{
  "workspace_id": "w1",
  "label": "proj-frontend"
}

>>> Tearing down cm2, the project's LAST task <<<
teardown cm2 complete (window fm-lab-herdr-proj-demo-1814532:w1:p3, worktree /home/friesent/.treehouse/frontend-941995/2/frontend)
>>> herdr workspace list after tearing down cm2 - the project workspace is GONE <<<

CONFIRMED: draining the project's last task closed its shared workspace.

=== Demo complete: no lab artifacts left behind ===

(Trimmed: per-run watcher-supervision and backlog-update advisory banners
emitted by fm-teardown.sh, unrelated to the herdr project-workspace behavior
under test, were omitted from the teardown sections above for readability.
The full raw output was reviewed and contained no other differences.)

Pipeline

Updates from git push no-mistakes

✅ **intent** - passed

✅ No issues found.

✅ **Rebase** - passed

✅ No issues found.

⚠️ **Review** - 1 info
  • ℹ️ tests/fm-backend-herdr-project-workspace-e2e.test.sh:4 - Three comments in this file (lines 4, 37, 175) still cross-reference a nonexistent 'docs/herdr-backend.md "Project workspace grouping"' heading. Commit 3ae8678 in this same branch was created specifically to fix this class of stale cross-reference (the actual heading is 'Presentation spaces', confirmed by grepping docs/herdr-backend.md's headings), but it only touched bin/backends/herdr.sh and bin/fm-spawn.sh, missing this test file which has the identical stale reference introduced earlier in the branch (fa0ba46). Comment-only, no functional impact.
✅ **Test** - passed

✅ No issues found.

  • bash tests/fm-backend-herdr.test.sh — full fake-CLI unit suite for bin/backends/herdr.sh (195 assertions, 0 failures), including the new/updated project-key-basename-vs-hash tests (test_project_key_basename_excludes_hash, test_project_workspace_record_snapshot_rejects_corruption's hash-bearing-label rejection case, test_project_container_ensure_reuses_and_self_heals_after_removal's clean-label assertions)
  • bash tests/fm-backend-herdr-project-workspace-e2e.test.sh — real herdr/jq/treehouse E2E test driving the actual bin/fm-spawn.sh and bin/fm-teardown.sh in an isolated lab session (all 9 assertions passed: same-project task sharing, cross-project isolation, decoy-label non-adoption, teardown/self-heal, unsafe-basename fallback)
  • Manual CLI verification: spawned two tasks (cm1, cm2) for the same scratch project under config/herdr-presentation-spaces=project against a real isolated herdr session, ran herdr workspace list/herdr tab list to confirm one shared workspace with a hash-free 'proj-frontend' label and one tab per task, inspected the persisted state/proj-<key>.herdr-workspace record to confirm the hash lives only in the project= field, then ran fm-teardown.sh on each task to confirm the workspace survives the first teardown and is removed after the last
✅ **Document** - passed

✅ No issues found.

✅ **Lint** - passed

✅ No issues found.

✅ **Push** - passed

✅ No issues found.

config/herdr-presentation-spaces now accepts "project" as a third value,
mutually exclusive with off/on: every task for one project shares that
project's own durable workspace instead of the per-home or disposable
per-task containers. The container is resolved from a persisted, live-
verified state/proj-<key>.herdr-workspace record rather than a label
search, self-healing with a fresh workspace once the record's workspace
is confirmed gone. It is gated behind the same Herdr 0.8.0 floor as
presentation spaces, with no on-style override.

Tab creation, teardown, and relaunch reuse existing primitives unchanged;
only the container-resolution step in fm-spawn.sh's herdr case arm is new.

Claude-Session: https://claude.ai/code/session_01CPaSAZkCLC1rnwHDCWNhnH
@friesentius friesentius changed the title feat(herdr): add project-grouped workspace topology for spawned tasks feat(herdr): add per-project workspace grouping topology Sep 11, 2026
… label

The basename-collision fix (fm_backend_herdr_project_key_for_path) correctly
disambiguates the persisted record filename and identity by appending a
short hash of the absolute project path, but that same combined key was also
feeding the visible herdr workspace --label and the record's own label=
field, so a project workspace showed e.g. "proj-frontend-a3f9c21c" instead
of a clean "proj-frontend".

Add fm_backend_herdr_project_key_basename to derive the plain, hash-free
basename, and thread it separately from the collision-resistant key through
fm_backend_herdr_project_workspace_ensure/_container_ensure and
fm-spawn.sh's caller: the record file/its project= field keep the full key,
while the herdr --label and the record's own label= field use only the
basename. Update the record snapshot's label verification, the affected
unit and e2e tests, and docs/herdr-backend.md accordingly.

Claude-Session: https://claude.ai/code/session_01CPaSAZkCLC1rnwHDCWNhnH
@friesentius friesentius changed the title feat(herdr): add per-project workspace grouping topology feat(herdr): group project tasks into a shared durable herdr workspace Sep 11, 2026
@friesentius
friesentius merged commit a7e4306 into main Sep 11, 2026
15 of 16 checks passed
friesentius pushed a commit that referenced this pull request Sep 21, 2026
…kunchenguid#3578)

* fix(bin): let verified harness ancestry outrank retained markers (#3)

* fix(bin): let a structural harness ancestor outrank a retained marker

bin/fm-harness.sh treated a verified environment marker as unconditionally
authoritative, so a Codex session started from an environment that had retained
CLAUDECODE=1 detected as claude. Session start then emitted Claude's Stop-owned
supervision protocol to a Codex primary, and every turn end was blocked for
missing Claude recovery.

The defect is the precedence boundary, not any one harness. codex, opencode,
kimi, and muse publish no identity marker at all, so with markers winning
outright any retained CLAUDECODE renamed them; the Cursor-before-Claude ordering
was a point patch on the same class of problem, and the launch-time marker
clearing only ever covered sessions fm-spawn started.

Markers and ancestry are now separate evidence layers that detect_own arbitrates:

- no ancestry match, or no marker: the single available layer answers, unchanged;
- same harness family: the marker's finer verdict stands, so a launch-selected
  pi-signed is not flattened to pi by an ancestry walk that can only see the
  shared launcher name;
- different harness with a structural (command-name) ancestor: ancestry wins,
  because only ancestry proves who owns the process tree;
- different harness with only a bare-interpreter script-path match: the marker
  wins, since a harness-shaped path in some node process's arguments is weaker
  evidence than a harness publishing its own identity.

The correction is symmetric: a retained CURSOR_AGENT no longer renames a claude
worker nested under cursor either.

Adds fm-harness.sh ancestry [<pid>], ancestry evidence with no marker layer, so
a real harness process can be asked what the walk makes of it.

tests/fm-harness-precedence.test.sh is the portable regression, built from real
renamed processes with no harness installed. Every case drives the two layers
apart and asserts each alone as well as the combination, so no case can pass
vacuously; it also pins Codex's real two-process install topology, since the fix
depends on the native binary being what a tool subprocess meets first. The
opt-in drift guard gains the matching live half: each installed harness's real
running process must still be identified by the ancestry walk, and it fails
naming the harness and version when a release changes that name.

Documentation follows the corrected contract in the script header, the
harness-adapters detection section, the codex, opencode, kimi, and cursor
references, and a dated verification record.

* fix(tests): drop the unused argument pass-through in the shim-topology helper

bin/fm-lint.sh refused the branch: run_shim declared a `[ancestry]` argument and
forwarded "$@", but every call site that varies the environment or passes the
ancestry subcommand invokes the shim entry point directly, so the helper is only
ever called with no arguments (ShellCheck SC2120/SC2119).

Behavior is unchanged: with no arguments "$@" expanded to nothing.

* fix(bin): examine the top of the process chain instead of assuming init

harness_ancestry stopped as soon as the next pid was 1, on the assumption that
pid 1 is always init and can never be a harness.
Inside a PID namespace that assumption inverts: the harness itself is pid 1, so
the walk never examined the one process that proves who owns the tree, reported
no ancestry at all, and handed the verdict straight back to a retained marker.

A real Codex session under `codex sandbox`, holding CLAUDECODE=1 and
CLAUDE_CODE_ENTRYPOINT=cli, is exactly that shape: it resolved claude and
rendered Claude's Stop-owned supervision protocol even with the marker-vs-ancestry
precedence boundary in place.
The same probe now resolves codex and renders the Codex foreground checkpoint.

A host's real pid 1 (init, systemd, launchd) matches no harness name, so
examining it costs one ps call and can introduce no false positive; the walk
still stops once that top process has been read, and a non-numeric or zero ppid
still ends it.

tests/fm-harness-precedence.test.sh pins the namespace shape with a fake ps that
reports every process as bash with ppid 1 and pid 1 as the harness.
The case asserts the marker still answers alone when pid 1 is host-shaped, so it
cannot pass vacuously, and it fails against the previous stop condition.

* docs(verification): record the real-Codex retained-marker evidence

The existing record proved the precedence boundary with the portable regression
and recorded each installed harness's process name behind the ancestry walk, but
it had no evidence from a real Codex process actually holding a retained Claude
marker, which is the failure the boundary exists for.

Adds the dated before/after result from codex-cli 0.152.0 under `codex sandbox`,
with the exact command and the decisive verdict and rendered protocol on each
side, and records the second boundary that shape exposed: the walk must examine
the top of the process chain, because inside a PID namespace the harness is pid 1.
Refreshes the portable regression's observed output for the case it gained.

* no-mistakes(review): blind ancestry in marker-pinned harness tests

* no-mistakes(review): blind ancestry in the Pi guard-routing test

* no-mistakes(review): classify precedence suite, dedupe ps stub, soften claims

* no-mistakes(review): model the spawn-and-wait Codex shim topology

* no-mistakes(document): correct stale muse marker-clearing detection claims

* no-mistakes: apply CI fixes

* fix(bin): examine the top of the chain in the lock and nudge walks too

The pid-1 defect corrected in bin/fm-harness.sh survived unchanged in the two
other harness-ancestry walks, on the exact topology the branch verified against
a real Codex process.

bin/fm-session-lock-lib.sh's fm_harness_ancestry_pids stopped as soon as the next
pid was 1, so a firstmate whose harness is pid 1 of its own PID namespace could
not find that harness at all and did not recognize its own session lock.
bin/fm-sessionstart-nudge.sh carried the same stop plus a blanket rejection of a
lock pid of 1, so the same session was told to run session start again on every
turn.

Both walks now compare the top process before stopping, matching the shape used
in bin/fm-harness.sh.
For the lock walk this is safe because fm_harness_process_matches rejects a
host's real pid 1.
For the nudge, `kill -0` still gates the lock pid, and on a host an unprivileged
`kill -0 1` fails, so a lock file that wrongly names pid 1 leaves the hook silent
rather than acting on init.

Each walk gains one regression case. The lock case drives a deterministic process
table whose pid 1 is the harness and asserts a host-shaped pid 1 still finds
nothing, so it cannot pass vacuously. The nudge case needs a real PID namespace,
because the builtin `kill -0` gate cannot be reached through a fake ps, and it
first proves the same fixture nudges with no lock present; it skips explicitly
where unprivileged namespaces are unavailable.

* no-mistakes(review): assert comm-strength detection from subprocess vantage in drift guard

* fix(bin): verify the live harness guard at the strength the guarantee needs

The marker-versus-ancestry boundary this branch ships is a strength claim:
detect_own hands an args-strength verdict straight back to a retained foreign
marker, so a harness is only protected where the ancestry walk reaches it at
comm strength.

The installed-harness drift guard probed the pane process alone. Under an
interpreter shim the pane process IS the shim, whose own script path is args
strength, while the native binary that carries comm strength is its child. The
guard therefore observed args for Codex, passed, and would have kept passing if
a release stopped spawning that native child at all, while real sessions
silently regressed to the original bug.

fm-harness.sh gains `ancestry-subtree`, which asks the walk from the pane
process and every descendant of it, the vantage a tool subprocess actually
occupies. The guard now requires comm strength somewhere in that set and
requires every vantage to name the same harness.

This supersedes the preceding commit's in-guard leaf walk, which reached the
same vantage but left the logic inside the test file, where CI could not pin it
and nothing else could reuse it. A harness-dependent check needs both halves:
`tests/fm-harness-precedence.test.sh` now carries a portable case proving the
subtree probe reaches a strength the top-of-session probe cannot, mutation
checked twice, once against the pre-change script and once by disabling
descendant enumeration. The subtree walk also avoids depending on tty and
process-group semantics that differ between Linux and macOS.

Verified live: codex-cli 0.152.0 reports [args codex;comm codex] and Claude Code
2.1.257 reports [comm claude].

* no-mistakes(review): narrow drift guard to the upward vantage path

* no-mistakes(review): judge only comm-strength vantages in drift guard

* no-mistakes(document): drop duplicated rationale in detection precedence evidence

* no-mistakes(review): fix pid-1 nudge case vacuity and descent no-arg expansion

* no-mistakes(document): drop branch-relative phrasing in detection precedence evidence

* no-mistakes(review): guard remaining empty positional expansions in fm-harness

* no-mistakes(document): scope cursor marker-ordering claim to the marker layer

* no-mistakes(review): Prefer comm-strength leaves in equal-depth descent ties

* no-mistakes(document): Document comm-strength descent tie-break

---------

* no-mistakes(review): Blind ancestry in stale gemini/rovo marker-precedence tests

* no-mistakes(document): Add missing equal-depth-tie test line to precedence evidence transcript

* no-mistakes(review): Fix stale/vacuous agy precedence test, add agy to precedence suite and docs

* no-mistakes(document): Fix stale kimi.md marker doc missed by ancestry-precedence fix

---------

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
friesentius pushed a commit that referenced this pull request Sep 21, 2026
…kunchenguid#3578)

* fix(bin): let verified harness ancestry outrank retained markers (#3)

* fix(bin): let a structural harness ancestor outrank a retained marker

bin/fm-harness.sh treated a verified environment marker as unconditionally
authoritative, so a Codex session started from an environment that had retained
CLAUDECODE=1 detected as claude. Session start then emitted Claude's Stop-owned
supervision protocol to a Codex primary, and every turn end was blocked for
missing Claude recovery.

The defect is the precedence boundary, not any one harness. codex, opencode,
kimi, and muse publish no identity marker at all, so with markers winning
outright any retained CLAUDECODE renamed them; the Cursor-before-Claude ordering
was a point patch on the same class of problem, and the launch-time marker
clearing only ever covered sessions fm-spawn started.

Markers and ancestry are now separate evidence layers that detect_own arbitrates:

- no ancestry match, or no marker: the single available layer answers, unchanged;
- same harness family: the marker's finer verdict stands, so a launch-selected
  pi-signed is not flattened to pi by an ancestry walk that can only see the
  shared launcher name;
- different harness with a structural (command-name) ancestor: ancestry wins,
  because only ancestry proves who owns the process tree;
- different harness with only a bare-interpreter script-path match: the marker
  wins, since a harness-shaped path in some node process's arguments is weaker
  evidence than a harness publishing its own identity.

The correction is symmetric: a retained CURSOR_AGENT no longer renames a claude
worker nested under cursor either.

Adds fm-harness.sh ancestry [<pid>], ancestry evidence with no marker layer, so
a real harness process can be asked what the walk makes of it.

tests/fm-harness-precedence.test.sh is the portable regression, built from real
renamed processes with no harness installed. Every case drives the two layers
apart and asserts each alone as well as the combination, so no case can pass
vacuously; it also pins Codex's real two-process install topology, since the fix
depends on the native binary being what a tool subprocess meets first. The
opt-in drift guard gains the matching live half: each installed harness's real
running process must still be identified by the ancestry walk, and it fails
naming the harness and version when a release changes that name.

Documentation follows the corrected contract in the script header, the
harness-adapters detection section, the codex, opencode, kimi, and cursor
references, and a dated verification record.

* fix(tests): drop the unused argument pass-through in the shim-topology helper

bin/fm-lint.sh refused the branch: run_shim declared a `[ancestry]` argument and
forwarded "$@", but every call site that varies the environment or passes the
ancestry subcommand invokes the shim entry point directly, so the helper is only
ever called with no arguments (ShellCheck SC2120/SC2119).

Behavior is unchanged: with no arguments "$@" expanded to nothing.

* fix(bin): examine the top of the process chain instead of assuming init

harness_ancestry stopped as soon as the next pid was 1, on the assumption that
pid 1 is always init and can never be a harness.
Inside a PID namespace that assumption inverts: the harness itself is pid 1, so
the walk never examined the one process that proves who owns the tree, reported
no ancestry at all, and handed the verdict straight back to a retained marker.

A real Codex session under `codex sandbox`, holding CLAUDECODE=1 and
CLAUDE_CODE_ENTRYPOINT=cli, is exactly that shape: it resolved claude and
rendered Claude's Stop-owned supervision protocol even with the marker-vs-ancestry
precedence boundary in place.
The same probe now resolves codex and renders the Codex foreground checkpoint.

A host's real pid 1 (init, systemd, launchd) matches no harness name, so
examining it costs one ps call and can introduce no false positive; the walk
still stops once that top process has been read, and a non-numeric or zero ppid
still ends it.

tests/fm-harness-precedence.test.sh pins the namespace shape with a fake ps that
reports every process as bash with ppid 1 and pid 1 as the harness.
The case asserts the marker still answers alone when pid 1 is host-shaped, so it
cannot pass vacuously, and it fails against the previous stop condition.

* docs(verification): record the real-Codex retained-marker evidence

The existing record proved the precedence boundary with the portable regression
and recorded each installed harness's process name behind the ancestry walk, but
it had no evidence from a real Codex process actually holding a retained Claude
marker, which is the failure the boundary exists for.

Adds the dated before/after result from codex-cli 0.152.0 under `codex sandbox`,
with the exact command and the decisive verdict and rendered protocol on each
side, and records the second boundary that shape exposed: the walk must examine
the top of the process chain, because inside a PID namespace the harness is pid 1.
Refreshes the portable regression's observed output for the case it gained.

* no-mistakes(review): blind ancestry in marker-pinned harness tests

* no-mistakes(review): blind ancestry in the Pi guard-routing test

* no-mistakes(review): classify precedence suite, dedupe ps stub, soften claims

* no-mistakes(review): model the spawn-and-wait Codex shim topology

* no-mistakes(document): correct stale muse marker-clearing detection claims

* no-mistakes: apply CI fixes

* fix(bin): examine the top of the chain in the lock and nudge walks too

The pid-1 defect corrected in bin/fm-harness.sh survived unchanged in the two
other harness-ancestry walks, on the exact topology the branch verified against
a real Codex process.

bin/fm-session-lock-lib.sh's fm_harness_ancestry_pids stopped as soon as the next
pid was 1, so a firstmate whose harness is pid 1 of its own PID namespace could
not find that harness at all and did not recognize its own session lock.
bin/fm-sessionstart-nudge.sh carried the same stop plus a blanket rejection of a
lock pid of 1, so the same session was told to run session start again on every
turn.

Both walks now compare the top process before stopping, matching the shape used
in bin/fm-harness.sh.
For the lock walk this is safe because fm_harness_process_matches rejects a
host's real pid 1.
For the nudge, `kill -0` still gates the lock pid, and on a host an unprivileged
`kill -0 1` fails, so a lock file that wrongly names pid 1 leaves the hook silent
rather than acting on init.

Each walk gains one regression case. The lock case drives a deterministic process
table whose pid 1 is the harness and asserts a host-shaped pid 1 still finds
nothing, so it cannot pass vacuously. The nudge case needs a real PID namespace,
because the builtin `kill -0` gate cannot be reached through a fake ps, and it
first proves the same fixture nudges with no lock present; it skips explicitly
where unprivileged namespaces are unavailable.

* no-mistakes(review): assert comm-strength detection from subprocess vantage in drift guard

* fix(bin): verify the live harness guard at the strength the guarantee needs

The marker-versus-ancestry boundary this branch ships is a strength claim:
detect_own hands an args-strength verdict straight back to a retained foreign
marker, so a harness is only protected where the ancestry walk reaches it at
comm strength.

The installed-harness drift guard probed the pane process alone. Under an
interpreter shim the pane process IS the shim, whose own script path is args
strength, while the native binary that carries comm strength is its child. The
guard therefore observed args for Codex, passed, and would have kept passing if
a release stopped spawning that native child at all, while real sessions
silently regressed to the original bug.

fm-harness.sh gains `ancestry-subtree`, which asks the walk from the pane
process and every descendant of it, the vantage a tool subprocess actually
occupies. The guard now requires comm strength somewhere in that set and
requires every vantage to name the same harness.

This supersedes the preceding commit's in-guard leaf walk, which reached the
same vantage but left the logic inside the test file, where CI could not pin it
and nothing else could reuse it. A harness-dependent check needs both halves:
`tests/fm-harness-precedence.test.sh` now carries a portable case proving the
subtree probe reaches a strength the top-of-session probe cannot, mutation
checked twice, once against the pre-change script and once by disabling
descendant enumeration. The subtree walk also avoids depending on tty and
process-group semantics that differ between Linux and macOS.

Verified live: codex-cli 0.152.0 reports [args codex;comm codex] and Claude Code
2.1.257 reports [comm claude].

* no-mistakes(review): narrow drift guard to the upward vantage path

* no-mistakes(review): judge only comm-strength vantages in drift guard

* no-mistakes(document): drop duplicated rationale in detection precedence evidence

* no-mistakes(review): fix pid-1 nudge case vacuity and descent no-arg expansion

* no-mistakes(document): drop branch-relative phrasing in detection precedence evidence

* no-mistakes(review): guard remaining empty positional expansions in fm-harness

* no-mistakes(document): scope cursor marker-ordering claim to the marker layer

* no-mistakes(review): Prefer comm-strength leaves in equal-depth descent ties

* no-mistakes(document): Document comm-strength descent tie-break

---------

* no-mistakes(review): Blind ancestry in stale gemini/rovo marker-precedence tests

* no-mistakes(document): Add missing equal-depth-tie test line to precedence evidence transcript

* no-mistakes(review): Fix stale/vacuous agy precedence test, add agy to precedence suite and docs

* no-mistakes(document): Fix stale kimi.md marker doc missed by ancestry-precedence fix

---------

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
friesentius pushed a commit that referenced this pull request Sep 21, 2026
…kunchenguid#3578)

* fix(bin): let verified harness ancestry outrank retained markers (#3)

* fix(bin): let a structural harness ancestor outrank a retained marker

bin/fm-harness.sh treated a verified environment marker as unconditionally
authoritative, so a Codex session started from an environment that had retained
CLAUDECODE=1 detected as claude. Session start then emitted Claude's Stop-owned
supervision protocol to a Codex primary, and every turn end was blocked for
missing Claude recovery.

The defect is the precedence boundary, not any one harness. codex, opencode,
kimi, and muse publish no identity marker at all, so with markers winning
outright any retained CLAUDECODE renamed them; the Cursor-before-Claude ordering
was a point patch on the same class of problem, and the launch-time marker
clearing only ever covered sessions fm-spawn started.

Markers and ancestry are now separate evidence layers that detect_own arbitrates:

- no ancestry match, or no marker: the single available layer answers, unchanged;
- same harness family: the marker's finer verdict stands, so a launch-selected
  pi-signed is not flattened to pi by an ancestry walk that can only see the
  shared launcher name;
- different harness with a structural (command-name) ancestor: ancestry wins,
  because only ancestry proves who owns the process tree;
- different harness with only a bare-interpreter script-path match: the marker
  wins, since a harness-shaped path in some node process's arguments is weaker
  evidence than a harness publishing its own identity.

The correction is symmetric: a retained CURSOR_AGENT no longer renames a claude
worker nested under cursor either.

Adds fm-harness.sh ancestry [<pid>], ancestry evidence with no marker layer, so
a real harness process can be asked what the walk makes of it.

tests/fm-harness-precedence.test.sh is the portable regression, built from real
renamed processes with no harness installed. Every case drives the two layers
apart and asserts each alone as well as the combination, so no case can pass
vacuously; it also pins Codex's real two-process install topology, since the fix
depends on the native binary being what a tool subprocess meets first. The
opt-in drift guard gains the matching live half: each installed harness's real
running process must still be identified by the ancestry walk, and it fails
naming the harness and version when a release changes that name.

Documentation follows the corrected contract in the script header, the
harness-adapters detection section, the codex, opencode, kimi, and cursor
references, and a dated verification record.

* fix(tests): drop the unused argument pass-through in the shim-topology helper

bin/fm-lint.sh refused the branch: run_shim declared a `[ancestry]` argument and
forwarded "$@", but every call site that varies the environment or passes the
ancestry subcommand invokes the shim entry point directly, so the helper is only
ever called with no arguments (ShellCheck SC2120/SC2119).

Behavior is unchanged: with no arguments "$@" expanded to nothing.

* fix(bin): examine the top of the process chain instead of assuming init

harness_ancestry stopped as soon as the next pid was 1, on the assumption that
pid 1 is always init and can never be a harness.
Inside a PID namespace that assumption inverts: the harness itself is pid 1, so
the walk never examined the one process that proves who owns the tree, reported
no ancestry at all, and handed the verdict straight back to a retained marker.

A real Codex session under `codex sandbox`, holding CLAUDECODE=1 and
CLAUDE_CODE_ENTRYPOINT=cli, is exactly that shape: it resolved claude and
rendered Claude's Stop-owned supervision protocol even with the marker-vs-ancestry
precedence boundary in place.
The same probe now resolves codex and renders the Codex foreground checkpoint.

A host's real pid 1 (init, systemd, launchd) matches no harness name, so
examining it costs one ps call and can introduce no false positive; the walk
still stops once that top process has been read, and a non-numeric or zero ppid
still ends it.

tests/fm-harness-precedence.test.sh pins the namespace shape with a fake ps that
reports every process as bash with ppid 1 and pid 1 as the harness.
The case asserts the marker still answers alone when pid 1 is host-shaped, so it
cannot pass vacuously, and it fails against the previous stop condition.

* docs(verification): record the real-Codex retained-marker evidence

The existing record proved the precedence boundary with the portable regression
and recorded each installed harness's process name behind the ancestry walk, but
it had no evidence from a real Codex process actually holding a retained Claude
marker, which is the failure the boundary exists for.

Adds the dated before/after result from codex-cli 0.152.0 under `codex sandbox`,
with the exact command and the decisive verdict and rendered protocol on each
side, and records the second boundary that shape exposed: the walk must examine
the top of the process chain, because inside a PID namespace the harness is pid 1.
Refreshes the portable regression's observed output for the case it gained.

* no-mistakes(review): blind ancestry in marker-pinned harness tests

* no-mistakes(review): blind ancestry in the Pi guard-routing test

* no-mistakes(review): classify precedence suite, dedupe ps stub, soften claims

* no-mistakes(review): model the spawn-and-wait Codex shim topology

* no-mistakes(document): correct stale muse marker-clearing detection claims

* no-mistakes: apply CI fixes

* fix(bin): examine the top of the chain in the lock and nudge walks too

The pid-1 defect corrected in bin/fm-harness.sh survived unchanged in the two
other harness-ancestry walks, on the exact topology the branch verified against
a real Codex process.

bin/fm-session-lock-lib.sh's fm_harness_ancestry_pids stopped as soon as the next
pid was 1, so a firstmate whose harness is pid 1 of its own PID namespace could
not find that harness at all and did not recognize its own session lock.
bin/fm-sessionstart-nudge.sh carried the same stop plus a blanket rejection of a
lock pid of 1, so the same session was told to run session start again on every
turn.

Both walks now compare the top process before stopping, matching the shape used
in bin/fm-harness.sh.
For the lock walk this is safe because fm_harness_process_matches rejects a
host's real pid 1.
For the nudge, `kill -0` still gates the lock pid, and on a host an unprivileged
`kill -0 1` fails, so a lock file that wrongly names pid 1 leaves the hook silent
rather than acting on init.

Each walk gains one regression case. The lock case drives a deterministic process
table whose pid 1 is the harness and asserts a host-shaped pid 1 still finds
nothing, so it cannot pass vacuously. The nudge case needs a real PID namespace,
because the builtin `kill -0` gate cannot be reached through a fake ps, and it
first proves the same fixture nudges with no lock present; it skips explicitly
where unprivileged namespaces are unavailable.

* no-mistakes(review): assert comm-strength detection from subprocess vantage in drift guard

* fix(bin): verify the live harness guard at the strength the guarantee needs

The marker-versus-ancestry boundary this branch ships is a strength claim:
detect_own hands an args-strength verdict straight back to a retained foreign
marker, so a harness is only protected where the ancestry walk reaches it at
comm strength.

The installed-harness drift guard probed the pane process alone. Under an
interpreter shim the pane process IS the shim, whose own script path is args
strength, while the native binary that carries comm strength is its child. The
guard therefore observed args for Codex, passed, and would have kept passing if
a release stopped spawning that native child at all, while real sessions
silently regressed to the original bug.

fm-harness.sh gains `ancestry-subtree`, which asks the walk from the pane
process and every descendant of it, the vantage a tool subprocess actually
occupies. The guard now requires comm strength somewhere in that set and
requires every vantage to name the same harness.

This supersedes the preceding commit's in-guard leaf walk, which reached the
same vantage but left the logic inside the test file, where CI could not pin it
and nothing else could reuse it. A harness-dependent check needs both halves:
`tests/fm-harness-precedence.test.sh` now carries a portable case proving the
subtree probe reaches a strength the top-of-session probe cannot, mutation
checked twice, once against the pre-change script and once by disabling
descendant enumeration. The subtree walk also avoids depending on tty and
process-group semantics that differ between Linux and macOS.

Verified live: codex-cli 0.152.0 reports [args codex;comm codex] and Claude Code
2.1.257 reports [comm claude].

* no-mistakes(review): narrow drift guard to the upward vantage path

* no-mistakes(review): judge only comm-strength vantages in drift guard

* no-mistakes(document): drop duplicated rationale in detection precedence evidence

* no-mistakes(review): fix pid-1 nudge case vacuity and descent no-arg expansion

* no-mistakes(document): drop branch-relative phrasing in detection precedence evidence

* no-mistakes(review): guard remaining empty positional expansions in fm-harness

* no-mistakes(document): scope cursor marker-ordering claim to the marker layer

* no-mistakes(review): Prefer comm-strength leaves in equal-depth descent ties

* no-mistakes(document): Document comm-strength descent tie-break

---------

* no-mistakes(review): Blind ancestry in stale gemini/rovo marker-precedence tests

* no-mistakes(document): Add missing equal-depth-tie test line to precedence evidence transcript

* no-mistakes(review): Fix stale/vacuous agy precedence test, add agy to precedence suite and docs

* no-mistakes(document): Fix stale kimi.md marker doc missed by ancestry-precedence fix

---------

Co-authored-by: NewAiCoder <170579485+NewAiCoder@users.noreply.github.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant