Conversation
A daemon-hosted background session sits several harness-named hops below the session that launched it, with no non-harness process in between, so the whole chain reads as one contiguous harness run. Resolving this session's identity as the outermost pid of that run therefore records the LAUNCHING session's pid. When the daemon above the session later exits, that recorded pid stops being an ancestor, and because it still names a live harness the Stop-owned auto-arm reads it as a competing session and stays inert for the rest of the session. These cases build that real three-level process shape and drive the real bin/fm-lock.sh, the real Stop auto-arm, and the real turn-end guard through it, alongside the negative cases that must keep behaving as they do today: a live launching session that genuinely owns the home, a demonstrably dead owner, away mode, and an inherited session declaration.
Under a session-hosting daemon, two different arrangements produce an identical process chain: an async hook or tool call run for session X inside a daemon-hosted worker, and a background session launched BY X through that same daemon. No process-table fact separates them, so resolving identity as the outermost pid of the contiguous harness run records X in both - which is right for the worker and wrong for the background session. Prefer the harness's own declaration of which process is this session, accepted only when it names a live harness process inside this contiguous run. Membership is what makes an untrustworthy value harmless: a pid inherited from an unrelated session is not in this ancestry and is ignored, and a dead one cannot be recorded as a live owner. Everything else falls back to the outermost pid unchanged, which is what every harness that declares nothing keeps using. The self-ownership predicate is deliberately untouched. It already accepts the recorded pid anywhere in the ancestry; the defect was that the pid it was given only belonged to the session transiently.
The writer now depends on a vendor-controlled surface: if Claude Code stopped setting CLAUDE_PID, or started letting an inherited value through, identity would silently fall back to the launching session's pid and a background session's supervision would go inert mid-session again. A stub agent cannot see that. The guard launches a real session with a deliberately wrong CLAUDE_PID planted in its environment and asserts the value reported inside that session's own hook is its own live process instead, so it proves the declaration is authoritative rather than merely present. It self-skips without the opt-in, refuses a pass that checked nothing, and fails naming the harness and version. Also records the dated result and re-selects the families a change to the identity library must re-run.
A print-mode session runs its hook directly beneath itself, so it cannot show what the writer would have recorded without the declaration. A real background agent can: its hook sits two harness-named hops below the shared session-hosting daemon, and the outermost pid of that chain is the daemon itself - a process every background session in the home shares and which outlives any one of them. The guard now exercises both shapes, refuses to pass if the background chain turned out to be a single hop (which would prove nothing about the case the fix exists for), reports the pid the old resolution would have recorded, and stops the agent it started.
The live session-declaration guard printed the resolved `claude` executable path alongside its version, so the measurement recorded in docs/verification/supervision.md carried an absolute home directory. The version is the identifying fact for that evidence, so the note now prints only the version and the recorded block matches what a real run emits. The guard still resolves and requires an executable `claude`, so an absent harness continues to fail loudly instead of verifying nothing.
Owner
Author
|
Superseded by kunchenguid#3122, which targets the upstream repo. Same branch, same commits. Closing this one so there is a single review surface. |
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
The failure
A background agent could write a session lock naming a process that was not its own, then permanently fail to recognise that lock as its own once that process exited. From then on it behaved as though another session held the home.
The cause is that the lock writer identified "this session" by walking to the outermost process in the contiguous harness ancestry. Two arrangements produce ancestries the process table cannot tell apart:
The walk answers "outermost" in both cases. In the first that is right. In the second it records the shared daemon, which outlives no single session and eventually exits, stranding the agent below it.
The change
bin/fm-session-lock-lib.shnow prefers the harness's own declaration of which process is its session (CLAUDE_PID) over the ancestry walk, because the harness knows what the process table cannot infer.That declaration is trusted only when both hold:
An inherited value from an unrelated session, a stale pid, a live process outside this ancestry, and a malformed value are all inert: each falls through to the previous behaviour.
tests/fm-session-lock-ancestry.test.shcovers each of those rejection paths with the accepted and rejected answers driven to deliberately different pids, so an assertion cannot pass for both branches.The ancestry fallback is unchanged for every harness that declares nothing. For those, the walk stops at the first match, so outermost and innermost are the same process and the overreach never existed.
Measurement
Claude Code 2.1.241, a print-mode session and a real background agent, each launched with a deliberately wrong
CLAUDE_PID:Recorded in
docs/verification/supervision.md, refreshed by the opt-in guardtests/fm-session-lock-declaration-live-e2e.test.sh. That guard needs a real harness, so it is env-gated and self-skipping rather than part of the default suite.What to look at
bin/fm-session-lock-lib.sh— the two guards that make the declaration safe to trust.tests/fm-session-lock-ancestry.test.sh— the rejection-path coverage, and the three-level launcher/daemon/session shape that reproduces the original failure.bin/fm-test-run.sh— the changed-path to family map keeps its over-selecting contract for the lock library.