Skip to content

fix(lifecycle_guard): allow bootstrap of an EXISTING non-gateway plist; surface guard blocks in execute_code - #767

Merged
Kyzcreig merged 2 commits into
mainfrom
fix/lifecycle-guard-bootstrap-existing-plist
Sep 20, 2026
Merged

Kyzcreig merged 2 commits into
mainfrom
fix/lifecycle-guard-bootstrap-existing-plist

Conversation

@Kyzcreig

Copy link
Copy Markdown
Collaborator

The defects

Two, both reproduced on live code before any edit.

1. launchctl bootstrap <existing plist> was refused label-independently

contains_launchctl_submit_command handles submit and bootstrap label-independently because a NEW job's label is attacker-chosen (NousResearch#62891). That reasoning is correct for submit (pure text) and for bootstrap of a path that does not exist yet — but not for a plist already on disk: launchd reads the Label key out of that file, so the label is a readable fact, not a claim.

Measured inside the supervised gateway, reloading the fleetreview router after a plutil -replace edit:

sudo launchctl bootout system/ai.hermes.fleetreview-router
sudo launchctl bootstrap system /Library/LaunchDaemons/ai.hermes.fleetreview-router.plist

Refused, though ai.hermes.fleetreview-router is not a gateway label. The workaround was the deprecated, label-gated launchctl load -w.

2. execute_code turned a guard block into a silent success

The guard refuses a terminal() call by returning a dict. Inside the sandbox that is a value, so the common script shape r = terminal(cmd) ran to completion and execute_code reported status=success, exit_code=0, output='' — a silent block, strictly worse than the direct terminal tool, which at least prints why.

The change

_bootstrap_targets_readable_non_gateway_plist allows bootstrap when every plist argument is an existing readable regular file whose Label is not a gateway label and whose Program/ProgramArguments do not reference a gateway entrypoint. Everything unreadable stays blocked: missing path, directory, FIFO, oversized (256 KiB cap), unparseable, missing Label, any unexpanded shell value in any argument, and submit in every shape.

Both refusal returns now carry blocked_by, and the generated hermes_tools stub raises ToolCallBlocked on that marker.

Before/after on the same input, part 2:

status exit_code output
before success 0 ''
after error 1 refusal in the traceback

Verification

Base fork/main 9de4895, scripts/run_tests.sh, sandboxed HOME.

  • tests/cron/test_lifecycle_guard_bootstrap_existing_plist.py — 31 passed
  • tests/tools/test_execute_code_surfaces_blocks.py — 11 passed
  • full sweep (tests/cron/, approval, code-execution, terminal, gateway-restart) — 107 files, 2083 passed, 4 skipped

Mutation proof, part 1 — 9 mutants, one per guard

mutant verdict
M1 drop gateway-label check KILLED (2)
M2 drop argv-entrypoint check KILLED (1)
M3 drop split-argv check KILLED (1)
M4 drop missing-Label fail-closed KILLED (1)
M5 drop unexpanded-shell check KILLED (1)
M8 unreadable treated as safe KILLED (5)
M9 any- instead of every-plist KILLED (6)
M6 drop size bound SURVIVED — equivalent under overlap
M7 drop S_ISREG check SURVIVED — equivalent under overlap

M6/M7 are reported honestly, not waved through. The reader carries three overlapping bounds (st_size pre-check, bounded read loop, post-read length check); dropping any one still yields None, and only dropping all three goes red (verified: M6g KILLED). S_ISREG is likewise shadowed because os.read on a directory raises IsADirectoryError and a FIFO read returns empty. Both properties are therefore pinned directly at the reader — test_oversized_file_reader_returns_none, test_fifo_argument_blocked, test_directory_named_like_a_plist_blocked, plus a positive control — so the property is gated even though no single mutation of its three implementations can be killed.

M3 and M5 initially SURVIVED — real test gaps, not equivalents. Added test_gateway_entrypoint_split_across_argv_blocked (a bare /usr/local/bin/hermes gateway run launcher matches no marker token) and test_variable_in_domain_argument_blocked (launchctl bootstrap gui/$UID <plist>, the common spelling, where the variable is in the domain not the path). Both mutants then went red.

Mutation proof, part 2 — 6 mutants, both directions, 6/6 KILLED

stub not routed · helper never raises · helper raises on everything · marker key renamed · submit-refusal unstamped · lifecycle-refusal unstamped.

The raise-on-everything mutant is the load-bearing control: it reds test_ordinary_nonzero_exit_is_not_a_block, proving an exit 7 or a grep miss still returns normally rather than exploding.

Against the real on-disk plists

Self-identity ai.hermes.gateway:

allowed  bootstrap /Library/LaunchDaemons/ai.hermes.fleetreview-router.plist
allowed  bootout   system/ai.hermes.fleetreview-router
BLOCKED  bootstrap ~/Library/LaunchAgents/ai.hermes.gateway.plist
allowed  bootstrap ~/Library/LaunchAgents/ai.hermes.gateway-daedalus.plist
BLOCKED  launchctl submit

Residual risk

Bootstrap now reads the plist at guard time, so a TOCTOU swap between the guard's read and launchd's read is possible. It requires local write access to the plist path, which already implies the ability to edit the gateway's own plist — so it does not widen the attack surface. Every other input remains fail-closed.

Pre-existing flakes (not caused by this change)

Under 64-way parallelism, test_media_delivery_parity, test_script_claim_heartbeat and test_scheduler_provider flaked — different tests each run. All pass in isolation on this branch; git diff --name-only fork/main is empty for each file; and test_media_delivery_parity flakes identically on a pristine fork/main worktree.

What to check

  • The _plist_declares_gateway_job heuristic — is "gateway" in tokens and any("hermes" in token) the right breadth, or should it be tighter/looser?
  • Whether GATEWAY_LIFECYCLE_BLOCK_MARKER belongs in cron/lifecycle_guard.py or somewhere more neutral, given tools/ now imports it.
  • The TOCTOU judgement above.

…t; surface guard blocks in execute_code

Two defects, both reproduced on live code before any edit.

1. `launchctl bootstrap <existing plist>` was refused label-independently.

`contains_launchctl_submit_command` treats submit and bootstrap
label-independently because a NEW job's label is attacker-chosen (NousResearch#62891).
That holds for `submit` (pure text) and for bootstrap of a path that does
not exist yet, but not for a plist already on disk: launchd reads `Label`
out of that file, so the label is a readable fact, not a claim.

Measured inside the supervised gateway on the Studio: reloading the
fleetreview router after a `plutil -replace` edit

    sudo launchctl bootout system/ai.hermes.fleetreview-router
    sudo launchctl bootstrap system /Library/LaunchDaemons/ai.hermes.fleetreview-router.plist

was refused, though `ai.hermes.fleetreview-router` is not a gateway label.
The workaround was the deprecated, label-gated `launchctl load -w`.

`_bootstrap_targets_readable_non_gateway_plist` now allows bootstrap when
every plist argument is an existing readable regular file whose Label is
not a gateway label and whose Program/ProgramArguments do not reference a
gateway entrypoint. Everything unreadable stays blocked: missing path,
directory, FIFO, oversized (256 KiB cap), unparseable, missing Label, any
unexpanded shell value in any argument, and `submit` in every shape.

2. execute_code turned a guard block into a silent success.

The guard refuses a terminal() call by RETURNING a dict. Inside the
sandbox that is a value, so the common script shape `r = terminal(cmd)`
ran to completion and execute_code reported status=success, exit_code=0,
output='' — a silent block, worse than the direct terminal tool, which at
least prints why.

Both refusal returns now carry `blocked_by`, and the generated
hermes_tools stub raises ToolCallBlocked on that marker. Measured on the
same input, before: status=success exit_code=0 output=''. After:
status=error exit_code=1 with the refusal in the traceback.

Verification (fork/main 9de4895, scripts/run_tests.sh, sandboxed HOME):

* tests/cron/test_lifecycle_guard_bootstrap_existing_plist.py 31 passed
* tests/tools/test_execute_code_surfaces_blocks.py 11 passed
* full sweep 107 files, 2083 passed, 4 skipped

Mutation proof, part 1 (9 mutants, per-guard):
  M1 drop gateway-label check        KILLED (2)
  M2 drop argv-entrypoint check      KILLED (1)
  M3 drop split-argv check           KILLED (1)
  M4 drop missing-Label fail-closed  KILLED (1)
  M5 drop unexpanded-shell check     KILLED (1)
  M8 unreadable treated as safe      KILLED (5)
  M9 any- instead of every-plist     KILLED (6)
  M6 drop size bound                 SURVIVED - equivalent under overlap
  M7 drop S_ISREG check              SURVIVED - equivalent under overlap

M6/M7 are honestly reported, not waved through. The reader carries three
overlapping bounds (st_size pre-check, bounded read loop, post-read length
check); dropping any one still yields None, and only dropping all three
goes red (verified: M6g KILLED). S_ISREG is likewise shadowed because
os.read on a directory raises IsADirectoryError and a FIFO read returns
empty. Both properties are pinned directly at the reader
(test_oversized_file_reader_returns_none, test_fifo_argument_blocked,
test_directory_named_like_a_plist_blocked) plus a positive control, so the
property is gated even though no single mutation of its three
implementations can be killed.

M3 and M5 initially SURVIVED - real test gaps, not equivalents. Added
test_gateway_entrypoint_split_across_argv_blocked and
test_variable_in_domain_argument_blocked (`launchctl bootstrap gui/$UID
<plist>`, the common spelling); both mutants then went red.

Mutation proof, part 2 (6 mutants, both directions): stub not routed,
helper never raises, helper raises on EVERYTHING, marker key renamed,
submit-refusal unstamped, lifecycle-refusal unstamped - 6/6 KILLED. The
raise-on-everything mutant is the load-bearing control: it reds
test_ordinary_nonzero_exit_is_not_a_block, so an `exit 7` or a grep miss
still returns normally.

Against the real on-disk plists, self-identity ai.hermes.gateway:
  allowed  bootstrap /Library/LaunchDaemons/ai.hermes.fleetreview-router.plist
  allowed  bootout system/ai.hermes.fleetreview-router
  BLOCKED  bootstrap ~/Library/LaunchAgents/ai.hermes.gateway.plist
  allowed  bootstrap ~/Library/LaunchAgents/ai.hermes.gateway-daedalus.plist
  BLOCKED  launchctl submit

Residual risk: bootstrap now reads the plist at guard time, so a
TOCTOU swap between the guard's read and launchd's read is possible. This
requires local write access to the plist path, which already implies the
ability to edit the gateway's own plist, so it does not widen the
attack surface. Every other input remains fail-closed.

Two pre-existing flakes surfaced under 64-way parallelism
(test_media_delivery_parity, test_script_claim_heartbeat,
test_scheduler_provider). Different tests each run; all pass in isolation
on this branch; `git diff --name-only fork/main` is empty for each file,
and test_media_delivery_parity flakes identically on a pristine fork/main
worktree. Not caused by this change.
…and is blocked

Review round 1 finding. The new allow-path asked "is this plist a gateway
job?" (Label / entrypoint argv) but never "does loading this plist EXECUTE
a gateway-lifecycle command?" — re-opening the NousResearch#62891 laundering shape with
a file instead of a submit line, and needing no root:

  Label=ai.hermes.helper
  ProgramArguments=["/bin/sh","-c","launchctl kickstart -k system/<self>"]
  launchctl bootstrap gui/501 $TMPDIR/ai.hermes.helper.plist   -> ALLOWED

_plist_declares_gateway_job now also runs the flattened Program/
ProgramArguments through the existing lifecycle scanner, per token (sh -c
payloads, referenced scripts) and joined (a command split across argv
words). The scan is re-entrant (a plist can bootstrap a plist), so it is
depth-bounded per thread and fails CLOSED at the bound.

Verified:
- incident shape (ai.hermes.fleetreview-router, argv /usr/bin/python3
  /opt/router.py) stays ALLOWED.
- 4 new witnesses: inline sh -c, referenced script, direct argv, mutual
  plist cycle -> all BLOCKED.
- mutation matrix 6/6 killed, each reddening its own named witness:
  drop-check 4 red; drop-per-token 1; drop-joined 1; bound-disabled 1
  (130 plist reads vs 3); depth-not-restored 5; bound-fails-open 1.
- 35 tests in the file, 220 lifecycle/guard tests, 42 safe_command tests
  all pass.
@Kyzcreig
Kyzcreig added this pull request to the merge queue Sep 20, 2026
@github-merge-queue
github-merge-queue Bot removed this pull request from the merge queue due to failed status checks Sep 20, 2026
@Kyzcreig

Copy link
Copy Markdown
Collaborator Author

Ejected from the merge queue 12:12 PT: merge_group run 35529883245 red on slice 3/16 only (runner studio-ang-ventures-hermes-agent-1). Two asyncio TimeoutErrors in tests/gateway/test_session_hygiene.py — outside this PR's diff, green on the PR's own run — while the Studio was under heavy local load (Apollo was running ~150 s real-clock suites there at the same minute). Wall-clock flake on a loaded self-hosted runner; re-enqueueing. Same run is also the first live proof of #766's split: slices 1–10 self-hosted, 11–16 ubuntu-latest, all green. — Apollo

@Kyzcreig
Kyzcreig added this pull request to the merge queue Sep 20, 2026
Merged via the queue into main with commit 02be1bf Sep 20, 2026
95 of 97 checks passed
@Kyzcreig
Kyzcreig deleted the fix/lifecycle-guard-bootstrap-existing-plist branch September 20, 2026 20:16
@Kyzcreig
Kyzcreig restored the fix/lifecycle-guard-bootstrap-existing-plist branch September 21, 2026 10:32
Kyzcreig added a commit that referenced this pull request Sep 25, 2026
…009b, batch 3)

#278 and lifecycle-guard bundles (#319+#594, ssh trio, #920+#933+#1017, #740, #767) re-ported onto upstream/main.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant