Conversation
teknium1
left a comment
There was a problem hiding this comment.
Thanks for preserving the direct-launch fallback and attempting real user-manager verification. The underlying premise remains live: current main directly launches workers at hermes_cli/kanban_db.py:8968-8989.
Problems
hermes_cli/kanban_db.py:9243startssystemd-run --user --scopewithPopen, but:9273verifies and later returns that launcher's PID. The installedsystemd-run(1)documentation describes scope mode as synchronous and the scoped command as manager-owned; thereforeproc.pidis not a demonstrated scoped worker PID. This conflicts with current reclaim/timeout lifecycle code, which signals the persisted PID (hermes_cli/kanban_db.py:6902-6928,:7094-7106).
Suggested changes
- Track a manager-reported scoped worker PID, or make scope-unit operations authoritative for liveness and termination; do not persist the
systemd-runwrapper asworker_pid. - Extend the native lifecycle test to assert the persisted PID is in the scope and that reclaim/timeout actually stops the scoped worker.
Automated hermes-sweeper review.
|
Ready for upstream re-review at exact head 56ee10a. The actionable scoped-worker PID/lifecycle finding is addressed with focused tests and live scope proof; details are in the original review thread. |
56ee10a to
b3b692d
Compare
|
Ready for fresh upstream re-review at exact head |
|
Fresh exact-head repair validation for @teknium1: the blocking review was submitted against |
b3b692d to
4eecc64
Compare
|
Current-main recut pushed at exact head |
|
Fresh exact-head repair validation for @teknium1: the blocking scoped-worker PID finding is addressed in the current PR candidate at exact head The implementation reads the transient scope's manager-reported Fresh proof on exact head:
The fork branch remains pushed at exact head |
Summary
systemd --userscopes when the dispatcher is itself running under a same-UID user serviceworker_pidand direct-launch event contract while adding truthful scope receipt fieldsThis is a current-main successor to the verified per-worker systemd-scope portion of #63073. It intentionally does not include the fleet-cap portion, which overlaps later scheduler work in #69942. Thanks to the original #63073 contributor for identifying and implementing the systemd-scope direction.
Why
A worker launched directly by a long-lived dispatcher remains in the dispatcher's service cgroup. If the worker or one of its descendants survives the Python task lifecycle, process ownership and cleanup become ambiguous. A verified transient user scope gives each worker run a manager-owned cgroup boundary without changing behavior on unsupported or unverified hosts.
The scope path is deliberately gated: Linux only, same-UID user manager, dispatcher hosted by its own
.service, and exact post-launch PID/cgroup verification. Any inability to prove those conditions preserves the direct launch behavior.Verification
mainatd5e135a51353c2dbc489d5c2583158b22d8efd7b; exact headb3b692dab6f2f41790ef9783d655f73c0f27ec03scripts/run_tests.sh tests/hermes_cli/test_kanban_worker_systemd_scope.py tests/hermes_cli/test_kanban_worker_spawn_toolsets.py tests/hermes_cli/test_kanban_worker_terminal_cwd.py tests/hermes_cli/test_kanban_worker_session_source.py tests/hermes_cli/test_kanban_boards.py— 42 passedruff check hermes_cli/kanban_db.py tests/hermes_cli/test_kanban_worker_systemd_scope.py— cleanpython -m py_compile hermes_cli/kanban_db.py— cleangit diff --check— cleanBoundaries
No deploy, service restart, live-runtime mutation, release-ref movement, merge, or auto-merge is included or implied.