Repository navigation
fix(process): negotiate the systemd properties a transient scope accepts - #102508
JoaoMarcos44 wants to merge 1 commit into
Conversation
`_systemd_run_user_scope_available()` probes the scope backend with the exact property set the worker spawn will use, and reads any non-zero exit as "systemd-run --user --scope is unavailable". One of those properties, `OOMPolicy=`, only became valid on *scope* units in systemd v253. Every manager older than that -- Ubuntu 22.04's 249, Debian 12 and RHEL 9's 252 -- answers the probe with `Unknown assignment: OOMPolicy=kill` and exits 1 before creating anything. On those hosts the probe therefore reports a working scope backend as missing, caches that verdict, and `restart_safe_gateway_child_argv()` fails closed on every dispatch. A hardening knob nobody asked about becomes a total cron and Kanban worker outage, and the operator sees only the verdict. Probe the desired set, and when systemd names a property it cannot parse, drop that property and probe again. The surviving set is cached and `_build_systemd_scope_argv()` spawns with exactly it, so the spawn can no longer ask for something the probe never proved. Managers that do accept `OOMPolicy=kill` keep it: an unsupported knob costs that knob, not the isolation every dispatch depends on. Failures that are not a property rejection -- a missing user D-Bus session above all -- are left alone: they are real unavailability, and retrying with a smaller set would only spin. That case still fails closed, and the refusal now quotes the systemd error that produced it instead of only the verdict. NousResearch#102357 identified `OOMPolicy=` as the rejected property from a live host; this change keeps that property wherever the manager supports it instead of removing it fleet-wide. Closes NousResearch#102486 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
PR #102508 — Negotiate the systemd properties a transient scope acceptsVerdict: Looks good. A rejected hardening property (e.g. What the change does
Non-blocking
Tests: |
|
Thanks @JoaoMarcos44 — this is the most thorough of the cluster: the Closing in favour of #102357 (@gkd2323c, earliest for #102486). The property set today is |
Problem
On every systemd older than v253, each restart-safe worker dispatch fails:
_systemd_run_user_scope_available()probes the backend with the same property setthe worker spawn will use, and reads any non-zero exit as "no scope backend". One of
those properties,
OOMPolicy=, became valid on scope units only in systemdv253 ("Scope units now support
OOMPolicy=", systemd v253 NEWS). Ubuntu 22.04ships 249; Debian 12 and RHEL 9 ship 252. All of them answer the probe with
Unknown assignment: OOMPolicy=killand exit 1 before creating anything.The verdict is then cached, and
restart_safe_gateway_child_argv()fails closed onevery dispatch that goes through it —
cron/scheduler.py::_launch_external_cron_workerand
hermes_cli/kanban_db.py::_restart_safe_worker_argv. A hardening knob nobodyasked about becomes a total worker outage, and the operator sees only the verdict,
never the
Unknown assignmentline that caused it.Root cause
The probe asserts a hardcoded property set, so one property this manager cannot
parse is indistinguishable from a missing user bus. The same list is then written a
second time in
_build_systemd_scope_argv(), so the spawn can ask for something theprobe never proved.
Fix
Probe the desired set; when systemd names a property it cannot parse, drop that
property and probe again. The surviving set is cached and the scope builder spawns
with exactly it.
MemoryAccounting+MemoryMax;OOMPolicydropped, once, with a warningOOMPolicy=killkeptAn unsupported knob costs that knob, not the isolation every dispatch depends on.
Failures that are not a property rejection are deliberately left alone: a missing
user bus is real unavailability, and retrying with a smaller set would only spin. That
path still fails closed, preserving the #101940 restart-durability contract untouched.
%%{init: {'theme': 'dark', 'themeVariables': { 'primaryColor': '#8b0000', 'mainBkg': '#0a0204', 'primaryTextColor': '#ffccd5', 'primaryBorderColor': '#ff0038', 'lineColor': '#ff0038'}}}%% graph TD A[Cron / Kanban Dispatch] --> B[Scope Availability Probe] B --> C{systemd-run exit code} C -->|0| D[Cache Accepted Property Set] C -->|Unknown assignment: PROP| E[Drop Rejected Property] C -->|Other failure e.g. no user bus| F[Fail Closed + Quote systemd Error] E --> B D --> G[Build Scope With The Proven Set] G --> H[Restart-Safe Worker Running]Relationship to open work
Supersedes #102357. That PR found the same rejected property on a live host — the
diagnosis is theirs and it is correct. It removes
OOMPolicy=killstatically from bothargv sites, which also removes the OOM-kill semantics on every host where the property
is valid (systemd >= 253), leaves the next added property free to repeat the same
outage, and is currently red on
Python tests / Run testsbecausetests/tools/test_process_registry.py:1960still pinsOOMPolicy=kill. Its own reviewthread asks for exactly this shape: "If per-worker OOM kill matters, consider re-adding
it conditionally behind a support probe later." This PR is that probe. If maintainers
prefer the one-line removal, #102357 should land instead of this.
Not this PR's layer:
cannot be created). Untouched here — this PR only stops the probe from producing a
false "unavailable", which is the precondition that PR's reviewer flagged as needing
to land first.
5f93083221added the memory bound.Test plan
tests/tools/test_process_registry.py(4 new, in the existing POSIX-only class):test_probe_negotiates_away_a_rejected_scope_property—Unknown assignment: OOMPolicy=killon the first probe, success on the retry; asserts the cached set is("MemoryAccounting", "MemoryMax"), that the retry carries a fresh unit name, andthat the memory bound survived.
test_probe_does_not_negotiate_a_missing_user_bus—Failed to connect to busprobesexactly once and stays unavailable.
test_scope_argv_uses_the_negotiated_property_set— the builder emits the negotiatedset, not the wish list.
test_fail_closed_error_reports_why_the_probe_failed— the refusal names the systemderror.
Existing coverage is unchanged:
test_wraps_in_systemd_scope_when_supervisor_and_availablestill passes because the builder asks for the full set when no probe has run.
Local run (Windows dev box, so the POSIX-only class is skipped there — the four new
tests were executed against the same module through an identical standalone mirror,
6 passed):
pytest tests/tools/test_process_registry.py tests/cron/test_restart_safe_worker.py→ 88 passed, 6 failed, and those same 6 fail identically on unmodified
main(Windows-only baseline noise:
TestTerminateHostPidPosix, PTY, stdin-EOF).ruff checkclean on both files.Closes #102486