Repository navigation
fix(process): drop OOMPolicy from systemd-run --scope argv (breaks cron worker dispatch on Linux) - #102357
fix(process): drop OOMPolicy from systemd-run --scope argv (breaks cron worker dispatch on Linux)#102357gkd2323c wants to merge 1 commit into
Conversation
OOMPolicy is a service-unit property; systemd-run --user --scope rejects it with 'Unknown assignment: OOMPolicy=kill'. The availability probe therefore always failed in supervised Linux gateways, making restart-safe cron worker dispatch (systemd scope) permanently unavailable and falling back to in-cgroup workers. MemoryMax + MemoryAccounting remain; OOMPolicy adds nothing for a transient scope (no service manager to act on the OOM event).
PR #102357 — drop OOMPolicy from systemd-run --scope argvVerdict: Reasonable one-line-class fix (unsupported property fails the whole What the change does
Non-blocking
Tests: None; manual host verification is the appropriate gate. |
andrexibiza
left a comment
There was a problem hiding this comment.
Reviewed exact head 280fbe97e2ae8e1d9d957ab75abdd06e9e065582 against current main 63279301bcbdc185c1b07b98a9312eb0c862f26d.
The production change is the correct narrow carrier for #102486. It removes OOMPolicy=kill from both the required systemd-run --user --scope probe and the real worker-scope argv while leaving MemoryAccounting=yes, finite MemoryMax, and the fail-closed RuntimeError branches untouched. That separates rejection of an optional property from actual absence of the required process-isolation substrate. This should compose before #102431; that PR asks a different policy question about genuinely unavailable isolation and is not a substitute for fixing this false-negative probe.
BLOCKER — the exact-head acceptance contract still requires the deleted property, so this PR's own deterministic test is red.
tests/tools/test_process_registry.py::TestSystemdCgroupIsolation::test_wraps_in_systemd_scope_when_supervisor_and_available still contains:
assert "OOMPolicy=kill" in propertiesThis head necessarily violates that assertion. The exact-head Python tests / Run tests check is failed, which in turn leaves All required checks pass failed. The passing lint, E2E, Windows, macOS, Docker, Nix, supply-chain, and attribution checks do not supersede that red required check.
Please update the existing scope-argv test to retain the required assertions for MemoryAccounting=yes and finite MemoryMax, and explicitly assert that OOMPolicy=kill is absent. Also make test_systemd_run_user_scope_available_caches_after_probe inspect its captured probe argv and assert the same absence there; otherwise only the real builder site is protected and the probe can regress independently. These are host-independent argv tests and directly encode the capability boundary this fix establishes.
Graph / closure: add Fixes #102486 to the PR body so the affected-host report closes through this implementation carrier. Keep #102431 separate: genuine failure to establish systemd-run --user --scope must continue to fail closed unless that broader safety policy is resolved on its own merits.
Once the deterministic test contract is repaired and every check on the resulting exact head is green, I see no remaining blocker in this narrow change.
A systemd-supervised gateway (INVOCATION_ID set) with no user D-Bus session — containers, minimal LXCs, macOS-style supervisors — fails EVERY scheduled job at dispatch: restart_safe_gateway_child_argv() raises and run_one_job() records a failure. The only symptom is silently skipped executions (a missed nightly backup, dead watchdogs). Degrade to a direct external subprocess with a warning instead of raising, unless HERMES_GATEWAY_CHILD_REQUIRE_SCOPE=1 re-enables fail-closed. Degraded jobs keep process separation and the full NousResearch#101940 ownership handoff; only cgroup isolation is lost. The dispatch is a GatewayChildDispatch (in_process / scoped / degraded) so the degraded case can never collapse into the 'not managed, stay in-process' sentinel — the exact failure mode that would recreate the restart-interruption edge NousResearch#101940 closed. Also drops OOMPolicy=kill from scope argv: it is a service-unit property that transient scopes reject ('Unknown assignment'), so the probe always failed where scopes actually work (NousResearch#102357, by gkd2323c, incorporated here with credit). Co-authored-by: gkd2323c (probe fix from NousResearch#102357)
…MPolicy too Requested in the review of NousResearch#102357: the spawn builder was pinned, the probe was not, and the probe is where the rejection was cached as "unavailable".
…MPolicy too Requested in the review of #102357: the spawn builder was pinned, the probe was not, and the probe is where the rejection was cached as "unavailable".
|
Merged via #104152 (rebase) — your commit is on @andrexibiza — both requests from your review are in the merged head: the spawn test keeps its Verified before merge with an A/B probe of |
Problem
Since #70716's restart-safe worker isolation landed (and 83efdf5 widened the Linux path to supervised gateways via
not _IS_LINUX), every cron job dispatched from a supervised systemd gateway fails with:Root cause
_systemd_run_user_scope_available()probes withsystemd-run --user --scopepassing--property OOMPolicy=kill. ButOOMPolicyis a service-unit property — transient scopes reject it:The probe therefore always returns non-zero in any environment where
systemd-run --user --scopeactually works, the negative result is cached (TTL), andrestart_safe_gateway_child_argv()raises — failing every restart-safe cron worker dispatch.Same invalid property is in
_build_systemd_scope_argv(), so even if the probe were bypassed the real worker spawn would fail identically.Fix
Drop
OOMPolicy=killfrom both the probe and the scope argv builder.MemoryAccounting=yes+MemoryMaxremain; a transient scope has no service manager acting on the OOM event, so the property was a no-op intended for service units only.Verified: probe returns True after the change, cron dispatch completes (job history: failed at 01:19/01:22 → ok at 01:26).