Repository navigation
Conversation
The NixOS module produces the one topology the restart-safe cron worker cannot run in: a system service (`User=hermes`, `isSystemUser = true`) whose uid has no systemd user manager. `run_one_job` hands every fire to `_launch_external_cron_worker` → `restart_safe_gateway_child_argv`, which requires a transient `systemd-run --user --scope` and fails closed when it cannot get one: Restart-safe cron worker dispatch failed: cannot create restart-safe systemd scope for gateway child: systemd-run --user --scope is unavailable (usually no reachable user D-Bus session at /run/user/995/bus). On a system-level service install, run `sudo loginctl enable-linger <gateway-user>` and restart the gateway. So on a stock `services.hermes-agent.enable = true` host, no cron job runs at all — every fire is recorded as a failed execution before an agent starts, and the occurrence is consumed, so a weekly job does not retry for another week. The error names the remedy, but nothing in the module applies it, and the terminal-worker path with the same cgroup-isolation goal degrades gracefully instead (`process_registry.py`: "worker shares the gateway cgroup"), so nothing else on the host looks wrong. Three parts, because linger alone leaves a startup race: - `linger = lib.mkDefault true` on the created user: the declarative `loginctl enable-linger`, which gives the uid a user manager and with it /run/user/<uid>/bus. mkDefault, so an operator can still refuse it. - after/wants linger-users.service, the unit that runs enable-linger. - a bounded preStart wait for the socket. logind starts user@<uid>.service asynchronously, and `run_gateway` resolves XDG_RUNTIME_DIR / DBUS_SESSION_BUS_ADDRESS exactly once at startup (802f0f9), so a bus that appears after ExecStart is one the process never sees for its lifetime. Non-fatal after 10s: a gateway without cron beats no gateway. All three are gated on `lingerEnabled`, read back off `config` rather than assumed: with `createUser = false` the operator owns the user, nothing here knows whether they lingered it, and the unit must not block on a bus that may never arrive. Test: nix/checks.nix `cron-worker-user-scope` relates the unit's `User=` to that user's `linger`, requires the ordering and the wait when it lingers, and requires neither when the module does not own the user. Red on the parent commit for the first three arms. Needs nixpkgs >= 25.05 for `users.manageLingering`.
|
Checked the mechanism against the released tags and against current 1. The fail-closed behaviour is real on 0.21.1 — and still present in 0.21.2. At both 2. On current
So after 0.21.2 the symptom is "cron runs, without cgroup isolation, one warning per gateway process" rather than "no cron at all" — unless an operator sets 3. Two current-
|
|
Reopened — the close was mine and it was a mistake. A downstream PR in our own fleet repo referenced this one, and I read "fix merged" off that. The fix shipped in our tree, not in yours. Confirmed with Thanks for the review — it is more useful than the PR it landed on. I checked each point against the tree rather than taking it on trust, and all three hold:
Agreed on reusing Why the first cut duplicated the predicate in the unit: So, pushing next:
One open question, since your docs already take a position on it. |
|
Closing. This landed on Thanks for picking it up. |
What does this PR do?
Gives the NixOS gateway's service uid a systemd user manager, so the restart-safe cron worker can create the transient scope it requires.
The module produces the one topology that worker cannot run in: a system service (
User=hermes,isSystemUser = true,nix/nixosModules.nix:214,348) whose uid has no user manager and therefore no/run/user/<uid>/bus. Since 0.21.1,run_one_jobhands every gateway-dispatched fire torestart_safe_gateway_child_argv, which raises rather than degrading:So on a stock
services.hermes-agent.enable = truehost, no cron job runs at all. The error names the remedy; nothing in the module applies it. Details and the quiet-failure analysis are in #110628.The runtime half already exists:
_ensure_user_systemd_env()(hermes_cli/gateway.py:4535, from 802f0f9) adoptsXDG_RUNTIME_DIR/DBUS_SESSION_BUS_ADDRESSon this exact topology, and its test describes the fixture as faking "the on-disk stateloginctl enable-lingerleaves behind". This PR makes the module satisfy that precondition.Why not relax the fail-closed instead: the docstring states that falling back "would recreate the restart interruption this handoff exists to prevent", so we read it as deliberate and left it alone. #110628 raises the separate question for topologies where a user scope is structurally impossible.
Related Issue
Fixes #110628
Type of Change
Changes Made
nix/nixosModules.nix— three parts, because linger alone leaves a startup race:users.users.${cfg.user}.linger = lib.mkDefault truein thecreateUserblock. The declarativeloginctl enable-linger;mkDefaultso an operator can still refuse it.after/wantsonlinger-users.service, the unit that applies a declaredlinger.preStartwait for/run/user/<uid>/bus. Ordering is not sufficient on its own:loginctl enable-lingerreturns before logind has finished startinguser@<uid>.service, andrun_gatewayresolves the bus environment exactly once at startup — a bus that appears afterExecStartis one that process never sees for its whole lifetime. Bounded at 10s and non-fatal: a gateway without cron beats no gateway.lingerEnabled, read back offconfigrather than assumed. WithcreateUser = falsethe operator owns the user and nothing here knows whether they lingered it, so all three parts stay off and the unit does not block on a bus that may never arrive.nix/checks.nix— newcron-worker-user-scopecheck.Container mode (
container.enable = true) is untouched: the gateway runs inside the container, without the host unit'sINVOCATION_ID, so the guard does not fire there.Requires nixpkgs >= 25.05 for
users.manageLingering(noted at the call site).How to Test
The precondition, on a NixOS host before the change — reproduces the gateway's cgroup and uid without waiting for a fire:
sudo systemd-run -p User=hermes -p Group=hermes --pipe --wait --collect \ /run/current-system/sw/bin/systemd-run --user --scope --quiet --unit probe /bin/true # Failed to connect to bus: No medium foundAfter
nixos-rebuild switchplus onesystemctl restart hermes-agent(a switch that only flips linger does not restart the unit, and the running process keeps its busless environment), the same probe exits 0 and gateway-dispatched fires run.The check is eval-only, so it also runs from a non-Linux host:
On this commit it evaluates. With
nix/nixosModules.nixreverted to the parent it throws, on the three arms the fix addresses:The fourth arm (a gateway whose uid is not known to linger must not block on a bus) passes on both, which is the point — it guards the gate, not the fix.
Checklist
Code
pytest tests/ -q— N/A, no Python changed; verified with thenix evalreceipts abovenix/checks.nix, red on the parent commit)Documentation & Housekeeping
cli-config.yaml.example— N/A, no config keysCONTRIBUTING.md/AGENTS.md— N/A