Skip to content

fix(nix): linger the service uid so cron can create its worker scope - #110629

Closed
haizaar wants to merge 1 commit into
NousResearch:mainfrom
zarmory:fix/nixos-linger-cron-scope
Closed

haizaar wants to merge 1 commit into
NousResearch:mainfrom
zarmory:fix/nixos-linger-cron-scope

Conversation

@haizaar

@haizaar haizaar commented Sep 14, 2026

Copy link
Copy Markdown

What does this PR do?

Gives the NixOS gateway's service uid a systemd user manager, so the restart-safe cron worker can create the transient scope it requires.

The module produces the one topology that worker cannot run in: a system service (User=hermes, isSystemUser = true, nix/nixosModules.nix:214,348) whose uid has no user manager and therefore no /run/user/<uid>/bus. Since 0.21.1, run_one_job hands every gateway-dispatched fire to restart_safe_gateway_child_argv, which raises rather than degrading:

Restart-safe cron worker dispatch failed: cannot create restart-safe systemd scope for
gateway child: systemd-run --user --scope is unavailable (usually no reachable user D-Bus
session at /run/user/995/bus). On a system-level service install, run
`sudo loginctl enable-linger <gateway-user>` and restart the gateway.

So on a stock services.hermes-agent.enable = true host, no cron job runs at all. The error names the remedy; nothing in the module applies it. Details and the quiet-failure analysis are in #110628.

The runtime half already exists: _ensure_user_systemd_env() (hermes_cli/gateway.py:4535, from 802f0f9) adopts XDG_RUNTIME_DIR / DBUS_SESSION_BUS_ADDRESS on this exact topology, and its test describes the fixture as faking "the on-disk state loginctl enable-linger leaves behind". This PR makes the module satisfy that precondition.

Why not relax the fail-closed instead: the docstring states that falling back "would recreate the restart interruption this handoff exists to prevent", so we read it as deliberate and left it alone. #110628 raises the separate question for topologies where a user scope is structurally impossible.

Related Issue

Fixes #110628

Type of Change

  • 🐛 Bug fix (non-breaking change that fixes an issue)

Changes Made

nix/nixosModules.nix — three parts, because linger alone leaves a startup race:

  • users.users.${cfg.user}.linger = lib.mkDefault true in the createUser block. The declarative loginctl enable-linger; mkDefault so an operator can still refuse it.
  • after/wants on linger-users.service, the unit that applies a declared linger.
  • A bounded preStart wait for /run/user/<uid>/bus. Ordering is not sufficient on its own: loginctl enable-linger returns before logind has finished starting user@<uid>.service, and run_gateway resolves the bus environment exactly once at startup — a bus that appears after ExecStart is one that process never sees for its whole lifetime. Bounded at 10s and non-fatal: a gateway without cron beats no gateway.
  • New lingerEnabled, read back off config rather than assumed. With createUser = false the operator owns the user and nothing here knows whether they lingered it, so all three parts stay off and the unit does not block on a bus that may never arrive.

nix/checks.nix — new cron-worker-user-scope check.

Container mode (container.enable = true) is untouched: the gateway runs inside the container, without the host unit's INVOCATION_ID, so the guard does not fire there.

Requires nixpkgs >= 25.05 for users.manageLingering (noted at the call site).

How to Test

The precondition, on a NixOS host before the change — reproduces the gateway's cgroup and uid without waiting for a fire:

sudo systemd-run -p User=hermes -p Group=hermes --pipe --wait --collect \
  /run/current-system/sw/bin/systemd-run --user --scope --quiet --unit probe /bin/true
# Failed to connect to bus: No medium found

After nixos-rebuild switch plus one systemctl restart hermes-agent (a switch that only flips linger does not restart the unit, and the running process keeps its busless environment), the same probe exits 0 and gateway-dispatched fires run.

The check is eval-only, so it also runs from a non-Linux host:

nix eval .#checks.x86_64-linux.cron-worker-user-scope.drvPath

On this commit it evaluates. With nix/nixosModules.nix reverted to the parent it throws, on the three arms the fix addresses:

error: cron worker user scope check failed:
  - the uid the gateway runs as (hermes) must linger: `systemd-run --user --scope` has no
    user manager to ask without it, and cron dispatch fails closed
  - a lingering gateway must be ordered after linger-users.service, the unit that runs
    `loginctl enable-linger`
  - a lingering gateway must wait for its user bus before ExecStart: the bus environment
    is resolved once at startup, so a bus that appears later is one this process never sees

The fourth arm (a gateway whose uid is not known to linger must not block on a bus) passes on both, which is the point — it guards the gate, not the fix.

Checklist

Code

  • I've read the Contributing Guide
  • My commit messages follow Conventional Commits
  • I searched for existing PRs to make sure this isn't a duplicate
  • My PR contains only changes related to this fix
  • I've run pytest tests/ -q — N/A, no Python changed; verified with the nix eval receipts above
  • I've added tests for my changes (nix/checks.nix, red on the parent commit)
  • I've tested on my platform: NixOS 25.11 / systemd 260.1 (GCE), check evaluated from macOS 15

Documentation & Housekeeping

  • I've updated relevant documentation — the reasoning lives at the call sites; happy to add a NixOS deployment note if you'd like one
  • cli-config.yaml.example — N/A, no config keys
  • CONTRIBUTING.md / AGENTS.md — N/A
  • Cross-platform impact — N/A, NixOS-module-only; container mode and Home Manager untouched
  • Tool descriptions/schemas — N/A

The NixOS module produces the one topology the restart-safe cron worker
cannot run in: a system service (`User=hermes`, `isSystemUser = true`)
whose uid has no systemd user manager. `run_one_job` hands every fire to
`_launch_external_cron_worker` → `restart_safe_gateway_child_argv`, which
requires a transient `systemd-run --user --scope` and fails closed when it
cannot get one:

  Restart-safe cron worker dispatch failed: cannot create restart-safe
  systemd scope for gateway child: systemd-run --user --scope is
  unavailable (usually no reachable user D-Bus session at
  /run/user/995/bus). On a system-level service install, run
  `sudo loginctl enable-linger <gateway-user>` and restart the gateway.

So on a stock `services.hermes-agent.enable = true` host, no cron job runs
at all — every fire is recorded as a failed execution before an agent
starts, and the occurrence is consumed, so a weekly job does not retry for
another week. The error names the remedy, but nothing in the module applies
it, and the terminal-worker path with the same cgroup-isolation goal
degrades gracefully instead (`process_registry.py`: "worker shares the
gateway cgroup"), so nothing else on the host looks wrong.

Three parts, because linger alone leaves a startup race:

- `linger = lib.mkDefault true` on the created user: the declarative
  `loginctl enable-linger`, which gives the uid a user manager and with it
  /run/user/<uid>/bus. mkDefault, so an operator can still refuse it.
- after/wants linger-users.service, the unit that runs enable-linger.
- a bounded preStart wait for the socket. logind starts user@<uid>.service
  asynchronously, and `run_gateway` resolves XDG_RUNTIME_DIR /
  DBUS_SESSION_BUS_ADDRESS exactly once at startup (802f0f9), so a bus that
  appears after ExecStart is one the process never sees for its lifetime.
  Non-fatal after 10s: a gateway without cron beats no gateway.

All three are gated on `lingerEnabled`, read back off `config` rather than
assumed: with `createUser = false` the operator owns the user, nothing here
knows whether they lingered it, and the unit must not block on a bus that
may never arrive.

Test: nix/checks.nix `cron-worker-user-scope` relates the unit's `User=` to
that user's `linger`, requires the ordering and the wait when it lingers,
and requires neither when the module does not own the user. Red on the
parent commit for the first three arms.

Needs nixpkgs >= 25.05 for `users.manageLingering`.
@KeyArgo

KeyArgo commented Sep 14, 2026

Copy link
Copy Markdown

Checked the mechanism against the released tags and against current main; three things worth folding into this PR, one of them a direct answer to the open question at the end of #110628.

1. The fail-closed behaviour is real on 0.21.1 — and still present in 0.21.2. At both v2026.9.7 and v2026.9.11, restart_safe_gateway_child_argv (tools/process_registry.py) raises unconditionally when _systemd_run_user_scope_available() is false: no caller-supplied policy, no fallback. cron/scheduler.py::_launch_external_cron_worker had no escape hatch either — its docstring at both tags reads "failure to establish the transient scope raises: falling back would recreate the restart interruption this handoff exists to prevent". So on a NixOS system unit every gateway-dispatched fire is consumed as a failed execution, exactly as described.

2. On current main that failure mode has already been re-decided — which answers open question 1. The degrade-with-warning tradeoff now exists, and it is the default:

  • restart_safe_gateway_child_argv returns a GatewayChildDispatch whose mode is in_process | scoped | degraded; require_restart_safe_scope=True raises, False logs once per process (_warn_scope_degraded_once) and dispatches the job as a direct external subprocess. degraded is deliberately a distinct mode so it can never collapse back into in_process and recreate the restart interruption fix(cron): preserve active turns across gateway restarts (salvage #101877) #101940 closed.
  • cron passes cron.require_restart_safe_scope, which defaults to false (hermes_cli/config_defaults.py); kanban always passes True, so kanban still fails closed regardless.
  • Documented in website/docs/user-guide/features/cron.md under Restart-safe workers under systemd: hosts without a user session — "containers, minimal LXCs, a service user without linger" — degrade by default, and linger is named there as "the lasting fix".

So after 0.21.2 the symptom is "cron runs, without cgroup isolation, one warning per gateway process" rather than "no cron at all" — unless an operator sets cron.require_restart_safe_scope: true. The module change here is still the right lasting fix (degraded → isolated, and it satisfies the precondition the runtime half already expects, from 802f0f9); it is the severity framing that only holds up to and including 0.21.2.

3. Two current-main facts that support the preStart wait, plus a line-number correction:

@alt-glitch alt-glitch added type/bug Something isn't working P2 Medium — degraded but workaround exists area/nix Nix flake, NixOS module, container packaging comp/cron Cron scheduler and job management labels Sep 14, 2026
@haizaar haizaar closed this Sep 14, 2026
@haizaar haizaar reopened this Sep 14, 2026
@haizaar

haizaar commented Sep 14, 2026

Copy link
Copy Markdown
Author

Reopened — the close was mine and it was a mistake. A downstream PR in our own fleet repo referenced this one, and I read "fix merged" off that. The fix shipped in our tree, not in yours. Confirmed with git grep linger upstream/main -- nix/nixosModules.nix → nothing, so nothing here is implemented on main.

Thanks for the review — it is more useful than the PR it landed on. I checked each point against the tree rather than taking it on trust, and all three hold:

  • Degrade mode is main-only. git grep -c require_restart_safe_scope v2026.9.11 -- cron/scheduler.py → no match; on upstream/main it is cron/scheduler.py:3139-3148 reading cron.require_restart_safe_scope, defaulting false at hermes_cli/config_defaults.py:1715, with hermes_cli/kanban_db_dispatch.py:2163,2175 passing True unconditionally. So fail-closed is the behaviour through 0.21.2 inclusive, and the severity framing in [Bug]: NixOS module — cron never runs on a system service (restart-safe worker cannot create its user scope) #110628 holds exactly up to that tag and no further. Our fleet is on 0.21.1, which is where we hit it.
  • Line refs. _ensure_user_systemd_env is hermes_cli/gateway.py:2135 on main and _wait_for_target_user_bus is :2165. The :4535 in the issue is the 0.21.1 tree, as you say — worth stating plainly rather than leaving two numbers floating around.
  • cron.md does name linger as the lasting fix, and update_cmd_fleet.py:919-920 does share the precondition.

Agreed on reusing _wait_for_target_user_bus, and it changes the shape of the fix for the better. The wait belongs in run_gateway, next to the _ensure_user_systemd_env() call it exists to make succeed — not in a systemd unit. Same predicate, one implementation, and it covers every system-level install rather than only NixOS: a hand-written unit, Docker-with-systemd, or any INVOCATION_ID topology has the identical race, and none of them are reachable from this module. The fleet-update path you point at gets it for free.

Why the first cut duplicated the predicate in the unit: _wait_for_target_user_bus does not exist at v2026.9.7, which is the base our fork rides. That is a constraint of our tree, not a preference, and it should not shape a PR against main.

So, pushing next:

  • hermes_cli/gateway.py — wait for our own user bus before adopting it, reusing _wait_for_target_user_bus, under the existing INVOCATION_ID gate at 4739-4741. Bounded by the helper's own 5s.
  • nix/nixosModules.nix — keeps linger = mkDefault true and the linger-users.service ordering, drops the preStart loop.
  • tests/hermes_cli/test_gateway_user_bus_adoption.py — a late-appearing-bus case: the socket is absent at the first probe and present shortly after, and the gateway still ends up with DBUS_SESSION_BUS_ADDRESS set. Red before the change.
  • nix/checks.nix — the cron-worker-user-scope arms narrow to linger + ordering, since the wait is no longer the module's business.

One open question, since your docs already take a position on it. cron.md recommends XDG_RUNTIME_DIR / DBUS_SESSION_BUS_ADDRESS in the unit for system-level installs. With the wait-then-adopt path above, that stops being necessary — the process resolves both itself, whenever the bus arrives. Would you rather the module also set them declaratively (/run/user/%U, belt and braces, and it makes the unit self-describing), or keep adoption as the single path and leave the docs describing the manual case only? Happy either way; I would rather not ship two mechanisms for one job without you choosing which is canonical.

@haizaar

haizaar commented Oct 1, 2026

Copy link
Copy Markdown
Author

Closing. This landed on main as ec30eef (cherry-picked with authorship preserved), with the follow-up ea5757a, which skips lingering in container mode. It shipped in v2026.9.24 (v0.21.5). The patch-id matches this branch exactly, so there is nothing left here to merge. #110628 is already closed.

Thanks for picking it up.

@haizaar haizaar closed this Oct 1, 2026
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

area/nix Nix flake, NixOS module, container packaging comp/cron Cron scheduler and job management P2 Medium — degraded but workaround exists type/bug Something isn't working

Projects

None yet

Development

Successfully merging this pull request may close these issues.

[Bug]: NixOS module — cron never runs on a system service (restart-safe worker cannot create its user scope)

3 participants