Skip to content
19 changes: 19 additions & 0 deletions architectures/aws-pcs/assets/add-cng-p5.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -393,6 +393,25 @@ Resources:
MIME-Version: 1.0

runcmd:
# --- Protect running jobs from unattended-upgrades / needrestart ---
# apt-daily-upgrade updates base libraries (e.g. glibc). needrestart then
# auto-restarts every service linked against them. slurmd links libc, so it
# gets restarted mid-upgrade - which KILLS the jobs running under it (the
# step is torn down and the job requeues from scratch). Exclude slurmd
# from needrestart's automatic restart so a security upgrade can
# never take down a running job. Security packages still install; only the
# slurmd auto-restart is suppressed. (PCS runs the controller managed-side;
# slurmd is the only Slurm systemd service on login/compute nodes. qr(^slurmd)
# also covers the versioned units, e.g. slurmd-25.11.)
# runcmd runs under /bin/sh (dash) - keep this POSIX-clean, no bashisms.
- |
mkdir -p /etc/needrestart/conf.d
cat > /etc/needrestart/conf.d/90-pcs-slurm.conf <<'NRCONF'
# AWS PCS: never auto-restart slurmd - restarting it kills
# the jobs running under it. Managed by add-cng UserData.
$nrconf{override_rc} = { qr(^slurmd) => 0 };
NRCONF
echo "needrestart: slurmd excluded from auto-restart" | tee /var/log/pcs-needrestart-guard.log
# Post-install script (optional) - generic OnNodeConfigured-style hook.
# Downloads and runs a user-supplied script (e.g. Enroot/Pyxis install).
- |
Expand Down
19 changes: 19 additions & 0 deletions architectures/aws-pcs/assets/add-cng-p6-b200.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -380,6 +380,25 @@ Resources:
MIME-Version: 1.0

runcmd:
# --- Protect running jobs from unattended-upgrades / needrestart ---
# apt-daily-upgrade updates base libraries (e.g. glibc). needrestart then
# auto-restarts every service linked against them. slurmd links libc, so it
# gets restarted mid-upgrade - which KILLS the jobs running under it (the
# step is torn down and the job requeues from scratch). Exclude slurmd
# from needrestart's automatic restart so a security upgrade can
# never take down a running job. Security packages still install; only the
# slurmd auto-restart is suppressed. (PCS runs the controller managed-side;
# slurmd is the only Slurm systemd service on login/compute nodes. qr(^slurmd)
# also covers the versioned units, e.g. slurmd-25.11.)
# runcmd runs under /bin/sh (dash) - keep this POSIX-clean, no bashisms.
- |
mkdir -p /etc/needrestart/conf.d
cat > /etc/needrestart/conf.d/90-pcs-slurm.conf <<'NRCONF'
# AWS PCS: never auto-restart slurmd - restarting it kills
# the jobs running under it. Managed by add-cng UserData.
$nrconf{override_rc} = { qr(^slurmd) => 0 };
NRCONF
echo "needrestart: slurmd excluded from auto-restart" | tee /var/log/pcs-needrestart-guard.log
# Post-install script (optional) - generic OnNodeConfigured-style hook.
# Downloads and runs a user-supplied script (e.g. Enroot/Pyxis install).
- |
Expand Down
19 changes: 19 additions & 0 deletions architectures/aws-pcs/assets/add-cng-p6-b300.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -383,6 +383,25 @@ Resources:
MIME-Version: 1.0

runcmd:
# --- Protect running jobs from unattended-upgrades / needrestart ---
# apt-daily-upgrade updates base libraries (e.g. glibc). needrestart then
# auto-restarts every service linked against them. slurmd links libc, so it
# gets restarted mid-upgrade - which KILLS the jobs running under it (the
# step is torn down and the job requeues from scratch). Exclude slurmd
# from needrestart's automatic restart so a security upgrade can
# never take down a running job. Security packages still install; only the
# slurmd auto-restart is suppressed. (PCS runs the controller managed-side;
# slurmd is the only Slurm systemd service on login/compute nodes. qr(^slurmd)
# also covers the versioned units, e.g. slurmd-25.11.)
# runcmd runs under /bin/sh (dash) - keep this POSIX-clean, no bashisms.
- |
mkdir -p /etc/needrestart/conf.d
cat > /etc/needrestart/conf.d/90-pcs-slurm.conf <<'NRCONF'
# AWS PCS: never auto-restart slurmd - restarting it kills
# the jobs running under it. Managed by add-cng UserData.
$nrconf{override_rc} = { qr(^slurmd) => 0 };
NRCONF
echo "needrestart: slurmd excluded from auto-restart" | tee /var/log/pcs-needrestart-guard.log
# Post-install script (optional) - generic OnNodeConfigured-style hook.
# Downloads and runs a user-supplied script (e.g. Enroot/Pyxis install).
- |
Expand Down
19 changes: 19 additions & 0 deletions architectures/aws-pcs/assets/add-cng.yaml
Original file line number Diff line number Diff line change
Expand Up @@ -451,6 +451,25 @@ Resources:
MIME-Version: 1.0

runcmd:
# --- Protect running jobs from unattended-upgrades / needrestart ---
# apt-daily-upgrade updates base libraries (e.g. glibc). needrestart then
# auto-restarts every service linked against them. slurmd links libc, so it
# gets restarted mid-upgrade - which KILLS the jobs running under it (the
# step is torn down and the job requeues from scratch). Exclude slurmd
# from needrestart's automatic restart so a security upgrade can
# never take down a running job. Security packages still install; only the
# slurmd auto-restart is suppressed. (PCS runs the controller managed-side;
# slurmd is the only Slurm systemd service on login/compute nodes. qr(^slurmd)
# also covers the versioned units, e.g. slurmd-25.11.)
# runcmd runs under /bin/sh (dash) — keep this POSIX-clean, no bashisms.
- |
mkdir -p /etc/needrestart/conf.d
cat > /etc/needrestart/conf.d/90-pcs-slurm.conf <<'NRCONF'
# AWS PCS: never auto-restart slurmd - restarting it kills
# the jobs running under it. Managed by add-cng.yaml UserData.
$nrconf{override_rc} = { qr(^slurmd) => 0 };

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

qr(^slurmd) also prefix-matches slurmdbd — no change needed here (FYI)

qr(^slurmd) is an unanchored-prefix regex, so it also matches slurmdbd.service in addition to slurmd and the versioned slurmd-25.11. On PCS login/compute nodes this has no practical effect — as the comment right above correctly notes, slurmd is the only Slurm unit here (the controller and slurmdbd are managed PCS-side). I'd leave it exactly as written: tightening it to qr(^slurmd($|-|\.)) would only guard a unit that isn't present on these nodes. Flagging purely so the prefix-match is a known, deliberate property. It's the same qr(^name) => 0 idiom needrestart ships for dbus/display-managers.

NRCONF
echo "needrestart: slurmd excluded from auto-restart" | tee /var/log/pcs-needrestart-guard.log

Copy link
Copy Markdown
Collaborator

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

tee writes the success line unconditionally (nit)

The echo … | tee runs regardless of whether the preceding mkdir -p / cat > actually succeeded (no set -e, no && chain), so in the near-impossible failure case (e.g. a read-only /etc at first boot) the log would still read needrestart: slurmd excluded from auto-restart when it wasn't. Cosmetic only — the failure it would misreport isn't a realistic first-boot condition, so I wouldn't hold the PR for it. Mentioning it in case you'd prefer the log line to be strictly truthful (chain the write with && before the echo).

# Post-install script (optional) - generic OnNodeConfigured-style hook.
# Downloads and runs a user-supplied script (e.g. Enroot/Pyxis install).
# Accepts BOTH an s3:// URL (fetched with `aws s3 cp` using the instance
Expand Down
21 changes: 21 additions & 0 deletions architectures/aws-pcs/docs/OPERATIONS.md
Original file line number Diff line number Diff line change
Expand Up @@ -466,6 +466,27 @@ cleanest mitigation is upstream — recording it here so users seeing it know th
workaround and the next contributor doesn't waste time looking for a bug in this PR's
scripts.

### 6.2 `needrestart` restarting `slurmd` would stop running jobs — already handled

When an unattended security upgrade updates a base library `slurmd` links (e.g. glibc),
`needrestart` restarts `slurmd`, and that restart tears down the `slurmstepd` steps under
it — **stopping every job on the node**. It is reproducible, not random, and not a reboot
or a Slurm-package upgrade.

The compute-node-group templates already guard against this: each `add-cng*` UserData
writes a `needrestart` drop-in so `slurmd` is never auto-restarted (security updates still
install; `needrestart` only defers the `slurmd` restart):

```perl
# /etc/needrestart/conf.d/90-pcs-slurm.conf (written by add-cng* UserData)
$nrconf{override_rc} = { qr(^slurmd) => 0 };
```

`slurmd` is the only Slurm systemd service on these nodes (the controller is managed by
PCS); `qr(^slurmd)` also matches the versioned units (e.g. `slurmd-25.11`). The drop-in is
a standalone `conf.d/*.conf` naming only `slurmd`, so if a later DLAMI or Slurm unit
handles this differently it neither conflicts nor errors — at worst it becomes redundant.

## 7. Recommendations recap

For a new production deploy:
Expand Down