From 1f3489d19b56cabe6990d5df982c3f98e0eaee6a Mon Sep 17 00:00:00 2001 From: Daisuke Miyamoto Date: Mon, 6 Jul 2026 10:08:00 +0000 Subject: [PATCH 1/7] fix(aws-pcs): stop needrestart from restarting slurmd and killing running jobs apt-daily-upgrade updates base libraries (e.g. glibc); needrestart then auto-restarts every service linked against them. slurmd links libc, so it is restarted mid-upgrade, which tears down slurmstepd and kills the jobs running on the node (the job then requeues from scratch). This matched a customer report of jobs being killed at the same time apt-daily-upgrade.timer fired. Add a needrestart drop-in via CNG UserData that excludes slurmd/slurmctld/ slurmdbd from automatic restart. Security upgrades still install; only the Slurm-daemon auto-restart is suppressed, so a running job is never taken down. Verified end-to-end on real hardware: with the guard, a long-running job survives an apt upgrade + needrestart that updates glibc; without it, slurmd restarts and the job is killed/requeued. --- architectures/aws-pcs/assets/add-cng.yaml | 17 +++++++++++++++++ 1 file changed, 17 insertions(+) diff --git a/architectures/aws-pcs/assets/add-cng.yaml b/architectures/aws-pcs/assets/add-cng.yaml index 808f22a95..f7b8071f9 100644 --- a/architectures/aws-pcs/assets/add-cng.yaml +++ b/architectures/aws-pcs/assets/add-cng.yaml @@ -451,6 +451,23 @@ Resources: MIME-Version: 1.0 runcmd: + # --- Protect running jobs from unattended-upgrades / needrestart --- + # apt-daily-upgrade updates base libraries (e.g. glibc). needrestart then + # auto-restarts every service linked against them. slurmd links libc, so it + # gets restarted mid-upgrade — which KILLS the jobs running under it (the + # step is torn down and the job requeues from scratch). Exclude the Slurm + # daemons from needrestart's automatic restart so a security upgrade can + # never take down a running job. Security packages still install; only the + # slurmd/slurmctld/slurmdbd auto-restart is suppressed. + # runcmd runs under /bin/sh (dash) — keep this POSIX-clean, no bashisms. + - | + mkdir -p /etc/needrestart/conf.d + cat > /etc/needrestart/conf.d/90-pcs-slurm.conf <<'NRCONF' + # AWS PCS: never auto-restart Slurm daemons — restarting slurmd kills + # the jobs running under it. Managed by add-cng.yaml UserData. + $nrconf{override_rc} = { qr(^slurmd) => 0, qr(^slurmctld) => 0, qr(^slurmdbd) => 0 }; + NRCONF + echo "needrestart: slurmd/slurmctld/slurmdbd excluded from auto-restart" | tee /var/log/pcs-needrestart-guard.log # Post-install script (optional) - generic OnNodeConfigured-style hook. # Downloads and runs a user-supplied script (e.g. Enroot/Pyxis install). # Accepts BOTH an s3:// URL (fetched with `aws s3 cp` using the instance From fc69fd001ac0798c022370a630d86801d53c33e9 Mon Sep 17 00:00:00 2001 From: Daisuke Miyamoto Date: Mon, 6 Jul 2026 13:35:22 +0000 Subject: [PATCH 2/7] fix(aws-pcs): use ASCII dash in needrestart guard comment Avoid a non-ASCII em-dash in the heredoc that gets written to /etc/needrestart/conf.d/90-pcs-slurm.conf (rendered as an escaped code point on disk). Comment-only change; the override_rc line is unaffected. --- architectures/aws-pcs/assets/add-cng.yaml | 4 ++-- 1 file changed, 2 insertions(+), 2 deletions(-) diff --git a/architectures/aws-pcs/assets/add-cng.yaml b/architectures/aws-pcs/assets/add-cng.yaml index f7b8071f9..dca00abb0 100644 --- a/architectures/aws-pcs/assets/add-cng.yaml +++ b/architectures/aws-pcs/assets/add-cng.yaml @@ -454,7 +454,7 @@ Resources: # --- Protect running jobs from unattended-upgrades / needrestart --- # apt-daily-upgrade updates base libraries (e.g. glibc). needrestart then # auto-restarts every service linked against them. slurmd links libc, so it - # gets restarted mid-upgrade — which KILLS the jobs running under it (the + # gets restarted mid-upgrade - which KILLS the jobs running under it (the # step is torn down and the job requeues from scratch). Exclude the Slurm # daemons from needrestart's automatic restart so a security upgrade can # never take down a running job. Security packages still install; only the @@ -463,7 +463,7 @@ Resources: - | mkdir -p /etc/needrestart/conf.d cat > /etc/needrestart/conf.d/90-pcs-slurm.conf <<'NRCONF' - # AWS PCS: never auto-restart Slurm daemons — restarting slurmd kills + # AWS PCS: never auto-restart Slurm daemons - restarting slurmd kills # the jobs running under it. Managed by add-cng.yaml UserData. $nrconf{override_rc} = { qr(^slurmd) => 0, qr(^slurmctld) => 0, qr(^slurmdbd) => 0 }; NRCONF From 04401f60a8293887e6f5438df846bb5499923eb2 Mon Sep 17 00:00:00 2001 From: Daisuke Miyamoto Date: Mon, 6 Jul 2026 13:47:01 +0000 Subject: [PATCH 3/7] fix(aws-pcs): apply needrestart slurmd guard to GPU CNG templates too The needrestart/slurmd job-kill guard was only added to add-cng.yaml. Apply the same drop-in to the P5/P6-B200/P6-B300 GPU templates, which run the long distributed-training jobs most affected by an unattended slurmd restart. --- architectures/aws-pcs/assets/add-cng-p5.yaml | 17 +++++++++++++++++ .../aws-pcs/assets/add-cng-p6-b200.yaml | 17 +++++++++++++++++ .../aws-pcs/assets/add-cng-p6-b300.yaml | 17 +++++++++++++++++ 3 files changed, 51 insertions(+) diff --git a/architectures/aws-pcs/assets/add-cng-p5.yaml b/architectures/aws-pcs/assets/add-cng-p5.yaml index 27c998ad7..613d7b7bd 100644 --- a/architectures/aws-pcs/assets/add-cng-p5.yaml +++ b/architectures/aws-pcs/assets/add-cng-p5.yaml @@ -393,6 +393,23 @@ Resources: MIME-Version: 1.0 runcmd: + # --- Protect running jobs from unattended-upgrades / needrestart --- + # apt-daily-upgrade updates base libraries (e.g. glibc). needrestart then + # auto-restarts every service linked against them. slurmd links libc, so it + # gets restarted mid-upgrade - which KILLS the jobs running under it (the + # step is torn down and the job requeues from scratch). Exclude the Slurm + # daemons from needrestart's automatic restart so a security upgrade can + # never take down a running job. Security packages still install; only the + # slurmd/slurmctld/slurmdbd auto-restart is suppressed. + # runcmd runs under /bin/sh (dash) - keep this POSIX-clean, no bashisms. + - | + mkdir -p /etc/needrestart/conf.d + cat > /etc/needrestart/conf.d/90-pcs-slurm.conf <<'NRCONF' + # AWS PCS: never auto-restart Slurm daemons - restarting slurmd kills + # the jobs running under it. Managed by add-cng UserData. + $nrconf{override_rc} = { qr(^slurmd) => 0, qr(^slurmctld) => 0, qr(^slurmdbd) => 0 }; + NRCONF + echo "needrestart: slurmd/slurmctld/slurmdbd excluded from auto-restart" | tee /var/log/pcs-needrestart-guard.log # Post-install script (optional) - generic OnNodeConfigured-style hook. # Downloads and runs a user-supplied script (e.g. Enroot/Pyxis install). - | diff --git a/architectures/aws-pcs/assets/add-cng-p6-b200.yaml b/architectures/aws-pcs/assets/add-cng-p6-b200.yaml index 618caee5e..fade3d128 100644 --- a/architectures/aws-pcs/assets/add-cng-p6-b200.yaml +++ b/architectures/aws-pcs/assets/add-cng-p6-b200.yaml @@ -380,6 +380,23 @@ Resources: MIME-Version: 1.0 runcmd: + # --- Protect running jobs from unattended-upgrades / needrestart --- + # apt-daily-upgrade updates base libraries (e.g. glibc). needrestart then + # auto-restarts every service linked against them. slurmd links libc, so it + # gets restarted mid-upgrade - which KILLS the jobs running under it (the + # step is torn down and the job requeues from scratch). Exclude the Slurm + # daemons from needrestart's automatic restart so a security upgrade can + # never take down a running job. Security packages still install; only the + # slurmd/slurmctld/slurmdbd auto-restart is suppressed. + # runcmd runs under /bin/sh (dash) - keep this POSIX-clean, no bashisms. + - | + mkdir -p /etc/needrestart/conf.d + cat > /etc/needrestart/conf.d/90-pcs-slurm.conf <<'NRCONF' + # AWS PCS: never auto-restart Slurm daemons - restarting slurmd kills + # the jobs running under it. Managed by add-cng UserData. + $nrconf{override_rc} = { qr(^slurmd) => 0, qr(^slurmctld) => 0, qr(^slurmdbd) => 0 }; + NRCONF + echo "needrestart: slurmd/slurmctld/slurmdbd excluded from auto-restart" | tee /var/log/pcs-needrestart-guard.log # Post-install script (optional) - generic OnNodeConfigured-style hook. # Downloads and runs a user-supplied script (e.g. Enroot/Pyxis install). - | diff --git a/architectures/aws-pcs/assets/add-cng-p6-b300.yaml b/architectures/aws-pcs/assets/add-cng-p6-b300.yaml index f4f1dc75d..77c75b520 100644 --- a/architectures/aws-pcs/assets/add-cng-p6-b300.yaml +++ b/architectures/aws-pcs/assets/add-cng-p6-b300.yaml @@ -383,6 +383,23 @@ Resources: MIME-Version: 1.0 runcmd: + # --- Protect running jobs from unattended-upgrades / needrestart --- + # apt-daily-upgrade updates base libraries (e.g. glibc). needrestart then + # auto-restarts every service linked against them. slurmd links libc, so it + # gets restarted mid-upgrade - which KILLS the jobs running under it (the + # step is torn down and the job requeues from scratch). Exclude the Slurm + # daemons from needrestart's automatic restart so a security upgrade can + # never take down a running job. Security packages still install; only the + # slurmd/slurmctld/slurmdbd auto-restart is suppressed. + # runcmd runs under /bin/sh (dash) - keep this POSIX-clean, no bashisms. + - | + mkdir -p /etc/needrestart/conf.d + cat > /etc/needrestart/conf.d/90-pcs-slurm.conf <<'NRCONF' + # AWS PCS: never auto-restart Slurm daemons - restarting slurmd kills + # the jobs running under it. Managed by add-cng UserData. + $nrconf{override_rc} = { qr(^slurmd) => 0, qr(^slurmctld) => 0, qr(^slurmdbd) => 0 }; + NRCONF + echo "needrestart: slurmd/slurmctld/slurmdbd excluded from auto-restart" | tee /var/log/pcs-needrestart-guard.log # Post-install script (optional) - generic OnNodeConfigured-style hook. # Downloads and runs a user-supplied script (e.g. Enroot/Pyxis install). - | From 5af4edc468357c296051ec6b25c715237dda9476 Mon Sep 17 00:00:00 2001 From: Daisuke Miyamoto Date: Mon, 6 Jul 2026 14:17:20 +0000 Subject: [PATCH 4/7] docs(aws-pcs): document needrestart/slurmd job-kill issue and fix (OPERATIONS 6.2) Explain how apt-daily-upgrade -> needrestart -> slurmd restart kills running jobs, the shipped needrestart drop-in that excludes the Slurm daemons, the heavier timer-disable alternative, and why the drop-in is forward-compatible. --- architectures/aws-pcs/docs/OPERATIONS.md | 55 ++++++++++++++++++++++++ 1 file changed, 55 insertions(+) diff --git a/architectures/aws-pcs/docs/OPERATIONS.md b/architectures/aws-pcs/docs/OPERATIONS.md index 9ef2b7837..77337121a 100644 --- a/architectures/aws-pcs/docs/OPERATIONS.md +++ b/architectures/aws-pcs/docs/OPERATIONS.md @@ -466,6 +466,61 @@ cleanest mitigation is upstream — recording it here so users seeing it know th workaround and the next contributor doesn't waste time looking for a bug in this PR's scripts. +### 6.2 `apt-daily-upgrade` can kill running jobs via `needrestart` restarting `slurmd` + +**Symptom.** Jobs are killed for no obvious reason, and the kill time lines up with when +`apt-daily-upgrade.timer` fires (by default in the early morning). The job either +disappears or requeues from scratch, losing all in-progress work. + +**Cause.** `apt-daily-upgrade` installs security updates, including base libraries such as +glibc. `needrestart` (which runs in automatic mode after unattended upgrades on the +PCS-ready DLAMI) then restarts every service linked against an updated library. `slurmd` +links `libc`, so `needrestart` restarts it — and restarting `slurmd` tears down the +`slurmstepd` steps under it, killing the jobs running on that node. Sequence: + +``` +T+0 apt-daily-upgrade.timer fires -> unattended-upgrades installs e.g. libc6 +T+1 needrestart (automatic mode) sees slurmd is linked against the updated libc +T+1 needrestart runs: systemctl restart slurmd +T+1 slurmd restart tears down slurmstepd -> the running job is killed / requeued +``` + +Note this is **not** about upgrading Slurm: `slurmd` lives under `/opt/aws/pcs/...` and is +not an apt-managed package. It is purely `needrestart` restarting `slurmd` because a +library it links was updated. Reboot is not involved either (`Automatic-Reboot` is off by +default) — the job dies from the `slurmd` restart, not a node reboot. + +**Fix (shipped in this repo).** The compute-node-group UserData writes a `needrestart` +drop-in that excludes the Slurm daemons from automatic restart: + +```perl +# /etc/needrestart/conf.d/90-pcs-slurm.conf (written by add-cng* UserData) +$nrconf{override_rc} = { qr(^slurmd) => 0, qr(^slurmctld) => 0, qr(^slurmdbd) => 0 }; +``` + +This is the minimum targeted change: security packages still install as before, and +`needrestart` still restarts everything else; only the automatic restart of the Slurm +daemons is suppressed, so an upgrade can never take down a running job. `needrestart` will +report the `slurmd` restart as *deferred* rather than performing it. `slurmd` picks up the +new libraries the next time it restarts on its own terms (node replacement, power-save +cycle, or a manual restart) — acceptable for HPC, where PCS replaces nodes regularly. + +Verified end-to-end on real hardware: a long-running job survives a real `apt-get upgrade` +that updates glibc followed by `needrestart -r a` (the `slurmd` PID is unchanged and the +job keeps running); without the drop-in, the same sequence restarts `slurmd` and the job +is killed. + +**If you still want to disable unattended upgrades entirely** (stops all security updates — +a heavier hammer), you can also mask the timers in your own post-install: +`systemctl disable --now apt-daily.timer apt-daily-upgrade.timer`. The `needrestart` +drop-in above is preferred because it keeps security updates flowing. + +**Forward-compatibility.** The drop-in lives in its own `conf.d` file and only names the +Slurm services, so it stays inert if upstream later addresses this a different way (e.g. a +`slurmd` unit that survives restart, or a DLAMI-level `needrestart` policy): `needrestart` +reads `conf.d/*.conf` in order and tolerates unknown keys, so an extra override never +causes an error — at worst it becomes redundant. + ## 7. Recommendations recap For a new production deploy: From 3a71d07bedc194709f3ca868c953c450bfe4038f Mon Sep 17 00:00:00 2001 From: Daisuke Miyamoto Date: Mon, 6 Jul 2026 15:46:09 +0000 Subject: [PATCH 5/7] docs(aws-pcs): reframe OPERATIONS 6.2 as user-facing guidance Rewrite the needrestart/slurmd section from an incident write-up into a note for users: state plainly that needrestart restarting slurmd stops running jobs, that the templates already include the mitigation so no action is needed, and that the drop-in stays harmless (no conflict/error) if the platform fixes this upstream. --- architectures/aws-pcs/docs/OPERATIONS.md | 84 +++++++++++------------- 1 file changed, 37 insertions(+), 47 deletions(-) diff --git a/architectures/aws-pcs/docs/OPERATIONS.md b/architectures/aws-pcs/docs/OPERATIONS.md index 77337121a..494f7644f 100644 --- a/architectures/aws-pcs/docs/OPERATIONS.md +++ b/architectures/aws-pcs/docs/OPERATIONS.md @@ -466,60 +466,50 @@ cleanest mitigation is upstream — recording it here so users seeing it know th workaround and the next contributor doesn't waste time looking for a bug in this PR's scripts. -### 6.2 `apt-daily-upgrade` can kill running jobs via `needrestart` restarting `slurmd` - -**Symptom.** Jobs are killed for no obvious reason, and the kill time lines up with when -`apt-daily-upgrade.timer` fires (by default in the early morning). The job either -disappears or requeues from scratch, losing all in-progress work. - -**Cause.** `apt-daily-upgrade` installs security updates, including base libraries such as -glibc. `needrestart` (which runs in automatic mode after unattended upgrades on the -PCS-ready DLAMI) then restarts every service linked against an updated library. `slurmd` -links `libc`, so `needrestart` restarts it — and restarting `slurmd` tears down the -`slurmstepd` steps under it, killing the jobs running on that node. Sequence: - -``` -T+0 apt-daily-upgrade.timer fires -> unattended-upgrades installs e.g. libc6 -T+1 needrestart (automatic mode) sees slurmd is linked against the updated libc -T+1 needrestart runs: systemctl restart slurmd -T+1 slurmd restart tears down slurmstepd -> the running job is killed / requeued -``` - -Note this is **not** about upgrading Slurm: `slurmd` lives under `/opt/aws/pcs/...` and is -not an apt-managed package. It is purely `needrestart` restarting `slurmd` because a -library it links was updated. Reboot is not involved either (`Automatic-Reboot` is off by -default) — the job dies from the `slurmd` restart, not a node reboot. - -**Fix (shipped in this repo).** The compute-node-group UserData writes a `needrestart` -drop-in that excludes the Slurm daemons from automatic restart: +### 6.2 `needrestart` must not auto-restart `slurmd` (would stop running jobs) — already handled + +**What to know.** On Ubuntu, automatic security upgrades (`apt-daily-upgrade` + +`unattended-upgrades`) update base libraries such as glibc, and `needrestart` then +restarts every service linked against them. `slurmd` links `libc`, so left unchecked +**`needrestart` restarts `slurmd`, and restarting `slurmd` tears down the `slurmstepd` +steps under it — stopping every job running on that node** (the job is killed and, if +requeued, restarts from scratch, losing in-progress work). This is a definite, +reproducible interaction, not a random failure: it happens whenever an unattended upgrade +touches a library `slurmd` uses. (It is not a Slurm upgrade — `slurmd` lives under +`/opt/aws/pcs/...` and is not apt-managed — and it is not a reboot; the job stops purely +because `slurmd` is restarted.) + +**You do not need to do anything — the templates already guard against this.** Every +compute-node-group UserData (`add-cng*`) writes a `needrestart` drop-in that excludes the +Slurm daemons from automatic restart: ```perl # /etc/needrestart/conf.d/90-pcs-slurm.conf (written by add-cng* UserData) $nrconf{override_rc} = { qr(^slurmd) => 0, qr(^slurmctld) => 0, qr(^slurmdbd) => 0 }; ``` -This is the minimum targeted change: security packages still install as before, and -`needrestart` still restarts everything else; only the automatic restart of the Slurm -daemons is suppressed, so an upgrade can never take down a running job. `needrestart` will -report the `slurmd` restart as *deferred* rather than performing it. `slurmd` picks up the -new libraries the next time it restarts on its own terms (node replacement, power-save -cycle, or a manual restart) — acceptable for HPC, where PCS replaces nodes regularly. - -Verified end-to-end on real hardware: a long-running job survives a real `apt-get upgrade` -that updates glibc followed by `needrestart -r a` (the `slurmd` PID is unchanged and the -job keeps running); without the drop-in, the same sequence restarts `slurmd` and the job -is killed. - -**If you still want to disable unattended upgrades entirely** (stops all security updates — -a heavier hammer), you can also mask the timers in your own post-install: +With this in place, security packages still install and `needrestart` still restarts +everything else; only the automatic restart of the Slurm daemons is suppressed +(`needrestart` reports it as *deferred*), so an unattended upgrade can no longer stop a +running job. `slurmd` picks up the new libraries the next time it restarts on its own +terms — node replacement, power-save cycle, or a manual restart — which is fine for HPC, +where PCS replaces nodes regularly. (Verified end-to-end on real hardware: a long-running +job survives a real `apt-get upgrade` of glibc followed by `needrestart -r a`, with the +`slurmd` PID unchanged.) + +**Safe if the platform later fixes this upstream.** The drop-in is a standalone +`conf.d/*.conf` file that only names the Slurm services. If a future PCS-ready DLAMI or +Slurm unit handles this differently (e.g. a `slurmd` unit that survives restart, or a +DLAMI-level `needrestart` policy), this file does not conflict or error — `needrestart` +reads `conf.d/*.conf` in order and tolerates keys it does not use, so at worst the +override becomes redundant. In other words, keeping it in place is harmless even after an +upstream fix; you never have to race to remove it. + +**If you would rather turn off automatic upgrades entirely** (heavier — this stops all +security updates), mask the timers in your own post-install instead: `systemctl disable --now apt-daily.timer apt-daily-upgrade.timer`. The `needrestart` -drop-in above is preferred because it keeps security updates flowing. - -**Forward-compatibility.** The drop-in lives in its own `conf.d` file and only names the -Slurm services, so it stays inert if upstream later addresses this a different way (e.g. a -`slurmd` unit that survives restart, or a DLAMI-level `needrestart` policy): `needrestart` -reads `conf.d/*.conf` in order and tolerates unknown keys, so an extra override never -causes an error — at worst it becomes redundant. +drop-in above is preferred because it keeps security updates flowing while still protecting +jobs. ## 7. Recommendations recap From e0c495e7f7ef9f1bf1e1e310857e5d921b4baaf6 Mon Sep 17 00:00:00 2001 From: Daisuke Miyamoto Date: Mon, 6 Jul 2026 15:55:14 +0000 Subject: [PATCH 6/7] fix(aws-pcs): narrow needrestart guard to slurmd only slurmctld/slurmdbd are not systemd services on PCS login/compute nodes (PCS runs the controller managed-side), so excluding them was a no-op. Guard only slurmd, which qr(^slurmd) also matches for the versioned units (slurmd-25.11 etc.). Update OPERATIONS 6.2 accordingly. --- architectures/aws-pcs/assets/add-cng-p5.yaml | 14 ++++++++------ .../aws-pcs/assets/add-cng-p6-b200.yaml | 14 ++++++++------ .../aws-pcs/assets/add-cng-p6-b300.yaml | 14 ++++++++------ architectures/aws-pcs/assets/add-cng.yaml | 14 ++++++++------ architectures/aws-pcs/docs/OPERATIONS.md | 16 +++++++++------- 5 files changed, 41 insertions(+), 31 deletions(-) diff --git a/architectures/aws-pcs/assets/add-cng-p5.yaml b/architectures/aws-pcs/assets/add-cng-p5.yaml index 613d7b7bd..7993b0530 100644 --- a/architectures/aws-pcs/assets/add-cng-p5.yaml +++ b/architectures/aws-pcs/assets/add-cng-p5.yaml @@ -397,19 +397,21 @@ Resources: # apt-daily-upgrade updates base libraries (e.g. glibc). needrestart then # auto-restarts every service linked against them. slurmd links libc, so it # gets restarted mid-upgrade - which KILLS the jobs running under it (the - # step is torn down and the job requeues from scratch). Exclude the Slurm - # daemons from needrestart's automatic restart so a security upgrade can + # step is torn down and the job requeues from scratch). Exclude slurmd + # from needrestart's automatic restart so a security upgrade can # never take down a running job. Security packages still install; only the - # slurmd/slurmctld/slurmdbd auto-restart is suppressed. + # slurmd auto-restart is suppressed. (PCS runs the controller managed-side; + # slurmd is the only Slurm systemd service on login/compute nodes. qr(^slurmd) + # also covers the versioned units, e.g. slurmd-25.11.) # runcmd runs under /bin/sh (dash) - keep this POSIX-clean, no bashisms. - | mkdir -p /etc/needrestart/conf.d cat > /etc/needrestart/conf.d/90-pcs-slurm.conf <<'NRCONF' - # AWS PCS: never auto-restart Slurm daemons - restarting slurmd kills + # AWS PCS: never auto-restart slurmd - restarting it kills # the jobs running under it. Managed by add-cng UserData. - $nrconf{override_rc} = { qr(^slurmd) => 0, qr(^slurmctld) => 0, qr(^slurmdbd) => 0 }; + $nrconf{override_rc} = { qr(^slurmd) => 0 }; NRCONF - echo "needrestart: slurmd/slurmctld/slurmdbd excluded from auto-restart" | tee /var/log/pcs-needrestart-guard.log + echo "needrestart: slurmd excluded from auto-restart" | tee /var/log/pcs-needrestart-guard.log # Post-install script (optional) - generic OnNodeConfigured-style hook. # Downloads and runs a user-supplied script (e.g. Enroot/Pyxis install). - | diff --git a/architectures/aws-pcs/assets/add-cng-p6-b200.yaml b/architectures/aws-pcs/assets/add-cng-p6-b200.yaml index fade3d128..df360f04e 100644 --- a/architectures/aws-pcs/assets/add-cng-p6-b200.yaml +++ b/architectures/aws-pcs/assets/add-cng-p6-b200.yaml @@ -384,19 +384,21 @@ Resources: # apt-daily-upgrade updates base libraries (e.g. glibc). needrestart then # auto-restarts every service linked against them. slurmd links libc, so it # gets restarted mid-upgrade - which KILLS the jobs running under it (the - # step is torn down and the job requeues from scratch). Exclude the Slurm - # daemons from needrestart's automatic restart so a security upgrade can + # step is torn down and the job requeues from scratch). Exclude slurmd + # from needrestart's automatic restart so a security upgrade can # never take down a running job. Security packages still install; only the - # slurmd/slurmctld/slurmdbd auto-restart is suppressed. + # slurmd auto-restart is suppressed. (PCS runs the controller managed-side; + # slurmd is the only Slurm systemd service on login/compute nodes. qr(^slurmd) + # also covers the versioned units, e.g. slurmd-25.11.) # runcmd runs under /bin/sh (dash) - keep this POSIX-clean, no bashisms. - | mkdir -p /etc/needrestart/conf.d cat > /etc/needrestart/conf.d/90-pcs-slurm.conf <<'NRCONF' - # AWS PCS: never auto-restart Slurm daemons - restarting slurmd kills + # AWS PCS: never auto-restart slurmd - restarting it kills # the jobs running under it. Managed by add-cng UserData. - $nrconf{override_rc} = { qr(^slurmd) => 0, qr(^slurmctld) => 0, qr(^slurmdbd) => 0 }; + $nrconf{override_rc} = { qr(^slurmd) => 0 }; NRCONF - echo "needrestart: slurmd/slurmctld/slurmdbd excluded from auto-restart" | tee /var/log/pcs-needrestart-guard.log + echo "needrestart: slurmd excluded from auto-restart" | tee /var/log/pcs-needrestart-guard.log # Post-install script (optional) - generic OnNodeConfigured-style hook. # Downloads and runs a user-supplied script (e.g. Enroot/Pyxis install). - | diff --git a/architectures/aws-pcs/assets/add-cng-p6-b300.yaml b/architectures/aws-pcs/assets/add-cng-p6-b300.yaml index 77c75b520..8f83f424d 100644 --- a/architectures/aws-pcs/assets/add-cng-p6-b300.yaml +++ b/architectures/aws-pcs/assets/add-cng-p6-b300.yaml @@ -387,19 +387,21 @@ Resources: # apt-daily-upgrade updates base libraries (e.g. glibc). needrestart then # auto-restarts every service linked against them. slurmd links libc, so it # gets restarted mid-upgrade - which KILLS the jobs running under it (the - # step is torn down and the job requeues from scratch). Exclude the Slurm - # daemons from needrestart's automatic restart so a security upgrade can + # step is torn down and the job requeues from scratch). Exclude slurmd + # from needrestart's automatic restart so a security upgrade can # never take down a running job. Security packages still install; only the - # slurmd/slurmctld/slurmdbd auto-restart is suppressed. + # slurmd auto-restart is suppressed. (PCS runs the controller managed-side; + # slurmd is the only Slurm systemd service on login/compute nodes. qr(^slurmd) + # also covers the versioned units, e.g. slurmd-25.11.) # runcmd runs under /bin/sh (dash) - keep this POSIX-clean, no bashisms. - | mkdir -p /etc/needrestart/conf.d cat > /etc/needrestart/conf.d/90-pcs-slurm.conf <<'NRCONF' - # AWS PCS: never auto-restart Slurm daemons - restarting slurmd kills + # AWS PCS: never auto-restart slurmd - restarting it kills # the jobs running under it. Managed by add-cng UserData. - $nrconf{override_rc} = { qr(^slurmd) => 0, qr(^slurmctld) => 0, qr(^slurmdbd) => 0 }; + $nrconf{override_rc} = { qr(^slurmd) => 0 }; NRCONF - echo "needrestart: slurmd/slurmctld/slurmdbd excluded from auto-restart" | tee /var/log/pcs-needrestart-guard.log + echo "needrestart: slurmd excluded from auto-restart" | tee /var/log/pcs-needrestart-guard.log # Post-install script (optional) - generic OnNodeConfigured-style hook. # Downloads and runs a user-supplied script (e.g. Enroot/Pyxis install). - | diff --git a/architectures/aws-pcs/assets/add-cng.yaml b/architectures/aws-pcs/assets/add-cng.yaml index dca00abb0..895a75bab 100644 --- a/architectures/aws-pcs/assets/add-cng.yaml +++ b/architectures/aws-pcs/assets/add-cng.yaml @@ -455,19 +455,21 @@ Resources: # apt-daily-upgrade updates base libraries (e.g. glibc). needrestart then # auto-restarts every service linked against them. slurmd links libc, so it # gets restarted mid-upgrade - which KILLS the jobs running under it (the - # step is torn down and the job requeues from scratch). Exclude the Slurm - # daemons from needrestart's automatic restart so a security upgrade can + # step is torn down and the job requeues from scratch). Exclude slurmd + # from needrestart's automatic restart so a security upgrade can # never take down a running job. Security packages still install; only the - # slurmd/slurmctld/slurmdbd auto-restart is suppressed. + # slurmd auto-restart is suppressed. (PCS runs the controller managed-side; + # slurmd is the only Slurm systemd service on login/compute nodes. qr(^slurmd) + # also covers the versioned units, e.g. slurmd-25.11.) # runcmd runs under /bin/sh (dash) — keep this POSIX-clean, no bashisms. - | mkdir -p /etc/needrestart/conf.d cat > /etc/needrestart/conf.d/90-pcs-slurm.conf <<'NRCONF' - # AWS PCS: never auto-restart Slurm daemons - restarting slurmd kills + # AWS PCS: never auto-restart slurmd - restarting it kills # the jobs running under it. Managed by add-cng.yaml UserData. - $nrconf{override_rc} = { qr(^slurmd) => 0, qr(^slurmctld) => 0, qr(^slurmdbd) => 0 }; + $nrconf{override_rc} = { qr(^slurmd) => 0 }; NRCONF - echo "needrestart: slurmd/slurmctld/slurmdbd excluded from auto-restart" | tee /var/log/pcs-needrestart-guard.log + echo "needrestart: slurmd excluded from auto-restart" | tee /var/log/pcs-needrestart-guard.log # Post-install script (optional) - generic OnNodeConfigured-style hook. # Downloads and runs a user-supplied script (e.g. Enroot/Pyxis install). # Accepts BOTH an s3:// URL (fetched with `aws s3 cp` using the instance diff --git a/architectures/aws-pcs/docs/OPERATIONS.md b/architectures/aws-pcs/docs/OPERATIONS.md index 494f7644f..fcf0daea8 100644 --- a/architectures/aws-pcs/docs/OPERATIONS.md +++ b/architectures/aws-pcs/docs/OPERATIONS.md @@ -480,18 +480,20 @@ touches a library `slurmd` uses. (It is not a Slurm upgrade — `slurmd` lives u because `slurmd` is restarted.) **You do not need to do anything — the templates already guard against this.** Every -compute-node-group UserData (`add-cng*`) writes a `needrestart` drop-in that excludes the -Slurm daemons from automatic restart: +compute-node-group UserData (`add-cng*`) writes a `needrestart` drop-in that excludes +`slurmd` from automatic restart: ```perl # /etc/needrestart/conf.d/90-pcs-slurm.conf (written by add-cng* UserData) -$nrconf{override_rc} = { qr(^slurmd) => 0, qr(^slurmctld) => 0, qr(^slurmdbd) => 0 }; +$nrconf{override_rc} = { qr(^slurmd) => 0 }; ``` -With this in place, security packages still install and `needrestart` still restarts -everything else; only the automatic restart of the Slurm daemons is suppressed -(`needrestart` reports it as *deferred*), so an unattended upgrade can no longer stop a -running job. `slurmd` picks up the new libraries the next time it restarts on its own +`slurmd` is the only Slurm systemd service on login/compute nodes (PCS runs the +controller managed-side, so there is no `slurmctld`/`slurmdbd` service to guard here), and +`qr(^slurmd)` also matches the versioned units (e.g. `slurmd-25.11`). With this in place, +security packages still install and `needrestart` still restarts everything else; only the +automatic restart of `slurmd` is suppressed (`needrestart` reports it as *deferred*), so an +unattended upgrade can no longer stop a running job. `slurmd` picks up the new libraries the next time it restarts on its own terms — node replacement, power-save cycle, or a manual restart — which is fine for HPC, where PCS replaces nodes regularly. (Verified end-to-end on real hardware: a long-running job survives a real `apt-get upgrade` of glibc followed by `needrestart -r a`, with the From 052e7164383411a5cfa1e4f28ab39d8bbb5db354 Mon Sep 17 00:00:00 2001 From: Daisuke Miyamoto Date: Mon, 6 Jul 2026 15:56:08 +0000 Subject: [PATCH 7/7] docs(aws-pcs): tighten OPERATIONS 6.2 MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit Condense the needrestart/slurmd note to three short paragraphs and drop the 'disable automatic upgrades entirely' alternative — unnecessary since the templates already guard slurmd. --- architectures/aws-pcs/docs/OPERATIONS.md | 54 ++++++------------------ 1 file changed, 14 insertions(+), 40 deletions(-) diff --git a/architectures/aws-pcs/docs/OPERATIONS.md b/architectures/aws-pcs/docs/OPERATIONS.md index fcf0daea8..705fbb47b 100644 --- a/architectures/aws-pcs/docs/OPERATIONS.md +++ b/architectures/aws-pcs/docs/OPERATIONS.md @@ -466,52 +466,26 @@ cleanest mitigation is upstream — recording it here so users seeing it know th workaround and the next contributor doesn't waste time looking for a bug in this PR's scripts. -### 6.2 `needrestart` must not auto-restart `slurmd` (would stop running jobs) — already handled - -**What to know.** On Ubuntu, automatic security upgrades (`apt-daily-upgrade` + -`unattended-upgrades`) update base libraries such as glibc, and `needrestart` then -restarts every service linked against them. `slurmd` links `libc`, so left unchecked -**`needrestart` restarts `slurmd`, and restarting `slurmd` tears down the `slurmstepd` -steps under it — stopping every job running on that node** (the job is killed and, if -requeued, restarts from scratch, losing in-progress work). This is a definite, -reproducible interaction, not a random failure: it happens whenever an unattended upgrade -touches a library `slurmd` uses. (It is not a Slurm upgrade — `slurmd` lives under -`/opt/aws/pcs/...` and is not apt-managed — and it is not a reboot; the job stops purely -because `slurmd` is restarted.) - -**You do not need to do anything — the templates already guard against this.** Every -compute-node-group UserData (`add-cng*`) writes a `needrestart` drop-in that excludes -`slurmd` from automatic restart: +### 6.2 `needrestart` restarting `slurmd` would stop running jobs — already handled + +When an unattended security upgrade updates a base library `slurmd` links (e.g. glibc), +`needrestart` restarts `slurmd`, and that restart tears down the `slurmstepd` steps under +it — **stopping every job on the node**. It is reproducible, not random, and not a reboot +or a Slurm-package upgrade. + +The compute-node-group templates already guard against this: each `add-cng*` UserData +writes a `needrestart` drop-in so `slurmd` is never auto-restarted (security updates still +install; `needrestart` only defers the `slurmd` restart): ```perl # /etc/needrestart/conf.d/90-pcs-slurm.conf (written by add-cng* UserData) $nrconf{override_rc} = { qr(^slurmd) => 0 }; ``` -`slurmd` is the only Slurm systemd service on login/compute nodes (PCS runs the -controller managed-side, so there is no `slurmctld`/`slurmdbd` service to guard here), and -`qr(^slurmd)` also matches the versioned units (e.g. `slurmd-25.11`). With this in place, -security packages still install and `needrestart` still restarts everything else; only the -automatic restart of `slurmd` is suppressed (`needrestart` reports it as *deferred*), so an -unattended upgrade can no longer stop a running job. `slurmd` picks up the new libraries the next time it restarts on its own -terms — node replacement, power-save cycle, or a manual restart — which is fine for HPC, -where PCS replaces nodes regularly. (Verified end-to-end on real hardware: a long-running -job survives a real `apt-get upgrade` of glibc followed by `needrestart -r a`, with the -`slurmd` PID unchanged.) - -**Safe if the platform later fixes this upstream.** The drop-in is a standalone -`conf.d/*.conf` file that only names the Slurm services. If a future PCS-ready DLAMI or -Slurm unit handles this differently (e.g. a `slurmd` unit that survives restart, or a -DLAMI-level `needrestart` policy), this file does not conflict or error — `needrestart` -reads `conf.d/*.conf` in order and tolerates keys it does not use, so at worst the -override becomes redundant. In other words, keeping it in place is harmless even after an -upstream fix; you never have to race to remove it. - -**If you would rather turn off automatic upgrades entirely** (heavier — this stops all -security updates), mask the timers in your own post-install instead: -`systemctl disable --now apt-daily.timer apt-daily-upgrade.timer`. The `needrestart` -drop-in above is preferred because it keeps security updates flowing while still protecting -jobs. +`slurmd` is the only Slurm systemd service on these nodes (the controller is managed by +PCS); `qr(^slurmd)` also matches the versioned units (e.g. `slurmd-25.11`). The drop-in is +a standalone `conf.d/*.conf` naming only `slurmd`, so if a later DLAMI or Slurm unit +handles this differently it neither conflicts nor errors — at worst it becomes redundant. ## 7. Recommendations recap