Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
Show all changes
145 commits
Select commit Hold shift + click to select a range
4586076
update README for PCS
DaisukeMiyamoto May 24, 2026
1e9a385
add reference cluster with PCS
DaisukeMiyamoto May 24, 2026
6d64291
Update PCS README: add deployment options, 1-click deploy, and fix S3…
DaisukeMiyamoto May 25, 2026
0814d0a
Add unified multi-NIC template for P5/P6 instances and update documen…
DaisukeMiyamoto May 25, 2026
838a4e9
Improve PCS deployment: rename to Pseries, update docs and examples
DaisukeMiyamoto May 25, 2026
42d6bad
Add placement group for on-demand P5/P6 instances in add-cng-p5
DaisukeMiyamoto May 25, 2026
51ba0ca
Add launch button and refine PCS documentation
DaisukeMiyamoto May 25, 2026
4369607
add launch button image
DaisukeMiyamoto May 25, 2026
01fd261
fix typos
DaisukeMiyamoto May 25, 2026
71d0176
fix typos
DaisukeMiyamoto May 25, 2026
e0ba463
Upgrade Slurm support to 25.05/25.11, default to 25.11
DaisukeMiyamoto May 25, 2026
a006cd5
Switch default AMI to PCS-specific DLAMI image
DaisukeMiyamoto May 25, 2026
64dda9c
Change mount path of FSxL from /shared to /fsx
DaisukeMiyamoto May 25, 2026
8ef8ddc
Simplify DLAMI build to use PCS-ready base image with Enroot/Pyxis only
DaisukeMiyamoto May 25, 2026
b2fd5c2
fix typos
DaisukeMiyamoto May 25, 2026
2c7c9f3
fix typos
DaisukeMiyamoto May 25, 2026
2ac6c78
update readme in pcs
DaisukeMiyamoto May 25, 2026
3b57ebf
add BaseAMI parameter for ImageBuilder
DaisukeMiyamoto May 26, 2026
5ec64bf
use awsome-distributed-ai bucket
DaisukeMiyamoto May 26, 2026
6b422d3
fix typos
DaisukeMiyamoto May 26, 2026
cdb6430
update architecture
DaisukeMiyamoto May 26, 2026
9540189
update architecture image
DaisukeMiyamoto May 26, 2026
a54bb33
Address PR review feedback: fix permissions, add documentation
DaisukeMiyamoto May 26, 2026
a84d993
Refactor cluster deployment to use modular add-cng templates
DaisukeMiyamoto May 26, 2026
97a9da1
Use variables for AZ ID in deployment examples
DaisukeMiyamoto May 26, 2026
f112965
Add AWS ParallelCluster Monitoring reference
DaisukeMiyamoto May 26, 2026
ba40b34
Add AWS Parallel Computing Service description
DaisukeMiyamoto May 26, 2026
05a338f
Add AWS Parallel Computing Service to top-level README
DaisukeMiyamoto May 26, 2026
29a7556
Add AWS ParallelCluster Monitoring integration to PCS templates
DaisukeMiyamoto May 27, 2026
4cb017a
Add Ubuntu 24.04 compatibility workaround for monitoring stack
DaisukeMiyamoto May 27, 2026
183f360
Fix RoleName output in cluster.yaml
DaisukeMiyamoto May 27, 2026
81d52fa
Add monitoring stack integration for AWS PCS
DaisukeMiyamoto May 28, 2026
fb3da65
Address PR #1109 review feedback (Round 2)
DaisukeMiyamoto May 28, 2026
fa17f30
Merge main: PR #1109 review feedback addressed
DaisukeMiyamoto May 28, 2026
38644a5
Fix monitoring installation bug with robust pattern matching
DaisukeMiyamoto May 28, 2026
05e1acc
Add architectures/aws-pcs directory for new structure
DaisukeMiyamoto May 28, 2026
e0d8511
Merge branch 'main' into deploy-monitoring
DaisukeMiyamoto May 28, 2026
5640eec
Add tests directory to architectures/aws-pcs
DaisukeMiyamoto May 28, 2026
cd586cf
Remove old 1.architectures/9.aws-pcs directory
DaisukeMiyamoto May 28, 2026
0e46758
Merge branch 'main' into deploy-monitoring
DaisukeMiyamoto May 28, 2026
840d578
Update README.md to reference new architectures/aws-pcs path
DaisukeMiyamoto May 28, 2026
2fad95c
Merge branch 'main' into deploy-monitoring
DaisukeMiyamoto May 28, 2026
d69beeb
Add ClusterName tag to launch templates
DaisukeMiyamoto May 28, 2026
b62b675
Apply monitoring integration to architectures/aws-pcs
DaisukeMiyamoto May 28, 2026
ac5e2d8
Add monitoring usage documentation to README
DaisukeMiyamoto May 28, 2026
1d19306
Fix MonitoringIAMPolicy ClusterName parameter and add missing template
DaisukeMiyamoto May 28, 2026
92a8d9a
Integrate IAM resources into cluster.yaml to eliminate external depen…
DaisukeMiyamoto May 29, 2026
918aa5c
Add Prerequisites stack reuse and UserData Enroot/Pyxis installation
DaisukeMiyamoto May 29, 2026
eec0aea
Refactor PCS deployment templates and update documentation
DaisukeMiyamoto May 29, 2026
605e106
Fix shell variable escaping in UserData for Enroot/Pyxis installation
DaisukeMiyamoto May 29, 2026
ce0bb6e
Externalize Enroot/Pyxis installation to separate script
DaisukeMiyamoto May 29, 2026
04b40cf
Fix Enroot/Pyxis installation script URL to use GitHub raw URL
DaisukeMiyamoto May 29, 2026
493ddc8
Generalize Enroot/Pyxis UserData into a post-install script hook
DaisukeMiyamoto May 31, 2026
3d6eef7
Add monitoring stack test cluster for Ubuntu 24.04 verification
DaisukeMiyamoto May 31, 2026
aa855df
Pin aws-parallelcluster-monitoring to release tag v2.6.2
DaisukeMiyamoto May 31, 2026
604cc38
Update README to pin MonitoringVersion to v2.6.2
DaisukeMiyamoto May 31, 2026
0cb78e8
Fix Enroot/Pyxis install on GPU nodes: gpg no-tty and stable nvidia r…
DaisukeMiyamoto May 31, 2026
61f9efa
Add P5 GPU training validation tests and 300 GiB root volume option
DaisukeMiyamoto May 31, 2026
de357d3
Bump monitoring to v2.6.3 and drop Ubuntu ec2-user workarounds
DaisukeMiyamoto May 31, 2026
287643a
Merge remote-tracking branch 'upstream/main' into deploy-monitoring
DaisukeMiyamoto May 31, 2026
e18f0e4
Tag nodes Name=HeadNode/Compute for monitoring dashboards
DaisukeMiyamoto May 31, 2026
23d490e
Make RootVolumeSize consistent across CNG templates and deploy-all
DaisukeMiyamoto May 31, 2026
d8f0ec7
Remove duplicate MinValue in add-cng.yaml RootVolumeSize
DaisukeMiyamoto May 31, 2026
e02baf5
PCS monitoring: free the Name tag, identify node type via monitoring-…
DaisukeMiyamoto May 31, 2026
f03e002
Add P6-B200 and P6-B300 GPU compute node group templates
DaisukeMiyamoto Jun 1, 2026
581a07e
deploy-all: select P-series NIC template by instance type
DaisukeMiyamoto Jun 1, 2026
8b9db0d
P6 templates: use EFA 'Use case 2' NIC layout (ENA primary + efa-only…
DaisukeMiyamoto Jun 1, 2026
7c8055d
P6 templates: use InterfaceType efa (not efa-only) for PCS compatibility
DaisukeMiyamoto Jun 2, 2026
9eaaae3
CNG templates: track latest launch template version
DaisukeMiyamoto Jun 2, 2026
efa6d49
add-cng-p6-b200: use EFA on all 8 network cards
DaisukeMiyamoto Jun 2, 2026
c1bcb64
P6 templates: set family-specific CngName defaults
DaisukeMiyamoto Jun 2, 2026
3c3a343
PCS monitoring: retry post-install on transient failure; docs to /opt…
DaisukeMiyamoto Jun 3, 2026
dff9bb9
PCS docs/params: human-friendly 1-click, FSx Region deployment types,…
DaisukeMiyamoto Jun 3, 2026
65c62ac
PCS GPU CNG: auto-derive EFA NIC count from instance type; clarify Ca…
DaisukeMiyamoto Jun 3, 2026
4fa78a2
deploy-all params: generic PostInstallScriptUrl override note; move M…
DaisukeMiyamoto Jun 3, 2026
1b2934c
README: lead with one-click ML-training-ready; align params to consol…
DaisukeMiyamoto Jun 3, 2026
0b612aa
README: add a quick NCCL multi-node GPU test (enroot import + Pyxis s…
DaisukeMiyamoto Jun 3, 2026
665c18f
Bump default MonitoringVersion to v2.6.5 (PRs #48/#49 merged upstream)
DaisukeMiyamoto Jun 3, 2026
004c33a
README: reference monitoring v2.6.5 (PCS /opt + DCGM Docker-29.x fixe…
DaisukeMiyamoto Jun 3, 2026
e80f216
docs: split full parameter reference into PARAMETERS.md; slim README …
DaisukeMiyamoto Jun 3, 2026
a43cded
README: run NCCL enroot import on the login node; drop PR references …
DaisukeMiyamoto Jun 3, 2026
f513568
README: move Templates after Cleanup; order Configuration as containe…
DaisukeMiyamoto Jun 3, 2026
128bd38
deploy-all: optional Grafana public access via GrafanaPublicAccessCid…
DaisukeMiyamoto Jun 3, 2026
13e9b56
README: document Grafana access as two options (SSM port-forward vs G…
DaisukeMiyamoto Jun 3, 2026
2f42631
README: link Quick Start to the Grafana public-access option (Grafana…
DaisukeMiyamoto Jun 3, 2026
3095b13
README: fold Cleanup into Quick Start (Console or CLI); merge User Ma…
DaisukeMiyamoto Jun 3, 2026
64609f8
docs: move architecture-components.md and PARAMETERS.md into docs/; f…
DaisukeMiyamoto Jun 3, 2026
3e19b9d
tests: drop obsolete monitoring test assets; refresh tests/README; bu…
DaisukeMiyamoto Jun 3, 2026
8032358
tests: consolidate into a single guide; flatten scripts; cover full m…
DaisukeMiyamoto Jun 3, 2026
b470050
README: add OpenZFS (/home) throughput row to the FSx storage table
DaisukeMiyamoto Jun 3, 2026
9dfa120
tests/README: add a Verified configurations section
DaisukeMiyamoto Jun 3, 2026
671d328
templates: fix inaccurate top-level Descriptions
DaisukeMiyamoto Jun 3, 2026
f511134
tests/README: add the validated 2x p6-b200 row (NCCL ~654 GB/s/8 nics…
DaisukeMiyamoto Jun 3, 2026
8e0bd27
install-enroot-pyxis.sh: wait for apt/dpkg lock + retry apt (fix cold…
DaisukeMiyamoto Jun 3, 2026
5b9fa44
tests/README: flag B300 NCCL bandwidth as needs-larger-message/more-n…
DaisukeMiyamoto Jun 3, 2026
f4da4d9
docs: add ROADMAP.md (future implementation TODO) + link from README
DaisukeMiyamoto Jun 3, 2026
16c8bd2
docs/ROADMAP: reset to the agreed open items (Multi-AZ prereq, manage…
DaisukeMiyamoto Jun 3, 2026
e784873
add-cng templates: make AmiId optional (auto-resolve PCS-ready DLAMI …
DaisukeMiyamoto Jun 3, 2026
27ec243
docs/ROADMAP: add user-management backend (LDAP/AD) integration item
DaisukeMiyamoto Jun 3, 2026
7d0ce47
docs/ROADMAP: add P6e-GB200/GB300 (Grace-Blackwell, arm64) support item
DaisukeMiyamoto Jun 3, 2026
6a80a7f
README: warn against combining BuildAMI=true with the default PostIns…
DaisukeMiyamoto Jun 3, 2026
664f194
tests/README: note BuildAMI=true must pair with PostInstallScriptUrl=…
DaisukeMiyamoto Jun 3, 2026
6589368
Merge branch 'awslabs:main' into deploy-monitoring
DaisukeMiyamoto Jun 3, 2026
aadd9f4
tests: record BuildAMI=true deploy-all validation (no double-install)
DaisukeMiyamoto Jun 3, 2026
d52c7ee
aws-pcs: fix PostInstallScriptUrl org (awslabs) and make enroot cache…
DaisukeMiyamoto Jun 4, 2026
bb32202
aws-pcs: validate FSx throughput against deployment type (review #3/#4)
DaisukeMiyamoto Jun 4, 2026
49e6322
deploy-all: force PostInstallScriptUrl empty when BuildAMI=true (revi…
DaisukeMiyamoto Jun 4, 2026
a9572d1
install-enroot-pyxis.sh: make idempotent + fix slurmd PATH (review #6…
DaisukeMiyamoto Jun 4, 2026
2522c82
deploy-all: keep PostInstallScriptUrl as a generic hook, warn (not bl…
DaisukeMiyamoto Jun 4, 2026
caff091
docs: GrafanaPublicAccessCidr exposes unauthenticated Prometheus/Push…
DaisukeMiyamoto Jun 4, 2026
398af36
docs: note p6-b300 GPU metrics gap (DCGM 4.2.0 pin); track upstream (…
DaisukeMiyamoto Jun 4, 2026
83d84df
GPU add-cng: lock InstanceType with AllowedValues + document the per-…
DaisukeMiyamoto Jun 4, 2026
91c0774
tests: reuse canonical NCCL/FSDP assets instead of shipping copies (r…
DaisukeMiyamoto Jun 4, 2026
288a439
install-enroot-pyxis.sh: document that Slurm 24.11 is out of scope (r…
DaisukeMiyamoto Jun 4, 2026
1bbb4d2
docs: AMI-pinning tip for production + align double-install notes wit…
DaisukeMiyamoto Jun 4, 2026
f3387f7
deploy-all + CNG: add DcgmExporterImage param to override dcgm-export…
DaisukeMiyamoto Jun 4, 2026
469a0b6
tests/README + README: reflect B300 validation findings (review #9 fo…
DaisukeMiyamoto Jun 4, 2026
b4dbcc7
install-enroot-pyxis.sh: pin slurmd PATH to the cluster's Slurm versi…
DaisukeMiyamoto Jun 4, 2026
13b3781
Default MonitoringVersion to v2.9.1 (DCGM_EXPORTER_IMAGE support for …
DaisukeMiyamoto Jun 4, 2026
ffaef37
install-enroot-pyxis.sh: derive slurmd PATH from the slurmd unit, not…
DaisukeMiyamoto Jun 4, 2026
0f96116
cluster.yaml: gate MetricsType on Slurm 25.11+ (fix 25.05 cluster cre…
DaisukeMiyamoto Jun 4, 2026
20593a8
install-enroot-pyxis.sh: resolve slurmd PATH from PCS profile.d/unit …
DaisukeMiyamoto Jun 4, 2026
87ebf94
Pass cluster Slurm version to post-install via PCS_SLURM_VERSION env …
DaisukeMiyamoto Jun 4, 2026
6e2d37c
install-enroot-pyxis.sh: build Pyxis only for the cluster's Slurm ver…
DaisukeMiyamoto Jun 4, 2026
b89b9af
tests/README: add pre-merge full test matrix + Enroot/Pyxis regressio…
DaisukeMiyamoto Jun 4, 2026
7bdd26e
AMI build: single-version Pyxis (add SlurmVersion param, plumb from d…
DaisukeMiyamoto Jun 4, 2026
aef17d3
docs: add OPERATIONS.md and slim README/tests caveats to one-line poi…
DaisukeMiyamoto Jun 4, 2026
1e401f9
docs/PARAMETERS.md: add DcgmExporterImage; cross-reference OPERATIONS…
DaisukeMiyamoto Jun 4, 2026
776103f
docs/OPERATIONS.md: explain why single-version Pyxis is fine across c…
DaisukeMiyamoto Jun 4, 2026
9dd8e1c
docs/OPERATIONS.md: record the cgroup-v2 prolog race as a known issue
DaisukeMiyamoto Jun 5, 2026
310fc60
docs/ROADMAP.md: add Software stack section (Spack, Intel oneAPI, NVI…
DaisukeMiyamoto Jun 5, 2026
c3448ac
README: surface SlurmVersion in §4 + document the AMI build's pipelin…
DaisukeMiyamoto Jun 5, 2026
c4017e2
README: tighten Key Features (capacity scope, drop one-click duplicat…
DaisukeMiyamoto Jun 5, 2026
5cbc003
README §7: reuse canonical micro-benchmarks/nccl-tests sbatch instead…
DaisukeMiyamoto Jun 5, 2026
1106db8
Default DcgmExporterImage to a DCGM 4.5.2 digest covering all support…
DaisukeMiyamoto Jun 5, 2026
0af2621
Strip personal-fork references before merge
DaisukeMiyamoto Jun 5, 2026
5251c79
deploy-all: refresh DcgmExporterImage console label to match new default
DaisukeMiyamoto Jun 5, 2026
7dc3412
Fix broken ldap_server reference paths
DaisukeMiyamoto Jun 5, 2026
5bc200a
deploy-all: improve 1-click parameter UX -- split AMI build, reorder,…
DaisukeMiyamoto Jun 5, 2026
2c295cc
deploy-all: drop docs/ refs from descriptions; move RootVolumeSize to…
DaisukeMiyamoto Jun 5, 2026
709b85a
FSx parameter labels: canonical service name + filesystem mount path
DaisukeMiyamoto Jun 5, 2026
88da755
Decouple AMI build from deploy-all; expose AmiId as a top-level clust…
DaisukeMiyamoto Jun 5, 2026
c068374
tests: split AMI-build path into its own Test 8 (independent flow)
DaisukeMiyamoto Jun 5, 2026
39d02a2
Consistency sweep: 'PCS-Ready DLAMI' naming + drop stale AMI-build ph…
DaisukeMiyamoto Jun 5, 2026
2ad37b7
Update architectures/aws-pcs/assets/pcs-ml-cluster-deploy-all.yaml
KeitaW Jun 6, 2026
File filter

Filter by extension

Filter by extension

Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
671 changes: 431 additions & 240 deletions architectures/aws-pcs/README.md

Large diffs are not rendered by default.

239 changes: 218 additions & 21 deletions architectures/aws-pcs/assets/add-cng-p5.yaml
Original file line number Diff line number Diff line change
@@ -1,5 +1,23 @@
# Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved.
# SPDX-License-Identifier: MIT-0
#
# Why a separate add-cng template per GPU family (p5 / p6-b200 / p6-b300)?
# The templates are ~85% identical; the real difference is the NetworkInterfaces
# block, whose EFA topology differs per family and cannot be parameterized in
# plain CloudFormation:
# - p5 / p5e / p5en: card 0 = EFA, EFA cards use DeviceIndex 1, 32 or 16 cards
# (this template already serves all three via the Use32Interfaces condition).
# - p6-b200: card 0 = EFA, DeviceIndex 1, 8 cards.
# - p6-b300: card 0 = ENA-only (no EFA), EFA on cards 1-16 with
# DeviceIndex 0, 17 cards.
# A single template would need ~32 per-card !If wrappers plus card-0/DeviceIndex
# branching, which obscures the list and means editing the already-validated
# p5/p6-b200 blocks. Fn::ForEach (AWS::LanguageExtensions) could generate the list
# from a mapping, but that transform requires CAPABILITY_AUTO_EXPAND, which breaks
# the README/workshop one-click quick-create links. So each family keeps its own
# flat, hand-checkable NIC list. Consolidating via Fn::ForEach is tracked in
# docs/ROADMAP.md. InstanceType is locked with AllowedValues so a type whose NIC
# layout doesn't match this template's block can't be selected.

AWSTemplateFormatVersion: '2010-09-09'
Description: Adding Compute Node Group to PCS Cluster with multi-network interface support for P5 series (On-Demand or Capacity Block)
Expand All @@ -17,7 +35,6 @@ Metadata:
- CngName
- QueueName
- InstanceType
- NetworkInterfaceCount
- MaxCount
- MinCount
- Label:
Expand All @@ -31,16 +48,37 @@ Metadata:
- FSxLustreFilesystemId
- FSxLustreFilesystemMountName
- FSxOpenZFSFilesystemId
- Label:
default: Monitoring configuration
Parameters:
- DeployMonitoring
- MonitoringVersion
- MonitoringRepo
- MonitoringRole
- Label:
default: Post-install script (OnNodeConfigured)
Parameters:
- PostInstallScriptUrl
- PostInstallScriptArgs

Parameters:
CapacityReservationId:
Type: String
Description: (Optional) Capacity Reservation ID for Capacity Blocks. Leave empty for On-Demand instances.
Description: >-
(Optional) Capacity Block for ML reservation ID. When SET, the node group
launches with MarketType=capacity-block against this reservation. Leave
EMPTY for On-Demand. NOTE: this is for Capacity Blocks ONLY — do not put an
On-Demand Capacity Reservation (ODCR) ID here. An ODCR with "open" matching
is consumed automatically by On-Demand launches, so for ODCR leave this empty.
Default: ""

ClusterId:
Type: String
Description: Cluster ID
Description: Cluster ID (actual PCS cluster ID like pcs_xxxxx)

ClusterName:
Type: String
Description: User-friendly cluster name

CngName:
Type: String
Expand All @@ -54,16 +92,23 @@ Parameters:

InstanceType:
Type: String
Description: Instance Type for Compute Node Group (e.g., p5.48xlarge, p5e.48xlarge, p5en.48xlarge)
Description: >-
P5-family instance type. The EFA interface count is derived automatically:
p5en.48xlarge = 16 network cards, p5.48xlarge / p5e.48xlarge = 32.
Default: "p5.48xlarge"

NetworkInterfaceCount:
Type: String
Description: Number of network interfaces for P5 instances (16 for p5en.48xlarge, 32 for p5.48xlarge and p5e.48xlarge)
Default: "32"
AllowedValues:
- "16"
- "32"
- p5.48xlarge
- p5e.48xlarge
- p5en.48xlarge

RootVolumeSize:
Type: Number
Description: >-
Root EBS volume size (GiB) for compute nodes. The PCS-Ready DLAMI default (~75
GiB) is too small for pulling/importing large container images (Enroot/Pyxis
squashfs, Megatron ~20 GB). 300 GiB gives ample headroom.
Default: 300
MinValue: 75

MaxCount:
Type: Number
Expand All @@ -78,9 +123,14 @@ Parameters:
Description: Select the Subnet where the Compute Node will be launched

AmiId:
Type: AWS::EC2::Image::Id
Description: Enter the AMI ID for the Login Node and Compute Node (e.g. ami-0abcdef1234567890)
ConstraintDescription: Must be a valid AMI ID (ami-xxxxxxxxxxxxxxxxx)
Type: String
Description: >-
(Optional) AMI ID for the nodes. Leave empty to auto-resolve the latest
PCS-Ready DLAMI (Ubuntu 24.04 x86_64) from SSM Parameter Store. Set an explicit
ami-xxxx only to override (e.g. a custom AMI built with Enroot/Pyxis pre-baked).
Default: ""
AllowedPattern: '^$|^ami-[0-9a-f]{8,17}$'
ConstraintDescription: Must be empty or a valid AMI ID (ami-xxxxxxxxxxxxxxxxx)

ClusterSecurityGroupId:
Type: AWS::EC2::SecurityGroup::Id
Expand All @@ -102,14 +152,91 @@ Parameters:
Type: String
Description: Filesystem ID for FSx for OpenZFS

DeployMonitoring:
Type: String
Description: Deploy monitoring stack (Prometheus/Grafana on Login Node, DCGM exporter on Compute Node with GPU)
Default: 'false'
AllowedValues:
- 'true'
- 'false'

MonitoringVersion:
Type: String
Description: >-
aws-parallelcluster-monitoring git ref to install (release tag, branch, or
'latest'). Used both to fetch post-install.sh from MonitoringRepo at this
ref and as the ref argument passed to it. Pin to a release tag for upstream
stability in production; a branch name (e.g. a fork's dev branch) can be
used together with MonitoringRepo for testing unreleased changes.
Default: 'v2.9.1'

MonitoringRepo:
Type: String
Description: >-
GitHub "owner/repo" to fetch the monitoring stack from. Override with a fork
together with a branch in MonitoringVersion to test unreleased changes before
they merge upstream.
Default: 'aws-samples/aws-parallelcluster-monitoring'

DcgmExporterImage:
Type: String
Description: >-
dcgm-exporter container image used by the monitoring stack on GPU nodes. The
default is a DCGM 4.5.2 build pinned by DIGEST -- this covers Hopper / B200 /
B300 alike (the digest pull bypasses the Docker 29.x OCI-index failure that
breaks newer NVCR tags). To pin to the older monitoring-default 4.2.0 (covers
Hopper + B200 only) or any other build, set this to that image reference
(preferably also a digest). No effect on CPU nodes.
Default: 'nvcr.io/nvidia/k8s/dcgm-exporter@sha256:a7ad6547d4546eaf4dd5d6b4c0b4db4101e63ef7dc3cdff7f42b767d2c60b706'

MonitoringRole:
Type: String
Description: Monitoring role for this node (login=head node with Prometheus/Grafana, compute=compute node with exporters, none=no monitoring)
Default: 'none'
AllowedValues:
- 'login'
- 'compute'
- 'none'

PostInstallScriptUrl:
Type: String
Description: >-
Optional URL of a post-install script to download and run on each node at
first boot (PCS equivalent of ParallelCluster OnNodeConfigured). Leave
empty to skip. To install Enroot/Pyxis, point this at scripts/install-enroot-pyxis.sh.
Must be an HTTP(S) URL (e.g. a GitHub raw URL); S3 public hosting is not allowed.
Default: ''

PostInstallScriptArgs:
Type: String
Description: Space-separated arguments passed to the post-install script.
Default: ''

SlurmVersion:
Type: String
Description: >-
The cluster's Slurm version, exported to the post-install script as
PCS_SLURM_VERSION so it can put the matching Slurm CLI bin on slurmd's PATH
(the Pyxis PMI hook needs scontrol). Must match the version the cluster was
created with.
Default: '25.11'
AllowedValues:
- '25.05'
- '25.11'

Conditions:
Use32Interfaces:
Fn::Equals:
- Ref: NetworkInterfaceCount
- "32"
# EFA interface count is derived from the instance type: p5en.48xlarge has 16
# network cards, p5.48xlarge / p5e.48xlarge have 32. So "use 32 interfaces"
# means "not p5en". No separate NetworkInterfaceCount parameter is needed.
Use32Interfaces: !Not [!Equals [!Ref InstanceType, "p5en.48xlarge"]]
UseCapacityBlock: !Not [!Equals [!Ref CapacityReservationId, ""]]
UseOnDemand: !Equals [!Ref CapacityReservationId, ""]
CreateQueue: !Not [!Equals [!Ref QueueName, ""]]
# Emit the monitoring-role tag for both login and compute nodes (any role
# other than 'none'). The monitoring stack identifies login nodes by this
# tag (monitoring-role=login) and excludes them from EC2 service discovery.
IsMonitoringRoleSet: !Not [!Equals [!Ref MonitoringRole, 'none']]
HasAmiId: !Not [!Equals [!Ref AmiId, ""]]

Resources:

Expand All @@ -132,10 +259,16 @@ Resources:
IamInstanceProfileArn: !Ref IamProfileArn
CustomLaunchTemplate:
TemplateId: !Ref PCSLaunchTemplate
Version: 1
# Track the latest launch template version so template updates (e.g. NIC
# or UserData changes) actually reach the node group. Hardcoding Version: 1
# pins the node group to the original version and silently ignores updates.
Version: !GetAtt PCSLaunchTemplate.LatestVersionNumber
SubnetIds:
- !Ref SubnetId
AmiId: !Ref AmiId
AmiId: !If
- HasAmiId
- !Ref AmiId
- '{{resolve:ssm:/aws/service/pcs/ami/dlami-base-ubuntu2404/x86_64/latest/ami-id}}'
InstanceConfigs:
- InstanceType: !Ref InstanceType
PurchaseOption: !If [UseCapacityBlock, CAPACITY_BLOCK, !Ref "AWS::NoValue"]
Expand All @@ -156,6 +289,12 @@ Resources:

LaunchTemplateData:
InstanceType: !Ref InstanceType
BlockDeviceMappings:
- DeviceName: /dev/sda1
Ebs:
VolumeSize: !Ref RootVolumeSize
VolumeType: gp3
DeleteOnTermination: true
InstanceMarketOptions: !If
- UseCapacityBlock
- MarketType: capacity-block
Expand All @@ -174,12 +313,34 @@ Resources:
Tags:
- Key: HPCRecipes
Value: "true"
# Name is free for arbitrary operator use. The monitoring stack
# identifies login vs compute nodes via the monitoring-role tag
# (below), not Name, so dashboards work regardless of this value.
# Default to PCS-<cngname> (mirrors CngName) for a recognizable
# console name; retag freely afterward.
- Key: Name
Value: !Sub 'PCS-${CngName}'
- Key: CngName
Value: !Sub 'PCS-${CngName}'
- Key: ClusterName
Value: !Ref ClusterName
- Key: pcs-cluster-id
Value: !Ref ClusterId
# Node type for the monitoring stack. Emitted for both login and
# compute (any role != none). Prometheus drops monitoring-role=login
# from EC2 SD (the static login_node job scrapes it); everything else
# is reported as instance_name=Compute. See aws-parallelcluster-monitoring
# prometheus-pcs.yml and issue #45.
- !If
- IsMonitoringRoleSet
- Key: monitoring-role
Value: !Ref MonitoringRole
- !Ref AWS::NoValue
MetadataOptions:
HttpEndpoint: enabled
HttpPutResponseHopLimit: 4
HttpTokens: required
InstanceMetadataTags: enabled
UserData:
Fn::Base64: !Sub |
MIME-Version: 1.0
Expand All @@ -190,15 +351,51 @@ Resources:
MIME-Version: 1.0

runcmd:
# Post-install script (optional) - generic OnNodeConfigured-style hook.
# Downloads and runs a user-supplied script (e.g. Enroot/Pyxis install).
- |
if [ -n "${PostInstallScriptUrl}" ]; then
echo "Downloading post-install script from ${PostInstallScriptUrl}..." | tee /var/log/pcs-post-install.log
curl -fsSL "${PostInstallScriptUrl}" -o /tmp/pcs-post-install.sh
chmod +x /tmp/pcs-post-install.sh
echo "Executing post-install script..." | tee -a /var/log/pcs-post-install.log
# Tell the script this cluster's Slurm version (it can't discover it at
# first boot, before slurmd/profile.d/controller exist). See script header.
export PCS_SLURM_VERSION="${SlurmVersion}"
bash /tmp/pcs-post-install.sh ${PostInstallScriptArgs} >> /var/log/pcs-post-install.log 2>&1
echo "Post-install script completed (exit $?)" | tee -a /var/log/pcs-post-install.log
else
echo "No post-install script configured (PostInstallScriptUrl empty)" | tee /var/log/pcs-post-install.log
fi
# FSx Home directories mount
- mkdir -p /tmp/home
- rsync -aA /home/ /tmp/home
- echo "${FSxOpenZFSFilesystemId}.fsx.${AWS::Region}.amazonaws.com:/fsx/ /home nfs noatime,nfsvers=3,sync,nconnect=16,rsize=1048576,wsize=1048576,defaults 0 0" >> /etc/fstab
- mount -a -t nfs defaults
- if [ "enabled" == "$(sestatus | awk '/^SELinux status:/{print $3}')" ]; then setsebool -P use_nfs_home_dirs 1; fi
- if [ "enabled" = "$(sestatus | awk '/^SELinux status:/{print $3}')" ]; then setsebool -P use_nfs_home_dirs 1; fi
- rsync -aA --ignore-existing /tmp/home/ /home
- rm -rf /tmp/home/
# If provided, mount FSxL filesystem as /fsx
- if [ ! -z "${FSxLustreFilesystemId}" ]; then mkdir -p /fsx; mount -t lustre ${FSxLustreFilesystemId}.fsx.${AWS::Region}.amazonaws.com@tcp:/${FSxLustreFilesystemMountName} /fsx; chmod 1777 /fsx; fi
# Monitoring stack installation (optional, enabled on Login Node only)
- |
if [ "${DeployMonitoring}" = "true" ]; then
echo "Installing monitoring stack..."
# Fetch post-install.sh from MonitoringRepo at the MonitoringVersion ref, and
# pass both the ref and the repo to it so the tarball is pulled from the same
# source. This lets a fork + branch be used for testing unreleased changes.
# As of v2.6.3 the installer auto-detects the Ubuntu 'ubuntu' user and runs
# install.sh itself (PR #44), so no ec2-user shim or local-var patch is needed.
# Optional dcgm-exporter image override (e.g. a B300-capable DCGM
# >= 4.4.0 build by digest). The monitoring installer reads
# DCGM_EXPORTER_IMAGE from its environment and falls back to its
# default pin when empty, so leaving DcgmExporterImage empty keeps
# the stock behaviour. See aws-parallelcluster-monitoring #50.
export DCGM_EXPORTER_IMAGE="${DcgmExporterImage}"
curl -fsSL https://raw.githubusercontent.com/${MonitoringRepo}/${MonitoringVersion}/post-install.sh -o /tmp/post-install.sh
bash /tmp/post-install.sh ${MonitoringVersion} ${MonitoringRepo} 2>&1 | tee /var/log/monitoring-install.log
echo "Monitoring installation complete"
fi

--==MYBOUNDARY==
NetworkInterfaces:
Expand Down
Loading