diff --git a/architectures/aws-pcs/README.md b/architectures/aws-pcs/README.md index 52eda6d10..6c1700c59 100644 --- a/architectures/aws-pcs/README.md +++ b/architectures/aws-pcs/README.md @@ -8,32 +8,37 @@ This repository provides reference architectures and deployment templates for se - **One click to an ML-training-ready cluster**: a single CloudFormation stack gives you a complete, ready-to-train environment — Slurm scheduler, GPU compute with EFA, shared FSx storage, the Enroot/Pyxis container runtime, and monitoring — with only the Availability Zone to choose. Submit distributed training jobs minutes after launch. - **Container runtime included**: Enroot/Pyxis is set up automatically, so `srun --container-image=...` works out of the box for containerized training. -- **Monitoring built in**: Grafana + Prometheus run on the login node, with DCGM Exporter on GPU compute nodes feeding the pre-built GPU dashboards (`DeployMonitoring=true`, on by default). Reach Grafana privately via SSM port-forward, or open it to a trusted CIDR with `GrafanaPublicAccessCidr`. +- **Monitoring built in**: Grafana + Prometheus on the login node, with DCGM Exporter on GPU nodes feeding pre-built GPU dashboards (on by default). Reach Grafana privately via SSM port-forward, or open it to a trusted CIDR. See [§8.2 Monitoring](#82-monitoring). - **GPU-ready, multi-NIC EFA**: dedicated launch templates for the P5 and P6 families, selected automatically by instance type, for high-bandwidth multi-node training. -- **Broad capacity-purchase support**: covers the full range of EC2 capacity options out of the box — On-Demand, On-Demand Capacity Reservations (ODCR), and Capacity Blocks for ML — selected per node group. +- **Flexible capacity options**: On-Demand, "open" On-Demand Capacity Reservations (consumed automatically), and Capacity Blocks for ML — selected per node group. (Targeting a *specific* ODCR is on the [roadmap](./docs/ROADMAP.md).) - **High-performance storage**: FSx for Lustre (shared scratch, `/fsx`) and FSx for OpenZFS (home directories, `/home`). +- **Multi-user ready**: opt-in OpenLDAP directory on the login node with SSSD on every compute node, so a team shares one cluster with consistent users — pairs with Slurm accounting. See [§8.3 User Management](#83-user-management). +- **Access control built in**: ready-to-deploy least-privilege IAM policy stacks for cluster admins and users, and login-node SSH / Grafana access gated to a trusted CIDR. See [§8.4 IAM Permissions](#84-iam-permissions). - **Modular components**: compose individual stacks (network/storage prerequisites, cluster scheduler, per-family compute node groups) instead of the all-in-one nested stack when you want to reuse infrastructure across clusters or iterate on one piece at a time. -> Built on the AWS-managed **PCS-Ready DLAMI** (NVIDIA driver, CUDA, PCS agent, and -> Slurm 25.05/25.11 pre-installed), so no custom AMI build is required by default — -> the cluster comes up without an Image Builder step. For frequent scaling, you can -> pre-bake Enroot/Pyxis into a custom DLAMI with the standalone -> [`pcs-ready-dlami-with-enroot-pyxis.yaml`](#9-pre-baking-enrootpyxis-into-a-custom-ami-optional) -> template and pass the result as `AmiId`. - ## 2. Architecture ![AWS PCS diagram](./images/ml-pcs-architecture.png) A default deployment (`pcs-ml-cluster-deploy-all.yaml`) provisions: -- VPC with public/private subnets, NAT gateway, and S3 endpoint +- VPC with a public subnet and private subnets in up to 3 AZs, a NAT gateway (primary AZ), and an S3 endpoint - FSx for Lustre (`/fsx`, high-performance shared scratch) and FSx for OpenZFS (`/home`) - PCS cluster with the Slurm scheduler (25.05 or 25.11), on the PCS-Ready DLAMI -- Login node group (public subnet) with the monitoring stack (Prometheus + Grafana + Nginx) +- Login node group (public subnet) with the monitoring stack (Prometheus + Grafana + Nginx); SSH/Grafana can be opened to a trusted CIDR - CPU compute node group (private subnet); EFA can be enabled for HPC/MPI workloads - Optional GPU (P5/P6) node group with multi-NIC EFA, plus DCGM Exporter for the GPU dashboards - Enroot/Pyxis container runtime installed at first boot via `PostInstallScriptUrl` (or pre-baked into a custom AMI you build separately and pass as `AmiId`) +Optional add-ons (off by default): +- **Multi-user directory**: OpenLDAP on the login node + SSSD on every compute node (`DirectoryService`) +- **IAM policy stacks**: least-privilege cluster-admin / cluster-user policies you can deploy separately + +Every node runs on the AWS-managed **PCS-Ready DLAMI** (NVIDIA driver, CUDA, PCS agent, +and Slurm 25.05/25.11 pre-installed), so **no custom AMI build is required** — the cluster +comes up without an Image Builder step, and Enroot/Pyxis is layered on at first boot. +See [AMI and container runtime](#ami-and-container-runtime) for pinning the AMI and the +optional pre-bake path for faster scaling. + --- ## 3. Quick Start @@ -62,8 +67,8 @@ Add a GPU queue and tune storage/monitoring via the parameters below. Once it's up: - **Connect** to the login node via SSM Session Manager — see [Accessing the Cluster](#6-accessing-the-cluster). -- **Open the Grafana dashboards** (deployed by default) via SSM port forwarding — see [Accessing Grafana](#accessing-grafana). -- **Want to reach Grafana directly in a browser** (no port forwarding)? Set `GrafanaPublicAccessCidr` to a trusted CIDR at deploy time — see [Option B — Direct public access](#option-b--direct-public-access-opt-in-via-grafanapublicaccesscidr). +- **Open the Grafana dashboards** (deployed by default) via SSM port forwarding — see [Accessing the Grafana dashboards](#accessing-the-grafana-dashboards). +- **Want to reach Grafana directly in a browser** (no port forwarding)? Set `GrafanaAccessCidr` to a trusted CIDR at deploy time — see [Option B — Direct public access](#option-b--direct-public-access-opt-in-via-grafanaaccesscidr). Prefer step-by-step instructions? See the [AI/ML for AWS PCS Workshop](https://catalog.workshops.aws/ml-on-pcs/). @@ -81,22 +86,54 @@ are deleted with the stack. ## 4. Configuration -Defaults give the most common production setup — Enroot/Pyxis installed at first -boot via `PostInstallScriptUrl` + `DeployMonitoring=true` — so a default deploy -only needs the Availability Zone (`PrimarySubnetAZ`). The most-used parameters: +The defaults give a working ML-training cluster, so the only required parameter is +the Availability Zone (`PrimarySubnetAZ`); everything else is optional. The most-used +parameters are grouped below to match the deploy-all console's parameter groups. Storage +parameters are covered in [§8.1 Storage](#81-storage-fsx-deployment-types--sizing); for the +complete reference see [PARAMETERS.md](./docs/PARAMETERS.md). + +**1. Network Configuration** | Parameter | Default | Purpose | |---|---|---| | `PrimarySubnetAZ` | *(required)* | Availability Zone to deploy into — the one required parameter | -| `SlurmVersion` | `25.11` | Slurm version (`25.05` or `25.11`); 25.11 is needed for the Slurm OpenMetrics dashboards. Drives Pyxis build version too. See [OPERATIONS.md §1](./docs/OPERATIONS.md#1-slurm-version-selection) | -| `AmiId` | *(empty → SSM auto)* | Empty auto-resolves to the latest **PCS-Ready Deep Learning AMI** (Ubuntu 24.04) from SSM. Pin to a specific `ami-xxx` for production, or pass an AMI built by [`pcs-ready-dlami-with-enroot-pyxis.yaml`](#9-pre-baking-enrootpyxis-into-a-custom-ami-optional) | -| `DeployMonitoring` | `true` | Deploy Prometheus + Grafana on the login node, plus DCGM Exporter on GPU compute nodes | -| `DeployOnDemandCNG` | `true` | Deploy the `cpu1` CPU queue (`OnDemandInstanceType`, default `c6i.4xlarge`) | -| `OnDemandEnableEfa` | `false` | Set `true` for HPC/MPI workloads on EFA-capable CPU types (hpc6a/hpc7a/hpc6id/hpc8a, c7i.metal, etc.). Auto-creates a cluster placement group. See [EFA on CPU HPC instances](#efa-on-cpu-hpc-instances-ondemandenableefa) | -| `DeployPseriesCNG` | `false` | Deploy a GPU (P5/P6) queue — see [GPU compute](#gpu-compute-p5p6) | -| `PseriesInstanceType` | `p5.48xlarge` | GPU instance type; auto-selects the multi-NIC template + EFA count | +| `AdditionalSubnetAZ2` / `…AZ3` | *(empty)* | Add private subnets in up to 2 more AZs (primary + 2). NAT gateway stays in the primary AZ | + +**2. PCS Cluster Configuration** + +| Parameter | Default | Purpose | +|---|---|---| +| `SlurmVersion` | `25.11` | `25.05` or `25.11`; 25.11 is needed for the Slurm OpenMetrics dashboards. See [OPERATIONS.md §1](./docs/OPERATIONS.md#1-slurm-version-selection) | +| `AmiId` | *(empty → SSM auto)* | Empty auto-resolves the latest PCS-Ready DLAMI. See [AMI and container runtime](#ami-and-container-runtime) | +| `SSHAccessCidr` | *(empty)* | Open SSH/22 on the login node to a trusted CIDR (default: SSM only). See [§6](#6-accessing-the-cluster) | + +**3. On-Demand Compute Node Group** + +| Parameter | Default | Purpose | +|---|---|---| +| `DeployOnDemandCNG` | `true` | Deploy the On-Demand compute queue (the `cpu1` queue by default) | +| `OnDemandInstanceType` | `c6i.4xlarge` | Any On-Demand type — CPU, a single-NIC GPU (`g6.12xlarge`, [Example 1](#example-1-single-nic-gpu-queue-g6)), or an HPC type. Multi-NIC P5/P6 use the GPU queue instead | +| `OnDemandEfaInterfaceCount` | `0` | `1`/`2` enables EFA for HPC/MPI on EFA-capable CPU types. See [§8.6](#86-cpu-compute-node-group--advanced-settings) | + +**4. GPU Compute Node Group - P5/P6 Series (Optional)** + +| Parameter | Default | Purpose | +|---|---|---| +| `DeployPseriesCNG` | `false` | Deploy a multi-NIC GPU (P5/P6) queue | +| `PseriesInstanceType` | `p5.48xlarge` | Picks the matching template + EFA NIC count automatically. See [GPU compute](#gpu-compute-p5p6) for the accepted types | | `CapacityReservationId` | *(empty)* | Capacity **Block** ID for the GPU queue; empty for On-Demand/ODCR | +**5. Additional Cluster Configuration (Monitoring, Multi-User, Container Runtime)** + +| Parameter | Default | Purpose | +|---|---|---| +| `MonitoringStack` | `Prometheus-LoginNode` | Prometheus + Grafana on the login node, DCGM Exporter on GPU nodes. `none` disables it. See [§8.2](#82-monitoring) | +| `GrafanaAccessCidr` | *(empty)* | Open HTTPS/443 (Grafana) on the login node to a trusted CIDR (default: SSM port-forward only) | +| `DirectoryService` | `none` | `OpenLDAP-LoginNode` for a multi-user cluster. See [§8.3](#83-user-management) | +| `PostInstallScriptUrl` | *(empty → auto)* | First-boot script; empty auto-installs Enroot/Pyxis from your templates bucket. Rarely changed. See [PARAMETERS.md](./docs/PARAMETERS.md) | + +(Group 5 also has `MonitoringRepo` / `MonitoringVersion` / `DcgmExporterImage` — pinned defaults, rarely changed; see [PARAMETERS.md](./docs/PARAMETERS.md).) + **See [PARAMETERS.md](./docs/PARAMETERS.md) for the complete parameter reference** (all 7 console parameter groups, with every default). The concept guides below cover the choices that need the most thought. @@ -107,9 +144,9 @@ choices that need the most thought. **PCS-Ready DLAMI** (Ubuntu 24.04 x86_64) from SSM (`/aws/service/pcs/ami/dlami-base-ubuntu2404/x86_64/latest/ami-id`) — no AMI choice needed. Enroot 3.5.0 + Pyxis 0.20.0 are layered on at first boot via -[`scripts/install-enroot-pyxis.sh`](./scripts/install-enroot-pyxis.sh) +[`assets/scripts/install-enroot-pyxis.sh`](./assets/scripts/install-enroot-pyxis.sh) (~8–12 min boot). For **frequent scaling**, pre-bake Enroot/Pyxis into a custom DLAMI -once with [§9](#9-pre-baking-enrootpyxis-into-a-custom-ami-optional) and pass that +once with [§8.5](#85-pre-baking-enrootpyxis-into-a-custom-ami) and pass that `ami-xxx` as `AmiId` (~3 min boot, deterministic state). The post-install hook is idempotent — it no-ops on a pre-baked AMI. @@ -143,53 +180,8 @@ automatically. > group at `PseriesMinCount = PseriesMaxCount = ` so the reserved > nodes launch immediately, rather than scaling from 0. -### EFA on CPU HPC instances (`OnDemandEnableEfa`) - -For tightly-coupled HPC / MPI workloads on CPU-only HPC instances, set -`OnDemandEnableEfa=true` (deploy-all) or `EnableEfa=true` (modular `add-cng.yaml`). -The CPU compute node group then launches with EFA `NetworkInterfaces` and a cluster -placement group is auto-created. GPU CNGs (P5/P6) are unaffected — they use their -own dedicated multi-NIC EFA wiring per-family. - -| Instance type | EFA interfaces | Aggregate spec | Set `OnDemandEfaInterfaceCount` to | -|---|---:|---:|---:| -| `hpc6a.48xlarge` | 1 | 100 Gbps | **1** | -| `hpc7a.96xlarge` (and 12/24/48) | 2 | 300 Gbps | **2** | -| `hpc6id.32xlarge` | 2 | 200 Gbps | **2** | -| `hpc8a.96xlarge` | 2 | 300 Gbps | **2** | -| `c7i.metal-48xl` etc. | 1 | varies | **1** | - -(Mismatching `OnDemandEfaInterfaceCount` with the instance type's actual -`MaximumEfaInterfaces` fails at launch.) - -**Placement group:** auto-created per-CNG by the template. Override with -`OnDemandPlacementGroupName=` to share one PG across multiple -CNGs (e.g. heterogeneous tightly-coupled jobs that span CPU + GPU). - -**Multi-NIC bandwidth needs multiple MPI pairs.** A single MPI pair uses one -libfabric endpoint and only one NIC. Use `osu_mbw_mr -np 32 -N 16` (or your -application's natural multi-pair pattern) to actually exercise both NICs on -hpc7a/hpc8a. See [tests/README.md Test 9](./tests/README.md#test-9-efa-on-cpu-hpc-instances-hpc6a--hpc7a--hpc8a) -for the full benchmark setup and validated bandwidth numbers. - -### Storage: FSx deployment types (Region availability) - -**FSx deployment types are not available in every Region.** Defaults match the most -capable type; switch to a more widely available one if your Region needs it. - -| Filesystem | Parameter | Default | Other values | Notes | -|---|---|---|---|---| -| Lustre (`/fsx`) | `LustreDeploymentType` | `PERSISTENT_2` | `PERSISTENT_1` | `PERSISTENT_2` (throughput 125/250/500/1000, metadata config) isn't in every Region; `PERSISTENT_1` (50/100/200) is in more Regions | -| Lustre (`/fsx`) | `PerUnitStorageThroughput` | `250` | any valid number | Must be valid for the type: P2 = 125/250/500/1000, P1 = 50/100/200 | -| Lustre (`/fsx`) | `FSxLustreEnableEfa` | `false` | `true` | Enable EFA on the Lustre filesystem. **The headline feature is GPUDirect Storage (GDS) for P5 / P5e / P5en / P6-B200 GPU CNGs**, which DMAs file data straight into GPU memory (requires the NVIDIA `nvidia-fs` / cuFile stack on the client — see [`docs/ROADMAP.md`](./docs/ROADMAP.md#client-side-lustre-on-efa--gds-support)). EFA-capable CPU CNGs (`OnDemandEnableEfa=true`) get the EFA *transport* path to storage as a secondary benefit, useful when a single client is pushing past ~10 GBps. **PERSISTENT_2 SSD only** (a CFN Rule fails the stack at create time on PERSISTENT_1). **Requires a much higher minimum `Capacity`**: at `PerUnitStorageThroughput=250` the minimum is **19200 GiB** (16× the 1200 GiB default). Set `Capacity` accordingly when enabling this | -| OpenZFS (`/home`) | `OpenZFSDeploymentType` | `SINGLE_AZ_HA_2` | `SINGLE_AZ_HA_1`, `SINGLE_AZ_2`, `SINGLE_AZ_1` | `SINGLE_AZ_1` is in all Regions; HA/2 variants vary. `MULTI_AZ` excluded (needs a second subnet) | -| OpenZFS (`/home`) | `HomeThroughput` | `320` | any valid number | Throughput (MB/s). Valid values depend on the deployment type: `SINGLE_AZ_2`/`SINGLE_AZ_HA_2` = 160/320/640/1280/2560/3840/5120/7680/10240; `SINGLE_AZ_HA_1` = 128/256/512/1024/2048/3072/4096; `SINGLE_AZ_1` = 64/128/256/512/1024/2048/3072/4096 | - -Check support before deploying: -[Lustre Regions](https://docs.aws.amazon.com/fsx/latest/LustreGuide/using-fsx-lustre.html) · -[OpenZFS Regions](https://docs.aws.amazon.com/fsx/latest/OpenZFSGuide/available-aws-regions.html). -If a deploy fails at the FSx resource with an "unsupported deployment type" error, -switch these parameters to a type your Region supports. +Storage (FSx for Lustre `/fsx` + OpenZFS `/home`) has sensible defaults; deployment +types, throughput, and capacity are covered in [§8.1 Storage](#81-storage-fsx-deployment-types--sizing). --- @@ -257,7 +249,6 @@ aws cloudformation create-stack \ ParameterKey=OnDemandQueueName,ParameterValue=hpc \ ParameterKey=OnDemandInstanceType,ParameterValue=hpc7a.96xlarge \ ParameterKey=OnDemandMaxCount,ParameterValue=4 \ - ParameterKey=OnDemandEnableEfa,ParameterValue=true \ ParameterKey=OnDemandEfaInterfaceCount,ParameterValue=2 \ --capabilities CAPABILITY_IAM CAPABILITY_NAMED_IAM ``` @@ -265,14 +256,16 @@ Replaces the default `cpu1` queue with an `hpc` queue of EFA-enabled hpc7a.96xlarge nodes. The CNG launches in an auto-created cluster placement group. For other HPC types, set `OnDemandInstanceType` and the matching `OnDemandEfaInterfaceCount` from the table in -[EFA on CPU HPC instances](#efa-on-cpu-hpc-instances-ondemandenableefa) +[EFA on CPU HPC instances](#efa-on-cpu-hpc-instances-ondemandefainterfacecount) (hpc6a = `1`; hpc7a/hpc6id/hpc8a = `2`). --- ## 6. Accessing the Cluster -Connect to the login node via SSM Session Manager — no public SSH needed. +Connect to the login node via SSM Session Manager — no public SSH needed. If you do +want direct SSH, set `SSHAccessCidr` to a trusted CIDR at deploy time to open SSH/22 +on the login node to that CIDR. **Console:** [EC2 Console](https://console.aws.amazon.com/ec2/home#Instances:) → filter by `aws:pcs:compute-node-group-name = login` → **Connect** → **Session Manager**. @@ -310,18 +303,68 @@ sbatch --partition=gpu-p6b300 nccl-tests-container.sbatch In `nccl-all_reduce_perf_.out`, EFA is active when you see `NET/OFI Selected provider is efa ... (found N nics)`; a healthy run ends with -`# Out of bounds values : 0 OK` and a busbw that scales with message size -(e.g. ~751 GB/s at 64 GiB on 2× p6-b300; raise `-e` past the 16 GiB default to -saturate larger fabrics). +`# Out of bounds values : 0 OK` and a busbw that scales with message size. For +expected bandwidth numbers per instance family and tuning notes, see +[tests/compute-test.md](./tests/compute-test.md#test-6-nccl-multi-node-efa). + +Before a long run, it's also worth checking GPU/EFA/NVLink health with the +[GPU Cluster Health Check suite](./tests/gpu-healthcheck-test.md) (nvidia-smi, DCGM +diagnostics, EFA enumeration, NCCL thresholds). For a full training example, see the [PyTorch FSDP test case](../../3.test_cases/pytorch/FSDP); the full validation matrix is in [tests/README.md](./tests/README.md). --- -## 8. Monitoring +## 8. Advanced Features + +### 8.1 Storage: FSx deployment types & sizing + +The `/fsx` (Lustre) and `/home` (OpenZFS) filesystems have working defaults. Two things +are worth knowing before changing them: **deployment types are not available in every +Region**, and **starting small then expanding is usually faster to deploy**. + +> **Tip — deploy small, expand after.** Both FSx for Lustre and OpenZFS can be **grown +> after creation** (increase storage capacity, and for Lustre throughput, via an FSx +> update — no data migration). A large filesystem also takes longer to create, which +> lengthens the whole stack deploy. So for the first deploy, leave `Capacity` / +> `HomeCapacity` at (or near) the minimum to get the cluster up quickly, then expand the +> filesystem to the size you need once the stack is `CREATE_COMPLETE`. -With `DeployMonitoring=true` (default), an integrated monitoring stack based on +**FSx deployment types are not available in every Region.** Defaults match the most +capable type; switch to a more widely available one if your Region needs it. + +| Filesystem | Parameter | Default | Other values | Notes | +|---|---|---|---|---| +| Lustre (`/fsx`) | `LustreDeploymentType` | `PERSISTENT_2` | `PERSISTENT_1` | `PERSISTENT_2` (throughput 125/250/500/1000, metadata config) isn't in every Region; `PERSISTENT_1` (50/100/200) is in more Regions | +| Lustre (`/fsx`) | `PerUnitStorageThroughput` | `250` | any valid number | Must be valid for the type: P2 = 125/250/500/1000, P1 = 50/100/200 | +| OpenZFS (`/home`) | `OpenZFSDeploymentType` | `SINGLE_AZ_HA_2` | `SINGLE_AZ_HA_1`, `SINGLE_AZ_2`, `SINGLE_AZ_1` | `SINGLE_AZ_1` is in all Regions; HA/2 variants vary. `MULTI_AZ` excluded (needs a second subnet) | +| OpenZFS (`/home`) | `HomeThroughput` | `320` | any valid number | Throughput (MB/s). Valid values depend on the deployment type: `SINGLE_AZ_2`/`SINGLE_AZ_HA_2` = 160/320/640/1280/2560/3840/5120/7680/10240; `SINGLE_AZ_HA_1`/`SINGLE_AZ_1` = 64/128/256/512/1024/2048/3072/4096 | + +Check support before deploying: +[Lustre Regions](https://docs.aws.amazon.com/fsx/latest/LustreGuide/using-fsx-lustre.html) · +[OpenZFS Regions](https://docs.aws.amazon.com/fsx/latest/OpenZFSGuide/available-aws-regions.html). +If a deploy fails at the FSx resource with an "unsupported deployment type" error, +switch these parameters to a type your Region supports. + +#### FSx for Lustre over EFA (GPUDirect Storage) + +`FSxLustreEnableEfa` (default `false`) enables EFA on the `/fsx` Lustre filesystem. +**The headline feature is GPUDirect Storage (GDS) for P5 / P5e / P5en / P6-B200 GPU CNGs**, +which DMAs file data straight into GPU memory (requires the NVIDIA `nvidia-fs` / cuFile +stack on the client — see the Client-side Lustre-on-EFA + GDS item in +[`docs/ROADMAP.md`](./docs/ROADMAP.md)). EFA-capable CPU CNGs +(`OnDemandEfaInterfaceCount > 0`) get the EFA *transport* path to storage as a secondary +benefit, useful when a single client is pushing past ~10 GBps. + +Constraints when enabling this: +- **PERSISTENT_2 SSD only** — a CFN Rule fails the stack at create time on `PERSISTENT_1`. +- **Much higher minimum `Capacity`** — at `PerUnitStorageThroughput=250` the minimum is + **19200 GiB** (16× the 1200 GiB default). Set `Capacity` accordingly. + +### 8.2 Monitoring + +With `MonitoringStack=Prometheus-LoginNode` (default), an integrated monitoring stack based on [aws-parallelcluster-monitoring](https://github.com/aws-samples/aws-parallelcluster-monitoring) is installed automatically: @@ -332,31 +375,19 @@ is installed automatically: Metrics cover Slurm jobs, GPU (utilization/memory/temperature/power/ECC/NVLink via DCGM), node CPU/memory/disk/network, and CloudWatch (EC2/FSx/PCS). The stack installs on node-local `/opt` (not the shared `/home`). Pre-built Grafana dashboards are -provisioned automatically — see [Accessing Grafana](#accessing-grafana) for the list +provisioned automatically — see [Accessing the Grafana dashboards](#accessing-the-grafana-dashboards) for the list and a screenshot. -> **GPU metrics work out of the box across the supported GPU range** (Hopper / B200 / -> B300). `DcgmExporterImage` defaults to a DCGM 4.5.2 build pinned by digest, validated -> on 2× p6-b300 and on B200. The monitoring stack's own default (DCGM 4.2.0) tops out -> at B200 and can't pull newer NVCR tags on Docker 29.x — overriding via digest at the -> deploy-all level is what bridges that. Override `DcgmExporterImage` only if you need -> to pin to a different build; details: -> [OPERATIONS.md §3.1](./docs/OPERATIONS.md#31-dcgmexporterimage-the-default-and-when-to-change-it). - > **Prefer AWS-managed Prometheus/Grafana?** If you'd rather use Amazon Managed Service > for Prometheus + Amazon Managed Grafana instead of the self-hosted stack on the login > node, see [`4.validation_and_observability/4.prometheus-grafana`](../../4.validation_and_observability/4.prometheus-grafana). -**Monitoring-related parameters:** -- `DeployMonitoring` (default `true`) -- `MonitoringVersion` — [aws-parallelcluster-monitoring](https://github.com/aws-samples/aws-parallelcluster-monitoring) git ref (release tag, branch, or `latest`; default `v2.9.1`). Pinned to a tag so upstream changes can't break deployments unexpectedly. `v2.9.1` adds the `DCGM_EXPORTER_IMAGE` override (lets `DcgmExporterImage` enable B300 GPU metrics); `v2.6.4`+ carry the PCS fixes (node-local `/opt` install + Docker-29.x DCGM tag). -- `MonitoringRepo` — `owner/repo` to fetch from (default `aws-samples/aws-parallelcluster-monitoring`). Point at a fork + a branch in `MonitoringVersion` to test unreleased changes. -- `DcgmExporterImage` — dcgm-exporter image used on GPU nodes; defaults to a DCGM 4.5.2 build pinned by digest (covers Hopper/B200/B300). Override only if you need to pin to a different build (e.g. the older monitoring-default DCGM 4.2.0). - -> Node type is identified by the `monitoring-role` tag (`login`/`compute`), not the EC2 -> `Name` tag — the `Name` tag defaults to `PCS-` and is free for you to retag. +`MonitoringStack` toggles the stack (`Prometheus-LoginNode` / `none`). The defaults work +out of the box; the source repo/version (`MonitoringRepo` / `MonitoringVersion`) and the +DCGM exporter image (`DcgmExporterImage`) are pinned and rarely need changing — see +[PARAMETERS.md](./docs/PARAMETERS.md) if you do. -### Accessing Grafana +#### Accessing the Grafana dashboards Log in to Grafana as **`admin`**; the password is generated per cluster and stored in SSM Parameter Store. Retrieve it (with `CLUSTER_ID` from the stack's `ClusterId` output): @@ -368,7 +399,7 @@ aws ssm get-parameter --name "/pcs/${CLUSTER_ID}/grafana/admin-password" \ There are two ways to reach the UI. -#### Option A — SSM port forwarding (default, private) +##### Option A — SSM port forwarding (default, private) No public access required; works even when the login node has no inbound rules. @@ -387,9 +418,9 @@ aws ssm start-session --target $INSTANCE_ID \ Then open `https://localhost:8443/grafana/`. -#### Option B — Direct public access (opt-in, via `GrafanaPublicAccessCidr`) +##### Option B — Direct public access (opt-in, via `GrafanaAccessCidr`) -To browse Grafana directly without port forwarding, set **`GrafanaPublicAccessCidr`** at +To browse Grafana directly without port forwarding, set **`GrafanaAccessCidr`** at deploy time to a CIDR you trust (e.g. your office IP `203.0.113.4/32`). deploy-all then creates a **login-only security group** that opens HTTPS/**443** to that CIDR and attaches it to the login node, so you can open: @@ -423,7 +454,7 @@ Security notes: as you are done. - The certificate is self-signed, so browsers show a warning — proceed past it, or put an ALB + ACM certificate in front for a trusted cert. -- Leaving `GrafanaPublicAccessCidr` empty (the default) keeps monitoring private; use +- Leaving `GrafanaAccessCidr` empty (the default) keeps monitoring private; use Option A. --- @@ -438,72 +469,116 @@ node's model, instance type, utilization, temperature, power, and memory: For detailed validation steps and the full test matrix (monitoring, containers, CPU/GPU, NCCL, FSDP), see the [Test & Validation Guide](tests/README.md). +> **Note — DCGM exporter version (GPU metrics).** GPU metrics work out of the box across +> the supported range (Hopper / B200 / B300) without overriding `DcgmExporterImage`: +> it defaults to a **DCGM 4.5.2** build pinned by digest (validated on 2× p6-b300 and on +> B200). The monitoring stack's own default (DCGM 4.2.0) tops out at B200 and can't pull +> newer NVCR tags on Docker 29.x — pinning by digest at the deploy-all level is what +> bridges that, and `MonitoringVersion v2.9.1` is the first release carrying the +> `DCGM_EXPORTER_IMAGE` override that lets this through (`v2.6.4`+ carry the other PCS +> fixes: node-local `/opt` install + the Docker-29.x DCGM tag). Override `DcgmExporterImage` +> only to pin a different build; details: +> [OPERATIONS.md §3.1](./docs/OPERATIONS.md#31-dcgmexporterimage-the-default-and-when-to-change-it). + +> **Note — node-type tagging.** The monitoring stack identifies login vs compute nodes by +> the `monitoring-role` tag (`login`/`compute`), **not** the EC2 `Name` tag — so the `Name` +> tag (default `PCS-`) is free for you to retag without breaking dashboards. + --- -## 9. Pre-baking Enroot/Pyxis into a custom AMI (optional) +### 8.3 User Management -The all-in-one template installs Enroot/Pyxis at **first boot** via -`PostInstallScriptUrl`, which is fast to deploy and avoids an Image Builder step. -For **frequent scaling** in production, pre-baking Enroot/Pyxis into a custom AMI -drops node boot time from ~8–12 min to ~3 min and pins every node to a deterministic -state. This is a separate, standalone path: build the AMI once with -[`pcs-ready-dlami-with-enroot-pyxis.yaml`](./assets/pcs-ready-dlami-with-enroot-pyxis.yaml), -then pass the resulting `ami-xxx` as `AmiId` to the cluster. +By default, the cluster runs as a single `ubuntu` user. For **multi-user +clusters** (per-user Slurm accounting, isolated home directories, team-based +access control), set `DirectoryService=OpenLDAP-LoginNode` at deploy time. + +This installs an OpenLDAP directory on the login node with SSSD on all compute +nodes — users you add are immediately visible cluster-wide. See the +**[User Management Guide](./docs/USER-MANAGEMENT.md)** for step-by-step +operations (adding/removing users, Slurm accounting, SSH access). + +> **Constraint — single login node.** The OpenLDAP server runs on the (single) login +> node, so enabling `DirectoryService` means a one-login-node cluster. The user database +> lives on OpenZFS `/home`, so it survives a login-node replacement (data is recoverable), +> but the directory **service** is a single point of failure while the login node is down. -**Step 1: Build the AMI** (~30 min one-time, separate stack) +### 8.4 IAM Permissions + +The cluster has two human roles, each with a ready-to-deploy IAM policy stack: + +| Role | What they can do | Deploy | +|---|---|---| +| **Cluster admin** ([`cluster-admin-iam.yaml`](./assets/cluster-admin-iam.yaml)) | deploy / update / delete the cluster (CloudFormation, PCS, EC2, FSx, scoped IAM) | [![Launch](images/launch-stack.svg)](https://console.aws.amazon.com/cloudformation/home#/stacks/quickcreate?templateUrl=https://awsome-distributed-ai.s3.amazonaws.com/templates/cluster-admin-iam.yaml&stackName=pcs-cluster-admins) | +| **Cluster user** ([`cluster-user-iam.yaml`](./assets/cluster-user-iam.yaml)) | SSM session to the **login node only** + read-only status; cannot create, modify, or delete anything | [![Launch](images/launch-stack.svg)](https://console.aws.amazon.com/cloudformation/home#/stacks/quickcreate?templateUrl=https://awsome-distributed-ai.s3.amazonaws.com/templates/cluster-user-iam.yaml&stackName=pcs-cluster-users) | + +Each template creates the customer-managed policies and an IAM group, and can +attach existing users at deploy time. See the **[IAM Permissions Guide](./docs/IAM.md)** +for roles, deploy instructions, security considerations, and the verification matrix. + +### 8.5 Pre-baking Enroot/Pyxis into a custom AMI + +The all-in-one template installs Enroot/Pyxis at **first boot** via +`PostInstallScriptUrl` (no Image Builder step). For **frequent scaling** in production, +pre-baking Enroot/Pyxis into a custom AMI drops node boot from ~8–12 min to ~3 min and +pins every node to a deterministic state. It's a standalone path: build the AMI once with +[`pcs-ready-dlami-with-enroot-pyxis.yaml`](./assets/pcs-ready-dlami-with-enroot-pyxis.yaml) +(single-Slurm-version by design), then pass the resulting `ami-xxx` as `AmiId`. [![Launch](images/launch-stack.svg)](https://console.aws.amazon.com/cloudformation/home#/stacks/quickcreate?templateUrl=https://awsome-distributed-ai.s3.amazonaws.com/templates/pcs-ready-dlami-with-enroot-pyxis.yaml&stackName=pcs-dlami) -```bash -aws cloudformation create-stack \ - --stack-name pcs-dlami \ - --template-url https://awsome-distributed-ai.s3.amazonaws.com/templates/pcs-ready-dlami-with-enroot-pyxis.yaml \ - --parameters ParameterKey=SlurmVersion,ParameterValue=25.11 \ - --capabilities CAPABILITY_IAM CAPABILITY_NAMED_IAM -``` +See **[docs/CUSTOM-AMI.md](./docs/CUSTOM-AMI.md)** for the full build → read AMI ID → +deploy procedure and the optional scheduled-rebuild / lifecycle / SSM-publish features. -The AMI is **single-Slurm-version by design**: Pyxis is a SPANK plugin whose ABI is -locked to its compile-time Slurm version, so pass the same `SlurmVersion` you'll use -on the cluster. +### 8.6 CPU compute node group — advanced settings -**Step 2: Read the resulting AMI ID** from the stack output `DLAMIforPCSAmiId`: +The On-Demand CPU queue is configured with `OnDemandInstanceType` (default +`c6i.4xlarge`) plus `OnDemandQueueName` / `OnDemandCngName` / `OnDemandMinCount` / +`OnDemandMaxCount`. The settings below cover tightly-coupled HPC / MPI workloads. -```bash -AMI_ID=$(aws cloudformation describe-stacks \ - --stack-name pcs-dlami \ - --query 'Stacks[0].Outputs[?OutputKey==`DLAMIforPCSAmiId`].OutputValue' \ - --output text) -echo "$AMI_ID" # ami-0xxxxxxxxxxxxxxxx -``` +#### EFA on CPU HPC instances (`OnDemandEfaInterfaceCount`) -**Step 3: Pass it to the cluster** as `AmiId` and clear `PostInstallScriptUrl` for -the cleanest boot: +For tightly-coupled HPC / MPI workloads on CPU-only HPC instances, set +`OnDemandEfaInterfaceCount` (deploy-all) or `EfaInterfaceCount` (modular +`add-cng.yaml`) to the instance type's EFA interface count. `0` (default) = no +EFA; `1` or `2` = enable EFA with that many interfaces — the CPU compute node +group then launches with EFA `NetworkInterfaces` and a cluster placement group +is auto-created. GPU CNGs (P5/P6) are unaffected — they use their own dedicated +multi-NIC EFA wiring per-family. + +| Instance type | EFA interfaces | Aggregate spec | Set the count to | +|---|---:|---:|---:| +| (any non-HPC type, e.g. `c6i.4xlarge`) | — | — | **0** (no EFA; default) | +| `hpc6a.48xlarge` | 1 | 100 Gbps | **1** | +| `hpc7a.96xlarge` (and 12/24/48) | 2 | 300 Gbps | **2** | +| `hpc6id.32xlarge` | 2 | 200 Gbps | **2** | +| `hpc8a.96xlarge` | 2 | 300 Gbps | **2** | +| `c7i.metal-48xl` etc. | 1 | varies | **1** | -```bash -aws cloudformation create-stack \ - --stack-name pcs-ml-cluster \ - --template-url https://awsome-distributed-ai.s3.amazonaws.com/templates/pcs-ml-cluster-deploy-all.yaml \ - --parameters \ - ParameterKey=PrimarySubnetAZ,ParameterValue=us-east-1a \ - ParameterKey=AmiId,ParameterValue=$AMI_ID \ - ParameterKey=PostInstallScriptUrl,ParameterValue= \ - --capabilities CAPABILITY_IAM CAPABILITY_NAMED_IAM -``` +(Setting a count > 0 on a non-EFA type, or mismatching the count with the +instance type's actual `MaximumEfaInterfaces`, fails at launch.) + +**Placement group:** auto-created per-CNG by the template. Override with +`OnDemandPlacementGroupName=` to share one PG across multiple +CNGs (e.g. heterogeneous tightly-coupled jobs that span CPU + GPU). -The post-install hook is idempotent, so leaving `PostInstallScriptUrl` at its default -also works — the installer detects Enroot/Pyxis is already present and is a fast -no-op. Setting it empty just shaves a few seconds off boot. +**Multi-NIC bandwidth needs multiple MPI pairs.** A single MPI pair uses one +libfabric endpoint and only one NIC. Use `osu_mbw_mr -np 32 -N 16` (or your +application's natural multi-pair pattern) to actually exercise both NICs on +hpc7a/hpc8a. See [tests/README.md Test 9](./tests/README.md#test-9-efa-on-cpu-hpc-instances-hpc6a--hpc7a--hpc8a) +for the full benchmark setup and validated bandwidth numbers. -**Optional features of `pcs-ready-dlami-with-enroot-pyxis.yaml`** (defaults are off): -- `BuildSchedule=Weekly`/`Monthly` for scheduled rebuilds against a moving base AMI -- `EnableLifecyclePolicy=true` to deprecate older AMIs after `LifecycleDeprecateAfterWeeks` -- `PublishToSsm=true` to publish the latest AMI ID to an SSM parameter for downstream stacks +### 8.7 Deploying updated templates before they are published -For production deploys that pin the AMI explicitly per cluster, none of these are needed. +The Quick Start deploys from the public production bucket (`awsome-distributed-ai`), which +only holds already-published templates. If you need to test **template or script changes +that aren't published yet** — a fork, a feature branch, or a PR under review — host the +templates + boot scripts in an S3 bucket you control and point the deploy at it via +`S3BucketName` / `S3KeyPrefix`. The step-by-step procedure (upload, deploy, iterate, +clean up) is in [docs/DEPLOY-TESTING.md](./docs/DEPLOY-TESTING.md). --- -## 10. Templates +## 9. Templates All templates live in [`assets/`](./assets/). `pcs-ml-cluster-deploy-all.yaml` nests the others; you can also deploy each individually for more control (e.g. reuse a VPC/FSx @@ -512,22 +587,80 @@ parameter and default, see [PARAMETERS.md](./docs/PARAMETERS.md). | Template | Purpose | Deploy | |---|---|---| -| [`pcs-ml-cluster-deploy-all.yaml`](./assets/pcs-ml-cluster-deploy-all.yaml) | All-in-one: Prerequisites + (optional AMI) + Cluster + login/CPU/GPU CNGs | [🚀](https://console.aws.amazon.com/cloudformation/home#/stacks/quickcreate?templateUrl=https://awsome-distributed-ai.s3.amazonaws.com/templates/pcs-ml-cluster-deploy-all.yaml&stackName=pcs-ml-cluster) | -| [`ml-cluster-prerequisites.yaml`](./assets/ml-cluster-prerequisites.yaml) | VPC, subnets, security groups, FSx for Lustre + OpenZFS | [🚀](https://console.aws.amazon.com/cloudformation/home#/stacks/quickcreate?templateUrl=https://awsome-distributed-ai.s3.amazonaws.com/templates/ml-cluster-prerequisites.yaml&stackName=pcs-prerequisites) | -| [`cluster.yaml`](./assets/cluster.yaml) | PCS cluster core (Slurm scheduler only, no nodes) | [🚀](https://console.aws.amazon.com/cloudformation/home#/stacks/quickcreate?templateUrl=https://awsome-distributed-ai.s3.amazonaws.com/templates/cluster.yaml&stackName=pcs-cluster) | -| [`add-cng.yaml`](./assets/add-cng.yaml) | Compute node group — login nodes, CPU / single-NIC-GPU queues (C6i, G5, G6); switches to a multi-NIC EFA `NetworkInterfaces` block when `EnableEfa=true` (HPC types: hpc6a/hpc7a/hpc6id/hpc8a …) | [🚀](https://console.aws.amazon.com/cloudformation/home#/stacks/quickcreate?templateUrl=https://awsome-distributed-ai.s3.amazonaws.com/templates/add-cng.yaml&stackName=pcs-add-cng) | -| [`add-cng-p5.yaml`](./assets/add-cng-p5.yaml) | P5/P5e/P5en nodes (16/32 EFA interfaces, by type) | [🚀](https://console.aws.amazon.com/cloudformation/home#/stacks/quickcreate?templateUrl=https://awsome-distributed-ai.s3.amazonaws.com/templates/add-cng-p5.yaml&stackName=pcs-add-cng-p5) | -| [`add-cng-p6-b200.yaml`](./assets/add-cng-p6-b200.yaml) | P6-B200 nodes (8 EFA interfaces) | [🚀](https://console.aws.amazon.com/cloudformation/home#/stacks/quickcreate?templateUrl=https://awsome-distributed-ai.s3.amazonaws.com/templates/add-cng-p6-b200.yaml&stackName=pcs-add-cng-p6-b200) | -| [`add-cng-p6-b300.yaml`](./assets/add-cng-p6-b300.yaml) | P6-B300 nodes (16 EFA interfaces) | [🚀](https://console.aws.amazon.com/cloudformation/home#/stacks/quickcreate?templateUrl=https://awsome-distributed-ai.s3.amazonaws.com/templates/add-cng-p6-b300.yaml&stackName=pcs-add-cng-p6-b300) | -| [`pcs-ready-dlami-with-enroot-pyxis.yaml`](./assets/pcs-ready-dlami-with-enroot-pyxis.yaml) | EC2 Image Builder: bake Enroot 3.5.0 + Pyxis 0.20.0 into the PCS-Ready DLAMI | [🚀](https://console.aws.amazon.com/cloudformation/home#/stacks/quickcreate?templateUrl=https://awsome-distributed-ai.s3.amazonaws.com/templates/pcs-ready-dlami-with-enroot-pyxis.yaml&stackName=pcs-dlami) | +| [`pcs-ml-cluster-deploy-all.yaml`](./assets/pcs-ml-cluster-deploy-all.yaml) | All-in-one: Prerequisites + (optional AMI) + Cluster + login/CPU/GPU CNGs | [![Launch](images/launch-stack.svg)](https://console.aws.amazon.com/cloudformation/home#/stacks/quickcreate?templateUrl=https://awsome-distributed-ai.s3.amazonaws.com/templates/pcs-ml-cluster-deploy-all.yaml&stackName=pcs-ml-cluster) | +| [`ml-cluster-prerequisites.yaml`](./assets/ml-cluster-prerequisites.yaml) | VPC, subnets, security groups, FSx for Lustre + OpenZFS | [![Launch](images/launch-stack.svg)](https://console.aws.amazon.com/cloudformation/home#/stacks/quickcreate?templateUrl=https://awsome-distributed-ai.s3.amazonaws.com/templates/ml-cluster-prerequisites.yaml&stackName=pcs-prerequisites) | +| [`cluster.yaml`](./assets/cluster.yaml) | PCS cluster core (Slurm scheduler only, no nodes) | [![Launch](images/launch-stack.svg)](https://console.aws.amazon.com/cloudformation/home#/stacks/quickcreate?templateUrl=https://awsome-distributed-ai.s3.amazonaws.com/templates/cluster.yaml&stackName=pcs-cluster) | +| [`add-cng.yaml`](./assets/add-cng.yaml) | Compute node group — login nodes, CPU / single-NIC-GPU queues (C6i, G5, G6); switches to a multi-NIC EFA `NetworkInterfaces` block when `EfaInterfaceCount > 0` (HPC types: hpc6a/hpc7a/hpc6id/hpc8a …) | [![Launch](images/launch-stack.svg)](https://console.aws.amazon.com/cloudformation/home#/stacks/quickcreate?templateUrl=https://awsome-distributed-ai.s3.amazonaws.com/templates/add-cng.yaml&stackName=pcs-add-cng) | +| [`add-cng-p5.yaml`](./assets/add-cng-p5.yaml) | P5/P5e/P5en nodes (16/32 EFA interfaces, by type) | [![Launch](images/launch-stack.svg)](https://console.aws.amazon.com/cloudformation/home#/stacks/quickcreate?templateUrl=https://awsome-distributed-ai.s3.amazonaws.com/templates/add-cng-p5.yaml&stackName=pcs-add-cng-p5) | +| [`add-cng-p6-b200.yaml`](./assets/add-cng-p6-b200.yaml) | P6-B200 nodes (8 EFA interfaces) | [![Launch](images/launch-stack.svg)](https://console.aws.amazon.com/cloudformation/home#/stacks/quickcreate?templateUrl=https://awsome-distributed-ai.s3.amazonaws.com/templates/add-cng-p6-b200.yaml&stackName=pcs-add-cng-p6-b200) | +| [`add-cng-p6-b300.yaml`](./assets/add-cng-p6-b300.yaml) | P6-B300 nodes (17 interfaces: 16 EFA + 1 ENA) | [![Launch](images/launch-stack.svg)](https://console.aws.amazon.com/cloudformation/home#/stacks/quickcreate?templateUrl=https://awsome-distributed-ai.s3.amazonaws.com/templates/add-cng-p6-b300.yaml&stackName=pcs-add-cng-p6-b300) | +| [`pcs-ready-dlami-with-enroot-pyxis.yaml`](./assets/pcs-ready-dlami-with-enroot-pyxis.yaml) | EC2 Image Builder: bake Enroot 3.5.0 + Pyxis 0.20.0 into the PCS-Ready DLAMI | [![Launch](images/launch-stack.svg)](https://console.aws.amazon.com/cloudformation/home#/stacks/quickcreate?templateUrl=https://awsome-distributed-ai.s3.amazonaws.com/templates/pcs-ready-dlami-with-enroot-pyxis.yaml&stackName=pcs-dlami) | `add-cng*` templates create a Slurm queue only when `QueueName` is set (leave it empty for login nodes). The P-series templates need a `CapacityReservationId` when using a Capacity Block. +### Template nesting structure (deploy-all) + +``` +pcs-ml-cluster-deploy-all.yaml ← user deploys this +│ +├─► ml-cluster-prerequisites.yaml ← VPC, subnets, SG, FSx Lustre + OpenZFS +│ +├─► cluster.yaml ← PCS cluster (Slurm scheduler), IAM role +│ • IAM role grants: ssm:PutParameter /pcs//grafana/*, /pcs//ldap/* +│ +├─► add-cng.yaml (login) ← login node (MinCount=1) +│ • MonitoringRole=login → Prometheus/Grafana +│ • DirectoryService=OpenLDAP-LoginNode → slapd server +│ │ +│ └─ UserData fetches external scripts: +│ ├─ PostInstallScriptUrl (default: install-enroot-pyxis.sh) +│ ├─ MonitoringRepo/MonitoringVersion → post-install.sh +│ └─ setup-directory.sh server (when DirectoryRole=server) +│ +├─► add-cng.yaml (compute) ← CPU queue (dynamic scaling 0→N) +│ • MonitoringRole=compute → Node Exporter +│ • DirectoryService=OpenLDAP-LoginNode → SSSD client +│ • EfaInterfaceCount>0 → EFA NetworkInterfaces + PG +│ │ +│ └─ UserData fetches external scripts: +│ ├─ PostInstallScriptUrl (same as login) +│ ├─ MonitoringRepo/MonitoringVersion → post-install.sh +│ └─ setup-directory.sh client (when DirectoryRole=client) +│ +└─► add-cng-p5.yaml / add-cng-p6-b200.yaml ← GPU queue (optional) + / add-cng-p6-b300.yaml + • Multi-NIC EFA (16/32 cards, per-family) + • MonitoringRole=compute → DCGM Exporter + • Same external script pattern as compute CNG + +Boot scripts (fetched at first boot from S3: s3:///scripts/): + assets/scripts/install-enroot-pyxis.sh ← Enroot 3.5.0 + Pyxis 0.20.0 + assets/scripts/setup-directory.sh ← multi-user directory (server + client) + +External boot scripts (fetched from GitHub): + aws-parallelcluster-monitoring post-install.sh ← monitoring stack installer + (https://github.com/aws-samples/aws-parallelcluster-monitoring) + fetched from: ${MonitoringRepo} @ ${MonitoringVersion} + +Helper scripts (NOT run at boot — for admin use on the login node): + assets/scripts/ldap-add-user.sh ← add POSIX users to LDAP directory + +External references (runtime): + SSM /aws/service/pcs/ami/.../latest/ami-id ← PCS-Ready DLAMI (when AmiId is empty) + SSM /pcs//grafana/admin-password ← auto-generated Grafana password + SSM /pcs//ldap/admin-password ← auto-generated LDAP admin password + +Standalone (not nested): + pcs-ready-dlami-with-enroot-pyxis.yaml ← AMI builder (separate stack) + cluster-admin-iam.yaml ← IAM admin policies + group (separate stack) + cluster-user-iam.yaml ← IAM user policy + group (separate stack) +``` + --- -## 11. Testing and Validation +## 10. Testing and Validation Every template has been deployed end-to-end on real hardware: P5 / P5e / P5en / P6-B200 / P6-B300 (NCCL all-reduce + FSDP Llama-2 7B), HPC EFA on hpc6a / hpc7a / @@ -538,12 +671,16 @@ numbers is in **[tests/README.md](./tests/README.md)**. --- -## 12. Additional Resources +## 11. Additional Resources In this repo: - [Parameter reference](./docs/PARAMETERS.md) — every deploy-all parameter and default -- [Operations guide](./docs/OPERATIONS.md) — version trade-offs, AMI pinning, monitoring/DCGM, FSx coupling, production settings +- [Operations guide](./docs/OPERATIONS.md) — version trade-offs, AMI pinning, monitoring/DCGM, FSx coupling, Lustre tuning, production settings +- [User management guide](./docs/USER-MANAGEMENT.md) — multi-user setup with OpenLDAP (add/remove users, groups, Slurm accounting) +- [IAM permissions guide](./docs/IAM.md) — cluster admin / cluster user roles, policy deploy, security considerations +- [Deploy & testing procedures](./docs/DEPLOY-TESTING.md) — development deploy workflow with test S3 bucket - [Test & Validation Guide](./tests/README.md) — reproducible matrix with measured numbers +- [GPU Cluster Health Check](../../4.validation_and_observability/2.gpu-cluster-healthcheck) — comprehensive GPU/EFA/NVLink validation suite (lightweight + intensive modes, Slurm prolog integration) - [Roadmap / TODO](./docs/ROADMAP.md) External: diff --git a/architectures/aws-pcs/assets/add-cng-p5.yaml b/architectures/aws-pcs/assets/add-cng-p5.yaml index 0926e505d..ca7ebd261 100644 --- a/architectures/aws-pcs/assets/add-cng-p5.yaml +++ b/architectures/aws-pcs/assets/add-cng-p5.yaml @@ -203,8 +203,8 @@ Parameters: Description: >- Optional URL of a post-install script to download and run on each node at first boot (PCS equivalent of ParallelCluster OnNodeConfigured). Leave - empty to skip. To install Enroot/Pyxis, point this at scripts/install-enroot-pyxis.sh. - Must be an HTTP(S) URL (e.g. a GitHub raw URL); S3 public hosting is not allowed. + empty to skip. To install Enroot/Pyxis, point this at assets/scripts/install-enroot-pyxis.sh. + Accepts an s3:// URL (instance-role fetch; private bucket OK) or an http(s):// URL (curl; public only). For Enroot/Pyxis point at scripts/install-enroot-pyxis.sh in the templates bucket. Default: '' PostInstallScriptArgs: @@ -224,6 +224,49 @@ Parameters: - '25.05' - '25.11' + DirectoryService: + Type: String + Description: >- + Multi-user directory service. 'none' = single ubuntu user (default, + unchanged). 'OpenLDAP-LoginNode' = deploy slapd on the login node (DB on shared + /home/ldap-db), compute nodes configure SSSD as LDAP clients. Future + options (SimpleAD, ManagedAD) will be added as AllowedValues. + Default: 'none' + AllowedValues: + - 'none' + - 'OpenLDAP-LoginNode' + + DirectoryDomainSuffix: + Type: String + Description: >- + LDAP domain suffix (e.g. dc=cluster,dc=internal). Only used when + DirectoryService != none. + Default: 'dc=cluster,dc=internal' + + DirectoryRole: + Type: String + Description: >- + This CNG's role in the directory service. 'server' = install and run + the directory (slapd on login node). 'client' = configure SSSD to + resolve users from the directory server. 'none' = skip directory setup. + deploy-all sets this automatically (login=server, compute=client) when + DirectoryService != none. + Default: 'none' + AllowedValues: + - 'none' + - 'server' + - 'client' + + S3BucketName: + Type: String + Description: S3 bucket where scripts are stored (same as nested templates bucket) + Default: 'awsome-distributed-ai' + + S3KeyPrefix: + Type: String + Description: S3 key prefix for scripts (e.g. templates/) + Default: 'templates/' + Conditions: # EFA interface count is derived from the instance type: p5en.48xlarge has 16 # network cards, p5.48xlarge / p5e.48xlarge have 32. So "use 32 interfaces" @@ -237,6 +280,7 @@ Conditions: # tag (monitoring-role=login) and excludes them from EC2 service discovery. IsMonitoringRoleSet: !Not [!Equals [!Ref MonitoringRole, 'none']] HasAmiId: !Not [!Equals [!Ref AmiId, ""]] + DirectoryEnabled: !Not [!Equals [!Ref DirectoryRole, 'none']] Resources: @@ -336,6 +380,16 @@ Resources: - Key: monitoring-role Value: !Ref MonitoringRole - !Ref AWS::NoValue + # Directory (multi-user) role tag, independent of monitoring-role. + # Emitted for both server (login) and client (compute) nodes when + # DirectoryRole != none. Compute clients discover the OpenLDAP + # server by querying for directory-role=server (NOT monitoring-role, + # which is owned by the monitoring stack). See setup-directory.sh. + - !If + - DirectoryEnabled + - Key: directory-role + Value: !Ref DirectoryRole + - !Ref AWS::NoValue MetadataOptions: HttpEndpoint: enabled HttpPutResponseHopLimit: 4 @@ -356,7 +410,10 @@ Resources: - | if [ -n "${PostInstallScriptUrl}" ]; then echo "Downloading post-install script from ${PostInstallScriptUrl}..." | tee /var/log/pcs-post-install.log - curl -fsSL "${PostInstallScriptUrl}" -o /tmp/pcs-post-install.sh + case "${PostInstallScriptUrl}" in + s3://*) aws s3 cp "${PostInstallScriptUrl}" /tmp/pcs-post-install.sh >> /var/log/pcs-post-install.log 2>&1 ;; + *) curl -fsSL "${PostInstallScriptUrl}" -o /tmp/pcs-post-install.sh ;; + esac chmod +x /tmp/pcs-post-install.sh echo "Executing post-install script..." | tee -a /var/log/pcs-post-install.log # Tell the script this cluster's Slurm version (it can't discover it at @@ -396,6 +453,22 @@ Resources: bash /tmp/post-install.sh ${MonitoringVersion} ${MonitoringRepo} 2>&1 | tee /var/log/monitoring-install.log echo "Monitoring installation complete" fi + - | + # Multi-user directory setup (optional, DirectoryRole != none) + if [ "${DirectoryRole}" != "none" ]; then + export LDAP_DOMAIN_SUFFIX="${DirectoryDomainSuffix}" + export LDAP_DOMAIN=$(echo "${DirectoryDomainSuffix}" | sed 's/dc=//g; s/,/./g') + export LDAP_ADMIN_PASSWORD=$(openssl rand -base64 16) + export CLUSTER_ID="${ClusterId}" + export DIRECTORY_DNS_IPS="" + # Pass the script source so the server role can install the + # ldap-add-user.sh helper onto /usr/local/bin from the same bucket. + export S3_BUCKET="${S3BucketName}" + export S3_KEY_PREFIX="${S3KeyPrefix}" + aws s3 cp s3://${S3BucketName}/${S3KeyPrefix}scripts/setup-directory.sh /tmp/setup-directory.sh + chmod +x /tmp/setup-directory.sh + bash /tmp/setup-directory.sh "${DirectoryRole}" 2>&1 | tee /var/log/directory-setup.log + fi --==MYBOUNDARY== NetworkInterfaces: diff --git a/architectures/aws-pcs/assets/add-cng-p6-b200.yaml b/architectures/aws-pcs/assets/add-cng-p6-b200.yaml index 40570aaa9..947ab1b94 100644 --- a/architectures/aws-pcs/assets/add-cng-p6-b200.yaml +++ b/architectures/aws-pcs/assets/add-cng-p6-b200.yaml @@ -194,8 +194,8 @@ Parameters: Description: >- Optional URL of a post-install script to download and run on each node at first boot (PCS equivalent of ParallelCluster OnNodeConfigured). Leave - empty to skip. To install Enroot/Pyxis, point this at scripts/install-enroot-pyxis.sh. - Must be an HTTP(S) URL (e.g. a GitHub raw URL); S3 public hosting is not allowed. + empty to skip. To install Enroot/Pyxis, point this at assets/scripts/install-enroot-pyxis.sh. + Accepts an s3:// URL (instance-role fetch; private bucket OK) or an http(s):// URL (curl; public only). For Enroot/Pyxis point at scripts/install-enroot-pyxis.sh in the templates bucket. Default: '' PostInstallScriptArgs: @@ -215,6 +215,49 @@ Parameters: - '25.05' - '25.11' + DirectoryService: + Type: String + Description: >- + Multi-user directory service. 'none' = single ubuntu user (default, + unchanged). 'OpenLDAP-LoginNode' = deploy slapd on the login node (DB on shared + /home/ldap-db), compute nodes configure SSSD as LDAP clients. Future + options (SimpleAD, ManagedAD) will be added as AllowedValues. + Default: 'none' + AllowedValues: + - 'none' + - 'OpenLDAP-LoginNode' + + DirectoryDomainSuffix: + Type: String + Description: >- + LDAP domain suffix (e.g. dc=cluster,dc=internal). Only used when + DirectoryService != none. + Default: 'dc=cluster,dc=internal' + + DirectoryRole: + Type: String + Description: >- + This CNG's role in the directory service. 'server' = install and run + the directory (slapd on login node). 'client' = configure SSSD to + resolve users from the directory server. 'none' = skip directory setup. + deploy-all sets this automatically (login=server, compute=client) when + DirectoryService != none. + Default: 'none' + AllowedValues: + - 'none' + - 'server' + - 'client' + + S3BucketName: + Type: String + Description: S3 bucket where scripts are stored (same as nested templates bucket) + Default: 'awsome-distributed-ai' + + S3KeyPrefix: + Type: String + Description: S3 key prefix for scripts (e.g. templates/) + Default: 'templates/' + Conditions: UseCapacityBlock: !Not [!Equals [!Ref CapacityReservationId, ""]] UseOnDemand: !Equals [!Ref CapacityReservationId, ""] @@ -224,6 +267,7 @@ Conditions: # tag (monitoring-role=login) and excludes them from EC2 service discovery. IsMonitoringRoleSet: !Not [!Equals [!Ref MonitoringRole, 'none']] HasAmiId: !Not [!Equals [!Ref AmiId, ""]] + DirectoryEnabled: !Not [!Equals [!Ref DirectoryRole, 'none']] Resources: @@ -323,6 +367,16 @@ Resources: - Key: monitoring-role Value: !Ref MonitoringRole - !Ref AWS::NoValue + # Directory (multi-user) role tag, independent of monitoring-role. + # Emitted for both server (login) and client (compute) nodes when + # DirectoryRole != none. Compute clients discover the OpenLDAP + # server by querying for directory-role=server (NOT monitoring-role, + # which is owned by the monitoring stack). See setup-directory.sh. + - !If + - DirectoryEnabled + - Key: directory-role + Value: !Ref DirectoryRole + - !Ref AWS::NoValue MetadataOptions: HttpEndpoint: enabled HttpPutResponseHopLimit: 4 @@ -343,7 +397,10 @@ Resources: - | if [ -n "${PostInstallScriptUrl}" ]; then echo "Downloading post-install script from ${PostInstallScriptUrl}..." | tee /var/log/pcs-post-install.log - curl -fsSL "${PostInstallScriptUrl}" -o /tmp/pcs-post-install.sh + case "${PostInstallScriptUrl}" in + s3://*) aws s3 cp "${PostInstallScriptUrl}" /tmp/pcs-post-install.sh >> /var/log/pcs-post-install.log 2>&1 ;; + *) curl -fsSL "${PostInstallScriptUrl}" -o /tmp/pcs-post-install.sh ;; + esac chmod +x /tmp/pcs-post-install.sh echo "Executing post-install script..." | tee -a /var/log/pcs-post-install.log # Tell the script this cluster's Slurm version (it can't discover it at @@ -397,6 +454,22 @@ Resources: done echo "Monitoring installation complete (exit $rc)" >> /var/log/monitoring-install.log fi + - | + # Multi-user directory setup (optional, DirectoryRole != none) + if [ "${DirectoryRole}" != "none" ]; then + export LDAP_DOMAIN_SUFFIX="${DirectoryDomainSuffix}" + export LDAP_DOMAIN=$(echo "${DirectoryDomainSuffix}" | sed 's/dc=//g; s/,/./g') + export LDAP_ADMIN_PASSWORD=$(openssl rand -base64 16) + export CLUSTER_ID="${ClusterId}" + export DIRECTORY_DNS_IPS="" + # Pass the script source so the server role can install the + # ldap-add-user.sh helper onto /usr/local/bin from the same bucket. + export S3_BUCKET="${S3BucketName}" + export S3_KEY_PREFIX="${S3KeyPrefix}" + aws s3 cp s3://${S3BucketName}/${S3KeyPrefix}scripts/setup-directory.sh /tmp/setup-directory.sh + chmod +x /tmp/setup-directory.sh + bash /tmp/setup-directory.sh "${DirectoryRole}" 2>&1 | tee /var/log/directory-setup.log + fi --==MYBOUNDARY== NetworkInterfaces: diff --git a/architectures/aws-pcs/assets/add-cng-p6-b300.yaml b/architectures/aws-pcs/assets/add-cng-p6-b300.yaml index 218f03783..a623394be 100644 --- a/architectures/aws-pcs/assets/add-cng-p6-b300.yaml +++ b/architectures/aws-pcs/assets/add-cng-p6-b300.yaml @@ -197,8 +197,8 @@ Parameters: Description: >- Optional URL of a post-install script to download and run on each node at first boot (PCS equivalent of ParallelCluster OnNodeConfigured). Leave - empty to skip. To install Enroot/Pyxis, point this at scripts/install-enroot-pyxis.sh. - Must be an HTTP(S) URL (e.g. a GitHub raw URL); S3 public hosting is not allowed. + empty to skip. To install Enroot/Pyxis, point this at assets/scripts/install-enroot-pyxis.sh. + Accepts an s3:// URL (instance-role fetch; private bucket OK) or an http(s):// URL (curl; public only). For Enroot/Pyxis point at scripts/install-enroot-pyxis.sh in the templates bucket. Default: '' PostInstallScriptArgs: @@ -218,6 +218,49 @@ Parameters: - '25.05' - '25.11' + DirectoryService: + Type: String + Description: >- + Multi-user directory service. 'none' = single ubuntu user (default, + unchanged). 'OpenLDAP-LoginNode' = deploy slapd on the login node (DB on shared + /home/ldap-db), compute nodes configure SSSD as LDAP clients. Future + options (SimpleAD, ManagedAD) will be added as AllowedValues. + Default: 'none' + AllowedValues: + - 'none' + - 'OpenLDAP-LoginNode' + + DirectoryDomainSuffix: + Type: String + Description: >- + LDAP domain suffix (e.g. dc=cluster,dc=internal). Only used when + DirectoryService != none. + Default: 'dc=cluster,dc=internal' + + DirectoryRole: + Type: String + Description: >- + This CNG's role in the directory service. 'server' = install and run + the directory (slapd on login node). 'client' = configure SSSD to + resolve users from the directory server. 'none' = skip directory setup. + deploy-all sets this automatically (login=server, compute=client) when + DirectoryService != none. + Default: 'none' + AllowedValues: + - 'none' + - 'server' + - 'client' + + S3BucketName: + Type: String + Description: S3 bucket where scripts are stored (same as nested templates bucket) + Default: 'awsome-distributed-ai' + + S3KeyPrefix: + Type: String + Description: S3 key prefix for scripts (e.g. templates/) + Default: 'templates/' + Conditions: UseCapacityBlock: !Not [!Equals [!Ref CapacityReservationId, ""]] UseOnDemand: !Equals [!Ref CapacityReservationId, ""] @@ -227,6 +270,7 @@ Conditions: # tag (monitoring-role=login) and excludes them from EC2 service discovery. IsMonitoringRoleSet: !Not [!Equals [!Ref MonitoringRole, 'none']] HasAmiId: !Not [!Equals [!Ref AmiId, ""]] + DirectoryEnabled: !Not [!Equals [!Ref DirectoryRole, 'none']] Resources: @@ -326,6 +370,16 @@ Resources: - Key: monitoring-role Value: !Ref MonitoringRole - !Ref AWS::NoValue + # Directory (multi-user) role tag, independent of monitoring-role. + # Emitted for both server (login) and client (compute) nodes when + # DirectoryRole != none. Compute clients discover the OpenLDAP + # server by querying for directory-role=server (NOT monitoring-role, + # which is owned by the monitoring stack). See setup-directory.sh. + - !If + - DirectoryEnabled + - Key: directory-role + Value: !Ref DirectoryRole + - !Ref AWS::NoValue MetadataOptions: HttpEndpoint: enabled HttpPutResponseHopLimit: 4 @@ -346,7 +400,10 @@ Resources: - | if [ -n "${PostInstallScriptUrl}" ]; then echo "Downloading post-install script from ${PostInstallScriptUrl}..." | tee /var/log/pcs-post-install.log - curl -fsSL "${PostInstallScriptUrl}" -o /tmp/pcs-post-install.sh + case "${PostInstallScriptUrl}" in + s3://*) aws s3 cp "${PostInstallScriptUrl}" /tmp/pcs-post-install.sh >> /var/log/pcs-post-install.log 2>&1 ;; + *) curl -fsSL "${PostInstallScriptUrl}" -o /tmp/pcs-post-install.sh ;; + esac chmod +x /tmp/pcs-post-install.sh echo "Executing post-install script..." | tee -a /var/log/pcs-post-install.log # Tell the script this cluster's Slurm version (it can't discover it at @@ -400,6 +457,22 @@ Resources: done echo "Monitoring installation complete (exit $rc)" >> /var/log/monitoring-install.log fi + - | + # Multi-user directory setup (optional, DirectoryRole != none) + if [ "${DirectoryRole}" != "none" ]; then + export LDAP_DOMAIN_SUFFIX="${DirectoryDomainSuffix}" + export LDAP_DOMAIN=$(echo "${DirectoryDomainSuffix}" | sed 's/dc=//g; s/,/./g') + export LDAP_ADMIN_PASSWORD=$(openssl rand -base64 16) + export CLUSTER_ID="${ClusterId}" + export DIRECTORY_DNS_IPS="" + # Pass the script source so the server role can install the + # ldap-add-user.sh helper onto /usr/local/bin from the same bucket. + export S3_BUCKET="${S3BucketName}" + export S3_KEY_PREFIX="${S3KeyPrefix}" + aws s3 cp s3://${S3BucketName}/${S3KeyPrefix}scripts/setup-directory.sh /tmp/setup-directory.sh + chmod +x /tmp/setup-directory.sh + bash /tmp/setup-directory.sh "${DirectoryRole}" 2>&1 | tee /var/log/directory-setup.log + fi --==MYBOUNDARY== NetworkInterfaces: diff --git a/architectures/aws-pcs/assets/add-cng.yaml b/architectures/aws-pcs/assets/add-cng.yaml index 37900f8b8..98b155c3b 100644 --- a/architectures/aws-pcs/assets/add-cng.yaml +++ b/architectures/aws-pcs/assets/add-cng.yaml @@ -9,7 +9,7 @@ Description: > (DeployMonitoring/MonitoringRole). A Slurm queue is created only when QueueName is set (leave empty for login nodes). - EFA support (EnableEfa=true) is available for CPU HPC instance types + EFA support (EfaInterfaceCount > 0) is available for CPU HPC instance types (hpc8a/hpc7a/hpc6a etc., 1-2 EFA interfaces). When enabled the LaunchTemplate switches to a NetworkInterfaces block with InterfaceType=efa, and a cluster placement group is auto-created (override with PlacementGroupName to reuse an @@ -41,7 +41,6 @@ Metadata: - Label: default: EFA (CPU HPC instances) Parameters: - - EnableEfa - EfaInterfaceCount - PlacementGroupName - Label: @@ -56,6 +55,12 @@ Metadata: Parameters: - PostInstallScriptUrl - PostInstallScriptArgs + - Label: + default: Multi-user directory (optional) + Parameters: + - DirectoryService + - DirectoryRole + - DirectoryDomainSuffix Parameters: ClusterId: @@ -189,8 +194,11 @@ Parameters: Description: >- Optional URL of a post-install script to download and run on each node at first boot (PCS equivalent of ParallelCluster OnNodeConfigured). Leave - empty to skip. To install Enroot/Pyxis, point this at scripts/install-enroot-pyxis.sh. - Must be an HTTP(S) URL (e.g. a GitHub raw URL); S3 public hosting is not allowed. + empty to skip. Accepts an s3:// URL (fetched with the instance role — works + with a private bucket) or an http(s):// URL (fetched with curl — public + only, e.g. a GitHub raw URL). To install Enroot/Pyxis, point this at + scripts/install-enroot-pyxis.sh in your templates bucket + (s3:///scripts/install-enroot-pyxis.sh). Default: '' PostInstallScriptArgs: @@ -210,38 +218,73 @@ Parameters: - '25.05' - '25.11' - EnableEfa: + DirectoryService: Type: String Description: >- - Enable Elastic Fabric Adapter (EFA) on the compute nodes for tightly-coupled - HPC/MPI workloads. When true, the LaunchTemplate switches to a - NetworkInterfaces block with InterfaceType=efa and a cluster placement group - is wired in (auto-created unless PlacementGroupName is supplied). Requires - an EFA-capable instance type (e.g. hpc8a.96xlarge, hpc7a.96xlarge, - hpc6a.48xlarge). For multi-NIC GPU families (P5/P6) use the - dedicated add-cng-p5/p6-* templates instead. - Default: 'false' + Multi-user directory service. 'none' = single ubuntu user (default, + unchanged). 'OpenLDAP-LoginNode' = deploy slapd on the login node (DB on shared + /home/ldap-db), compute nodes configure SSSD as LDAP clients. Future + options (SimpleAD, ManagedAD) will be added as AllowedValues. + Default: 'none' AllowedValues: - - 'true' - - 'false' + - 'none' + - 'OpenLDAP-LoginNode' + + DirectoryDomainSuffix: + Type: String + Description: >- + LDAP domain suffix (e.g. dc=cluster,dc=internal). Only used when + DirectoryService != none. + Default: 'dc=cluster,dc=internal' + + DirectoryRole: + Type: String + Description: >- + This CNG's role in the directory service. 'server' = install and run + the directory (slapd on login node). 'client' = configure SSSD to + resolve users from the directory server. 'none' = skip directory setup. + deploy-all sets this automatically (login=server, compute=client) when + DirectoryService != none. + Default: 'none' + AllowedValues: + - 'none' + - 'server' + - 'client' + + S3BucketName: + Type: String + Description: S3 bucket where scripts are stored (same as nested templates bucket) + Default: 'awsome-distributed-ai' + + S3KeyPrefix: + Type: String + Description: S3 key prefix for scripts (e.g. templates/) + Default: 'templates/' EfaInterfaceCount: Type: Number Description: >- - Number of EFA network interfaces to attach when EnableEfa=true. Use the - MaximumEfaInterfaces value reported by the instance type: hpc8a/hpc7a/hpc6id - = 2; hpc6a/c7i.metal = 1. Mismatching this with the instance type's - capability fails at launch. Ignored when EnableEfa=false. - Default: 1 - AllowedValues: [1, 2] + Number of Elastic Fabric Adapter (EFA) network interfaces on the compute + nodes. 0 (default) = no EFA (standard ENA networking). 1 or 2 = enable EFA + with that many interfaces: the LaunchTemplate switches to a NetworkInterfaces + block with InterfaceType=efa and a cluster placement group is wired in + (auto-created unless PlacementGroupName is supplied). Set the count to the + instance type's MaximumEfaInterfaces — hpc8a.96xlarge / hpc7a.* / + hpc6id.32xlarge = 2; hpc6a.48xlarge / c7i.metal = 1. EFA requires an + EFA-capable type; a non-EFA type (e.g. c6i.4xlarge) fails to launch with + count > 0. For multi-NIC GPU families (P5/P6) use the dedicated + add-cng-p5/p6-* templates instead. + Default: 0 + AllowedValues: [0, 1, 2] PlacementGroupName: Type: String Description: >- (Optional) Name of an existing cluster placement group to launch nodes into. Leave empty to auto-create a per-CNG cluster placement group when - EnableEfa=true. Specify a value to share one placement group across multiple - CNGs (e.g. a heterogeneous tightly-coupled job). Ignored when EnableEfa=false. + EfaInterfaceCount > 0. Specify a value to share one placement group across + multiple CNGs (e.g. a heterogeneous tightly-coupled job). Ignored when + EfaInterfaceCount = 0. Default: '' Conditions: @@ -252,7 +295,8 @@ Conditions: IsMonitoringRoleSet: !Not [!Equals [!Ref MonitoringRole, 'none']] HasExtraSecurityGroup: !Not [!Equals [!Ref ExtraSecurityGroupId, ""]] HasAmiId: !Not [!Equals [!Ref AmiId, ""]] - EfaEnabled: !Equals [!Ref EnableEfa, 'true'] + # EFA is enabled when the interface count is non-zero (0 = standard ENA). + EfaEnabled: !Not [!Equals [!Ref EfaInterfaceCount, 0]] HasUserPlacementGroup: !Not [!Equals [!Ref PlacementGroupName, ""]] # Auto-create a per-CNG cluster placement group only when EFA is on AND the # user did not provide one. Override PlacementGroupName to reuse an existing @@ -260,9 +304,8 @@ Conditions: CreatePlacementGroup: !And - !Condition EfaEnabled - !Not [!Condition HasUserPlacementGroup] - Has2ndEfaInterface: !And - - !Condition EfaEnabled - - !Equals [!Ref EfaInterfaceCount, 2] + Has2ndEfaInterface: !Equals [!Ref EfaInterfaceCount, 2] + DirectoryEnabled: !Not [!Equals [!Ref DirectoryRole, 'none']] Resources: @@ -349,6 +392,16 @@ Resources: - Key: monitoring-role Value: !Ref MonitoringRole - !Ref AWS::NoValue + # Directory (multi-user) role tag, independent of monitoring-role. + # Emitted for both server (login) and client (compute) nodes when + # DirectoryRole != none. Compute clients discover the OpenLDAP + # server by querying for directory-role=server (NOT monitoring-role, + # which is owned by the monitoring stack). See setup-directory.sh. + - !If + - DirectoryEnabled + - Key: directory-role + Value: !Ref DirectoryRole + - !Ref AWS::NoValue MetadataOptions: HttpEndpoint: enabled HttpPutResponseHopLimit: 4 @@ -413,10 +466,18 @@ Resources: runcmd: # Post-install script (optional) - generic OnNodeConfigured-style hook. # Downloads and runs a user-supplied script (e.g. Enroot/Pyxis install). + # Accepts BOTH an s3:// URL (fetched with `aws s3 cp` using the instance + # role — works with a PRIVATE bucket, the default for Enroot/Pyxis) and an + # http(s):// URL (fetched anonymously with curl — for public GitHub raw or + # any user-hosted script). Pick by scheme so dev accounts without public S3 + # still install Enroot/Pyxis from their own private templates bucket. - | if [ -n "${PostInstallScriptUrl}" ]; then echo "Downloading post-install script from ${PostInstallScriptUrl}..." | tee /var/log/pcs-post-install.log - curl -fsSL "${PostInstallScriptUrl}" -o /tmp/pcs-post-install.sh + case "${PostInstallScriptUrl}" in + s3://*) aws s3 cp "${PostInstallScriptUrl}" /tmp/pcs-post-install.sh >> /var/log/pcs-post-install.log 2>&1 ;; + *) curl -fsSL "${PostInstallScriptUrl}" -o /tmp/pcs-post-install.sh ;; + esac chmod +x /tmp/pcs-post-install.sh echo "Executing post-install script..." | tee -a /var/log/pcs-post-install.log # Tell the script this cluster's Slurm version (it can't discover it at @@ -470,6 +531,22 @@ Resources: done echo "Monitoring installation complete (exit $rc)" >> /var/log/monitoring-install.log fi + - | + # Multi-user directory setup (optional, DirectoryRole != none) + if [ "${DirectoryRole}" != "none" ]; then + export LDAP_DOMAIN_SUFFIX="${DirectoryDomainSuffix}" + export LDAP_DOMAIN=$(echo "${DirectoryDomainSuffix}" | sed 's/dc=//g; s/,/./g') + export LDAP_ADMIN_PASSWORD=$(openssl rand -base64 16) + export CLUSTER_ID="${ClusterId}" + export DIRECTORY_DNS_IPS="" + # Pass the script source so the server role can install the + # ldap-add-user.sh helper onto /usr/local/bin from the same bucket. + export S3_BUCKET="${S3BucketName}" + export S3_KEY_PREFIX="${S3KeyPrefix}" + aws s3 cp s3://${S3BucketName}/${S3KeyPrefix}scripts/setup-directory.sh /tmp/setup-directory.sh + chmod +x /tmp/setup-directory.sh + bash /tmp/setup-directory.sh "${DirectoryRole}" 2>&1 | tee /var/log/directory-setup.log + fi --==MYBOUNDARY== diff --git a/architectures/aws-pcs/assets/cluster-admin-iam.yaml b/architectures/aws-pcs/assets/cluster-admin-iam.yaml new file mode 100644 index 000000000..58af14d14 --- /dev/null +++ b/architectures/aws-pcs/assets/cluster-admin-iam.yaml @@ -0,0 +1,418 @@ +# Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. +# SPDX-License-Identifier: MIT-0 +AWSTemplateFormatVersion: "2010-09-09" +Description: | + Creates two customer-managed IAM policies that together grant the + permissions a deploying principal needs to create / update / delete an + AWS PCS reference cluster from the templates in architectures/aws-pcs/. + Optionally attaches them to an IAM Group whose members become PCS cluster + admins. The policy set is split into two because it exceeds the IAM + per-policy 6144-char limit: a "core" policy for everyday cluster CRUD, + and an optional "Image Builder" policy for the standalone DLAMI builder + template (pcs-ready-dlami-with-enroot-pyxis.yaml). +Metadata: + AWS::CloudFormation::Interface: + ParameterGroups: + - Label: { default: "Group + IAM users (optional)" } + Parameters: [GroupName, AttachUsers, AttachImageBuilderPolicy] + ParameterLabels: + GroupName: { default: "Name of the IAM group to create" } + AttachUsers: { default: "Existing IAM users to add to the group (comma-separated, optional)" } + AttachImageBuilderPolicy: { default: "Also attach the Image Builder admin policy" } +Parameters: + GroupName: + Type: String + Default: PCSClusterAdmins + Description: IAM group name for cluster administrators + AttachUsers: + Type: CommaDelimitedList + Default: "" + Description: Comma-separated list of existing IAM user names to add to the group. Leave empty to skip. + AttachImageBuilderPolicy: + Type: String + Default: "false" + AllowedValues: ["true", "false"] + Description: Attach the optional Image Builder admin policy (needed only when deploying pcs-ready-dlami-with-enroot-pyxis.yaml). +Conditions: + HasUsers: !Not [!Equals [!Join ["", !Ref AttachUsers], ""]] + AttachImageBuilder: !Equals [!Ref AttachImageBuilderPolicy, "true"] +Resources: + PCSClusterAdminPolicy: + Type: AWS::IAM::ManagedPolicy + Properties: + ManagedPolicyName: !Sub "${AWS::StackName}-PCSClusterAdmin-core" + Description: "Core CFN+EC2+FSx+PCS+IAM(scoped)+SSM+KMS+Secrets+Logs permissions for AWS PCS cluster CRUD" + PolicyDocument: + { + "Version": "2012-10-17", + "Statement": [ + { + "Sid": "CloudFormationStackLifecycle", + "Effect": "Allow", + "Action": [ + "cloudformation:CreateStack", + "cloudformation:UpdateStack", + "cloudformation:DeleteStack", + "cloudformation:CancelUpdateStack", + "cloudformation:ContinueUpdateRollback", + "cloudformation:CreateChangeSet", + "cloudformation:ExecuteChangeSet", + "cloudformation:DescribeChangeSet", + "cloudformation:DeleteChangeSet", + "cloudformation:DescribeStacks", + "cloudformation:DescribeStackEvents", + "cloudformation:DescribeStackResource", + "cloudformation:DescribeStackResources", + "cloudformation:ListStackResources", + "cloudformation:GetTemplate", + "cloudformation:GetTemplateSummary", + "cloudformation:TagResource", + "cloudformation:UntagResource" + ], + "Resource": "*" + }, + { + "Sid": "CloudFormationDescribeStackless", + "Effect": "Allow", + "Action": [ + "cloudformation:ValidateTemplate", + "cloudformation:ListStacks", + "cloudformation:ListExports", + "cloudformation:ListImports" + ], + "Resource": "*" + }, + { + "Sid": "EC2NetworkingAndComputeLifecycle", + "Effect": "Allow", + "Action": [ + "ec2:CreateVpc", + "ec2:DeleteVpc", + "ec2:ModifyVpcAttribute", + "ec2:AssociateVpcCidrBlock", + "ec2:DisassociateVpcCidrBlock", + "ec2:CreateSubnet", + "ec2:DeleteSubnet", + "ec2:ModifySubnetAttribute", + "ec2:CreateInternetGateway", + "ec2:DeleteInternetGateway", + "ec2:AttachInternetGateway", + "ec2:DetachInternetGateway", + "ec2:CreateNatGateway", + "ec2:DeleteNatGateway", + "ec2:AllocateAddress", + "ec2:ReleaseAddress", + "ec2:AssociateAddress", + "ec2:DisassociateAddress", + "ec2:CreateRouteTable", + "ec2:DeleteRouteTable", + "ec2:CreateRoute", + "ec2:DeleteRoute", + "ec2:ReplaceRoute", + "ec2:AssociateRouteTable", + "ec2:DisassociateRouteTable", + "ec2:CreateSecurityGroup", + "ec2:DeleteSecurityGroup", + "ec2:AuthorizeSecurityGroupIngress", + "ec2:AuthorizeSecurityGroupEgress", + "ec2:RevokeSecurityGroupIngress", + "ec2:RevokeSecurityGroupEgress", + "ec2:UpdateSecurityGroupRuleDescriptionsIngress", + "ec2:UpdateSecurityGroupRuleDescriptionsEgress", + "ec2:CreateVpcEndpoint", + "ec2:ModifyVpcEndpoint", + "ec2:DeleteVpcEndpoints", + "ec2:CreateLaunchTemplate", + "ec2:CreateLaunchTemplateVersion", + "ec2:ModifyLaunchTemplate", + "ec2:DeleteLaunchTemplate", + "ec2:DeleteLaunchTemplateVersions", + "ec2:CreatePlacementGroup", + "ec2:DeletePlacementGroup", + "ec2:CreateNetworkInterface", + "ec2:DeleteNetworkInterface", + "ec2:ModifyNetworkInterfaceAttribute", + "ec2:AttachNetworkInterface", + "ec2:DetachNetworkInterface", + "ec2:CreateTags", + "ec2:DeleteTags", + "ec2:RunInstances", + "ec2:CreateFleet" + ], + "Resource": "*" + }, + { + "Sid": "EC2DescribeForPCS", + "Effect": "Allow", + "Action": [ + "ec2:Describe*", + "ec2:GetSecurityGroupsForVpc" + ], + "Resource": "*" + }, + { + "Sid": "FSxLifecycle", + "Effect": "Allow", + "Action": [ + "fsx:CreateFileSystem", + "fsx:UpdateFileSystem", + "fsx:DeleteFileSystem", + "fsx:DescribeFileSystems", + "fsx:DescribeBackups", + "fsx:CreateBackup", + "fsx:DeleteBackup", + "fsx:TagResource", + "fsx:UntagResource", + "fsx:ListTagsForResource" + ], + "Resource": "*" + }, + { + "Sid": "PCSFullAccess", + "Effect": "Allow", + "Action": [ + "pcs:*" + ], + "Resource": "*" + }, + { + "Sid": "IAMRoleAndInstanceProfileLifecycle", + "Effect": "Allow", + "Action": [ + "iam:CreateRole", + "iam:DeleteRole", + "iam:GetRole", + "iam:UpdateAssumeRolePolicy", + "iam:UpdateRole", + "iam:UpdateRoleDescription", + "iam:PutRolePolicy", + "iam:DeleteRolePolicy", + "iam:GetRolePolicy", + "iam:ListRolePolicies", + "iam:AttachRolePolicy", + "iam:DetachRolePolicy", + "iam:ListAttachedRolePolicies", + "iam:CreateInstanceProfile", + "iam:DeleteInstanceProfile", + "iam:GetInstanceProfile", + "iam:ListInstanceProfilesForRole", + "iam:AddRoleToInstanceProfile", + "iam:RemoveRoleFromInstanceProfile", + "iam:TagRole", + "iam:UntagRole", + "iam:TagInstanceProfile", + "iam:UntagInstanceProfile" + ], + "Resource": [ + "arn:aws:iam::*:role/*PCS*", + "arn:aws:iam::*:role/*pcs*", + "arn:aws:iam::*:role/*ImageBuilder*", + "arn:aws:iam::*:instance-profile/*PCS*", + "arn:aws:iam::*:instance-profile/*pcs*", + "arn:aws:iam::*:instance-profile/*ImageBuilder*" + ] + }, + { + "Sid": "IAMPassRoleToPCSAndEC2", + "Effect": "Allow", + "Action": "iam:PassRole", + "Resource": [ + "arn:aws:iam::*:role/*PCS*", + "arn:aws:iam::*:role/*pcs*" + ], + "Condition": { + "StringEquals": { + "iam:PassedToService": [ + "ec2.amazonaws.com", + "pcs.amazonaws.com" + ] + } + } + }, + { + "Sid": "ServiceLinkedRolesOneTime", + "Effect": "Allow", + "Action": "iam:CreateServiceLinkedRole", + "Resource": "arn:aws:iam::*:role/aws-service-role/*", + "Condition": { + "StringLike": { + "iam:AWSServiceName": [ + "pcs.amazonaws.com", + "spot.amazonaws.com", + "spotfleet.amazonaws.com", + "fsx.amazonaws.com", + "s3.data-source.lustre.fsx.amazonaws.com", + "imagebuilder.amazonaws.com" + ] + } + } + }, + { + "Sid": "SSMParameterLifecycle", + "Effect": "Allow", + "Action": [ + "ssm:PutParameter", + "ssm:GetParameter", + "ssm:GetParameters", + "ssm:GetParametersByPath", + "ssm:DeleteParameter", + "ssm:DeleteParameters", + "ssm:DescribeParameters", + "ssm:AddTagsToResource", + "ssm:RemoveTagsFromResource", + "ssm:ListTagsForResource", + "ssm:LabelParameterVersion" + ], + "Resource": [ + "arn:aws:ssm:*:*:parameter/pcs/*", + "arn:aws:ssm:*:*:parameter/aws/service/*" + ] + }, + { + "Sid": "KMSAccessForDefaultEncryption", + "Effect": "Allow", + "Action": [ + "kms:Decrypt", + "kms:Encrypt", + "kms:GenerateDataKey", + "kms:GenerateDataKeyWithoutPlaintext", + "kms:CreateGrant", + "kms:DescribeKey" + ], + "Resource": "*" + }, + { + "Sid": "SecretsManagerForPCS", + "Effect": "Allow", + "Action": [ + "secretsmanager:CreateSecret", + "secretsmanager:UpdateSecret", + "secretsmanager:RotateSecret", + "secretsmanager:DescribeSecret", + "secretsmanager:DeleteSecret", + "secretsmanager:TagResource", + "secretsmanager:UntagResource" + ], + "Resource": "arn:aws:secretsmanager:*:*:secret:pcs!*" + }, + { + "Sid": "CloudWatchLogsForMonitoring", + "Effect": "Allow", + "Action": [ + "logs:CreateLogGroup", + "logs:DeleteLogGroup", + "logs:DescribeLogGroups", + "logs:PutRetentionPolicy", + "logs:TagResource", + "logs:UntagResource", + "logs:CreateDelivery", + "logs:DeleteDelivery", + "logs:GetDelivery", + "logs:PutDeliverySource", + "logs:DeleteDeliverySource", + "logs:PutDeliveryDestination", + "logs:DeleteDeliveryDestination" + ], + "Resource": "*" + }, + { + "Sid": "PCSVendedLogsDelivery", + "Effect": "Allow", + "Action": [ + "pcs:AllowVendedLogDeliveryForResource" + ], + "Resource": "*" + } + ] + } + PCSClusterAdminImageBuilderPolicy: + Type: AWS::IAM::ManagedPolicy + Condition: AttachImageBuilder + Properties: + ManagedPolicyName: !Sub "${AWS::StackName}-PCSClusterAdmin-imagebuilder" + Description: "Image Builder permissions for the optional standalone DLAMI builder template" + PolicyDocument: + { + "Version": "2012-10-17", + "Statement": [ + { + "Sid": "IAMPassRoleToImageBuilder", + "Effect": "Allow", + "Action": "iam:PassRole", + "Resource": "arn:aws:iam::*:role/*ImageBuilder*", + "Condition": { + "StringEquals": { + "iam:PassedToService": "imagebuilder.amazonaws.com" + } + } + }, + { + "Sid": "ImageBuilderLifecycleOptional", + "Effect": "Allow", + "Action": [ + "imagebuilder:CreateImagePipeline", + "imagebuilder:UpdateImagePipeline", + "imagebuilder:DeleteImagePipeline", + "imagebuilder:GetImagePipeline", + "imagebuilder:ListImagePipelines", + "imagebuilder:CreateImageRecipe", + "imagebuilder:DeleteImageRecipe", + "imagebuilder:GetImageRecipe", + "imagebuilder:ListImageRecipes", + "imagebuilder:CreateComponent", + "imagebuilder:DeleteComponent", + "imagebuilder:GetComponent", + "imagebuilder:ListComponents", + "imagebuilder:CreateInfrastructureConfiguration", + "imagebuilder:UpdateInfrastructureConfiguration", + "imagebuilder:DeleteInfrastructureConfiguration", + "imagebuilder:GetInfrastructureConfiguration", + "imagebuilder:ListInfrastructureConfigurations", + "imagebuilder:CreateDistributionConfiguration", + "imagebuilder:UpdateDistributionConfiguration", + "imagebuilder:DeleteDistributionConfiguration", + "imagebuilder:GetDistributionConfiguration", + "imagebuilder:ListDistributionConfigurations", + "imagebuilder:CreateImage", + "imagebuilder:DeleteImage", + "imagebuilder:GetImage", + "imagebuilder:ListImages", + "imagebuilder:StartImagePipelineExecution", + "imagebuilder:CancelImageCreation", + "imagebuilder:CreateLifecyclePolicy", + "imagebuilder:UpdateLifecyclePolicy", + "imagebuilder:DeleteLifecyclePolicy", + "imagebuilder:GetLifecyclePolicy", + "imagebuilder:ListLifecyclePolicies", + "imagebuilder:TagResource", + "imagebuilder:UntagResource", + "imagebuilder:ListTagsForResource" + ], + "Resource": "*" + } + ] + } + PCSClusterAdminGroup: + Type: AWS::IAM::Group + Properties: + GroupName: !Ref GroupName + ManagedPolicyArns: + - !Ref PCSClusterAdminPolicy + - !If [AttachImageBuilder, !Ref PCSClusterAdminImageBuilderPolicy, !Ref "AWS::NoValue"] + GroupMembership: + Type: AWS::IAM::UserToGroupAddition + Condition: HasUsers + Properties: + GroupName: !Ref PCSClusterAdminGroup + Users: !Ref AttachUsers +Outputs: + GroupName: + Description: IAM group with cluster-admin permissions + Value: !Ref PCSClusterAdminGroup + CorePolicyArn: + Description: ARN of the core admin managed policy + Value: !Ref PCSClusterAdminPolicy + ImageBuilderPolicyArn: + Condition: AttachImageBuilder + Description: ARN of the optional Image Builder admin managed policy + Value: !Ref PCSClusterAdminImageBuilderPolicy diff --git a/architectures/aws-pcs/assets/cluster-user-iam.yaml b/architectures/aws-pcs/assets/cluster-user-iam.yaml new file mode 100644 index 000000000..b02ca519f --- /dev/null +++ b/architectures/aws-pcs/assets/cluster-user-iam.yaml @@ -0,0 +1,139 @@ +# Copyright Amazon.com, Inc. or its affiliates. All Rights Reserved. +# SPDX-License-Identifier: MIT-0 +AWSTemplateFormatVersion: "2010-09-09" +Description: | + Creates a customer-managed IAM policy + group for end users of an + already-deployed AWS PCS reference cluster. Members can: + - Discover the login node (ec2:DescribeInstances) and read CFN stack outputs + - Read PCS cluster / queue / compute-node-group status (pcs:Get*/List*) + - Open SSM Session Manager / SSH-over-SSM / port-forward to LOGIN nodes + - Retrieve the Grafana admin password from Parameter Store + - Terminate their own SSM sessions + They cannot create, modify, or delete cluster resources, and cannot open + shells on compute nodes (the ssm:resourceTag Name=PCS-login* scope blocks that. Note: the Name tag is operator-mutable; if you rename login nodes, update the policy accordingly). +Metadata: + AWS::CloudFormation::Interface: + ParameterGroups: + - Label: { default: "Group + IAM users (optional)" } + Parameters: [GroupName, AttachUsers] + ParameterLabels: + GroupName: { default: "Name of the IAM group to create" } + AttachUsers: { default: "Existing IAM users to add to the group (comma-separated, optional)" } +Parameters: + GroupName: + Type: String + Default: PCSClusterUsers + Description: IAM group name for cluster end users + AttachUsers: + Type: CommaDelimitedList + Default: "" + Description: Comma-separated list of existing IAM user names to add to the group. Leave empty to skip. +Conditions: + HasUsers: !Not [!Equals [!Join ["", !Ref AttachUsers], ""]] +Resources: + PCSClusterUserPolicy: + Type: AWS::IAM::ManagedPolicy + Properties: + ManagedPolicyName: !Sub "${AWS::StackName}-PCSClusterUser" + Description: "Read-only + SSM Session Manager (login-node only) for end users of an already-deployed AWS PCS cluster" + PolicyDocument: + { + "Version": "2012-10-17", + "Statement": [ + { + "Sid": "DiscoverClusterAndLoginNode", + "Effect": "Allow", + "Action": [ + "ec2:DescribeInstances", + "ec2:DescribeInstanceStatus", + "ec2:DescribeTags", + "cloudformation:DescribeStacks", + "cloudformation:ListStacks", + "cloudformation:DescribeStackResources", + "cloudformation:ListStackResources" + ], + "Resource": "*" + }, + { + "Sid": "PCSReadOnly", + "Effect": "Allow", + "Action": [ + "pcs:GetCluster", + "pcs:ListClusters", + "pcs:GetComputeNodeGroup", + "pcs:ListComputeNodeGroups", + "pcs:GetQueue", + "pcs:ListQueues", + "pcs:ListTagsForResource" + ], + "Resource": "*" + }, + { + "Sid": "GrafanaPasswordRead", + "Effect": "Allow", + "Action": [ + "ssm:GetParameter", + "ssm:GetParameters" + ], + "Resource": "arn:aws:ssm:*:*:parameter/pcs/*/grafana/*" + }, + { + "Sid": "SSMSessionToLoginNode", + "Effect": "Allow", + "Action": [ + "ssm:StartSession" + ], + "Resource": [ + "arn:aws:ec2:*:*:instance/*", + "arn:aws:ssm:*:*:document/AWS-StartSSHSession", + "arn:aws:ssm:*:*:document/AWS-StartPortForwardingSession", + "arn:aws:ssm:*:*:document/AWS-StartPortForwardingSessionToRemoteHost", + "arn:aws:ssm:*:*:document/SSM-SessionManagerRunShell" + ], + "Condition": { + "StringLike": { + "ssm:resourceTag/Name": "PCS-login*" + } + } + }, + { + "Sid": "SSMSessionDocumentLookup", + "Effect": "Allow", + "Action": [ + "ssm:DescribeSessions", + "ssm:GetConnectionStatus", + "ssm:DescribeInstanceProperties", + "ssm:DescribeInstanceInformation" + ], + "Resource": "*" + }, + { + "Sid": "SSMSessionTerminateOwn", + "Effect": "Allow", + "Action": [ + "ssm:TerminateSession", + "ssm:ResumeSession" + ], + "Resource": "arn:aws:ssm:*:*:session/${aws:username}-*" + } + ] + } + PCSClusterUserGroup: + Type: AWS::IAM::Group + Properties: + GroupName: !Ref GroupName + ManagedPolicyArns: + - !Ref PCSClusterUserPolicy + GroupMembership: + Type: AWS::IAM::UserToGroupAddition + Condition: HasUsers + Properties: + GroupName: !Ref PCSClusterUserGroup + Users: !Ref AttachUsers +Outputs: + GroupName: + Description: IAM group with cluster-user (read-only + SSM-to-login) permissions + Value: !Ref PCSClusterUserGroup + PolicyArn: + Description: ARN of the user managed policy + Value: !Ref PCSClusterUserPolicy diff --git a/architectures/aws-pcs/assets/cluster.yaml b/architectures/aws-pcs/assets/cluster.yaml index f4d1dfad3..032830e1e 100644 --- a/architectures/aws-pcs/assets/cluster.yaml +++ b/architectures/aws-pcs/assets/cluster.yaml @@ -159,16 +159,55 @@ Resources: ManagedPolicyArns: - !Sub "arn:${AWS::Partition}:iam::aws:policy/AmazonSSMManagedInstanceCore" - !Sub "arn:${AWS::Partition}:iam::aws:policy/AmazonS3ReadOnlyAccess" + # Instance-role permissions are inline (PutRolePolicy), NOT a separate + # AWS::IAM::ManagedPolicy — so deploying this stack needs only + # iam:CreateRole + iam:PutRolePolicy (which the cluster-admin policy + # grants), never iam:CreatePolicy. Split by feature so monitoring and + # multi-user don't depend on each other: Policies: - - PolicyName: PcsRegisterInstancePolicy + # Always present (every cluster): register with PCS + read boot scripts + # from S3 + read this cluster's own SSM parameters (grafana password, + # ldap admin password) + decrypt SSM SecureStrings. None of this is + # monitoring- or directory-specific; it is the baseline an instance + # needs to register, fetch its post-install/directory scripts, and + # store/read the per-cluster secrets it generates. + - PolicyName: PcsInstanceBasePolicy PolicyDocument: Version: "2012-10-17" Statement: - - Effect: Allow + - Sid: PcsRegisterInstance + Effect: Allow Action: - pcs:RegisterComputeNodeGroupInstance Resource: "*" + - Sid: S3GetScripts + Effect: Allow + Action: + - s3:GetObject + Resource: + - arn:aws:s3:::awsome-distributed-ai/templates/scripts/* + - arn:aws:s3:::*/templates/scripts/* + + - Sid: SSMClusterParametersForThisClusterOnly + Effect: Allow + Action: + - ssm:GetParameter + - ssm:PutParameter + - ssm:AddTagsToResource + Resource: + - !Sub 'arn:${AWS::Partition}:ssm:${AWS::Region}:${AWS::AccountId}:parameter/pcs/${PCSCluster.Id}/grafana/*' + - !Sub 'arn:${AWS::Partition}:ssm:${AWS::Region}:${AWS::AccountId}:parameter/pcs/${PCSCluster.Id}/ldap/*' + + - Sid: KMSForSSMSecureString + Effect: Allow + Action: + - kms:Decrypt + Resource: '*' + Condition: + StringEquals: + kms:ViaService: !Sub 'ssm.${AWS::Region}.amazonaws.com' + # IAM Instance Profile PcsInstanceProfile: Type: AWS::IAM::InstanceProfile @@ -179,12 +218,17 @@ Resources: Roles: - !Ref PcsInstanceIamRole - # Monitoring IAM Policy (optional, attached to the instance role) - MonitoringPolicy: - Type: AWS::IAM::ManagedPolicy + # Monitoring instance permissions — inline policy added to the instance role + # only when monitoring is enabled. Inline (RoleName + PutRolePolicy via an + # AWS::IAM::Policy) instead of a ManagedPolicy, so the deployer never needs + # iam:CreatePolicy. Holds ONLY monitoring-specific reads (Prometheus service + # discovery + cost/CloudWatch metrics); multi-user/SSM/S3 perms live in the + # always-on base policy above. + MonitoringInstancePolicy: + Type: AWS::IAM::Policy Condition: IsMonitoringEnabled Properties: - ManagedPolicyName: !Sub '${PCSCluster.Id}-monitoring-policy' + PolicyName: !Sub '${PCSCluster.Id}-monitoring' Roles: - !Ref PcsInstanceIamRole PolicyDocument: @@ -225,23 +269,6 @@ Resources: - pricing:DescribeServices Resource: '*' - - Sid: SSMGrafanaPasswordForThisClusterOnly - Effect: Allow - Action: - - ssm:GetParameter - - ssm:PutParameter - - ssm:AddTagsToResource - Resource: !Sub 'arn:${AWS::Partition}:ssm:${AWS::Region}:${AWS::AccountId}:parameter/pcs/${PCSCluster.Id}/grafana/*' - - - Sid: KMSForSSMSecureString - Effect: Allow - Action: - - kms:Decrypt - Resource: '*' - Condition: - StringEquals: - kms:ViaService: !Sub 'ssm.${AWS::Region}.amazonaws.com' - - Sid: CloudWatchGetMetricData Effect: Allow Action: diff --git a/architectures/aws-pcs/assets/ml-cluster-prerequisites.yaml b/architectures/aws-pcs/assets/ml-cluster-prerequisites.yaml index 7d2719a1c..6b346430e 100644 --- a/architectures/aws-pcs/assets/ml-cluster-prerequisites.yaml +++ b/architectures/aws-pcs/assets/ml-cluster-prerequisites.yaml @@ -27,6 +27,8 @@ Metadata: default: Availability Zone configuration for the subnets Parameters: - PrimarySubnetAZ + - AdditionalSubnetAZ2 + - AdditionalSubnetAZ3 - Label: default: FSx for Lustre (/fsx) storage size Parameters: @@ -43,6 +45,10 @@ Metadata: default: Enable EFA + GDS on FSx for Lustre (PERSISTENT_2 only) PrimarySubnetAZ: default: Availability zone id to deploy the primary subnets + AdditionalSubnetAZ2: + default: "(Optional) 2nd AZ for an additional private subnet" + AdditionalSubnetAZ3: + default: "(Optional) 3rd AZ for an additional private subnet" CreateS3Endpoint: default: Create an S3 endpoint @@ -60,6 +66,23 @@ Parameters: Description: Availability zone id in which the public subnet and primary private subnet will be created. Type: AWS::EC2::AvailabilityZone::Name + AdditionalSubnetAZ2: + Description: >- + (Optional) Availability Zone for a 2nd private subnet. Empty (default) = + single-AZ. Set to create an additional private subnet for multi-AZ FSx + (OpenZFS MULTI_AZ), Simple AD, or higher-availability workloads. The NAT + gateway stays in the primary AZ (shared), so this subnet reaches the + internet via cross-AZ routing. + Type: String + Default: "" + + AdditionalSubnetAZ3: + Description: >- + (Optional) Availability Zone for a 3rd private subnet. Empty (default) = + not created. Requires AdditionalSubnetAZ2 to also be set. + Type: String + Default: "" + CreateS3Endpoint: AllowedValues: - 'true' @@ -120,7 +143,7 @@ Parameters: The headline feature is GPUDirect Storage (GDS) for P5/P5e/P5en/P6-B200 GPU clients, which DMAs file data straight into GPU memory (the NVIDIA nvidia-fs/cuFile stack must be installed on the client — currently a - ROADMAP follow-up). EFA-capable CPU CNGs (OnDemandEnableEfa=true) get + ROADMAP follow-up). EFA-capable CPU CNGs (OnDemandEfaInterfaceCount > 0) get the EFA transport path to storage as a secondary benefit, useful when a single client is pushing past ~10 GBps. **EFA-enabled Lustre has a much larger minimum storage capacity than the non-EFA default**: at @@ -224,6 +247,11 @@ Rules: Conditions: S3EndpointCondition: !Equals [!Ref 'CreateS3Endpoint', 'true'] + # Additional private subnets (multi-AZ). Each is created only when its AZ + # parameter is set. AZ3 is meant to be used together with AZ2 (a 3rd subnet + # without a 2nd is unusual but technically allowed — both are independent). + HasSubnetAZ2: !Not [!Equals [!Ref AdditionalSubnetAZ2, ""]] + HasSubnetAZ3: !Not [!Equals [!Ref AdditionalSubnetAZ3, ""]] # PERSISTENT_2 is the only Lustre deployment type that supports a metadata # configuration (and the EFA/Intelligent-Tiering features); PERSISTENT_1 does # not, so the MetadataConfiguration block is only emitted for PERSISTENT_2. @@ -256,7 +284,7 @@ Resources: CidrBlock: !FindInMap [Networking, VPC, CIDR0] Tags: - Key: Name - Value: ML Cluster VPC + Value: !Ref VPCName VpcCidrBlock: Type: AWS::EC2::VPCCidrBlock @@ -334,18 +362,46 @@ Resources: - Key: Name Value: !Join [ ' ', [ !Ref VPCName, 'Public Subnet -', !Ref PrimarySubnetAZ ] ] - # Create the primary private subnet + # Private subnets are carved from CIDR1 (10.1.0.0/16) split into four /18 + # blocks, so up to three private AZs (primary + 2 additional) can coexist + # without overlap. The primary uses index 0; additional subnets use 1 and 2. PrimaryPrivateSubnet: Type: AWS::EC2::Subnet DependsOn: [VpcCidrBlock] Properties: VpcId: !Ref VPC - CidrBlock: !Select [ 0, !Cidr [ !FindInMap [Networking, VPC, CIDR1], 2, 15 ]] + CidrBlock: !Select [ 0, !Cidr [ !FindInMap [Networking, VPC, CIDR1], 4, 14 ]] AvailabilityZone: !Ref PrimarySubnetAZ Tags: - Key: Name Value: !Join [ ' ', [ !Ref VPCName, 'Private Subnet -', !Ref PrimarySubnetAZ ] ] + # Additional private subnet 2 (optional, multi-AZ) + AdditionalPrivateSubnet2: + Type: AWS::EC2::Subnet + Condition: HasSubnetAZ2 + DependsOn: [VpcCidrBlock] + Properties: + VpcId: !Ref VPC + CidrBlock: !Select [ 1, !Cidr [ !FindInMap [Networking, VPC, CIDR1], 4, 14 ]] + AvailabilityZone: !Ref AdditionalSubnetAZ2 + Tags: + - Key: Name + Value: !Join [ ' ', [ !Ref VPCName, 'Private Subnet -', !Ref AdditionalSubnetAZ2 ] ] + + # Additional private subnet 3 (optional, multi-AZ) + AdditionalPrivateSubnet3: + Type: AWS::EC2::Subnet + Condition: HasSubnetAZ3 + DependsOn: [VpcCidrBlock] + Properties: + VpcId: !Ref VPC + CidrBlock: !Select [ 2, !Cidr [ !FindInMap [Networking, VPC, CIDR1], 4, 14 ]] + AvailabilityZone: !Ref AdditionalSubnetAZ3 + Tags: + - Key: Name + Value: !Join [ ' ', [ !Ref VPCName, 'Private Subnet -', !Ref AdditionalSubnetAZ3 ] ] + # Create and set the public route table PublicRouteTable: Type: AWS::EC2::RouteTable @@ -386,6 +442,23 @@ Resources: SubnetId: !Ref PrimaryPrivateSubnet RouteTableId: !Ref PrivateRouteTable + # Additional private subnets share the same private route table, so they reach + # the internet through the primary AZ's single NAT gateway (cross-AZ). This is + # intentional — no per-AZ NAT (cost/HA trade-off documented in the AZ params). + AdditionalPrivateSubnet2RTAssociation: + Type: AWS::EC2::SubnetRouteTableAssociation + Condition: HasSubnetAZ2 + Properties: + SubnetId: !Ref AdditionalPrivateSubnet2 + RouteTableId: !Ref PrivateRouteTable + + AdditionalPrivateSubnet3RTAssociation: + Type: AWS::EC2::SubnetRouteTableAssociation + Condition: HasSubnetAZ3 + Properties: + SubnetId: !Ref AdditionalPrivateSubnet3 + RouteTableId: !Ref PrivateRouteTable + # S3 endpoint S3Endpoint: Condition: S3EndpointCondition @@ -495,6 +568,18 @@ Outputs: Description: ID of the primary private subnet (alias for compatibility) Export: Name: !Sub ${AWS::StackName}-PrivateSubnet + AdditionalPrivateSubnet2: + Condition: HasSubnetAZ2 + Value: !Ref AdditionalPrivateSubnet2 + Description: ID of the 2nd additional private subnet (multi-AZ) + Export: + Name: !Sub ${AWS::StackName}-AdditionalPrivateSubnet2 + AdditionalPrivateSubnet3: + Condition: HasSubnetAZ3 + Value: !Ref AdditionalPrivateSubnet3 + Description: ID of the 3rd additional private subnet (multi-AZ) + Export: + Name: !Sub ${AWS::StackName}-AdditionalPrivateSubnet3 SecurityGroup: Value: !Ref SecurityGroup Description: SecurityGroup for Batch diff --git a/architectures/aws-pcs/assets/pcs-ml-cluster-deploy-all.yaml b/architectures/aws-pcs/assets/pcs-ml-cluster-deploy-all.yaml index a2d764dd9..a8e54e6e4 100644 --- a/architectures/aws-pcs/assets/pcs-ml-cluster-deploy-all.yaml +++ b/architectures/aws-pcs/assets/pcs-ml-cluster-deploy-all.yaml @@ -16,7 +16,8 @@ Metadata: default: "1. Network Configuration" Parameters: - PrimarySubnetAZ - - VPCName + - AdditionalSubnetAZ2 + - AdditionalSubnetAZ3 - CreateS3Endpoint - Label: default: "2. PCS Cluster Configuration" @@ -25,17 +26,11 @@ Metadata: - LoginNodeInstanceType - RootVolumeSize - AmiId - - DeployMonitoring - - GrafanaPublicAccessCidr + - SSHAccessCidr - ManagedAccounting - AccountingPolicyEnforcement - Label: - default: "3. Container Runtime (Post-install Script)" - Parameters: - - PostInstallScriptUrl - - PostInstallScriptArgs - - Label: - default: "4. On-Demand Compute Node Group" + default: "3. On-Demand Compute Node Group" Parameters: - DeployOnDemandCNG - OnDemandInstanceType @@ -43,11 +38,10 @@ Metadata: - OnDemandMaxCount - OnDemandCngName - OnDemandQueueName - - OnDemandEnableEfa - OnDemandEfaInterfaceCount - OnDemandPlacementGroupName - Label: - default: "5. GPU Compute Node Group - P5/P6 Series (Optional)" + default: "4. GPU Compute Node Group - P5/P6 Series (Optional)" Parameters: - DeployPseriesCNG - PseriesInstanceType @@ -57,7 +51,19 @@ Metadata: - PseriesCngName - PseriesQueueName - Label: - default: "6. FSx for Lustre (/fsx) + FSx for OpenZFS (/home) (Advanced)" + default: "5. Additional Cluster Configuration (Monitoring, Multi-User, Container Runtime)" + Parameters: + - MonitoringStack + - GrafanaAccessCidr + - MonitoringRepo + - MonitoringVersion + - DcgmExporterImage + - DirectoryService + - DirectoryDomainSuffix + - PostInstallScriptUrl + - PostInstallScriptArgs + - Label: + default: "6. FSx Storage (/fsx and /home)" Parameters: - Capacity - LustreDeploymentType @@ -69,36 +75,37 @@ Metadata: - HomeThroughput - OpenZFSDeploymentType - Label: - default: "7. Developer / Advanced (Nested Templates & Monitoring Source)" + default: "7. Developer / Advanced (do not change for normal use)" Parameters: - S3BucketName - S3KeyPrefix - - MonitoringVersion - - MonitoringRepo - - DcgmExporterImage ParameterLabels: # Friendly display names shown in the CloudFormation console 1-click UI, # so users see plain-language labels instead of raw CamelCase parameter keys. PrimarySubnetAZ: default: "Availability Zone (required)" - VPCName: - default: "VPC name" + AdditionalSubnetAZ2: + default: "(Optional) 2nd AZ for an additional private subnet" + AdditionalSubnetAZ3: + default: "(Optional) 3rd AZ for an additional private subnet" CreateS3Endpoint: default: "Create S3 VPC endpoint" LoginNodeInstanceType: default: "Login node instance type" SlurmVersion: default: "Slurm version" - DeployMonitoring: - default: "Deploy monitoring (Prometheus/Grafana)" - GrafanaPublicAccessCidr: - default: "Grafana public access CIDR (empty = SSM only)" + SSHAccessCidr: + default: "SSH access CIDR for login node (empty = SSM only)" + MonitoringStack: + default: "Monitoring stack (none or Prometheus-LoginNode)" + GrafanaAccessCidr: + default: "Grafana public access CIDR (empty = SSM port-forward only)" MonitoringVersion: default: "Monitoring version (git ref)" MonitoringRepo: default: "Monitoring repo (owner/repo)" DcgmExporterImage: - default: "dcgm-exporter image (default DCGM 4.5.2 by digest; covers H100/B200/B300)" + default: "dcgm-exporter image (default covers H100/B200/B300)" ManagedAccounting: default: "Slurm managed accounting" AccountingPolicyEnforcement: @@ -110,7 +117,7 @@ Metadata: RootVolumeSize: default: "Root volume size (GiB)" AmiId: - default: "AMI ID (empty = latest PCS-Ready Deep Learning AMI from SSM; pin in production)" + default: "AMI ID (empty = SSM auto-resolve; pin in production)" DeployOnDemandCNG: default: "Deploy CPU (on-demand) queue" OnDemandInstanceType: @@ -123,10 +130,8 @@ Metadata: default: "CPU node group name" OnDemandQueueName: default: "CPU queue name" - OnDemandEnableEfa: - default: "Enable EFA on CPU queue (for HPC instance types)" OnDemandEfaInterfaceCount: - default: "EFA interface count (1 for hpc6a, 2 for hpc8a/hpc7a/hpc6id)" + default: "EFA interfaces on CPU queue (0 = off, 1 or 2 for HPC types)" OnDemandPlacementGroupName: default: "Existing placement group name (empty = auto-create when EFA on)" DeployPseriesCNG: @@ -179,11 +184,6 @@ Parameters: Default: "templates/" # Network Configuration - VPCName: - Description: Name of your VPC - Default: 'ML-Cluster-VPC' - Type: String - PrimarySubnetAZ: Description: >- Availability Zone to deploy the cluster into. The only required choice -- @@ -191,6 +191,21 @@ Parameters: FSx) are both placed here. Type: AWS::EC2::AvailabilityZone::Name + AdditionalSubnetAZ2: + Description: >- + (Optional) 2nd Availability Zone for an additional private subnet. Empty + (default) = single-AZ. Enables multi-AZ FSx (OpenZFS MULTI_AZ) and future + multi-AZ workloads. The NAT gateway stays in the primary AZ. + Type: String + Default: "" + + AdditionalSubnetAZ3: + Description: >- + (Optional) 3rd Availability Zone for an additional private subnet. Empty + (default) = not created. Requires AdditionalSubnetAZ2 to also be set. + Type: String + Default: "" + CreateS3Endpoint: AllowedValues: - 'true' @@ -249,7 +264,7 @@ Parameters: headline feature is GPUDirect Storage (GDS) for P5/P5e/P5en/P6-B200 GPU clients, which DMAs file data straight into GPU memory (the NVIDIA nvidia-fs/cuFile stack must be installed on the client -- currently a - ROADMAP follow-up). EFA-capable CPU CNGs (OnDemandEnableEfa=true) get + ROADMAP follow-up). EFA-capable CPU CNGs (OnDemandEfaInterfaceCount > 0) get the EFA transport path to storage as a secondary benefit, useful when a single client is pushing past ~10 GBps. **EFA-enabled Lustre has a much higher minimum Capacity than non-EFA**: at PerUnitStorageThroughput=250 @@ -262,6 +277,29 @@ Parameters: - 'true' - 'false' + # Multi-user directory + DirectoryService: + Type: String + Description: >- + Multi-user directory service. 'none' = single ubuntu user (default, + unchanged). 'OpenLDAP-LoginNode' = deploy slapd on the login node (DB on shared + /home/ldap-db), compute nodes configure SSSD as LDAP clients. Future + options (SimpleAD, ManagedAD) will be added here. NOTE: OpenLDAP-LoginNode + runs the directory server on a SINGLE login node — keep the login node + group at 1 instance (do not scale it) while this is enabled. For + multi-login-node / HA directories, use the future SimpleAD/ManagedAD options. + Default: 'none' + AllowedValues: + - 'none' + - 'OpenLDAP-LoginNode' + + DirectoryDomainSuffix: + Type: String + Description: >- + LDAP domain suffix (e.g. dc=cluster,dc=internal). Only used when + DirectoryService != none. + Default: 'dc=cluster,dc=internal' + HomeCapacity: Description: "Home directories storage capacity in GiB" Type: Number @@ -311,13 +349,15 @@ Parameters: PostInstallScriptUrl: Type: String Description: >- - HTTP(S) URL of a script run on every node at first boot (PCS equivalent of - ParallelCluster OnNodeConfigured). Default installs Enroot/Pyxis. Set empty - to skip, or point at any other HTTP(S) script for custom first-boot setup. - Idempotent: a no-op if Enroot/Pyxis is already pre-baked into AmiId, so - leaving the default works whether or not you supply a custom AMI. - S3 hosting is not allowed. - Default: 'https://raw.githubusercontent.com/awslabs/awsome-distributed-ai/main/architectures/aws-pcs/scripts/install-enroot-pyxis.sh' + URL of a script run on every node at first boot (PCS equivalent of + ParallelCluster OnNodeConfigured). Accepts an s3:// URL (fetched with the + instance role — works with a private bucket) or an http(s):// URL (fetched + with curl — public only). Leave EMPTY (default) to auto-install Enroot/Pyxis + from the templates bucket (s3:///scripts/install-enroot-pyxis.sh), + so it works in dev accounts without public S3. Set to your own s3://or + https:// script for custom first-boot setup, or to a single space to skip. + Idempotent: a no-op if Enroot/Pyxis is already pre-baked into AmiId. + Default: '' PostInstallScriptArgs: Type: String @@ -355,22 +395,34 @@ Parameters: Default: "m6i.4xlarge" # Monitoring Configuration - DeployMonitoring: + SSHAccessCidr: Type: String - Description: Deploy monitoring stack (Prometheus, Grafana, DCGM) on Login Node - Default: 'true' + Description: >- + (Optional) CIDR allowed to reach the login node on SSH/22. Empty (default) + = SSH over SSM only. Set to your office IP or VPN range for direct SSH + access (multi-user clusters, VS Code Remote, scp). + Default: "" + AllowedPattern: '^$|^((25[0-5]|2[0-4]\d|1\d\d|[1-9]?\d)\.){3}(25[0-5]|2[0-4]\d|1\d\d|[1-9]?\d)/(3[0-2]|[12]?\d)$' + ConstraintDescription: Must be empty or a valid IPv4 CIDR (e.g. 203.0.113.4/32) + + MonitoringStack: + Type: String + Description: >- + Monitoring stack to deploy. 'Prometheus-LoginNode' (default) installs + Prometheus + Grafana + DCGM Exporter on the login node. 'none' skips + monitoring entirely. Future options (AMP-AMG, CloudWatch) will be added. + Default: 'Prometheus-LoginNode' AllowedValues: - - 'true' - - 'false' + - 'Prometheus-LoginNode' + - 'none' - GrafanaPublicAccessCidr: + GrafanaAccessCidr: Type: String Description: >- - (Optional) CIDR allowed to reach the login node on HTTPS/443. Empty (default) - = SSM port-forward only (recommended). When set, also exposes the - UNauthenticated /prometheus/, /pushgateway/, /slurmexporter/ proxy paths - (not just password-gated Grafana). Use the tightest CIDR you can; clear it - when done. + (Optional) CIDR allowed to reach the login node on HTTPS/443 (Grafana). + Empty (default) = SSM port-forward only. When set, also exposes the + unauthenticated /prometheus/, /pushgateway/, /slurmexporter/ proxy paths. + Use the tightest CIDR you can; clear it when done. Default: "" AllowedPattern: '^$|^((25[0-5]|2[0-4]\d|1\d\d|[1-9]?\d)\.){3}(25[0-5]|2[0-4]\d|1\d\d|[1-9]?\d)/(3[0-2]|[12]?\d)$' ConstraintDescription: Must be empty or a valid IPv4 CIDR (e.g. 203.0.113.4/32) @@ -439,37 +491,29 @@ Parameters: # EFA on the on-demand CPU CNG (for HPC instances like hpc8a/hpc7a/hpc6a). # GPU CNGs (P5/P6) have their own multi-NIC EFA wiring in the dedicated - # add-cng-p5/p6-* templates; these knobs do NOT affect those. - OnDemandEnableEfa: - Type: String - Description: >- - Enable Elastic Fabric Adapter (EFA) on the on-demand (CPU) compute node - group. Use with EFA-capable HPC instance types (hpc8a.96xlarge, - hpc7a.96xlarge, hpc6a.48xlarge, etc.). Default false keeps the - standard single-NIC LaunchTemplate. No effect on the P-series GPU CNG. - Default: 'false' - AllowedValues: - - 'true' - - 'false' - + # add-cng-p5/p6-* templates; this knob does NOT affect those. OnDemandEfaInterfaceCount: Type: Number Description: >- - Number of EFA interfaces to attach when OnDemandEnableEfa=true. Use 2 for - hpc8a/hpc7a/hpc6id; use 1 for hpc6a/c7i.metal. Mismatching this with - the instance type's MaximumEfaInterfaces fails at launch. Ignored when - OnDemandEnableEfa=false. - Default: 1 - AllowedValues: [1, 2] + Number of Elastic Fabric Adapter (EFA) interfaces on the on-demand (CPU) + compute node group. 0 (default) = no EFA (standard ENA networking). 1 or 2 + = enable EFA with that many interfaces (switches to a NetworkInterfaces + LaunchTemplate + a cluster placement group). Set the count to the instance + type's MaximumEfaInterfaces: hpc8a.96xlarge / hpc7a.* / hpc6id.32xlarge = 2; + hpc6a.48xlarge / c7i.metal = 1. EFA requires an EFA-capable type — a non-EFA + type (e.g. the default c6i.4xlarge) fails to launch with count > 0. No + effect on the P-series GPU CNG. + Default: 0 + AllowedValues: [0, 1, 2] OnDemandPlacementGroupName: Type: String Description: >- (Optional, advanced) Existing cluster placement group name to launch the on-demand CPU nodes into. Leave empty (default) to auto-create a per-CNG - cluster placement group when OnDemandEnableEfa=true. Set to share one PG - across multiple CNGs (e.g. heterogeneous tightly-coupled jobs). Ignored - when OnDemandEnableEfa=false. + cluster placement group when OnDemandEfaInterfaceCount > 0. Set to share one + PG across multiple CNGs (e.g. heterogeneous tightly-coupled jobs). Ignored + when OnDemandEfaInterfaceCount = 0. Default: '' # P-series Compute Node Group @@ -586,10 +630,21 @@ Rules: Conditions: UseLocalTemplates: !Equals [!Ref S3BucketName, 'local'] + # When PostInstallScriptUrl is left empty, default it to the Enroot/Pyxis + # installer in the templates bucket (s3://, fetched via the instance role — + # works with a private bucket). A non-empty value (s3:// or https://) is used + # verbatim; a single space disables the post-install hook. + PostInstallScriptUrlIsDefault: !Equals [!Ref PostInstallScriptUrl, ''] DeployOnDemandCNGEnabled: !Equals [!Ref DeployOnDemandCNG, 'true'] + DirectoryEnabled: !Not [!Equals [!Ref DirectoryService, 'none']] DeployPseriesCNGEnabled: !Equals [!Ref DeployPseriesCNG, 'true'] # Open Grafana (login node, HTTPS/443) to a CIDR only when one is provided. - GrafanaPublicAccess: !Not [!Equals [!Ref GrafanaPublicAccessCidr, ""]] + SSHAccessEnabled: !Not [!Equals [!Ref SSHAccessCidr, ""]] + GrafanaAccessEnabled: !Not [!Equals [!Ref GrafanaAccessCidr, ""]] + LoginAccessEnabled: !Or + - !Condition SSHAccessEnabled + - !Condition GrafanaAccessEnabled + MonitoringEnabled: !Not [!Equals [!Ref MonitoringStack, 'none']] # Select the right multi-NIC EFA template from the P-series instance type. # p6-b200/p6-b300 have different network-card counts than P5, so each family # gets its own nested stack; anything else falls back to the P5 template. @@ -616,8 +671,10 @@ Resources: - ml-cluster-prerequisites.yaml - !Sub 'https://${S3BucketName}.s3.amazonaws.com/${S3KeyPrefix}ml-cluster-prerequisites.yaml' Parameters: - VPCName: !Ref VPCName + VPCName: !Sub '${AWS::StackName}-VPC' PrimarySubnetAZ: !Ref PrimarySubnetAZ + AdditionalSubnetAZ2: !Ref AdditionalSubnetAZ2 + AdditionalSubnetAZ3: !Ref AdditionalSubnetAZ3 CreateS3Endpoint: !Ref CreateS3Endpoint Capacity: !Ref Capacity LustreDeploymentType: !Ref LustreDeploymentType @@ -649,7 +706,7 @@ Resources: SlurmVersion: !Ref SlurmVersion ManagedAccounting: !Ref ManagedAccounting AccountingPolicyEnforcement: !Ref AccountingPolicyEnforcement - DeployMonitoring: !Ref DeployMonitoring + DeployMonitoring: !If [MonitoringEnabled, "true", "false"] Tags: - Key: Name Value: !Sub '${AWS::StackName}-Cluster' @@ -657,25 +714,36 @@ Resources: ############################################################################# # 4. Login Node Group Stack ############################################################################# - # Optional login-only security group that exposes Grafana (HTTPS/443) to a - # specific CIDR. Created only when GrafanaPublicAccessCidr is set. It is - # attached ONLY to the login node (via ExtraSecurityGroupId), so compute nodes - # and FSx — which share the cluster security group — are not exposed. - GrafanaPublicSecurityGroup: + # Optional login-only security group for SSH and/or Grafana access from a + # specific CIDR. Created when SSHAccessCidr and/or GrafanaAccessCidr is set. + # Attached ONLY to the login node (via ExtraSecurityGroupId) — compute nodes + # and FSx (which share the cluster security group) are NOT exposed. + LoginAccessSecurityGroup: Type: AWS::EC2::SecurityGroup - Condition: GrafanaPublicAccess + Condition: LoginAccessEnabled Properties: - GroupDescription: Allow inbound HTTPS to Grafana on the login node + GroupDescription: Allow inbound SSH/HTTPS to the login node from trusted CIDRs VpcId: !GetAtt PrerequisitesStack.Outputs.VPC SecurityGroupIngress: - - IpProtocol: tcp - FromPort: 443 - ToPort: 443 - CidrIp: !Ref GrafanaPublicAccessCidr - Description: Grafana HTTPS from the allowed CIDR + - !If + - SSHAccessEnabled + - IpProtocol: tcp + FromPort: 22 + ToPort: 22 + CidrIp: !Ref SSHAccessCidr + Description: SSH from SSHAccessCidr + - !Ref AWS::NoValue + - !If + - GrafanaAccessEnabled + - IpProtocol: tcp + FromPort: 443 + ToPort: 443 + CidrIp: !Ref GrafanaAccessCidr + Description: Grafana HTTPS from GrafanaAccessCidr + - !Ref AWS::NoValue Tags: - Key: Name - Value: !Sub '${AWS::StackName}-grafana-public' + Value: !Sub '${AWS::StackName}-login-access' LoginNodeGroupStack: Type: AWS::CloudFormation::Stack @@ -698,21 +766,29 @@ Resources: AmiId: !Ref AmiId ClusterSecurityGroupId: !GetAtt PrerequisitesStack.Outputs.SecurityGroup ExtraSecurityGroupId: !If - - GrafanaPublicAccess - - !Ref GrafanaPublicSecurityGroup + - LoginAccessEnabled + - !Ref LoginAccessSecurityGroup - "" IamProfileArn: !GetAtt ClusterStack.Outputs.InstanceProfileArn FSxLustreFilesystemId: !GetAtt PrerequisitesStack.Outputs.FSxLustreFilesystemId FSxLustreFilesystemMountName: !GetAtt PrerequisitesStack.Outputs.FSxLustreFilesystemMountname FSxOpenZFSFilesystemId: !GetAtt PrerequisitesStack.Outputs.FSxOFilesystemId - DeployMonitoring: !Ref DeployMonitoring + DeployMonitoring: !If [MonitoringEnabled, "true", "false"] MonitoringVersion: !Ref MonitoringVersion MonitoringRepo: !Ref MonitoringRepo DcgmExporterImage: !Ref DcgmExporterImage SlurmVersion: !Ref SlurmVersion MonitoringRole: login - PostInstallScriptUrl: !Ref PostInstallScriptUrl + PostInstallScriptUrl: !If + - PostInstallScriptUrlIsDefault + - !Sub 's3://${S3BucketName}/${S3KeyPrefix}scripts/install-enroot-pyxis.sh' + - !Ref PostInstallScriptUrl PostInstallScriptArgs: !Ref PostInstallScriptArgs + DirectoryService: !Ref DirectoryService + DirectoryRole: !If [DirectoryEnabled, 'server', 'none'] + DirectoryDomainSuffix: !Ref DirectoryDomainSuffix + S3BucketName: !Ref S3BucketName + S3KeyPrefix: !Ref S3KeyPrefix Tags: - Key: Name Value: !Sub '${AWS::StackName}-LoginNodeGroup' @@ -745,20 +821,27 @@ Resources: FSxLustreFilesystemId: !GetAtt PrerequisitesStack.Outputs.FSxLustreFilesystemId FSxLustreFilesystemMountName: !GetAtt PrerequisitesStack.Outputs.FSxLustreFilesystemMountname FSxOpenZFSFilesystemId: !GetAtt PrerequisitesStack.Outputs.FSxOFilesystemId - DeployMonitoring: !Ref DeployMonitoring + DeployMonitoring: !If [MonitoringEnabled, "true", "false"] MonitoringVersion: !Ref MonitoringVersion MonitoringRepo: !Ref MonitoringRepo DcgmExporterImage: !Ref DcgmExporterImage SlurmVersion: !Ref SlurmVersion MonitoringRole: compute - PostInstallScriptUrl: !Ref PostInstallScriptUrl + PostInstallScriptUrl: !If + - PostInstallScriptUrlIsDefault + - !Sub 's3://${S3BucketName}/${S3KeyPrefix}scripts/install-enroot-pyxis.sh' + - !Ref PostInstallScriptUrl PostInstallScriptArgs: !Ref PostInstallScriptArgs - # EFA on CPU HPC instances (hpc8a/hpc7a/hpc6a). Forwarded to add-cng.yaml, - # which auto-creates a per-CNG cluster placement group when EnableEfa=true and - # PlacementGroupName is empty. - EnableEfa: !Ref OnDemandEnableEfa + # EFA on CPU HPC instances (hpc8a/hpc7a/hpc6a). add-cng.yaml enables EFA + # when EfaInterfaceCount > 0 and auto-creates a per-CNG cluster placement + # group when PlacementGroupName is empty. EfaInterfaceCount: !Ref OnDemandEfaInterfaceCount PlacementGroupName: !Ref OnDemandPlacementGroupName + DirectoryService: !Ref DirectoryService + DirectoryRole: !If [DirectoryEnabled, 'client', 'none'] + DirectoryDomainSuffix: !Ref DirectoryDomainSuffix + S3BucketName: !Ref S3BucketName + S3KeyPrefix: !Ref S3KeyPrefix Tags: - Key: Name Value: !Sub '${AWS::StackName}-OnDemandCNG' @@ -792,14 +875,22 @@ Resources: FSxLustreFilesystemMountName: !GetAtt PrerequisitesStack.Outputs.FSxLustreFilesystemMountname FSxOpenZFSFilesystemId: !GetAtt PrerequisitesStack.Outputs.FSxOFilesystemId CapacityReservationId: !Ref CapacityReservationId - DeployMonitoring: !Ref DeployMonitoring + DeployMonitoring: !If [MonitoringEnabled, "true", "false"] MonitoringVersion: !Ref MonitoringVersion MonitoringRepo: !Ref MonitoringRepo DcgmExporterImage: !Ref DcgmExporterImage SlurmVersion: !Ref SlurmVersion MonitoringRole: compute - PostInstallScriptUrl: !Ref PostInstallScriptUrl + PostInstallScriptUrl: !If + - PostInstallScriptUrlIsDefault + - !Sub 's3://${S3BucketName}/${S3KeyPrefix}scripts/install-enroot-pyxis.sh' + - !Ref PostInstallScriptUrl PostInstallScriptArgs: !Ref PostInstallScriptArgs + DirectoryService: !Ref DirectoryService + DirectoryRole: !If [DirectoryEnabled, 'client', 'none'] + DirectoryDomainSuffix: !Ref DirectoryDomainSuffix + S3BucketName: !Ref S3BucketName + S3KeyPrefix: !Ref S3KeyPrefix Tags: - Key: Name Value: !Sub '${AWS::StackName}-PseriesCNG' @@ -833,14 +924,22 @@ Resources: FSxLustreFilesystemMountName: !GetAtt PrerequisitesStack.Outputs.FSxLustreFilesystemMountname FSxOpenZFSFilesystemId: !GetAtt PrerequisitesStack.Outputs.FSxOFilesystemId CapacityReservationId: !Ref CapacityReservationId - DeployMonitoring: !Ref DeployMonitoring + DeployMonitoring: !If [MonitoringEnabled, "true", "false"] MonitoringVersion: !Ref MonitoringVersion MonitoringRepo: !Ref MonitoringRepo DcgmExporterImage: !Ref DcgmExporterImage SlurmVersion: !Ref SlurmVersion MonitoringRole: compute - PostInstallScriptUrl: !Ref PostInstallScriptUrl + PostInstallScriptUrl: !If + - PostInstallScriptUrlIsDefault + - !Sub 's3://${S3BucketName}/${S3KeyPrefix}scripts/install-enroot-pyxis.sh' + - !Ref PostInstallScriptUrl PostInstallScriptArgs: !Ref PostInstallScriptArgs + DirectoryService: !Ref DirectoryService + DirectoryRole: !If [DirectoryEnabled, 'client', 'none'] + DirectoryDomainSuffix: !Ref DirectoryDomainSuffix + S3BucketName: !Ref S3BucketName + S3KeyPrefix: !Ref S3KeyPrefix Tags: - Key: Name Value: !Sub '${AWS::StackName}-PseriesCNG' @@ -874,14 +973,22 @@ Resources: FSxLustreFilesystemMountName: !GetAtt PrerequisitesStack.Outputs.FSxLustreFilesystemMountname FSxOpenZFSFilesystemId: !GetAtt PrerequisitesStack.Outputs.FSxOFilesystemId CapacityReservationId: !Ref CapacityReservationId - DeployMonitoring: !Ref DeployMonitoring + DeployMonitoring: !If [MonitoringEnabled, "true", "false"] MonitoringVersion: !Ref MonitoringVersion MonitoringRepo: !Ref MonitoringRepo DcgmExporterImage: !Ref DcgmExporterImage SlurmVersion: !Ref SlurmVersion MonitoringRole: compute - PostInstallScriptUrl: !Ref PostInstallScriptUrl + PostInstallScriptUrl: !If + - PostInstallScriptUrlIsDefault + - !Sub 's3://${S3BucketName}/${S3KeyPrefix}scripts/install-enroot-pyxis.sh' + - !Ref PostInstallScriptUrl PostInstallScriptArgs: !Ref PostInstallScriptArgs + DirectoryService: !Ref DirectoryService + DirectoryRole: !If [DirectoryEnabled, 'client', 'none'] + DirectoryDomainSuffix: !Ref DirectoryDomainSuffix + S3BucketName: !Ref S3BucketName + S3KeyPrefix: !Ref S3KeyPrefix Tags: - Key: Name Value: !Sub '${AWS::StackName}-PseriesCNG' diff --git a/architectures/aws-pcs/scripts/install-enroot-pyxis.sh b/architectures/aws-pcs/assets/scripts/install-enroot-pyxis.sh similarity index 100% rename from architectures/aws-pcs/scripts/install-enroot-pyxis.sh rename to architectures/aws-pcs/assets/scripts/install-enroot-pyxis.sh diff --git a/architectures/aws-pcs/assets/scripts/ldap-add-user.sh b/architectures/aws-pcs/assets/scripts/ldap-add-user.sh new file mode 100755 index 000000000..6843a5817 --- /dev/null +++ b/architectures/aws-pcs/assets/scripts/ldap-add-user.sh @@ -0,0 +1,70 @@ +#!/bin/bash +# ldap-add-user.sh — Helper to add a POSIX user to the OpenLDAP directory. +# Run on the login node as root (or with LDAP admin credentials). +# +# Usage: ./ldap-add-user.sh [uid] [gid] [ssh-pub-key] +# +# Example: +# ./ldap-add-user.sh alice 10001 3000 +# ./ldap-add-user.sh bob 10002 3000 "ssh-rsa AAAA..." + +set -euo pipefail + +USERNAME="${1:?Usage: $0 [uid] [gid] [ssh-pub-key]}" +USER_UID="${2:-$((10000 + RANDOM % 50000))}" +USER_GID="${3:-3000}" +SSH_PUBKEY="${4:-}" + +# Auto-detect LDAP config from sssd.conf or environment +LDAP_DOMAIN_SUFFIX="${LDAP_DOMAIN_SUFFIX:-$(grep ldap_search_base /etc/sssd/sssd.conf 2>/dev/null | awk -F= '{print $2}' | tr -d ' ' || echo 'dc=cluster,dc=internal')}" +LDAP_ADMIN_DN="cn=admin,${LDAP_DOMAIN_SUFFIX}" + +echo "Adding user: ${USERNAME} (uid=${USER_UID}, gid=${USER_GID})" +echo "LDAP base: ${LDAP_DOMAIN_SUFFIX}" + +# Prompt for admin password if not set +if [ -z "${LDAP_ADMIN_PASSWORD:-}" ]; then + read -sp "LDAP admin password: " LDAP_ADMIN_PASSWORD + echo +fi + +# Create user entry +LDIF=$(cat <&1 + +# Set a random initial password (user should change via ldappasswd) +INITIAL_PW=$(openssl rand -base64 12) +ldappasswd -x -H ldap://localhost -D "${LDAP_ADMIN_DN}" -w "${LDAP_ADMIN_PASSWORD}" \ + -s "${INITIAL_PW}" "uid=${USERNAME},ou=People,${LDAP_DOMAIN_SUFFIX}" + +echo "" +echo "User '${USERNAME}' created successfully." +echo " UID: ${USER_UID}" +echo " GID: ${USER_GID}" +echo " Home: /home/${USERNAME} (auto-created on first login via pam_mkhomedir)" +echo " Initial password: ${INITIAL_PW}" +echo "" +echo "To add to Slurm accounting:" +echo " sacctmgr add user ${USERNAME} account=default" diff --git a/architectures/aws-pcs/assets/scripts/setup-directory.sh b/architectures/aws-pcs/assets/scripts/setup-directory.sh new file mode 100755 index 000000000..f7850a72c --- /dev/null +++ b/architectures/aws-pcs/assets/scripts/setup-directory.sh @@ -0,0 +1,319 @@ +#!/bin/bash +# setup-directory.sh — Multi-user directory setup for PCS reference architecture. +# +# Single script for both login node (LDAP server) and compute nodes (SSSD client). +# Called from CNG UserData when DirectoryService != "none". +# +# Usage: setup-directory.sh +# role = "server" (login node — installs slapd, creates base OUs) +# "client" (compute node — installs SSSD, points at LDAP server) +# +# IMPORTANT — single login node only when DirectoryService=OpenLDAP-LoginNode: +# The OpenLDAP server runs on THE login node and its DB lives on shared +# /home (OpenZFS). This design assumes exactly ONE login node (server). Do +# NOT scale the login node group above MinCount=MaxCount=1 with the directory +# enabled: multiple login nodes would each tag themselves directory-role=server +# (clients would pick an arbitrary one) and each run a slapd against the same +# /home/ldap-db MDB files (concurrent-open corruption). For HA multi-server +# directory, use a managed backend (Simple AD / Managed AD) instead — that is +# the planned DirectoryService=SimpleAD / ManagedAD extension path. +# +# Environment variables (from UserData): +# LDAP_DOMAIN_SUFFIX — e.g. "dc=cluster,dc=internal" +# LDAP_DOMAIN — e.g. "cluster.internal" (derived from suffix) +# LDAP_ADMIN_PASSWORD — auto-generated by UserData (server role only) +# CLUSTER_ID — PCS cluster ID (for SSM parameter path) +# LDAP_SERVER_URI — e.g. "ldap://10.1.x.x" (client role; empty = discover) +# DIRECTORY_DNS_IPS — comma-separated DNS IPs (future SimpleAD; empty = discover login IP) +# S3_BUCKET — bucket holding the scripts (server role; to install the +# ldap-add-user.sh helper onto /usr/local/bin) +# S3_KEY_PREFIX — key prefix for the scripts (default "templates/") + +set -euo pipefail + +ROLE="${1:-client}" +LDAP_DOMAIN_SUFFIX="${LDAP_DOMAIN_SUFFIX:-dc=cluster,dc=internal}" +LDAP_DOMAIN="${LDAP_DOMAIN:-cluster.internal}" +LDAP_DB_DIR="/home/ldap-db" + +export DEBIAN_FRONTEND=noninteractive + +# AWS CLI calls below (ssm put-parameter, ec2 describe-instances) pass no +# --region: on EC2 the CLI resolves both credentials and region from the +# instance's IMDS credential provider (it obtains the IMDSv2 token itself, even +# under HttpTokens=required). Reading placement/region with a plain curl would +# need a manual token and return empty without one, so we deliberately let the +# CLI handle it instead of passing a (possibly empty) --region. + +echo "[directory] Role: ${ROLE}" +echo "[directory] Domain: ${LDAP_DOMAIN} (${LDAP_DOMAIN_SUFFIX})" + +############################################################################### +# apt helpers — at first boot cloud-init's own unattended-upgrades/apt run can +# still hold the dpkg lock, so a bare `apt-get install` fails with +# "Could not get lock /var/lib/dpkg/lock-frontend". Wait for the lock to clear, +# then run apt-get with built-in retries. Mirrors install-enroot-pyxis.sh. +############################################################################### +wait_for_apt_lock() { + local waited=0 + while fuser /var/lib/dpkg/lock-frontend >/dev/null 2>&1 \ + || fuser /var/lib/dpkg/lock >/dev/null 2>&1 \ + || fuser /var/lib/apt/lists/lock >/dev/null 2>&1; do + if [ "$waited" -ge 300 ]; then + echo "[directory] WARNING: dpkg lock still held after ${waited}s; proceeding anyway." + break + fi + echo "[directory] Waiting for apt/dpkg lock to clear (${waited}s)..." + sleep 10 + waited=$((waited + 10)) + done +} + +apt_get() { + wait_for_apt_lock + local attempt + for attempt in 1 2 3; do + if DEBIAN_FRONTEND=noninteractive apt-get "$@"; then + return 0 + fi + echo "[directory] apt-get $* failed (attempt ${attempt}/3); waiting for lock and retrying..." + wait_for_apt_lock + sleep 5 + done + echo "[directory] ERROR: apt-get $* failed after 3 attempts." + return 1 +} + +############################################################################### +# SERVER role — login node: install slapd + configure +############################################################################### +setup_server() { + LDAP_ADMIN_PASSWORD="${LDAP_ADMIN_PASSWORD:?LDAP_ADMIN_PASSWORD must be set}" + CLUSTER_ID="${CLUSTER_ID:-unknown}" + + echo "[directory-server] Running apt-get update..." + apt_get update -qq + + echo "[directory-server] Installing slapd + ldap-utils..." + debconf-set-selections </dev/null)" ]; then + cp -a /var/lib/ldap/* "${LDAP_DB_DIR}/" + fi + chown -R openldap:openldap "${LDAP_DB_DIR}" + fi + + # Update slapd DB directory in cn=config + if [ -d /etc/ldap/slapd.d ]; then + MDB_LDIF=$(find /etc/ldap/slapd.d -name "olcDatabase*mdb*" -o -name "olcDatabase*hdb*" | head -1) + if [ -n "$MDB_LDIF" ] && grep -q "olcDbDirectory" "$MDB_LDIF"; then + sed -i "s|olcDbDirectory:.*|olcDbDirectory: ${LDAP_DB_DIR}|" "$MDB_LDIF" + fi + fi + + # AppArmor: allow slapd to access /home/ldap-db + if [ -f /etc/apparmor.d/usr.sbin.slapd ]; then + if ! grep -q "${LDAP_DB_DIR}" /etc/apparmor.d/usr.sbin.slapd; then + sed -i "/\/var\/lib\/ldap\/ r,/a\\ ${LDAP_DB_DIR}/ r," /etc/apparmor.d/usr.sbin.slapd + sed -i "/\/var\/lib\/ldap\/\*\* rwk,/a\\ ${LDAP_DB_DIR}/** rwk," /etc/apparmor.d/usr.sbin.slapd + apparmor_parser -r /etc/apparmor.d/usr.sbin.slapd 2>/dev/null || true + fi + fi + + chown -R openldap:openldap "${LDAP_DB_DIR}" + systemctl start slapd + systemctl enable slapd + + # Wait for slapd + for i in $(seq 1 10); do + ldapsearch -x -H ldap://localhost -b "" -s base namingContexts >/dev/null 2>&1 && break + sleep 1 + done + + echo "[directory-server] Creating base OUs..." + ldapadd -x -H ldap://localhost -D "cn=admin,${LDAP_DOMAIN_SUFFIX}" -w "${LDAP_ADMIN_PASSWORD}" </dev/null || true +dn: ou=People,${LDAP_DOMAIN_SUFFIX} +objectClass: organizationalUnit +ou: People +EOF + ldapadd -x -H ldap://localhost -D "cn=admin,${LDAP_DOMAIN_SUFFIX}" -w "${LDAP_ADMIN_PASSWORD}" </dev/null || true +dn: ou=Groups,${LDAP_DOMAIN_SUFFIX} +objectClass: organizationalUnit +ou: Groups +EOF + # Default group + ldapadd -x -H ldap://localhost -D "cn=admin,${LDAP_DOMAIN_SUFFIX}" -w "${LDAP_ADMIN_PASSWORD}" </dev/null || true +dn: cn=clusterusers,ou=Groups,${LDAP_DOMAIN_SUFFIX} +objectClass: posixGroup +cn: clusterusers +gidNumber: 3000 +EOF + + # Persist admin password — prefer SSM (access-controlled, auditable). + # No explicit --region: the AWS CLI resolves it from the instance's IMDS + # credential provider (reading placement/region directly would need an + # IMDSv2 token). + if aws ssm put-parameter \ + --name "/pcs/${CLUSTER_ID}/ldap/admin-password" \ + --value "${LDAP_ADMIN_PASSWORD}" \ + --type SecureString \ + --overwrite 2>/dev/null; then + echo "[directory-server] Admin password stored in SSM: /pcs/${CLUSTER_ID}/ldap/admin-password" + else + echo "${LDAP_ADMIN_PASSWORD}" > "${LDAP_DB_DIR}/.admin-password" + chmod 600 "${LDAP_DB_DIR}/.admin-password" + chown openldap:openldap "${LDAP_DB_DIR}/.admin-password" + echo "[directory-server] WARNING: SSM put failed. Password saved to ${LDAP_DB_DIR}/.admin-password" + fi + + # Install the ldap-add-user helper alongside this script so admins have it + # on PATH (/usr/local/bin) on the login node. Fetched from the same S3 + # location this script came from (S3_BUCKET/S3_KEY_PREFIX passed by UserData); + # falls back to a no-op with a hint if those are unset. + if [ -n "${S3_BUCKET:-}" ]; then + if aws s3 cp "s3://${S3_BUCKET}/${S3_KEY_PREFIX:-templates/}scripts/ldap-add-user.sh" \ + /usr/local/bin/ldap-add-user.sh 2>/dev/null; then + chmod 755 /usr/local/bin/ldap-add-user.sh + echo "[directory-server] Installed helper: /usr/local/bin/ldap-add-user.sh" + else + echo "[directory-server] WARNING: could not fetch ldap-add-user.sh from s3://${S3_BUCKET}/${S3_KEY_PREFIX:-templates/}scripts/ — add users with raw ldapadd (see USER-MANAGEMENT.md)." + fi + else + echo "[directory-server] NOTE: S3_BUCKET unset; skipping ldap-add-user.sh install. Add users with raw ldapadd (see USER-MANAGEMENT.md)." + fi + + # Also configure SSSD on the login node itself (so getent works locally) + setup_client_internal "ldap://localhost" + + echo "[directory-server] OpenLDAP server ready." + echo "[directory-server] Admin DN: cn=admin,${LDAP_DOMAIN_SUFFIX}" +} + +############################################################################### +# CLIENT role — compute node: install SSSD + configure LDAP provider +############################################################################### +setup_client() { + # Determine LDAP server URI + local server_uri="${LDAP_SERVER_URI:-}" + + if [ -z "$server_uri" ]; then + # Discover login node IP + if [ -n "${DIRECTORY_DNS_IPS:-}" ]; then + # Future: SimpleAD — use provided DNS IPs + server_uri="ldap://${DIRECTORY_DNS_IPS%%,*}" + else + # OpenLDAP-LoginNode — discover the directory server's private IP via + # its EC2 tags (pcs-cluster-id + directory-role=server). The + # directory-role tag is owned by the multi-user feature and is + # independent of the monitoring stack's monitoring-role tag. The AWS + # CLI auto-resolves the region from the instance's IMDS credential + # provider, so no explicit --region (which would need an IMDSv2 token + # to read placement/region) is required. + if [ -z "${CLUSTER_ID:-}" ]; then + echo "[directory-client] ERROR: CLUSTER_ID is empty; cannot discover directory server. LDAP client not configured." + return 1 + fi + local login_ip + login_ip=$(aws ec2 describe-instances \ + --filters "Name=tag:pcs-cluster-id,Values=${CLUSTER_ID}" \ + "Name=tag:directory-role,Values=server" \ + "Name=instance-state-name,Values=running" \ + --query 'Reservations[0].Instances[0].PrivateIpAddress' \ + --output text 2>/dev/null || echo "") + if [ -z "$login_ip" ] || [ "$login_ip" = "None" ]; then + echo "[directory-client] ERROR: could not discover directory server IP (CLUSTER_ID=${CLUSTER_ID}, directory-role=server). Check the compute node's IAM role has ec2:DescribeInstances. LDAP client not configured." + return 1 + fi + server_uri="ldap://${login_ip}" + fi + fi + + setup_client_internal "$server_uri" +} + +############################################################################### +# Internal: configure SSSD (shared by server + client roles) +############################################################################### +setup_client_internal() { + local server_uri="$1" + + echo "[directory-client] Running apt-get update..." + apt_get update -qq + + echo "[directory-client] Installing SSSD..." + # sssd-tools provides sss_cache, needed to invalidate the SSSD cache after a + # user is deleted/modified in LDAP (otherwise the change is not visible until + # the cache entry's TTL expires — see USER-MANAGEMENT.md "Deleting a user"). + apt_get install -y sssd sssd-tools libpam-sss libnss-sss ldap-utils + + echo "[directory-client] Configuring SSSD (server: ${server_uri})..." + cat > /etc/sssd/sssd.conf <> /etc/pam.d/common-session + fi + + # Configure NSS + sed -i 's/^passwd:.*/passwd: files systemd sss/' /etc/nsswitch.conf + sed -i 's/^group:.*/group: files systemd sss/' /etc/nsswitch.conf + sed -i 's/^shadow:.*/shadow: files sss/' /etc/nsswitch.conf + + systemctl enable sssd + systemctl restart sssd + + echo "[directory-client] SSSD configured. Server: ${server_uri}" +} + +############################################################################### +# Main +############################################################################### +case "$ROLE" in + server) + setup_server + ;; + client) + setup_client + ;; + *) + echo "Usage: $0 " + exit 1 + ;; +esac diff --git a/architectures/aws-pcs/docs/CUSTOM-AMI.md b/architectures/aws-pcs/docs/CUSTOM-AMI.md new file mode 100644 index 000000000..1f9c53da8 --- /dev/null +++ b/architectures/aws-pcs/docs/CUSTOM-AMI.md @@ -0,0 +1,70 @@ +# Pre-baking Enroot/Pyxis into a custom AMI + +The all-in-one template installs Enroot/Pyxis at **first boot** via +`PostInstallScriptUrl`, which is fast to deploy and avoids an Image Builder step. For +**frequent scaling** in production, pre-baking Enroot/Pyxis into a custom AMI drops node +boot time from ~8–12 min to ~3 min and pins every node to a deterministic state. + +This is a separate, standalone path: build the AMI once with +[`pcs-ready-dlami-with-enroot-pyxis.yaml`](../assets/pcs-ready-dlami-with-enroot-pyxis.yaml), +then pass the resulting `ami-xxx` as `AmiId` to the cluster. + +## Step 1: Build the AMI (~30 min one-time, separate stack) + +[![Launch](../images/launch-stack.svg)](https://console.aws.amazon.com/cloudformation/home#/stacks/quickcreate?templateUrl=https://awsome-distributed-ai.s3.amazonaws.com/templates/pcs-ready-dlami-with-enroot-pyxis.yaml&stackName=pcs-dlami) + +```bash +aws cloudformation create-stack \ + --stack-name pcs-dlami \ + --template-url https://awsome-distributed-ai.s3.amazonaws.com/templates/pcs-ready-dlami-with-enroot-pyxis.yaml \ + --parameters ParameterKey=SlurmVersion,ParameterValue=25.11 \ + --capabilities CAPABILITY_IAM CAPABILITY_NAMED_IAM +``` + +The AMI is **single-Slurm-version by design**: Pyxis is a SPANK plugin whose ABI is +locked to its compile-time Slurm version, so pass the same `SlurmVersion` you'll use on +the cluster. + +## Step 2: Read the resulting AMI ID + +From the stack output `DLAMIforPCSAmiId`: + +```bash +AMI_ID=$(aws cloudformation describe-stacks \ + --stack-name pcs-dlami \ + --query 'Stacks[0].Outputs[?OutputKey==`DLAMIforPCSAmiId`].OutputValue' \ + --output text) +echo "$AMI_ID" # ami-0xxxxxxxxxxxxxxxx +``` + +## Step 3: Pass it to the cluster as `AmiId` + +Optionally skip the boot-time Enroot/Pyxis install (it's already baked in) by setting +`PostInstallScriptUrl` to a single space: + +```bash +aws cloudformation create-stack \ + --stack-name pcs-ml-cluster \ + --template-url https://awsome-distributed-ai.s3.amazonaws.com/templates/pcs-ml-cluster-deploy-all.yaml \ + --parameters \ + ParameterKey=PrimarySubnetAZ,ParameterValue=us-east-1a \ + ParameterKey=AmiId,ParameterValue=$AMI_ID \ + ParameterKey=PostInstallScriptUrl,ParameterValue=' ' \ + --capabilities CAPABILITY_IAM CAPABILITY_NAMED_IAM +``` + +Leaving `PostInstallScriptUrl` at its default (empty → auto-install Enroot/Pyxis from the +templates bucket) also works on a pre-baked AMI: the installer detects Enroot/Pyxis is +already present and is a fast idempotent no-op. Passing a single space skips the +download+check entirely, shaving a few seconds off boot. + +## Optional features of `pcs-ready-dlami-with-enroot-pyxis.yaml` + +Defaults are off: +- `BuildSchedule=Weekly`/`Monthly` for scheduled rebuilds against a moving base AMI +- `EnableLifecyclePolicy=true` to deprecate older AMIs after `LifecycleDeprecateAfterWeeks` +- `PublishToSsm=true` to publish the latest AMI ID to an SSM parameter for downstream stacks + +For production deploys that pin the AMI explicitly per cluster, none of these are needed. + +> The build path is validated by [tests/infra-test.md Test 8](../tests/infra-test.md#test-8-pre-baked-ami-build-standalone-dlami-template). diff --git a/architectures/aws-pcs/docs/DEPLOY-TESTING.md b/architectures/aws-pcs/docs/DEPLOY-TESTING.md new file mode 100644 index 000000000..4c495b498 --- /dev/null +++ b/architectures/aws-pcs/docs/DEPLOY-TESTING.md @@ -0,0 +1,176 @@ +# Deploying updated templates before they are published + +The one-click Quick Start in the [README](../README.md#3-quick-start) deploys from the +**public production bucket** (`awsome-distributed-ai`), which only has the templates that +have already been merged and published. This guide is for the other case: **you have +template/script changes that are not yet in the public bucket** (e.g. a fork, a feature +branch, or a PR under review) and you want to deploy and test them. + +The approach is the same in every case: **host the templates + boot scripts in an S3 +bucket you control, then point the deploy at that bucket** via the `S3BucketName` / +`S3KeyPrefix` parameters. The nested stacks and the first-boot scripts are all fetched +from `s3:///...`, so overriding those two parameters redirects +the entire deploy to your copy. + +## 1. Prerequisites + +- AWS CLI configured with credentials for your account. +- An S3 bucket you own, in any Region (the bucket is reached by its global name, so it + does not need to be in the deploy Region). It can be **private** — the templates are + fetched by your CLI/CloudFormation, and the boot scripts are fetched by the node + instance role, so no public access is required. +- A local checkout of the branch/fork with the changes you want to test. + +This guide uses these placeholders — substitute your own: + +```bash +BUCKET=my-pcs-templates # an S3 bucket you control +PREFIX=templates/ # key prefix (keep the trailing slash) +REGION=us-east-1 # the Region to deploy the cluster into +AZ=us-east-1a # an Availability Zone in $REGION +``` + +## 2. Upload the templates + scripts to your bucket + +Run from the repo's `architectures/aws-pcs` directory (so `assets/` is the source): + +```bash +aws s3 sync assets/ "s3://${BUCKET}/${PREFIX}" \ + --exclude "*" --include "*.yaml" --include "*.sh" +``` + +This uploads both the CloudFormation templates (`*.yaml`) and the boot scripts (`*.sh`) +in one command. The scripts live under `assets/scripts/`, so they land at +`s3://${BUCKET}/${PREFIX}scripts/` — which is exactly where the default +`PostInstallScriptUrl` looks (`s3:///scripts/install-enroot-pyxis.sh`). +Re-run this sync after every change you want to test. + +## 3. Deploy pointing at your bucket + +Pass `S3BucketName` (and `S3KeyPrefix` if you changed it from the `templates/` default). +That single override is what makes the nested stacks **and** the first-boot Enroot/Pyxis +installer come from your copy instead of the public bucket: + +```bash +aws cloudformation create-stack \ + --stack-name pcs-test \ + --template-url "https://${BUCKET}.s3.amazonaws.com/${PREFIX}pcs-ml-cluster-deploy-all.yaml" \ + --parameters \ + ParameterKey=PrimarySubnetAZ,ParameterValue=${AZ} \ + ParameterKey=S3BucketName,ParameterValue=${BUCKET} \ + ParameterKey=S3KeyPrefix,ParameterValue=${PREFIX} \ + --capabilities CAPABILITY_IAM CAPABILITY_NAMED_IAM \ + --region ${REGION} +``` + +> **Why this matters.** If you deploy the updated top-level template but leave +> `S3BucketName` at its default, the nested stacks and boot scripts are still pulled from +> the **public** bucket — so you'd be testing your top-level change against the old +> published nested templates/scripts. Always override `S3BucketName` to your bucket when +> testing unpublished changes. + +Add any parameters you're testing on top — for example multi-user: + +```bash + ParameterKey=DirectoryService,ParameterValue=OpenLDAP-LoginNode \ +``` + +or an EFA-capable CPU queue: + +```bash + ParameterKey=OnDemandInstanceType,ParameterValue=hpc8a.96xlarge \ + ParameterKey=OnDemandEfaInterfaceCount,ParameterValue=2 \ +``` + +See [PARAMETERS.md](./PARAMETERS.md) for the full list. + +## 4. Monitor progress + +```bash +aws cloudformation describe-stacks --stack-name pcs-test --region ${REGION} \ + --query 'Stacks[0].StackStatus' --output text + +aws cloudformation describe-stack-events --stack-name pcs-test --region ${REGION} \ + --query 'StackEvents[?ResourceStatus!=`CREATE_IN_PROGRESS`].[Timestamp,LogicalResourceId,ResourceStatus]' \ + --output text | head -20 +``` + +Typical timeline (~25-30 min): Prerequisites (VPC + FSx) ~10 min → Cluster ~5 min → +Login + compute node groups (parallel) ~5 min → node boot + cloud-init (post-install, +monitoring, directory) ~3-8 min. + +## 5. Connect and verify + +Connect to the login node over SSM (no public IP or SSH key needed): + +```bash +CLUSTER_ID=$(aws cloudformation describe-stacks --stack-name pcs-test --region ${REGION} \ + --query 'Stacks[0].Outputs[?OutputKey==`ClusterId`].OutputValue' --output text) +LOGIN_ID=$(aws ec2 describe-instances --region ${REGION} \ + --filters "Name=tag:pcs-cluster-id,Values=$CLUSTER_ID" \ + "Name=tag:monitoring-role,Values=login" \ + "Name=instance-state-name,Values=running" \ + --query 'Reservations[0].Instances[0].InstanceId' --output text) +aws ssm start-session --target $LOGIN_ID --region ${REGION} +``` + +Then `sudo su - ubuntu` and run the checks relevant to your change. A few quick ones: + +```bash +# Slurm sees the queues / nodes (adjust the version if you set SlurmVersion=25.05) +export PATH=/opt/aws/pcs/scheduler/slurm-25.11/bin:$PATH; sinfo -N; squeue + +# Container runtime came from your bucket (the log shows the s3:// it fetched) +grep s3:// /var/log/pcs-post-install.log + +# Monitoring containers (when MonitoringStack=Prometheus-LoginNode) +sudo docker ps --format "table {{.Names}}\t{{.Status}}" +``` + +For the full reproducible test matrix (monitoring, container runtime, CPU/GPU, NCCL, +FSDP, multi-user, GPU health), see [tests/README.md](../tests/README.md). + +## 6. Iterate + +After editing a template or script, re-run the **sync** (step 2), then +`update-stack` (same parameters as `create-stack`) — or delete and recreate if the change +can't be applied in place (e.g. a subnet CIDR change): + +```bash +aws cloudformation update-stack --stack-name pcs-test \ + --template-url "https://${BUCKET}.s3.amazonaws.com/${PREFIX}pcs-ml-cluster-deploy-all.yaml" \ + --parameters ParameterKey=PrimarySubnetAZ,UsePreviousValue=true \ + ParameterKey=S3BucketName,UsePreviousValue=true \ + ... \ + --capabilities CAPABILITY_IAM CAPABILITY_NAMED_IAM --region ${REGION} +``` + +> A stack update only re-runs first-boot scripts on **newly launched** nodes; existing +> nodes keep what they booted with. To re-test a boot-script change, let the node group +> scale a fresh node (or replace the affected nodes). + +## 7. Cleanup + +```bash +aws cloudformation delete-stack --stack-name pcs-test --region ${REGION} +``` + +Nested stacks (and FSx — back up data first) are deleted automatically. If a CNG stack +hits `DELETE_FAILED` (a PCS timing dependency), delete the PCS compute node groups first, +then retry: + +```bash +for cng in $(aws pcs list-compute-node-groups --cluster-identifier $CLUSTER_ID --region ${REGION} \ + --query 'computeNodeGroups[].id' --output text); do + aws pcs delete-compute-node-group --cluster-identifier $CLUSTER_ID \ + --compute-node-group-identifier $cng --region ${REGION} +done +sleep 60 +aws cloudformation delete-stack --stack-name pcs-test --region ${REGION} +``` + +## After the change is published + +Once the change is merged and the maintainer has synced `assets/` to the public bucket +(`awsome-distributed-ai`), no override is needed — the defaults already point there, and +the [README Quick Start](../README.md#3-quick-start) works as written. diff --git a/architectures/aws-pcs/docs/IAM.md b/architectures/aws-pcs/docs/IAM.md new file mode 100644 index 000000000..68aa924cb --- /dev/null +++ b/architectures/aws-pcs/docs/IAM.md @@ -0,0 +1,161 @@ +# IAM Permissions Guide + +The cluster distinguishes **two human roles** with very different +responsibilities, and ships a ready-to-deploy IAM policy stack for each: + +| Role | Who | What they can do | Template | +|---|---|---|---| +| **Cluster admin** | The person who deploys/updates/deletes the cluster | Full CRUD on the infrastructure: CloudFormation, PCS, EC2 (VPC/SG/launch templates/placement groups/NAT/EIP), FSx, scoped IAM, SSM Parameter Store, KMS, Secrets Manager, and (optionally) Image Builder | [`cluster-admin-iam.yaml`](../assets/cluster-admin-iam.yaml) | +| **Cluster user** | Engineers who just run jobs on an existing cluster | SSM session **to the login node only**, port-forward Grafana, read the Grafana password, read PCS cluster/queue status. **Cannot create, modify, or delete anything**, and cannot open shells on compute nodes | [`cluster-user-iam.yaml`](../assets/cluster-user-iam.yaml) | + +Splitting the roles means the deployer's broad permissions never have to be +handed to every engineer who just wants to `srun`, and the user role can be +given out widely — it can't accidentally delete the cluster or get a shell on a +compute node. + +--- + +## Deploying the policies + +Each template creates the customer-managed IAM policies, an IAM group with them +attached, and (optionally) adds existing IAM users to that group. Deploy both +as CloudFormation stacks: + +| Stack | Creates | Deploy | +|---|---|---| +| **Cluster admin** | `-PCSClusterAdmin-core` (always) + `-PCSClusterAdmin-imagebuilder` (if opted in) managed policies + an IAM group | [![Launch](../images/launch-stack.svg)](https://console.aws.amazon.com/cloudformation/home#/stacks/quickcreate?templateUrl=https://awsome-distributed-ai.s3.amazonaws.com/templates/cluster-admin-iam.yaml&stackName=pcs-cluster-admins) | +| **Cluster user** | `-PCSClusterUser` managed policy + an IAM group | [![Launch](../images/launch-stack.svg)](https://console.aws.amazon.com/cloudformation/home#/stacks/quickcreate?templateUrl=https://awsome-distributed-ai.s3.amazonaws.com/templates/cluster-user-iam.yaml&stackName=pcs-cluster-users) | + +Or from the CLI (use `--template-body` against a local checkout for a pre-merge +sandbox test): + +```bash +# Admin: create the policies + group, attach existing users, include Image Builder perms +aws cloudformation create-stack \ + --stack-name pcs-cluster-admins \ + --template-body file://architectures/aws-pcs/assets/cluster-admin-iam.yaml \ + --parameters ParameterKey=AttachUsers,ParameterValue=alice,bob \ + ParameterKey=AttachImageBuilderPolicy,ParameterValue=true \ + --capabilities CAPABILITY_IAM CAPABILITY_NAMED_IAM + +# User: create the policy + group, attach existing users +aws cloudformation create-stack \ + --stack-name pcs-cluster-users \ + --template-body file://architectures/aws-pcs/assets/cluster-user-iam.yaml \ + --parameters ParameterKey=AttachUsers,ParameterValue=carol,dave \ + --capabilities CAPABILITY_IAM CAPABILITY_NAMED_IAM +``` + +Both templates take an `AttachUsers` parameter (comma-separated existing IAM +user names) so you can wire up group membership at deploy time, or leave it +empty and add users to the group later. The admin template's +`AttachImageBuilderPolicy` defaults to `false`; set it `true` only when the +admin will also deploy the standalone DLAMI builder +(`pcs-ready-dlami-with-enroot-pyxis.yaml`). + +### What the cluster user can do once attached + +```bash +# Find the login node and open a session +INSTANCE_ID=$(aws ec2 describe-instances \ + --filters "Name=tag:Name,Values=PCS-login" "Name=instance-state-name,Values=running" \ + --query 'Reservations[0].Instances[0].InstanceId' --output text) +aws ssm start-session --target $INSTANCE_ID + +# Port-forward Grafana (443 -> 8443), then open https://localhost:8443/grafana/ +aws ssm start-session --target $INSTANCE_ID \ + --document-name AWS-StartPortForwardingSession \ + --parameters '{"portNumber":["443"],"localPortNumber":["8443"]}' + +# Read the Grafana admin password +aws ssm get-parameter --name "/pcs//grafana/admin-password" \ + --with-decryption --query 'Parameter.Value' --output text +``` + +--- + +## Considerations + +These are **sample, slightly-broader-than-strict-least-privilege** policies, +derived from the AWS-published +[minimum permissions for an AWS PCS service administrator](https://docs.aws.amazon.com/pcs/latest/userguide/security-min-permissions.html) +plus the extra permissions the all-in-one template needs because it provisions +VPC + FSx + IAM roles itself (the AWS reference policy assumes those already +exist). Review and tighten before production use. + +**Login-node access is scoped by the `Name` tag.** The user policy conditions +`ssm:StartSession` on `ssm:resourceTag/Name` matching `PCS-login*`. PCS does not +emit a dedicated "is this a login node" tag, so the templates set +`Name=PCS-login` on the login node and `Name=PCS-` on compute nodes +(`PCS-cpu1`, `PCS-hpc8a`, …) — the most stable signal available. **The `Name` +tag is operator-mutable**: if you re-tag a login node, update the policy +condition to match (or fork the templates to add a dedicated `IsLoginNode=true` +tag and key off that). + +**Combined CRUD is intentional, not a mistake.** The admin policy covers +create + update + delete in one policy because (1) CFN rollback on a failed +Create requires Delete actions, (2) UpdateStack is operationally a superset of +Create (it may replace resources), and (3) drift detection during Update calls +Describe across every service. If you want a read-only variant, reduce the same +actions to `*:Describe*` / `*:Get*` / `*:List*` for an auditor role. + +**The admin policy is split into core + Image Builder** because the combined +document (~7.4 KB) exceeds the IAM 6,144-character per-policy limit. The +~5.8 KB core covers a normal deploy; the ~1.7 KB Image Builder add-on is only +needed for the standalone DLAMI builder. The CFN template attaches both to the +group when `AttachImageBuilderPolicy=true`. + +**There is no `AmazonPCSFullAccess` managed policy** as of January 2026 — AWS +publishes only `AWSPCSComputeNodePolicy` (for compute instances) and +`AWSPCSServiceRolePolicy` (the service-linked role). The PCS portion of the +admin policy must therefore be customer-managed. + +**Pairing with AWS-managed policies.** For a smaller customer-managed surface +you can attach AWS-managed policies for parts of the stack and trim the matching +statements: `AWSCloudFormationFullAccess`, `AmazonFSxFullAccess`, +`AWSImageBuilderFullAccess` are reasonable fits. Avoid `AmazonEC2FullAccess` — +it is materially overprivileged (e.g. EBS public-share); prefer the +customer-managed EC2 statements in the template. + +### Not covered by these policies + +- **The compute instance role** (passed to EC2 by `cluster.yaml`) — provisioned + by the templates themselves; use the AWS-managed `AWSPCSComputeNodePolicy`. +- **The Image Builder build instance role** — use the AWS-managed + `EC2InstanceProfileForImageBuilder` / + `EC2InstanceProfileForImageBuilderECRContainerBuilds`. +- **Fine-grained per-cluster scoping** — both policies use `Resource: "*"` for + many EC2/VPC actions because resource-level scoping there is limited. This is + a deliberate sample-grade choice. + +### Refining to least-privilege via CloudTrail + +To generate a tighter policy from real usage: + +1. Deploy a representative cluster in a sandbox account with broad permissions on + the deploying principal (so nothing fails for spurious IAM reasons). +2. Let the full lifecycle run — deploy, then `delete-stack` — so CloudTrail + captures every API call. +3. Generate a policy from CloudTrail with IAM Access Analyzer + (`aws accessanalyzer start-policy-generation` → `get-generated-policy`), then + diff against the template's statements. Access Analyzer output is usually + *narrower* on actions but leaves `Resource: "*"`; the template's resource ARNs + are usually the keepers. + +The same approach works for the user policy — exercise the user workflows +(start a session, port-forward Grafana, terminate it) in a sandbox, then narrow. + +--- + +## Verifying the policies + +To confirm the admin policy can deploy a cluster end-to-end and the user policy is +correctly constrained (login-only SSM, no LDAP-password access), see the reproducible +procedure in [tests/iam-test.md](../tests/iam-test.md). + +> **Note on the template-source bucket.** The admin policy grants no `s3:GetObject`, +> because the production templates live in the public `awsome-distributed-ai` +> bucket (CFN fetches `--template-url` anonymously). If you host the templates in +> a **private** bucket, the deploying principal additionally needs `s3:GetObject` +> on that bucket — grant it separately; it is intentionally out of the +> cluster-admin policy. diff --git a/architectures/aws-pcs/docs/OPERATIONS.md b/architectures/aws-pcs/docs/OPERATIONS.md index 04fe044ce..bcd43e933 100644 --- a/architectures/aws-pcs/docs/OPERATIONS.md +++ b/architectures/aws-pcs/docs/OPERATIONS.md @@ -45,9 +45,11 @@ as a separate stack, then pass its output to the cluster. the default `PostInstallScriptUrl` delivers. - **Pre-baked AMI** — build the AMI separately (see the README's *Pre-baking Enroot/Pyxis into a custom AMI* section), then pass its `ami-xxx` as - the cluster's `AmiId`. Use `PostInstallScriptUrl=""` for the cleanest boot (the - installer is idempotent so leaving the default is a fast no-op, but skipping the - download saves a few seconds on every node launch). + the cluster's `AmiId`. Set `PostInstallScriptUrl=' '` (a single space) for the + cleanest boot — that skips the Enroot/Pyxis install entirely. (Leaving it at the + default — empty, which auto-installs from the templates bucket — also works on a + pre-baked AMI: the installer is idempotent and detects Enroot/Pyxis is already + present, a fast no-op; the single space just avoids the download+check.) ### 2.1 The AMI is single-Slurm-version, by design @@ -135,7 +137,7 @@ want — leave it alone. ### 3.2 Public Grafana exposure -`GrafanaPublicAccessCidr` opens **TCP/443 on the login node** to a CIDR via a +`GrafanaAccessCidr` opens **TCP/443 on the login node** to a CIDR via a login-only security group. The nginx in front of Grafana also proxies **`/prometheus/`, `/pushgateway/`, `/slurmexporter/` without authentication** — opening the CIDR exposes those too, not just password-gated Grafana. Use the tightest CIDR you @@ -168,8 +170,8 @@ on the deployment type**: |---|---|---| | Lustre | `PERSISTENT_2` (default) | 125 / 250 / 500 / 1000 MB/s/TiB | | Lustre | `PERSISTENT_1` (older Regions) | 50 / 100 / 200 MB/s/TiB | -| OpenZFS | `SINGLE_AZ_HA_2` (default) | 160 / 320 / 640 / 1280 / 2560 / 3840 / 5120 / 7680 MB/s | -| OpenZFS | `SINGLE_AZ_HA_1` / `SINGLE_AZ_2` / `SINGLE_AZ_1` | 64 / 128 / 192 / 256 / 384 / 512 / 768 / 1024 MB/s | +| OpenZFS | `SINGLE_AZ_2` / `SINGLE_AZ_HA_2` (default) | 160 / 320 / 640 / 1280 / 2560 / 3840 / 5120 / 7680 / 10240 MB/s | +| OpenZFS | `SINGLE_AZ_HA_1` / `SINGLE_AZ_1` | 64 / 128 / 256 / 512 / 1024 / 2048 / 3072 / 4096 MB/s | The templates enforce the valid pair via CloudFormation `Rules` so a mismatch fails at stack-create time with a clear message instead of deep in the nested FSx stack. If you @@ -425,5 +427,5 @@ For a new production deploy: and pass its output as `AmiId` (~3 min boot vs ~6 min). Match `SlurmVersion` between the AMI build stack and the cluster stack. - Default `DcgmExporterImage` covers H100/B200/B300; override only to pin a different build -- Minimum-CIDR `GrafanaPublicAccessCidr` if used at all; otherwise empty (SSM port-forward) +- Minimum-CIDR `GrafanaAccessCidr` if used at all; otherwise empty (SSM port-forward) - Throughput values that match the chosen FSx deployment types diff --git a/architectures/aws-pcs/docs/PARAMETERS.md b/architectures/aws-pcs/docs/PARAMETERS.md index f801f2549..fea9bc97d 100644 --- a/architectures/aws-pcs/docs/PARAMETERS.md +++ b/architectures/aws-pcs/docs/PARAMETERS.md @@ -1,13 +1,13 @@ # Parameter Reference — `pcs-ml-cluster-deploy-all.yaml` -Full parameter list for the all-in-one deployment template, grouped to match the -sections shown in the CloudFormation console (the console also shows friendly labels via -`AWS::CloudFormation::Interface`). Defaults give the most common production setup — -the latest PCS-Ready Deep Learning AMI auto-resolved from SSM, Enroot/Pyxis installed -at first boot via `PostInstallScriptUrl`, monitoring enabled — so a default deploy -only needs the Availability Zone (`PrimarySubnetAZ`). To pre-bake Enroot/Pyxis into +Full parameter list for the all-in-one deployment template. The sections and their order +match the CloudFormation console's parameter groups exactly (the console shows friendly +labels via `AWS::CloudFormation::Interface`). Defaults give the most common production +setup — the latest PCS-Ready Deep Learning AMI auto-resolved from SSM, Enroot/Pyxis +installed at first boot via `PostInstallScriptUrl`, monitoring enabled — so a default +deploy only needs the Availability Zone (`PrimarySubnetAZ`). To pre-bake Enroot/Pyxis into a custom AMI for faster boots, build it separately with -[`pcs-ready-dlami-with-enroot-pyxis.yaml`](../README.md#9-pre-baking-enrootpyxis-into-a-custom-ami-optional) +[`pcs-ready-dlami-with-enroot-pyxis.yaml`](../README.md#85-pre-baking-enrootpyxis-into-a-custom-ami) and pass its output as `AmiId`. For conceptual guidance (GPU instance/EFA selection, FSx Region availability, container @@ -17,10 +17,17 @@ runtime options), see the [README](../README.md#4-configuration). | Parameter | Default | Purpose | |---|---|---| -| `PrimarySubnetAZ` | *(required)* | Availability Zone to deploy into — the one required parameter | -| `VPCName` | `ML-Cluster-VPC` | Name for the created VPC | +| `PrimarySubnetAZ` | *(required)* | Availability Zone to deploy into — the one required parameter. Holds the public subnet (login node), the primary private subnet (compute, FSx), and the single NAT gateway | +| `AdditionalSubnetAZ2` | *(empty)* | (Optional) 2nd AZ for an additional private subnet. Empty = single-AZ. Enables multi-AZ layouts (e.g. OpenZFS `MULTI_AZ`). Shares the primary AZ's NAT gateway (cross-AZ egress, no per-AZ NAT) | +| `AdditionalSubnetAZ3` | *(empty)* | (Optional) 3rd AZ for an additional private subnet. Requires `AdditionalSubnetAZ2` to also be set. Max 3 private AZs total | | `CreateS3Endpoint` | `true` | Create an S3 VPC endpoint | +The VPC name is fixed to `${StackName}-VPC` (derived from the stack name, so +multiple deployments in one account get unique VPC names automatically) — there +is no `VPCName` parameter on the all-in-one template. The standalone +`ml-cluster-prerequisites.yaml` still accepts a `VPCName` parameter if you deploy +it directly. + ## 2. PCS Cluster Configuration | Parameter | Default | Purpose | @@ -28,20 +35,12 @@ runtime options), see the [README](../README.md#4-configuration). | `SlurmVersion` | `25.11` | Slurm version (`25.05` or `25.11`). Drives which monitoring you get (Slurm OpenMetrics is 25.11+ only) and is also threaded into the CNG UserData so the right-version Pyxis is installed; see [OPERATIONS.md §1](./OPERATIONS.md#1-slurm-version-selection) | | `LoginNodeInstanceType` | `m6i.4xlarge` | Login node instance type | | `RootVolumeSize` | `300` | Root EBS volume size (GiB) on every node (login + compute); 300 leaves room for large container images (Megatron `.sqsh` ~20 GB) | -| `AmiId` | *(empty → SSM auto-resolve)* | AMI ID for every node group. **Empty (default) auto-resolves to the latest PCS-Ready Deep Learning AMI** (Ubuntu 24.04, x86_64) from SSM (`/aws/service/pcs/ami/dlami-base-ubuntu2404/x86_64/latest/ami-id`). For production, **pin to a specific `ami-xxx`** so a later scale-out cannot drift onto a newer base. Use a custom AMI built off the PCS-Ready DLAMI base (e.g. via [`pcs-ready-dlami-with-enroot-pyxis.yaml`](../README.md#9-pre-baking-enrootpyxis-into-a-custom-ami-optional)) when you want Enroot/Pyxis pre-baked or other customizations. See [OPERATIONS.md §4](./OPERATIONS.md#4-ami-selection-amiid--pin-in-production) | -| `DeployMonitoring` | `true` | Deploy Prometheus/Grafana/DCGM on the login node | -| `GrafanaPublicAccessCidr` | *(empty)* | When set to a CIDR, opens HTTPS/443 on the login node to that CIDR via a login-only security group. Empty = SSM port-forward only. **443 also exposes the unauthenticated `/prometheus/`, `/pushgateway/`, `/slurmexporter/` proxy paths**, not just the password-gated Grafana. Use the tightest CIDR you can; `0.0.0.0/0` is accepted for short-lived PoC/workshop use but exposes those endpoints to the whole internet | +| `AmiId` | *(empty → SSM auto-resolve)* | AMI ID for every node group. **Empty (default) auto-resolves to the latest PCS-Ready Deep Learning AMI** (Ubuntu 24.04, x86_64) from SSM (`/aws/service/pcs/ami/dlami-base-ubuntu2404/x86_64/latest/ami-id`). For production, **pin to a specific `ami-xxx`** so a later scale-out cannot drift onto a newer base. Use a custom AMI built off the PCS-Ready DLAMI base (e.g. via [`pcs-ready-dlami-with-enroot-pyxis.yaml`](../README.md#85-pre-baking-enrootpyxis-into-a-custom-ami)) when you want Enroot/Pyxis pre-baked or other customizations. See [OPERATIONS.md §4](./OPERATIONS.md#4-ami-selection-amiid--pin-in-production) | +| `SSHAccessCidr` | *(empty)* | When set to a CIDR, opens SSH/22 on the login node to that CIDR via a login-only security group (attached to the login node only, never compute). Empty (default) = SSH over SSM only. Set to your office/VPN range for direct `ssh`/`scp`/VS Code Remote (common for multi-user clusters) | | `ManagedAccounting` | `disabled` | Enable Slurm managed accounting (requires Slurm 24.11+) | | `AccountingPolicyEnforcement` | `none` | Slurm accounting policy enforcement (`none` or `associations,limits,safe`) | -## 3. Container Runtime (Post-install Script) - -| Parameter | Default | Purpose | -|---|---|---| -| `PostInstallScriptUrl` | Enroot/Pyxis installer | HTTP(S) script run on every node at first boot (PCS equivalent of ParallelCluster `OnNodeConfigured`). Empty = skip; or override with any other HTTP(S) script. Idempotent: a no-op if Enroot/Pyxis is already pre-baked into `AmiId` | -| `PostInstallScriptArgs` | *(empty)* | Arguments passed to the post-install script | - -## 4. On-Demand Compute Node Group (CPU) +## 3. On-Demand Compute Node Group (CPU) | Parameter | Default | Purpose | |---|---|---| @@ -51,11 +50,10 @@ runtime options), see the [README](../README.md#4-configuration). | `OnDemandMaxCount` | `4` | CPU queue maximum nodes | | `OnDemandCngName` | `cpu1` | CPU node-group name | | `OnDemandQueueName` | `cpu1` | CPU Slurm queue name | -| `OnDemandEnableEfa` | `false` | Enable EFA on the CPU CNG (HPC/MPI workloads on hpc6a/hpc7a/hpc6id/hpc8a, c7i.metal, etc.). Switches the CNG's LaunchTemplate to a `NetworkInterfaces` block with `InterfaceType=efa` and wires in a cluster placement group. No effect on the GPU CNG. See [README §EFA on CPU HPC instances](../README.md#efa-on-cpu-hpc-instances-ondemandenableefa) | -| `OnDemandEfaInterfaceCount` | `1` | Number of EFA interfaces. Match the instance type's `MaximumEfaInterfaces`: hpc8a/hpc7a/hpc6id = `2`; hpc6a/c7i.metal = `1`. Mismatching fails at launch. Ignored when `OnDemandEnableEfa=false` | -| `OnDemandPlacementGroupName` | *(empty)* | Existing cluster placement group name to launch nodes into. Empty + `OnDemandEnableEfa=true` auto-creates a per-CNG cluster placement group; supplying a name reuses an existing one (e.g. shared across CPU + GPU CNGs for heterogeneous tightly-coupled jobs). Ignored when `OnDemandEnableEfa=false` | +| `OnDemandEfaInterfaceCount` | `0` | EFA interfaces on the CPU CNG. **`0` (default) = no EFA** (standard ENA). `1` or `2` = enable EFA with that many interfaces (switches the LaunchTemplate to a `NetworkInterfaces` block with `InterfaceType=efa` + a cluster placement group). Set the count to the instance type's `MaximumEfaInterfaces`: `hpc8a.96xlarge`/`hpc7a.*`/`hpc6id.32xlarge`=2; `hpc6a.48xlarge`/`c7i.metal`=1. **EFA needs an EFA-capable type** — a non-EFA type (e.g. the default `c6i.4xlarge`) fails to launch with count > 0. No effect on the GPU CNG. See [README §8.6 CPU compute node group](../README.md#86-cpu-compute-node-group--advanced-settings) | +| `OnDemandPlacementGroupName` | *(empty)* | Existing cluster placement group name to launch nodes into. Empty + `OnDemandEfaInterfaceCount > 0` auto-creates a per-CNG cluster placement group; supplying a name reuses an existing one (e.g. shared across CPU + GPU CNGs for heterogeneous tightly-coupled jobs). Ignored when `OnDemandEfaInterfaceCount = 0` | -## 5. GPU Compute Node Group — P5/P6 (Optional) +## 4. GPU Compute Node Group — P5/P6 (Optional) See [GPU compute](../README.md#gpu-compute-p5p6) for instance/EFA/capacity guidance. @@ -69,19 +67,34 @@ See [GPU compute](../README.md#gpu-compute-p5p6) for instance/EFA/capacity guida | `PseriesCngName` | `gpu-p5` | GPU node-group name | | `PseriesQueueName` | `gpu-p5` | GPU Slurm queue name | -## 6. FSx for Lustre (`/fsx`) + FSx for OpenZFS (`/home`) (Advanced) +## 5. Additional Cluster Configuration (Monitoring, Multi-User, Container Runtime) -See [Storage: FSx deployment types](../README.md#storage-fsx-deployment-types-region-availability) for Region availability. +| Parameter | Default | Purpose | +|---|---|---| +| `MonitoringStack` | `Prometheus-LoginNode` | Monitoring stack to deploy. `Prometheus-LoginNode` = self-hosted Prometheus + Grafana + DCGM Exporter on the login node. `none` = no monitoring. (Renamed from the old boolean `DeployMonitoring`; `-` enum, extensible to future `AMP-AMG`/`CloudWatch`) | +| `GrafanaAccessCidr` | *(empty)* | When set to a CIDR, opens HTTPS/443 (Grafana) on the login node to that CIDR via the login-only security group. Empty = SSM port-forward only. **443 also exposes the unauthenticated `/prometheus/`, `/pushgateway/`, `/slurmexporter/` proxy paths**, not just the password-gated Grafana. Use the tightest CIDR you can. (Renamed from `GrafanaPublicAccessCidr`) | +| `MonitoringRepo` | `aws-samples/aws-parallelcluster-monitoring` | GitHub `owner/repo` for the monitoring stack; override with a fork + a branch in `MonitoringVersion` to test unreleased changes | +| `MonitoringVersion` | `v2.9.1` | [aws-parallelcluster-monitoring](https://github.com/aws-samples/aws-parallelcluster-monitoring) git ref (release tag, branch, or `latest`). `v2.9.1` adds the `DCGM_EXPORTER_IMAGE` override (needed for B300 GPU metrics) and brings Grafana 13; `v2.6.4`+ carry the PCS `/opt` install + Docker-29.x DCGM fixes. Pin to a tag for stability. Migration notes: [OPERATIONS.md §3](./OPERATIONS.md#3-monitoring-monitoringversion) | +| `DcgmExporterImage` | DCGM 4.5.2 by digest | `dcgm-exporter` image used on GPU nodes. Defaults to a DCGM 4.5.2 build pinned by digest (`nvcr.io/nvidia/k8s/dcgm-exporter@sha256:a7ad6547...`) covering Hopper / B200 / B300. The digest pull bypasses the Docker-29.x OCI-index failure on newer NVCR tags. Override (any image reference, ideally also a digest) to pin to a different build — e.g. the monitoring stack's older default 4.2.0. No effect on CPU nodes. See [OPERATIONS.md §3.1](./OPERATIONS.md#31-dcgmexporterimage-the-default-and-when-to-change-it) | +| `DirectoryService` | `none` | Multi-user directory. `none` = single `ubuntu` user. `OpenLDAP-LoginNode` = slapd on the login node (DB on shared `/home/ldap-db`) + SSSD on all compute nodes. **Single login node only** — keep the login node group at 1 instance while enabled. See [USER-MANAGEMENT.md](./USER-MANAGEMENT.md) | +| `DirectoryDomainSuffix` | `dc=cluster,dc=internal` | LDAP domain suffix. Only used when `DirectoryService != none` | +| `PostInstallScriptUrl` | *(empty → auto)* | Script run on every node at first boot (PCS equivalent of ParallelCluster `OnNodeConfigured`). **Empty (default) auto-installs Enroot/Pyxis** from `s3:///scripts/install-enroot-pyxis.sh` (fetched with the instance role, so it works with a **private** bucket — no public S3 needed). Accepts an `s3://` URL (instance-role fetch) or an `http(s)://` URL (curl, public only, e.g. GitHub raw). Set to a single space to skip. Idempotent: a no-op if Enroot/Pyxis is already pre-baked into `AmiId` | +| `PostInstallScriptArgs` | *(empty)* | Arguments passed to the post-install script. Normally left empty — most users never touch the container-runtime parameters | + +## 6. FSx Storage (`/fsx` and `/home`) + +See [README §8.1 Storage](../README.md#81-storage-fsx-deployment-types--sizing) for Region +availability and the "deploy small, expand after" tip. | Parameter | Default | Purpose | |---|---|---| -| `Capacity` | `1200` | FSx for Lustre (`/fsx`) capacity (GiB; 1200 or increments of 2400) | +| `Capacity` | `1200` | FSx for Lustre (`/fsx`) capacity (GiB; 1200 or increments of 2400). Can be increased after creation, so start small for a faster first deploy | | `LustreDeploymentType` | `PERSISTENT_2` | FSx for Lustre (`/fsx`) deployment type (`PERSISTENT_2` / `PERSISTENT_1`) — Region-dependent | | `PerUnitStorageThroughput` | `250` | FSx for Lustre (`/fsx`) throughput (MB/s/TiB); valid values depend on the deployment type | | `Compression` | `LZ4` | FSx for Lustre (`/fsx`) data compression (`LZ4` / `NONE`) | | `LustreVersion` | `2.15` | FSx for Lustre (`/fsx`) software version (`2.15` / `2.12`) | -| `FSxLustreEnableEfa` | `false` | Enable EFA on the FSx for Lustre filesystem. **The headline feature is GPUDirect Storage (GDS) for P5/P5e/P5en/P6-B200 GPU clients**, which DMAs file data straight into GPU memory (requires the NVIDIA `nvidia-fs` / cuFile stack on the client — tracked as a follow-up in [docs/ROADMAP.md](./ROADMAP.md#client-side-lustre-on-efa--gds-support)). EFA-capable CPU CNGs (`OnDemandEnableEfa=true`) get the EFA *transport* path to storage as a secondary benefit, useful when a single client is pushing past ~10 GBps. **PERSISTENT_2 SSD only** — a CFN Rule on the prerequisites template fails the stack at create time when combined with PERSISTENT_1 (rather than silently ignoring the opt-in). **Requires a much larger `Capacity` than non-EFA**: at `PerUnitStorageThroughput=250` the minimum is **19200 GiB** (16× the 1200 GiB non-EFA default). The full minimum-capacity matrix per throughput tier is in the [FSx for Lustre User Guide](https://docs.aws.amazon.com/fsx/latest/LustreGuide/efa.html). The FSx side rejects undersized capacity at stack-create time with a clear error | -| `HomeCapacity` | `512` | FSx for OpenZFS (`/home`) capacity (GiB) | +| `FSxLustreEnableEfa` | `false` | Enable EFA on the FSx for Lustre filesystem. **The headline feature is GPUDirect Storage (GDS) for P5/P5e/P5en/P6-B200 GPU clients**, which DMAs file data straight into GPU memory (requires the NVIDIA `nvidia-fs` / cuFile stack on the client — tracked as a follow-up in [docs/ROADMAP.md](./ROADMAP.md)). EFA-capable CPU CNGs (`OnDemandEfaInterfaceCount > 0`) get the EFA *transport* path to storage as a secondary benefit, useful when a single client is pushing past ~10 GBps. **PERSISTENT_2 SSD only** — a CFN Rule on the prerequisites template fails the stack at create time when combined with PERSISTENT_1 (rather than silently ignoring the opt-in). **Requires a much larger `Capacity` than non-EFA**: at `PerUnitStorageThroughput=250` the minimum is **19200 GiB** (16× the 1200 GiB non-EFA default). The full minimum-capacity matrix per throughput tier is in the [FSx for Lustre User Guide](https://docs.aws.amazon.com/fsx/latest/LustreGuide/efa.html). The FSx side rejects undersized capacity at stack-create time with a clear error | +| `HomeCapacity` | `512` | FSx for OpenZFS (`/home`) capacity (GiB). Can be increased after creation | | `HomeThroughput` | `320` | FSx for OpenZFS (`/home`) throughput (MB/s) | | `OpenZFSDeploymentType` | `SINGLE_AZ_HA_2` | FSx for OpenZFS (`/home`) deployment type (`SINGLE_AZ_HA_2` / `SINGLE_AZ_HA_1` / `SINGLE_AZ_2` / `SINGLE_AZ_1`) — Region-dependent | @@ -91,6 +104,3 @@ See [Storage: FSx deployment types](../README.md#storage-fsx-deployment-types-re |---|---|---| | `S3BucketName` | `awsome-distributed-ai` | S3 bucket the nested templates are fetched from | | `S3KeyPrefix` | `templates/` | S3 key prefix for the nested templates | -| `MonitoringVersion` | `v2.9.1` | [aws-parallelcluster-monitoring](https://github.com/aws-samples/aws-parallelcluster-monitoring) git ref (release tag, branch, or `latest`). `v2.9.1` adds the `DCGM_EXPORTER_IMAGE` override (needed for B300 GPU metrics) and brings Grafana 13; `v2.6.4`+ carry the PCS `/opt` install + Docker-29.x DCGM fixes. Pin to a tag for stability. Migration notes: [OPERATIONS.md §3](./OPERATIONS.md#3-monitoring-monitoringversion) | -| `MonitoringRepo` | `aws-samples/aws-parallelcluster-monitoring` | GitHub `owner/repo` for the monitoring stack; override with a fork + a branch in `MonitoringVersion` to test unreleased changes | -| `DcgmExporterImage` | DCGM 4.5.2 by digest | `dcgm-exporter` image used on GPU nodes. Defaults to a DCGM 4.5.2 build pinned by digest (`nvcr.io/nvidia/k8s/dcgm-exporter@sha256:a7ad6547...`) covering Hopper / B200 / B300. The digest pull bypasses the Docker-29.x OCI-index failure on newer NVCR tags. Override (any image reference, ideally also a digest) to pin to a different build — e.g. the monitoring stack's older default 4.2.0. No effect on CPU nodes. See [OPERATIONS.md §3.1](./OPERATIONS.md#31-dcgmexporterimage-the-default-and-when-to-change-it) | diff --git a/architectures/aws-pcs/docs/ROADMAP.md b/architectures/aws-pcs/docs/ROADMAP.md index d66e3d431..b86022142 100644 --- a/architectures/aws-pcs/docs/ROADMAP.md +++ b/architectures/aws-pcs/docs/ROADMAP.md @@ -8,10 +8,27 @@ Priority: 🔴 high · 🟡 medium · 🟢 low ## Templates & deployment -- [ ] 🟡 **Multi-AZ support in the prerequisites stack.** `ml-cluster-prerequisites.yaml` - currently creates a single private subnet, so `OpenZFSDeploymentType` excludes the - `MULTI_AZ` types. Add a second private subnet (and the related routing) to enable - Multi-AZ FSx and higher-availability deployments. +- [x] 🟡 **Multi-AZ support in the prerequisites stack.** `ml-cluster-prerequisites.yaml` + now supports up to 3 private-subnet AZs via `AdditionalSubnetAZ2`/`AdditionalSubnetAZ3` + (CIDR1 split into four /18 blocks; additional subnets share the primary AZ's single + NAT gateway). This unblocks `OpenZFSDeploymentType=MULTI_AZ` and higher-availability + layouts. *(Note: OpenZFS MULTI_AZ wiring of the 2nd subnet into the FSx resource is a + follow-up; the subnets + routing are in place.)* +- [ ] 🟡 **Targeted ODCR support for GPU node groups.** Today `CapacityReservationId` + on `add-cng-p5`/`add-cng-p6-b200`/`add-cng-p6-b300` is **Capacity Block for ML only** — + setting it forces `MarketType=capacity-block` and drops the placement group, so a + *targeted* On-Demand Capacity Reservation (ODCR) cannot be consumed (only "open" ODCRs, + via the empty/On-Demand path, work). Add a `CapacityReservationType` enum + (`none` | `capacity-block` | `targeted-odcr`) and branch the launch template: + `targeted-odcr` sets `CapacityReservationTarget` **without** `MarketType=capacity-block` + and **keeps** the placement group (On-Demand billing against the reservation). + `none`/`capacity-block` stay equivalent to today (backward compatible). Replaces the + current "do not put an ODCR ID here" caveat. Verification can be done **without GPU + capacity**: (1) static — create the GPU CNGs with `Min/MaxCount=0` and assert the + generated launch template's `CapacityReservationSpecification`/`InstanceMarketOptions`/ + `Placement` per type; (2) dynamic — the branch logic is instance-family-independent, so + exercise actual targeted-ODCR consumption (`InstanceLifecycle` empty = On-Demand, reserved + count decrements) on a cheap type (c6i/g5). - [ ] 🟢 **Trainium (Trn) validation.** Validate the templates on Trainium instances (e.g. trn1/trn2) — node group, EFA/networking, and a sample training run. - [ ] 🟡 **Graviton (arm64) CPU CNG support — `hpc7g` / `c7gn`.** EFA-capable arm64 @@ -21,7 +38,7 @@ Priority: 🔴 high · 🟡 medium · 🟢 low `/aws/service/pcs/ami/dlami-base-ubuntu2404/arm64/latest/ami-id` (verified via `aws ssm get-parameters-by-path`), so this is well-defined as a follow-up: branch `AmiId` resolution by the CNG's instance architecture (or expose an `arm64` toggle), - add an arm64 Enroot/Pyxis first-boot path (`scripts/install-enroot-pyxis.sh` is x86 + add an arm64 Enroot/Pyxis first-boot path (`assets/scripts/install-enroot-pyxis.sh` is x86 only today), and validate hpc7g + c7gn end-to-end on real hardware. - [ ] 🟡 **P6e-GB200 / P6e-GB300 (Grace-Blackwell) support.** Add node-group templates for the GB200/GB300 NVL instances (e.g. p6e-gb200.36xlarge). These are Grace (arm64) CPUs @@ -89,11 +106,13 @@ Priority: 🔴 high · 🟡 medium · 🟢 low ## User management -- [ ] 🟡 **Integrate a user-management backend (LDAP/AD).** Provide a way to manage cluster - users centrally instead of the single `ubuntu` user — e.g. integrate an LDAP/OpenLDAP or - AWS Managed Microsoft AD directory (see `1.architectures/6.ldap_server`) so login/compute - nodes authenticate against a shared directory (multi-user clusters, per-user home dirs, - Slurm accounting per user). +- [x] 🟡 **Integrate a user-management backend (LDAP/AD).** Done for OpenLDAP: + `DirectoryService=OpenLDAP-LoginNode` runs slapd on the login node (DB on shared + `/home/ldap-db`) with SSSD on all compute nodes (CPU + GPU). Users added via + `ldap-add-user` resolve cluster-wide; home dirs auto-create; Slurm sees LDAP users + transparently. See `docs/USER-MANAGEMENT.md`. *(Follow-up: managed-directory options + `DirectoryService=SimpleAD`/`ManagedAD` for multi-login-node / HA — the param enum is + already extensible.)* ## Monitoring @@ -101,6 +120,14 @@ Priority: 🔴 high · 🟡 medium · 🟢 low Prometheus + Amazon Managed Grafana as an alternative to the self-hosted stack on the login node (see `4.validation_and_observability/4.prometheus-grafana`), so users can use a managed backend instead of running the containers themselves. +- [x] 🟡 **Rename `DeployMonitoring` → `MonitoringStack` (enum).** Done in deploy-all: + `MonitoringStack: none | Prometheus-LoginNode` (default `Prometheus-LoginNode`), + aligning with the `DirectoryService` `-` pattern. `AMP-AMG`/`CloudWatch` + remain as future AllowedValues for the managed-monitoring item above. deploy-all + converts to the nested templates' `DeployMonitoring=true/false` internally, so + add-cng*.yaml are unchanged. **Breaking change** at the deploy-all interface + (bundled into the major-update PR alongside `GrafanaPublicAccessCidr`→`GrafanaAccessCidr` + and `SSHAccessCidr`). ## Testing / docs diff --git a/architectures/aws-pcs/docs/USER-MANAGEMENT.md b/architectures/aws-pcs/docs/USER-MANAGEMENT.md new file mode 100644 index 000000000..bd7850656 --- /dev/null +++ b/architectures/aws-pcs/docs/USER-MANAGEMENT.md @@ -0,0 +1,564 @@ +# User Management Guide + +This guide is written for **cluster administrators who may not be familiar with +LDAP**. It covers the day-to-day operations of managing users on a PCS +reference architecture cluster with `DirectoryService=OpenLDAP-LoginNode`. + +By default, the cluster runs as a single `ubuntu` user. When multi-user is +enabled, an OpenLDAP directory runs on the login node and provides centralized +POSIX user accounts visible on all nodes via SSSD. + +--- + +## Quick reference (common tasks) + +Run on the login node. `$ADMIN_PW` is the LDAP admin password from SSM (see +[Getting the admin password](#getting-the-admin-password)); the `LDAP_*` env +vars must be passed **inline to `sudo`** (`sudo VAR=... cmd`) — `sudo -E` alone +drops them under the default `env_reset`/`secure_path`. + +> **Slurm path matches `SlurmVersion`.** The `sacctmgr` / `sbatch` paths below use +> `slurm-25.11` (the default). If you deployed with `SlurmVersion=25.05`, replace +> `slurm-25.11` with `slurm-25.05` in every path (or just run `export +> PATH=/opt/aws/pcs/scheduler/slurm-$(ls /opt/aws/pcs/scheduler | sed 's/slurm-//')/bin:$PATH` +> once and drop the absolute path). + +| Task | Command (run on the login node) | +|---|---| +| Add a user | `sudo LDAP_ADMIN_PASSWORD="$ADMIN_PW" ldap-add-user.sh alice 10001 3000` | +| List all users | `ldapsearch -x -H ldap://localhost -b ou=People,dc=cluster,dc=internal uid` | +| Delete a user | `ldapdelete -x -H ldap://localhost -D cn=admin,dc=cluster,dc=internal -w "$ADMIN_PW" uid=alice,ou=People,dc=cluster,dc=internal` then `sudo sss_cache -E` | +| Reset a user's password | `ldappasswd -x -H ldap://localhost -D cn=admin,dc=cluster,dc=internal -w "$ADMIN_PW" -s NEWPASS uid=alice,ou=People,dc=cluster,dc=internal` | +| Add user to Slurm accounting | `sudo /opt/aws/pcs/scheduler/slurm-25.11/bin/sacctmgr -i add user alice Account=ml-team` (root = accounting admin) | +| Verify user on compute node | `srun -N1 -n1 -p cpu1 id alice` | + +`ldap-add-user.sh` is a helper script installed on the login node at +`/usr/local/bin/` (it wraps `ldapadd`+`ldappasswd` so you don't need the LDAP +syntax). List/delete/reset use the raw `ldap*` tools directly — the full +commands are documented in each section below. + +--- + +## How it works (overview) + +``` +┌─────────────────────────────────────────────────────────┐ +│ Login Node │ +│ │ +│ ┌──────────┐ ┌──────────┐ ┌──────────────┐ │ +│ │ slapd │────►│ SSSD │────►│ NSS / PAM │ │ +│ │ (OpenLDAP│ │ (cache) │ │ (getent, │ │ +│ │ server) │ │ │ │ login, su) │ │ +│ └──────────┘ └──────────┘ └──────────────┘ │ +│ │ │ +│ DB: /home/ldap-db/ (shared OpenZFS) │ +└───────┼──────────────────────────────────────────────────┘ + │ ldap://login-ip:389 + ▼ +┌─────────────────────────────────────────────────────────┐ +│ Compute Node │ +│ │ +│ ┌──────────┐ ┌──────────────┐ │ +│ │ SSSD │────►│ NSS / PAM │ │ +│ │ (client)│ │ (getent, │ │ +│ │ │ │ srun user) │ │ +│ └──────────┘ └──────────────┘ │ +└─────────────────────────────────────────────────────────┘ +``` + +**Key points:** +- Users are stored in the LDAP database on the login node +- The database lives on shared `/home` (OpenZFS NFS) — it survives login node restart/replacement +- Every node (login + compute) runs SSSD which queries LDAP for user info +- When you add a user in LDAP, they become visible on all nodes within seconds +- Home directories are auto-created at first login (shared `/home` on OpenZFS) +- Slurm sees LDAP users transparently — no Slurm configuration needed for user resolution + +> ⚠️ **Single login node only.** `OpenLDAP-LoginNode` runs the directory server +> on **one** login node, so keep the login node group at `MinCount=MaxCount=1` +> while the directory is enabled. Compute clients discover the server by its +> `directory-role=server` tag and the slapd database is a single MDB on shared +> `/home`; running two login nodes would give clients an ambiguous server and +> have two `slapd` processes open the same database files concurrently +> (corruption risk). If you need multiple login nodes or a highly-available +> directory, use a managed backend (the planned `SimpleAD` / `ManagedAD` +> `DirectoryService` options) rather than the login-node OpenLDAP. + +### How a compute node finds the LDAP server (tag-based discovery) + +This part is **not obvious**, so it's worth spelling out. A compute node does +**not** receive the login node's IP as a parameter — PCS launches the login and +compute node groups independently, and the login node's private IP isn't known +at template-synthesis time (and it changes if the login node is replaced). +Instead, discovery happens **at compute-node boot**, by EC2 tag lookup: + +1. When the directory is enabled, the login node group tags its instance + `directory-role=server` (alongside `pcs-cluster-id=`). The + compute node groups tag themselves `directory-role=client`. This + `directory-role` tag is **dedicated to the directory feature** — it is + deliberately *separate* from the monitoring stack's `monitoring-role` tag, so + the two features don't depend on each other. +2. On first boot, each compute node runs `setup-directory.sh client`, which + calls `aws ec2 describe-instances` filtering for + `tag:pcs-cluster-id=` + `tag:directory-role=server` + + `instance-state-name=running`, and reads the matching instance's + `PrivateIpAddress`. The `pcs-cluster-id` filter scopes the lookup to **this + cluster only**, so multiple PCS clusters can share one VPC without their + compute nodes finding the wrong cluster's LDAP server. (`CLUSTER_ID` is + passed from `${ClusterId}` in UserData; the script aborts client setup if it + is empty, rather than risk matching another cluster's server.) +3. That IP becomes the SSSD `ldap_uri` (`ldap://`). SSSD on the + compute node then resolves users from the login node's slapd. + +Implications to be aware of: + +- **Compute nodes need `ec2:DescribeInstances`** in their instance role (the + cluster IAM role already grants it). Without it, discovery fails and the node + boots without LDAP (check `/var/log/directory-setup.log` for + `could not discover directory server IP`). +- **The login (server) node must be running before a compute node boots** for + discovery to succeed. In normal deploy order the login node group comes up + first; a compute node that scales up later simply queries the + already-running server. +- **If the login node is replaced**, its new instance re-tags itself + `directory-role=server` and re-attaches to the same `/home/ldap-db`, so newly + booting compute nodes discover the new IP automatically. Already-running + compute nodes keep their cached `ldap_uri`; they pick up the new IP on their + next boot (or after an SSSD reconfigure). +- **An explicit override exists**: set `LDAP_SERVER_URI` (or `DIRECTORY_DNS_IPS`, + for the future managed-directory path) in the client's environment to skip the + tag lookup entirely — used by the SimpleAD/ManagedAD extension and handy for + debugging. + +--- + +## Enabling multi-user + +### Option 1: deploy-all (recommended) + +```bash +aws cloudformation create-stack \ + --stack-name my-cluster \ + --template-url https://awsome-distributed-ai.s3.amazonaws.com/templates/pcs-ml-cluster-deploy-all.yaml \ + --parameters \ + ParameterKey=PrimarySubnetAZ,ParameterValue=us-east-2b \ + ParameterKey=DirectoryService,ParameterValue=OpenLDAP-LoginNode \ + --capabilities CAPABILITY_IAM CAPABILITY_NAMED_IAM +``` + +That's it. The login node will have slapd running and compute nodes will be +configured as LDAP clients automatically at first boot. + +### Option 2: modular deployment + +Pass these to your add-cng.yaml stacks: +- Login CNG: `DirectoryService=OpenLDAP-LoginNode`, `DirectoryRole=server` +- Compute CNG: `DirectoryService=OpenLDAP-LoginNode`, `DirectoryRole=client` + +### Parameters + +| Parameter | Default | Description | +|---|---|---| +| `DirectoryService` | `none` | Set to `OpenLDAP-LoginNode` to enable multi-user | +| `DirectoryRole` | `none` | Auto-set by deploy-all: `server` for login, `client` for compute | +| `DirectoryDomainSuffix` | `dc=cluster,dc=internal` | LDAP base DN (change only if you need a different domain) | + +--- + +## Day-to-day operations + +All commands below run on the **login node** as root (`sudo`). + +### Getting the admin password + +The LDAP admin password is auto-generated at cluster creation and stored in +AWS Systems Manager Parameter Store: + +```bash +CLUSTER_ID= + +aws ssm get-parameter \ + --name "/pcs/${CLUSTER_ID}/ldap/admin-password" \ + --with-decryption \ + --query 'Parameter.Value' \ + --output text +``` + +> **If SSM is empty** (instance role lacked the permission at first boot): +> ```bash +> sudo cat /home/ldap-db/.admin-password +> ``` + +Store this password somewhere safe — you'll need it for all user management +operations. + +--- + +### Adding a user + +**Using the helper script** (recommended): + +```bash +# Usage: ldap-add-user.sh [ssh-public-key] +sudo LDAP_ADMIN_PASSWORD="" ldap-add-user.sh alice 10001 3000 +``` + +This creates the user with: +- Username: `alice` +- UID: `10001` (pick a unique number in range 10001–59999) +- GID: `3000` (= `clusterusers` group, the default) +- Home directory: `/home/alice` (auto-created on first login) +- Shell: `/bin/bash` +- A random initial password (printed to stdout) + +**With an SSH key** (user can log in immediately): + +```bash +sudo LDAP_ADMIN_PASSWORD="" ldap-add-user.sh alice 10001 3000 "ssh-rsa AAAA... alice@laptop" +``` + +**Verifying the user was created:** + +```bash +# On login node +getent passwd alice +# Expected: alice:*:10001:3000:alice:/home/alice:/bin/bash + +id alice +# Expected: uid=10001(alice) gid=3000(clusterusers) groups=3000(clusterusers) +``` + +--- + +### Adding multiple users (batch) + +Create a file `users.txt`: +``` +alice 10001 3000 ssh-rsa AAAA... +bob 10002 3000 ssh-rsa BBBB... +carol 10003 3000 +``` + +Then: +```bash +while read name uid gid key; do + sudo LDAP_ADMIN_PASSWORD="" ldap-add-user.sh "$name" "$uid" "$gid" "$key" +done < users.txt +``` + +--- + +### Listing all users + +```bash +# Simple list +ldapsearch -x -H ldap://localhost -b "ou=People,dc=cluster,dc=internal" \ + "(objectClass=posixAccount)" uid uidNumber | grep -E "^uid:|^uidNumber:" + +# Or just use getent (shows all LDAP users + system users) +getent passwd | awk -F: '$3 >= 10000 {print $1, $3, $6}' +``` + +--- + +### Deleting a user + +```bash +ADMIN_PW="" +ldapdelete -x -H ldap://localhost \ + -D "cn=admin,dc=cluster,dc=internal" \ + -w "$ADMIN_PW" \ + "uid=alice,ou=People,dc=cluster,dc=internal" +``` + +**Invalidate the SSSD cache** so the deletion takes effect immediately instead +of lingering until the cache entry's TTL expires. SSSD caches lookups (so users +still resolve during a brief LDAP outage), which means a freshly-deleted user +keeps resolving via `getent` until the cache is cleared. Run on the login node +**and** any running compute node: +```bash +sudo sss_cache -E # invalidate all cached entries (needs the sssd-tools package, pre-installed) +# across compute nodes: +srun -N -n bash -c 'sudo sss_cache -E' +``` + +Also remove from Slurm accounting: +```bash +sudo /opt/aws/pcs/scheduler/slurm-25.11/bin/sacctmgr -i remove user alice +``` + +The user's home directory (`/home/alice`) is NOT deleted automatically. +Remove it manually if needed: +```bash +sudo rm -rf /home/alice +``` + +--- + +### Resetting a user's password + +```bash +ADMIN_PW="" +NEW_PW="temporary-password-123" + +ldappasswd -x -H ldap://localhost \ + -D "cn=admin,dc=cluster,dc=internal" \ + -w "$ADMIN_PW" \ + -s "$NEW_PW" \ + "uid=alice,ou=People,dc=cluster,dc=internal" + +echo "New password for alice: $NEW_PW" +``` + +Tell the user to change it after login: +```bash +# User runs this after logging in +ldappasswd -x -H ldap://localhost \ + -D "uid=alice,ou=People,dc=cluster,dc=internal" \ + -W -s "my-new-password" \ + "uid=alice,ou=People,dc=cluster,dc=internal" +``` + +--- + +### Creating groups + +```bash +ADMIN_PW="" + +ldapadd -x -H ldap://localhost \ + -D "cn=admin,dc=cluster,dc=internal" \ + -w "$ADMIN_PW" << EOF +dn: cn=ml-team,ou=Groups,dc=cluster,dc=internal +objectClass: posixGroup +cn: ml-team +gidNumber: 3001 +memberUid: alice +memberUid: bob +EOF +``` + +### Adding a user to a group + +```bash +ldapmodify -x -H ldap://localhost \ + -D "cn=admin,dc=cluster,dc=internal" \ + -w "$ADMIN_PW" << EOF +dn: cn=ml-team,ou=Groups,dc=cluster,dc=internal +changetype: modify +add: memberUid +memberUid: carol +EOF +``` + +--- + +## Slurm accounting + +PCS manages the Slurm accounting database internally (enable it with +`ManagedAccounting=enabled` at deploy time). You just need to register users and +accounts. + +> **Run `sacctmgr` add/modify/remove as `root`.** In PCS managed accounting the +> Administrator is `root` — the default `ubuntu` user is not an accounting admin, +> so `sacctmgr -i add ...` as `ubuntu` fails with *"Only +> admins/operators/coordinators can add accounts"*. Use `sudo` with the full +> path (the Slurm bin dir isn't on root's `PATH` by default): + +```bash +SACCTMGR=/opt/aws/pcs/scheduler/slurm-25.11/bin/sacctmgr + +# Create a Slurm account (typically one per team or project) +sudo $SACCTMGR -i add account ml-team Description="ML Team" + +# Add LDAP users to the account +sudo $SACCTMGR -i add user alice Account=ml-team +sudo $SACCTMGR -i add user bob Account=ml-team + +# Verify (read-only — works as any user with the Slurm bin on PATH) +sudo $SACCTMGR show user alice bob format=User,Account,DefaultAccount +``` + +Read-only `sacct` / `sreport` / `sacctmgr show ...` work as `ubuntu`; only the +mutating `sacctmgr` verbs need `root`. + +> **Note:** if `AccountingPolicyEnforcement=none` (the default), users can +> submit jobs even without being registered in `sacctmgr`. Registration is +> needed for fairshare/priority and for `sacct` history to show the user name. + +--- + +## Verifying users on compute nodes + +After adding a user, verify they're visible on compute nodes: + +```bash +export PATH=/opt/aws/pcs/scheduler/slurm-25.11/bin:$PATH + +# Single node +srun -N 1 -n 1 -p cpu1 bash -c 'getent passwd alice; id alice' + +# All nodes +srun -N 4 -n 4 -p cpu1 bash -c 'echo "$(hostname): $(id alice)"' +``` + +If a user isn't visible yet (SSSD cache delay, typically <5 sec): +```bash +srun -N 1 -n 1 -p cpu1 bash -c 'sudo sss_cache -E; sleep 2; getent passwd alice' +``` + +--- + +## Running jobs as a specific user + +Users log in to the login node and submit jobs normally: + +```bash +# User 'alice' logs in via SSH and runs: +srun -p cpu1 -N 1 -n 1 bash -c 'whoami; hostname' +sbatch --partition=cpu1 my-training.sbatch +``` + +The job runs as `alice` (uid=10001) on the compute node. The user's home +directory `/home/alice` is visible on the compute node (shared OpenZFS). + +--- + +## Troubleshooting + +### "User not found" on compute node + +```bash +# Check SSSD is running on compute +srun -N 1 -n 1 bash -c 'systemctl status sssd | head -3' + +# Check LDAP connectivity from compute +srun -N 1 -n 1 bash -c 'ldapsearch -x -H ldap:// -b dc=cluster,dc=internal uid=alice' + +# Force cache refresh +srun -N 1 -n 1 bash -c 'sudo sss_cache -E; sudo systemctl restart sssd' +``` + +### slapd not running on login node + +```bash +sudo systemctl status slapd +sudo journalctl -u slapd -n 20 +# Check install log +cat /var/log/directory-setup.log +``` + +### "Invalid credentials" when running ldap commands + +You're using the wrong admin password. Retrieve it from SSM or the fallback +file (see [Getting the admin password](#getting-the-admin-password)). + +### Home directory not created + +```bash +# Check pam_mkhomedir is configured +grep pam_mkhomedir /etc/pam.d/common-session +# Expected: session optional pam_mkhomedir.so skel=/etc/skel umask=0022 + +# Manually create (should auto-create on next login) +sudo mkdir -p /home/alice +sudo chown alice:clusterusers /home/alice +``` + +### New compute node doesn't resolve users + +New compute nodes boot with the latest LaunchTemplate version, which includes +SSSD client setup. If a node was launched before `DirectoryService` was enabled +(e.g. during a stack update), it won't have SSSD. Terminate the node and let +PCS replace it with a new one. + +--- + +## UID/GID conventions + +| Range | Purpose | +|---|---| +| 0–999 | System users (do not use) | +| 1000 | `ubuntu` (DLAMI default user) | +| 3000 | `clusterusers` group (default GID for new users) | +| 3001+ | Additional groups (create as needed) | +| 10001–59999 | LDAP user UIDs | + +**Always specify UIDs explicitly** when creating users. This ensures +consistency across all nodes and NFS mounts. Do not rely on auto-increment. + +--- + +## Data persistence and backup + +| Data | Location | Survives node replacement? | Survives stack delete? | +|---|---|---|---| +| LDAP database | `/home/ldap-db/` (shared OpenZFS) | ✅ | ❌ (FSx deleted) | +| User home directories | `/home//` (shared OpenZFS) | ✅ | ❌ (FSx deleted) | +| Admin password | SSM Parameter Store | ✅ | ✅ | + +### Backup + +```bash +# Export LDAP database to a file (run periodically via cron) +sudo slapcat -l /home/ldap-backup-$(date +%Y%m%d).ldif +``` + +### Restore (on a fresh login node) + +```bash +sudo systemctl stop slapd +sudo slapadd -l /home/ldap-backup-YYYYMMDD.ldif +sudo chown -R openldap:openldap /home/ldap-db +sudo systemctl start slapd +``` + +--- + +## Access methods + +| Method | Best for | Setup required | +|---|---|---| +| **Direct SSH** (port 22) | Multi-user teams, VS Code/JupyterLab | SG rule opening port 22 to a CIDR | +| **SSH over SSM** | Security-sensitive environments | IAM credentials + SSM plugin per user | +| **SSM Session Manager** | Admin-only access | IAM credentials only | + +For multi-user clusters, **Direct SSH** is recommended. Users connect with +their SSH key that was added during user creation: + +```bash +ssh alice@ +``` + +--- + +## Template structure + +``` +deploy-all.yaml +├─► cluster.yaml (IAM role with ssm:PutParameter for /pcs//ldap/*) +├─► add-cng.yaml (login) → DirectoryRole=server → setup-directory.sh server +│ (installs slapd + configures SSSD locally) +└─► add-cng.yaml (compute) → DirectoryRole=client → setup-directory.sh client + (installs SSSD, discovers login node IP) +``` + +--- + +## Upgrading to AWS Simple AD (future) + +If you outgrow OpenLDAP (need HA, >50 users, Kerberos), the +`DirectoryService` parameter is designed for extension: + +```yaml +DirectoryService: SimpleAD # future AllowedValue +``` + +Migration path: +1. Export users: `slapcat > users.ldif` +2. Deploy Simple AD (separate stack, requires 2 AZs) +3. Import users +4. Redeploy cluster with `DirectoryService=SimpleAD` +5. Decommission slapd + +See [docs/ROADMAP.md](./ROADMAP.md) for tracking. diff --git a/architectures/aws-pcs/docs/architecture-components.md b/architectures/aws-pcs/docs/architecture-components.md deleted file mode 100644 index 3ca5e8af0..000000000 --- a/architectures/aws-pcs/docs/architecture-components.md +++ /dev/null @@ -1,122 +0,0 @@ -# AWS PCS Architecture Components for Diagram - -## Main Components (Required for Diagram) - -### Network Layer -- **VPC** with dual CIDR blocks (10.0.0.0/16 + 10.1.0.0/16) -- **Public Subnet** (NAT Gateway placement) -- **Private Subnet** (compute nodes, FSx placement) -- **S3 VPC Endpoint** (gateway endpoint for template/data access) - -### Storage Layer -- **FSx for Lustre** (scratch storage, high-speed I/O for ML training) - - PERSISTENT_2 deployment type - - Configurable throughput (125-1000 MB/s/TiB) - - LZ4 compression -- **FSx for OpenZFS** (home directories, persistent user data) - - Single-AZ HA deployment - - NFS exports with no_root_squash - - Automatic daily backups - -### Compute Layer -- **AWS PCS Cluster** (Slurm scheduler) - - Head node (managed service, not visible) - - Slurm versions: 25.05, 25.11 -- **Login Node Group** (public subnet) - - SSH/SSM access point - - Container tooling (Enroot/Pyxis) - - Monitoring stack (Prometheus, Grafana) -- **On-Demand Compute Node Group** (private subnet) - - CPU-based workloads - - Auto-scaling with Slurm - - Optional Enroot/Pyxis via UserData -- **GPU Compute Node Group - P5/P6** (private subnet) - - 32x EFA network interfaces - - Multi-node ML training - - Optional Enroot/Pyxis via UserData - - Dynamic scaling (MinCount=0) - -### Optional Components -- **EC2 Image Builder** (custom AMI with Enroot/Pyxis pre-installed) - - Scheduled builds (manual, weekly, monthly) - - Published to SSM Parameter Store -- **Monitoring Stack** (on login node) - - Prometheus (metrics collection) - - Grafana (visualization) - - DCGM Exporter (GPU metrics) - - Slurm OpenMetrics - -## Key Architectural Features - -### Dual Deployment Modes -1. **Custom AMI Mode** (Production) - - ImageBuilder pre-installs Enroot/Pyxis (~20-30 min build) - - Fast node boot (~3 minutes) - - Best for production clusters - -2. **UserData Installation Mode** (Testing) - - Enroot/Pyxis installed on first boot (~8-12 minutes) - - No AMI build required - - Best for rapid iteration and testing - -### Prerequisites Stack Reuse -- Network and storage resources can be deployed once (pcs-shared-infra) -- Multiple clusters reference shared infrastructure via CloudFormation Exports/Imports -- Reduces deployment time from 20-30 minutes to 5-10 minutes for subsequent clusters - -### Dual Tagging Strategy -- **ClusterName tag**: User-friendly name (CloudFormation stack name) -- **pcs-cluster-id tag**: Actual PCS cluster ID (e.g., pcs_i3bddqwdrp) -- Both applied to all instances (login nodes, compute nodes) -- Used by monitoring stack for IMDS metadata and SSM parameter paths - -## Data Flow - -1. **User Access**: - - Users → SSH/SSM → Login Node (public subnet) - -2. **Job Submission**: - - Login Node → Slurm Scheduler (PCS managed) - - Scheduler → Compute Nodes (private subnet) - -3. **Storage Access**: - - Compute Nodes → FSx for Lustre (scratch, training data) - - Compute Nodes → FSx for OpenZFS (home directories) - - Compute Nodes → S3 (via VPC endpoint, checkpoints/results) - -4. **Container Workflow**: - - Users → Login Node (Enroot/Pyxis commands) - - Slurm jobs → Pull container images (Docker Hub, NVIDIA NGC) - - Compute Nodes → Run containerized training jobs - -5. **Monitoring**: - - Compute Nodes → Prometheus (metrics push) - - Users → Grafana dashboard (via Login Node) - - Slurm → OpenMetrics exporter → Prometheus - -## Security Features - -- **Network Isolation**: Compute nodes in private subnet, NAT Gateway for outbound only -- **Security Group**: All-to-all EFA communication within cluster -- **IAM Roles**: Integrated instance profiles for S3/SSM access -- **SSM Session Manager**: SSH alternative for login node access - -## Diagram Layout Recommendations - -- **Top Layer**: User access (SSH/SSM) → Login Node -- **Middle Layer**: Slurm scheduler (PCS managed, abstract as service icon) -- **Compute Layer**: On-Demand + GPU node groups (separate swim lanes) -- **Storage Layer**: FSx for Lustre + OpenZFS (bottom or side panel) -- **S3 VPC Endpoint**: Connection from private subnet to S3 service -- **Optional Panel**: ImageBuilder + Monitoring stack (dashed border) - -## Components to Exclude from Diagram - -- Internet Gateway (generic networking) -- NAT Gateway (generic networking) -- Route Tables (generic networking) -- Elastic IP (generic networking) -- Security Group rules (detail level) -- IAM roles/policies (non-visual) -- CloudFormation stacks (deployment mechanism) -- SSM parameters (configuration storage) diff --git a/architectures/aws-pcs/tests/README.md b/architectures/aws-pcs/tests/README.md index 02d920552..1bbc375c5 100644 --- a/architectures/aws-pcs/tests/README.md +++ b/architectures/aws-pcs/tests/README.md @@ -1,917 +1,151 @@ # AWS PCS — Test & Validation Guide -This directory is a single guide for validating an AWS PCS cluster deployed from the -templates in [`../assets`](../assets). Each test below lists **what to run** and the -**expected result**. +Test procedures for validating an AWS PCS cluster deployed from the templates +in [`../assets`](../assets). -For non-test operational guidance (Slurm version trade-offs, AMI single-version rule, -`MonitoringVersion` migration, the `DcgmExporterImage` default and when to change it, -AMI pinning, FSx deployment-type ↔ throughput coupling), see -[`../docs/OPERATIONS.md`](../docs/OPERATIONS.md). - -Rather than ship its own copies, this guide reuses the repository's canonical benchmark -and training assets and documents only the **PCS-specific deltas** (queue/partition names, -running the Enroot import on the login node, putting caches on `/fsx`): - -| Stage | Canonical asset to use | PCS-specific delta documented here | -|---|---|---| -| NCCL `all_reduce` over EFA | [`micro-benchmarks/nccl-tests/slurm/nccl-tests-container.sbatch`](../../../micro-benchmarks/nccl-tests) | [Test 6](#test-6-nccl-multi-node-efa) — import on the login node, partition name | -| FSDP Llama-2 7B training | [`3.test_cases/pytorch/FSDP`](../../../3.test_cases/pytorch/FSDP) | [Test 7](#test-7-fsdp-sample-training) — venv **or** Enroot container; cache on `/fsx`, 2 nodes | -| GPU sanity (nvidia-smi) | (one `srun` line, no script) | [Test 5](#test-5-p5p6-gpu-multi-nic) | - -All Slurm commands run as the **`ubuntu`** user from the login node (SSM session or -SSH). Slurm binaries are only on `PATH` in a login shell — over SSM/SSH wrap commands -as `bash -lc "sinfo; squeue"`. - ---- - -## Pre-merge full test matrix (run before opening / updating the PR) - -Run this **complete set** before merging template or script changes. Several real bugs -only appeared in specific combinations (a 25.05-only `MetricsType` rejection, a Pyxis -SPANK plugin built for the wrong Slurm version, a first-boot post-install failure), so -do **not** assume a single 25.11 GPU run covers everything. - -| # | Dimension | What to cover | Why it matters | -|---|---|---|---| -| 1 | **Every Slurm version** | Deploy one cluster per `SlurmVersion` `AllowedValues` value (**25.05 and 25.11**) and run a Pyxis `--container-image` job on each | `MetricsType` is 25.11-only (25.05 cluster create fails if it's set unconditionally); the Pyxis SPANK plugin is ABI-locked to its Slurm version, so a wrong-version `spank_pyxis.so` stops slurmd. **Any change to `scripts/install-enroot-pyxis.sh` MUST be retested on all supported versions.** | -| 2 | **First-boot container runtime install** | Default `PostInstallScriptUrl` runs `install-enroot-pyxis.sh` at first boot; verify Enroot/Pyxis lands and a `--container-image` job works → [Test 2](#test-2-enrootpyxis-container-runtime-first-boot-install) | The default path used by every cluster that doesn't override `AmiId`. Bugs in `install-enroot-pyxis.sh` only show up on a clean first boot. | -| 3 | **First-boot from a clean deploy** | Validate on a **freshly deployed** cluster, not a node you hand-patched | Post-install runs during cloud-init *before* slurmd/profile.d/controller exist; bugs there (e.g. version detection, `set -e` aborts) only show on a clean first boot, not after a live re-run. | -| 4 | **CPU queue** | `DeployOnDemandCNG=true`; Pyxis container job on `cpu1` | Baseline; also the cheapest way to exercise items 1–3 without GPU capacity. | -| 5 | **Each GPU family** (as capacity allows) | `p5`/`p5e`/`p5en`, `p6-b200`, `p6-b300`: nvidia-smi, NCCL all_reduce, FSDP | EFA NIC layout and the dcgm-exporter image differ per family (see notes). | -| 6 | **Monitoring** | 6 login containers up, all Prometheus targets healthy, GPU dashboards populate on every supported GPU family with the default `DcgmExporterImage` (DCGM 4.5.2 by digest) | — | -| 7 | **Template lint** | `aws cloudformation validate-template` on every edited `assets/*.yaml` | Catches structural errors before a deploy round-trip. | -| 8 | **Pre-baked AMI build path** (when touched) | If `pcs-ready-dlami-with-enroot-pyxis.yaml` or any code it bakes (`scripts/install-enroot-pyxis.sh`) changes: build an AMI per supported `SlurmVersion`, then deploy a cluster with `AmiId=` + `PostInstallScriptUrl=""` and run a container job → [Test 8](#test-8-pre-baked-ami-build-standalone-dlami-template) | Independent path: the cluster stack does NOT run Image Builder, so an `install-enroot-pyxis.sh` fix is only in the AMI after a rebuild. The AMI is single-Slurm-version by design — `SlurmVersion` on the DLAMI stack must match the cluster's `SlurmVersion` (the SPANK plugin is ABI-locked). **Skip this row only if neither the AMI build template nor the install script changed.** | -| 9 | **CPU EFA path** (when touched) | If `add-cng.yaml`'s EFA wiring (`EnableEfa` / `EfaInterfaceCount` / `PlacementGroupName`) or the deploy-all forwarding (`OnDemand{EnableEfa,EfaInterfaceCount,PlacementGroupName}`) changes: deploy with `OnDemandEnableEfa=true` on at least one EFA-capable HPC type (e.g. hpc7a.96xlarge, 2 NICs), verify `lspci`/`fi_info` show the EFA NICs, and run a 2-node OSU `osu_mbw_mr` → [Test 9](#test-9-efa-on-cpu-hpc-instances-hpc6a--hpc7a--hpc8a) | EFA enables a different LaunchTemplate shape (`NetworkInterfaces` block with `InterfaceType=efa`, mutually exclusive with `SecurityGroupIds`); a regression here only shows on a real EFA deploy, not template-validate. **Skip this row only if no EFA-related wiring was touched.** | -| 10 | **FSx storage health + performance** | (A) Both filesystems mount with correct options; read/write sanity; FSx-side params match CFN. (B) Performance regression/improvement test: stat IOPS, sequential read/write BW, multi-node concurrent stat, flock correctness → [Test 10](#test-10-fsx-storage-health-and-performance) | Run Part A after every deploy; run Part B before/after any Lustre-related change (mount options, lctl tunables, stripe, FSxLustreEnableEfa, Lustre version). A regression >10% on any metric should block the change. | - -Tests 1–10 below are the per-item how-to. The single-cluster shortcut (one deploy that -covers monitoring + CPU + one GPU family) is fine for iterating; rows 8 and 9 are -separate paths run only when their inputs change. The full matrix above is the bar -for **merge**. - ---- - -## Coverage matrix - -What this guide covers, and the template/parameter that exercises it: - -| Dimension | Options to test | How | -|---|---|---| -| **Monitoring** | enabled (default) | `DeployMonitoring=true` → [Test 1](#test-1-monitoring-stack) | -| **Container runtime — default first-boot path** | UserData install on every node | Default `PostInstallScriptUrl` (no AMI override) → [Test 2](#test-2-enrootpyxis-container-runtime-first-boot-install) | -| **CPU nodes** | `c6i`/`c7i` etc. | `DeployOnDemandCNG=true` → [Test 3](#test-3-cpu-queue) | -| **Single-NIC GPU** | `g5`/`g6` | On-Demand CNG with a G-series type → [Test 4](#test-4-g-series-gpu-single-nic) | -| **Multi-NIC GPU** | `p5`/`p5e`/`p5en`, `p6-b200`, `p6-b300` | `DeployPseriesCNG=true` + `PseriesInstanceType` → [Test 5](#test-5-p5p6-gpu-multi-nic) | -| **NCCL / EFA** | 2-node all_reduce | [Test 6](#test-6-nccl-multi-node-efa) | -| **Sample training** | FSDP Llama-2 7B | [Test 7](#test-7-fsdp-sample-training) | -| **Container runtime — pre-baked AMI path** | Standalone DLAMI build + cluster pinned to its output | `pcs-ready-dlami-with-enroot-pyxis.yaml` (separate stack) → cluster with `AmiId=ami-xxx` + `PostInstallScriptUrl=""` → [Test 8](#test-8-pre-baked-ami-build-standalone-dlami-template) | -| **EFA on CPU HPC instances** | hpc6a (1 NIC), hpc7a / hpc8a (2 NIC) | `OnDemandEnableEfa=true` + `OnDemandEfaInterfaceCount=1\|2` (deploy-all) or `EnableEfa=true` + `EfaInterfaceCount=1\|2` (modular `add-cng.yaml`); auto-creates a per-CNG cluster placement group, override with `OnDemandPlacementGroupName` / `PlacementGroupName` to share → [Test 9](#test-9-efa-on-cpu-hpc-instances-hpc6a--hpc7a--hpc8a) | -| **FSx storage health + performance** | Lustre + OpenZFS mounts, read/write, FSx-side parameters honored; performance regression/improvement test for any Lustre-related change | Default deploy already mounts both filesystems; verify + benchmark → [Test 10](#test-10-fsx-storage-health-and-performance) | - -A single `pcs-ml-cluster-deploy-all.yaml` deploy with `DeployMonitoring=true`, -`DeployOnDemandCNG=true`, and `DeployPseriesCNG=true` exercises Tests 1–7 in one cluster. -Test 8 is a **separate** flow (independent stack for the AMI, then a cluster pinned to -its output); run it when `pcs-ready-dlami-with-enroot-pyxis.yaml` or -`scripts/install-enroot-pyxis.sh` change. See [the README](../README.md#5-usage-examples) -for deploy commands; set `PseriesInstanceType` to the GPU family you want to validate. - ---- - -## Verified configurations - -Key configurations validated on real hardware with these templates (representative -results; exact bandwidth/throughput vary with NCCL/EFA versions and message size): - -| Config | Region | Capacity | Monitoring | NCCL all_reduce (2-node peak busbw) | FSDP Llama-2 7B (2-node) | -|---|---|---|---|---|---| -| **2× p6-b200.48xlarge** (16× B200) | us-west-2 | Capacity Block | ✅ v2.9.1, 16 GPUs in Grafana | **~654 GB/s** @16 GiB (EFA, `found 8 nics`, `#wrong 0`) | **~223 TFLOPS/GPU, ~86k tok/s** | -| **2× p6-b300.48xlarge** (16× B300) | us-west-2 | Capacity Block | ✅ v2.9.1, 16 B300 GPUs in Grafana with the default `DcgmExporterImage` (DCGM 4.5.2 by digest) ‡ | **~751 GB/s** @64 GiB (EFA, `found 16 nics`, `#wrong 0`) † | **~205 TFLOPS/GPU, ~79k tok/s** (venv); **~193 TFLOPS/GPU** (container) | -| **2× p5.48xlarge** (16× H100) | us-east-2 | Capacity Block | ✅ | **~480 GB/s** (EFA, `found 32 nics`, `#wrong 0`) | ~60 TFLOPS/GPU | -| **Slurm 25.05, CPU + PostInstall** | us-west-2 | On-Demand | ✅ v2.9.1 (no `slurm_openmetrics` job §) | first-boot Pyxis OK, `srun --container-image=ubuntu:22.04` clean | n/a | -| **Slurm 25.11, CPU + PostInstall** | us-west-2 | On-Demand | ✅ v2.9.1 incl. Slurm OpenMetrics | first-boot Pyxis OK | n/a | -| **Login + CPU (`c6i`)** | us-west-2 / us-east-* | On-Demand | ✅ all targets up | n/a | n/a | -| **Grafana public access** (login-only SG) | us-west-2 | — | ✅ reachable at `https:///grafana/` from the allowed CIDR | — | — | - -### HPC EFA on CPU instances (`OnDemandEnableEfa=true`) - -OSU MPI micro-benchmarks 7.4 on 2 nodes, Slurm 25.11, AWS Open MPI 4.1.7 -(`/opt/amazon/openmpi`), libfabric provider `efa`. Tuned env: `FI_PROVIDER=efa`, -huge page on (default), `FI_EFA_FORK_SAFE=1`, `OMPI_MCA_pml=cm`, `OMPI_MCA_mtl=ofi`, -`OMPI_MCA_mtl_ofi_provider_include=efa`. All deploys used `OnDemandEnableEfa=true` -with the matching `OnDemandEfaInterfaceCount`; the cluster placement group was -auto-created per-CNG by `add-cng.yaml`. - -| Instance | NICs | Spec aggregate ¶ | `osu_latency` 1B | `osu_bw` peak (1 pair) | `osu_bibw` peak (1 pair) | `osu_mbw_mr` peak (16 pair × 2 nodes) | `osu_allreduce` 1B (32 ranks) | -|---|---:|---:|---:|---:|---:|---:|---:| -| **hpc6a.48xlarge** | 1 | 100 Gbps | 13.78 µs | 96.3 Gbps (12.04 GB/s) | 152 Gbps †† | **97.6 Gbps** (97.6% of spec) | 27.5 µs | -| **hpc7a.96xlarge** | 2 | 300 Gbps | 14.27 µs | 95.6 Gbps (11.95 GB/s) | 156 Gbps | **263 Gbps** (87.6% of spec) | 23.4 µs | -| **hpc8a.96xlarge** | 2 | 300 Gbps | **10.31 µs** ★ | **210 Gbps** (26.27 GB/s) | **363 Gbps** †† | **341 Gbps** ‡‡ | **17.4 µs** ★ | - -¶ AWS docs ([HPC instance specs](https://docs.aws.amazon.com/ec2/latest/instancetypes/hpc.html)) -"Baseline / Burst bandwidth (Gbps)" column. The notes also state "you must attach at -least 2 ENIs, to separate network cards, to achieve [aggregate] throughput" with a -per-ENI cap of 150 Gbps for hpc7a and 170 Gbps for hpc6id; **per-ENI cap for hpc8a is -not stated in AWS docs**. - -†† `osu_bibw` measures bidirectional bandwidth on a single pair (simultaneous -send+recv). It can exceed the uni-directional spec because it counts both -directions; not directly comparable to the "300 Gbps" aggregate spec. - -‡‡ hpc8a's `osu_mbw_mr` reading exceeds the docs' aggregate spec (300 Gbps). The -reading is reproducible across two independent runs (job 6 = 42.68 GB/s, job 7 = -43.85 GB/s), so it isn't a one-off artifact. Possible explanations include (a) AWS -docs spec is per-direction sustained while `osu_mbw_mr` reports forward-direction -peak, (b) Nitro v6 / EFAv4 efficiency on hpc8a leaves headroom above the published -number, (c) ack/control traffic in reverse direction inflates the receiver-side -counter view. **NIC-level Prometheus counters** (`node_amazonefa_tx_bytes` and -`rx_bytes`, available via the v2.7+ `efa-metrics.sh` textfile collector) can resolve -which is the case for a given run; in a 30s-update collector the wire-level peak -shows below the OSU-reported peak because the OSU phases for each message size are -shorter than the textfile sample window. - -★ hpc8a is fastest on every metric. The Nitro v6 / EFAv4 generation lift over hpc7a -(Nitro v4) shows up most clearly in latency (~30% better) and single-pair bandwidth -(~2× — single pair on hpc7a hits the 150 Gbps per-ENI cap, hpc8a does not appear to -hit one). - -**Tuning notes** (collected during these runs): -- **`osu_bw` (single pair) understates aggregate fabric bandwidth.** A single MPI - pair uses one libfabric endpoint; on instances with 2 NICs, the second NIC is idle. - Use `osu_mbw_mr -np 32 -N 16` for the realistic aggregate number; `osu_bw` only - matches the per-ENI cap. -- **Disabling huge pages costs throughput.** An earlier hpc7a run with - `FI_EFA_USE_HUGE_PAGE=0` showed 68 Gbps `osu_bw`; leaving the variable unset (= 1 - default) brought it to 95.6 Gbps on the same instance. -- **Open MPI 4.1.7 on the PCS-Ready DLAMI has no `pml=ofi`.** Set - `OMPI_MCA_pml=cm` and `OMPI_MCA_mtl=ofi` instead — the `cm` PML dispatches to the - OFI MTL. Setting `OMPI_MCA_pml=ofi` directly fails with "mca_pml_base_open() - failed → Returned 'Not found' (-13)". -- **EFA monitoring works with the default v2.9.1 stack.** The `efa-metrics.sh` - textfile collector emits `node_amazonefa_*` series for `tx_bytes`, `rx_bytes`, - `recv_bytes`, `send_bytes`, RDMA `read/write` bytes/work-requests, retransmits, - and unresponsive-remote-endpoint events. The Compute Node Details Grafana - dashboard has dedicated EFA panels for these. Sampling cadence is 30 sec - (textfile-collector timer), so short OSU sub-tests show below their wall-clock peak. - -### Stack creation times (measured, deploy-all) - -Wall-clock from `aws cloudformation create-stack` to top-level `CREATE_COMPLETE`, on a -warmed account in us-west-2 (no first-time-in-region provisioning). Useful for sizing -how long a deploy round-trip takes during a review cycle. - -| Configuration | First-boot path | Time | -|---|---|---| -| 25.05, CPU only | Default `PostInstallScriptUrl` (Enroot/Pyxis at first boot) | **~24m** | -| 25.11, CPU only | Default `PostInstallScriptUrl` (Enroot/Pyxis at first boot) | **~31m** | -| 25.11, GPU + CPU (Capacity Block) | Default `PostInstallScriptUrl` (Enroot/Pyxis at first boot) | **~44m** | -| 25.05/25.11, CPU only, **pre-baked AMI** | Custom AMI built **separately** via `pcs-ready-dlami-with-enroot-pyxis.yaml` (~30m one-time), then cluster deployed with `AmiId=` + `PostInstallScriptUrl=""` | AMI build ~30m + cluster ~25m | - -Notes: Prerequisites (VPC + dual FSx) is the long-pole on the cluster path (~20-25m); -the default first-boot path adds the per-boot Enroot/Pyxis install (~2-3m). Pre-baking -the AMI is a separate ~30m one-time job — its cost amortizes when you redeploy or -scale clusters that share the AMI. GPU adds the P-series CNG and the GPU node first boot. - -Notes: -- **Deploy path:** all of the above came up from `pcs-ml-cluster-deploy-all.yaml` with - the default first-boot Enroot/Pyxis installer + `DeployMonitoring=true`. -- **Container runtime:** validated via first-boot UserData install (the default); the - pre-baked-AMI path uses the same `install-enroot-pyxis.sh` baked into the DLAMI by - `pcs-ready-dlami-with-enroot-pyxis.yaml`. -- **Pre-baked AMI path validated** (us-west-2, CPU-only): build the AMI with the - standalone DLAMI template, take its `DLAMIforPCSAmiId` output, deploy the cluster - with `AmiId=` and `PostInstallScriptUrl=""`. ImageBuilder bakes Enroot 3.5.0 - + per-version Pyxis into the DLAMI; login/compute nodes boot ready, no first-boot - install. The AMI is single-Slurm-version on purpose — see - [docs/OPERATIONS.md §2](../docs/OPERATIONS.md#2-container-runtime-postinstall-vs-ami-build). -- **EFA interface count** is derived from the instance type (p5/p5e = 32, p5en = 16, - p6-b200 = 8, p6-b300 = 16-of-17); see [README GPU compute](../README.md#gpu-compute-p5p6). -- **FSDP runs both ways** — validated with a shared-`/fsx` **venv** (~200 TFLOPS/GPU) and - with an **Enroot/Pyxis container** (`CONTAINER_IMAGE=/fsx/pytorch-fsdp.sqsh`, ~193 - TFLOPS/GPU) on the same 2× p6-b300. See Test 7 for both. **FSDP loss** stays at ln(vocab) - in this smoke test — a known dataloader/vocab quirk of the test case, not a cluster issue. -- **† B300 NCCL bandwidth scales past 16 GiB.** A 2-node / **16 GiB** all_reduce reaches only - ~654 GB/s busbw — it doesn't saturate all 16 EFA cards. Re-measured with larger messages on - 2× p6-b300: **64 GiB → ~751 GB/s** busbw (`found 16 nics`, `#wrong 0`), still climbing, so - 16 GiB was indeed unsaturated. **128 GiB and 256 GiB OOM** (the all_reduce buffer exceeds - B300 GPU memory), so ~64 GiB is the practical max single-buffer size here; for a true peak, - scale to more nodes rather than larger buffers. -- **‡ B300 GPU metrics work with the default `DcgmExporterImage`** — a DCGM 4.5.2 - build pinned by digest (`nvcr.io/nvidia/k8s/dcgm-exporter@sha256:a7ad6547…`), - validated on 2× p6-b300. Override only if you need to pin to a different DCGM build; - see [docs/OPERATIONS.md §3.1](../docs/OPERATIONS.md#31-dcgmexporterimage-the-default-and-when-to-change-it). -- **§ Slurm OpenMetrics is 25.11+ only.** On 25.05 the Slurm dashboards stay empty (the - rest of monitoring works fine). See - [docs/OPERATIONS.md §1](../docs/OPERATIONS.md#1-slurm-version-selection). - ---- - -## Test 1: Monitoring stack - -With `DeployMonitoring=true` (default), Prometheus/Grafana/exporters install on the -login node and DCGM/node exporters on the compute nodes. - -```bash -# On the login node: -docker ps --format "table {{.Names}}\t{{.Status}}" # login: prometheus, grafana, nginx, cloudwatch-exporter, node-exporter, pushgateway -tail -5 /var/log/monitoring-install.log # ends "...complete (exit 0)" -ls -la /opt/aws-parallelcluster-monitoring # installed on node-local /opt (NOT /home) -curl -s http://localhost:9090/api/v1/targets | \ - python3 -c 'import sys,json;[print(t["labels"].get("instance"),t["health"]) for t in json.load(sys.stdin)["data"]["activeTargets"]]' -curl -s http://localhost:6817/metrics | head # Slurm OpenMetrics -``` - -**Expected:** the six login containers are `Up`; install log exits 0; the tree is under -`/opt`; all Prometheus targets `up`; Slurm OpenMetrics returns Prometheus-format text. -For dashboard access (SSM port-forward or public CIDR) see -[README §8 Monitoring](../README.md#8-monitoring). Use `MonitoringVersion=v2.9.1`+ on PCS. +For operational guidance (Slurm version trade-offs, AMI pinning, monitoring, +FSx tuning), see [`../docs/OPERATIONS.md`](../docs/OPERATIONS.md). --- -## Test 2: Enroot/Pyxis container runtime (first-boot install) - -This is the default path used by every cluster that doesn't override `AmiId`: -`PostInstallScriptUrl` runs `install-enroot-pyxis.sh` once on each node at first boot. -The pre-baked-AMI path is validated separately as [Test 8](#test-8-pre-baked-ami-build-standalone-dlami-template). - -Deploy `pcs-ml-cluster-deploy-all.yaml` with both `AmiId` and `PostInstallScriptUrl` -left at their defaults (so SSM auto-resolves the latest PCS-Ready DLAMI and -post-install runs the Enroot/Pyxis installer), then on any node: +## Pre-merge test matrix -```bash -which enroot # /usr/bin/enroot -ls /opt/aws/pcs/scheduler/slurm-*/lib/slurm/spank_pyxis.so # per-version Pyxis SPANK plugin -cat /etc/aws/pcs/scheduler/slurm-*/plugstack.conf.d/pyxis.conf # points at the matching .so -tail -1 /var/log/pcs-post-install.log # "...completed (exit 0)" -``` +Run this **complete set** before merging template or script changes. -**Expected:** `enroot` on `PATH`; a `spank_pyxis.so` under the **cluster's** Slurm version -dir, and the plugstack `pyxis.conf` referencing that exact path; post-install log exits 0. -The Test 1/6/7 container jobs are the functional proof that Pyxis works. - -> **⚠️ Regression-test rule for `scripts/install-enroot-pyxis.sh`.** This script has bitten -> us repeatedly in ways a single 25.11 GPU run does not catch. **Any change to it MUST be -> retested across the full matrix at the top of this guide**, specifically: -> - **All supported Slurm versions** (25.05 **and** 25.11). The Pyxis SPANK plugin is -> ABI-locked to its Slurm version — a plugin built for the wrong version stops slurmd from -> starting (`Incompatible Slurm plugin version`). The script builds Pyxis for the version -> passed in `PCS_SLURM_VERSION` and installs the `.so` to a per-version path; a regression -> here only shows on the *other* version. -> - **The pre-baked AMI path too** ([Test 8](#test-8-pre-baked-ami-build-standalone-dlami-template)). -> `pcs-ready-dlami-with-enroot-pyxis.yaml` carries its **own copy** of the Enroot/Pyxis -> steps in its Image Builder UserData — editing `install-enroot-pyxis.sh` does **not** -> change the AMI path until you rebuild. Build an AMI per supported `SlurmVersion`, -> deploy a cluster pinned to it (`AmiId=` + `PostInstallScriptUrl=""`), and -> run a container job. -> - **On a clean first boot**, not a hand-patched node — post-install runs before -> slurmd/profile.d/controller exist, and several bugs only appear there. +| # | Category | Tests | File | When to run | +|---|---|---|---|---| +| 0 | **Docs lint** | `bash tests/lint-docs.sh` — no stale/renamed param refs in docs, every deploy-all param documented, README anchors resolve (the docs counterpart of template-lint; runs in seconds, no AWS) | [`lint-docs.sh`](./lint-docs.sh) | Every PR (esp. param renames) | +| 1-3, 8 | **Infrastructure** | Monitoring stack, container runtime (first-boot + AMI build), template lint | [`infra-test.md`](./infra-test.md) | Every PR | +| 4-6 | **Compute** | CPU queue, GPU families (G/P5/P6), NCCL multi-node EFA | [`compute-test.md`](./compute-test.md) | Every PR | +| 7, 7b | **Training** | FSDP Llama-2 7B (HF-streamed) + Megatron-LM GPT-3 (TP/PP/DP, local data) | [`training-test.md`](./training-test.md) | GPU PRs | +| 9 | **HPC EFA** | EFA on CPU instances (hpc6a/hpc7a/hpc8a), OSU benchmarks | [`hpc-efa-test.md`](./hpc-efa-test.md) | EFA wiring changes | +| 10 | **Storage** | FSx health check + performance regression test (noatime benchmark) | [`storage-test.md`](./storage-test.md) | FSx / mount changes | +| 11-12 | **Multi-user** | OpenLDAP directory + Slurm managed accounting | [`multi-user-test.md`](./multi-user-test.md) | Directory / accounting changes | +| 13 | **GPU health** | GPU Cluster Health Check suite (DCGM, EFA, NVLink, NCCL thresholds) | [`gpu-healthcheck-test.md`](./gpu-healthcheck-test.md) | GPU CNG deploys | +| 14 | **IAM** | cluster-admin deploys+deletes (no `iam:CreatePolicy`); cluster-user is SSM-login-only and can't read the LDAP password | [`iam-test.md`](./iam-test.md) | IAM policy changes | --- -## Test 3: CPU queue - -Deployed by default as `cpu1` (`DeployOnDemandCNG=true`, `c6i.4xlarge`, 0–4 dynamic). +## Quick-start: single-cluster shortcut -```bash -sinfo # cpu1 partition present, nodes idle~ -srun --partition=cpu1 --nodes=1 hostname # a node powers up and runs -``` - -**Expected:** `cpu1` shows in `sinfo`; a dynamically-scaled node launches and the job -returns its hostname. +A single `pcs-ml-cluster-deploy-all.yaml` deploy with `MonitoringStack=Prometheus-LoginNode`, +`DeployOnDemandCNG=true`, and `DeployPseriesCNG=true` exercises Tests 1–7 in +one cluster. Tests 8-13 are separate paths run only when their inputs change. --- -## Test 4: G-series GPU (single NIC) +## Major-update PR — configurations run end-to-end on real hardware -Single-NIC GPU instances (`g5`/`g6`) use `add-cng.yaml` — deploy them as the On-Demand -CNG, e.g. `OnDemandInstanceType=g6.12xlarge`, `OnDemandQueueName=gpu-g6` (see -[README Example 2](../README.md#5-usage-examples)). +The major-update PR (IAM policies, multi-user OpenLDAP, `MonitoringStack` rename, +`SSHAccessCidr`/`GrafanaAccessCidr`, multi-AZ subnets, `OnDemandEfaInterfaceCount` +0/1/2 collapse, VPCName fixed to `${StackName}-VPC`, instance-role perms split +inline) was validated end-to-end in us-east-2 with a single `deploy-all` cluster: -```bash -srun --partition=gpu-g6 --nodes=1 --gres=gpu:1 \ - --container-image=docker://nvidia/cuda:12.4.1-base-ubuntu22.04 nvidia-smi ``` - -**Expected:** the container runs and `nvidia-smi` lists the node's GPU(s) — confirms -single-NIC GPU + Pyxis on a G-series queue. - ---- - -## Test 5: P5/P6 GPU (multi-NIC) - -Multi-NIC GPU node groups, selected automatically by `PseriesInstanceType` -(`p5.48xlarge`/`p5e`/`p5en` → `add-cng-p5.yaml`; `p6-b200.48xlarge` → `add-cng-p6-b200`; -`p6-b300.48xlarge` → `add-cng-p6-b300`). The EFA interface count is derived from the -type — no parameter to set. A one-line interactive `srun` is enough for a GPU sanity -check (no batch script needed); set `--partition` to your GPU queue: - -```bash -srun --partition=gpu-p6b200 --nodes=1 --gres=gpu:8 \ - --container-image=docker://nvidia/cuda:12.4.1-base-ubuntu22.04 nvidia-smi +PrimarySubnetAZ=us-east-2b AdditionalSubnetAZ2=us-east-2a AdditionalSubnetAZ3=us-east-2c +DirectoryService=OpenLDAP-LoginNode SSHAccessCidr=/32 +MonitoringStack=Prometheus-LoginNode ``` -**Expected:** `nvidia-smi` lists 8 GPUs of the expected model (H100 / H200 / B200 / -B300). This confirms the multi-NIC launch template booted and Pyxis works on the GPU -node. EFA itself is exercised by Test 6. - ---- - -## Test 6: NCCL multi-node (EFA) - -2-node × 8-GPU `all_reduce_perf` over EFA, using the repo's canonical launcher -[`micro-benchmarks/nccl-tests/slurm/nccl-tests-container.sbatch`](../../../micro-benchmarks/nccl-tests/slurm/nccl-tests-container.sbatch) -(it reads `$IMAGE`, default `/fsx/nccl-tests.sqsh`). Only two PCS-specific deltas: - -1. **Import the image on the login node** — `enroot import` builds its overlayfs on the - node-local root disk (the login node has a 300 GiB root via `RootVolumeSize`); FSx - Lustre can't host that overlay, so only the resulting `.sqsh` goes to shared `/fsx`. - Pin a specific image tag for reproducible numbers (don't use `latest`): - - ```bash - # On the login node (direct, not a batch job). enroot URI form is - # docker://[REGISTRY#]REPO:TAG — the registry needs a '#', or it 401s on Docker Hub. - TAG=cuda12.8.1-efa1.43.2-ofiv1.16.3-ncclv2.27.7-1-testsv2.16.9 - enroot import -o /fsx/nccl-tests.sqsh "docker://public.ecr.aws#hpc-cloud/nccl-tests:${TAG}" - ``` - -2. **Submit with your GPU queue** — set the partition to the queue you deployed - (`gpu-p5` / `gpu-p6b200` / `gpu-p6b300`); the canonical script defaults to 2 nodes, - 8 tasks/node: - - ```bash - cd /fsx && sbatch --partition=gpu-p6b200 \ - /fsx/awsome-distributed-ai/micro-benchmarks/nccl-tests/slurm/nccl-tests-container.sbatch - ``` - -**Expected** (in `nccl-all_reduce_perf_.out`): -- EFA is the provider: `NET/OFI Selected provider is efa, fabric is efa-direct (found N nics)` - (N = EFA interface count: 32 for p5/p5e, 16 for p5en, 8 for p6-b200, 16 for p6-b300). -- Correctness: `# Out of bounds values : 0 OK`, every size `#wrong: 0`. -- `busbw` rises with message size — on 2× p6-b300, ~654 GB/s at 16 GiB but **~751 GB/s at - 64 GiB** (16 cards aren't saturated at 16 GiB); ~480 GB/s on 2× p5. - -> **Sizing the sweep on B300.** The canonical script sweeps to `-e 16G`. On p6-b300 that -> under-reports peak bandwidth — raise it to `-e 64G` (edit the `all_reduce_perf` line) to -> see the cards saturate. **Don't go to 128 GiB/256 GiB**: the all_reduce buffer exceeds -> B300 GPU memory and the job is OOM-killed. For a higher peak, add nodes, not buffer size. - ---- - -## Test 7: FSDP sample training - -A short FSDP Llama-2 7B run, using the repo's canonical training case -[`3.test_cases/pytorch/FSDP`](../../../3.test_cases/pytorch/FSDP) and its -[`slurm/llama2_7b-training.sbatch`](../../../3.test_cases/pytorch/FSDP/slurm/llama2_7b-training.sbatch). -Follow that case's README; the only PCS-specific deltas are where things live on the -shared filesystems and the node count: - -1. **Build the venv on shared `/fsx`** so every compute node sees it (the canonical - `slurm/create_venv.sh` creates `./env` in place — run it from a `/fsx` checkout): - - ```bash - cd /fsx && git clone --depth 1 https://github.com/awslabs/awsome-distributed-ai.git - cd /fsx/awsome-distributed-ai/3.test_cases/pytorch/FSDP/slurm && bash create_venv.sh - export HF_TOKEN=hf_xxx # gated Llama-2 tokenizer - ``` - -2. **Keep the HuggingFace cache on `/fsx` (Lustre), not `/home` (NFS).** Concurrent rank - file-locking on NFS throws `OSError: [Errno 116] Stale file handle`; export - `HF_HOME=/fsx/.hf-cache` before submitting. - -3. **Submit 2 nodes** (the canonical sbatch defaults to 4) on your GPU queue. The venv must - be on `PATH` for `torchrun` to resolve on every node — point `PATH` at the shared venv via - `--export` (the canonical sbatch doesn't `activate` it): - - ```bash - sbatch --nodes=2 --partition=gpu-p6b200 \ - --export=ALL,PATH=/fsx/awsome-distributed-ai/3.test_cases/pytorch/FSDP/slurm/env/bin:$PATH,HF_HOME=/fsx/.hf-cache \ - llama2_7b-training.sbatch - ``` - -### Option B — run it in an Enroot/Pyxis container instead of the venv - -The same canonical sbatch switches to container mode when `CONTAINER_IMAGE` is set (it adds -`--container-image`/`--container-mounts` and runs `./train.py` inside). Build the image once -on the login node and submit with `CONTAINER_IMAGE` — no venv needed: - -```bash -# On the login node (300 GiB root disk + Docker), build + import to /fsx: -cd /fsx/awsome-distributed-ai/3.test_cases/pytorch/FSDP -sudo docker build -t fsdp:pytorch -f Dockerfile . -enroot import -o /fsx/pytorch-fsdp.sqsh dockerd://fsdp:pytorch - -# Submit (container mode; mounts $(pwd) into /fsx inside the container): -cd slurm && sbatch --nodes=2 --partition=gpu-p6b300 \ - --export=ALL,CONTAINER_IMAGE=/fsx/pytorch-fsdp.sqsh,HF_HOME=/fsx/.hf-cache,HF_TOKEN=hf_xxx \ - llama2_7b-training.sbatch -``` - -> If the Dockerfile's `FROM` tag (a `public.ecr.aws/hpc-cloud/nccl-tests` tag) has been -> rotated out of the registry, substitute a current tag from that repo before building. - -**Expected** (in `logs/llama2_7b-FSDP_.out`), either path: -- NCCL initializes over EFA (`found N nics`) and training logs ~100 steps + a validation - step, saving checkpoints under `./checkpoints`. -- Throughput per step, e.g. on 2× p6-b300 **~200 TFLOPS/GPU, ~77k tokens/s** (venv) / - **~193 TFLOPS/GPU** (container); ~60 TFLOPS on 2× p5/H100. (Loss is constant at ln(vocab) - in this smoke test — a known dataloader/vocab quirk of the test case, not a cluster problem.) - -> **Multi-NIC tip:** the canonical sbatch already sets `NCCL_SOCKET_IFNAME=^docker,lo,veth,eth` -> (NCCL auto-selects). Do **not** pin a single interface on P5/P6 — all NICs share one -> subnet and pinning one breaks the cross-node NCCL bootstrap ring. - ---- - -## Test 8: Pre-baked AMI build (standalone DLAMI template) - -**When to run:** when `pcs-ready-dlami-with-enroot-pyxis.yaml` or any code it bakes -in (`scripts/install-enroot-pyxis.sh`) changes — the cluster stack does NOT run -Image Builder, so a fix to the install script is only in the AMI after a rebuild. -Skip this test if you only touched the cluster templates. - -This is an **independent flow**, not a deploy-all parameter: build an AMI with the -standalone template, then deploy a cluster pinned to that AMI ID with -`PostInstallScriptUrl=""` so nothing else runs at boot. - -The AMI is **single-Slurm-version by design** (Pyxis SPANK plugin ABI is -version-locked) — so when you run this test, run it for **every supported -`SlurmVersion`** that the install-script change could affect (typically both 25.05 -and 25.11). - -### Step 1 — build the AMI (~30 min one-time per Slurm version) - -```bash -SLURM_VERSION=25.11 # repeat for 25.05 if relevant - -aws cloudformation create-stack \ - --stack-name pcs-dlami-${SLURM_VERSION/./} \ - --template-url https://midaisuk-llm-dev.s3.amazonaws.com/templates/pcs-ready-dlami-with-enroot-pyxis.yaml \ - --parameters ParameterKey=SlurmVersion,ParameterValue=${SLURM_VERSION} \ - --capabilities CAPABILITY_IAM CAPABILITY_NAMED_IAM \ - --profile claude --region us-west-2 - -aws cloudformation wait stack-create-complete \ - --stack-name pcs-dlami-${SLURM_VERSION/./} \ - --profile claude --region us-west-2 - -AMI_ID=$(aws cloudformation describe-stacks \ - --stack-name pcs-dlami-${SLURM_VERSION/./} \ - --query 'Stacks[0].Outputs[?OutputKey==`DLAMIforPCSAmiId`].OutputValue' \ - --output text --profile claude --region us-west-2) -echo "$AMI_ID" # ami-0xxxxxxxxxxxxxxxx -``` - -### Step 2 — deploy a cluster pinned to that AMI - -```bash -aws cloudformation create-stack \ - --stack-name pcs-amitest-${SLURM_VERSION/./} \ - --template-url https://midaisuk-llm-dev.s3.amazonaws.com/templates/pcs-ml-cluster-deploy-all.yaml \ - --parameters \ - ParameterKey=PrimarySubnetAZ,ParameterValue=us-west-2a \ - ParameterKey=SlurmVersion,ParameterValue=${SLURM_VERSION} \ - ParameterKey=AmiId,ParameterValue=${AMI_ID} \ - ParameterKey=PostInstallScriptUrl,ParameterValue= \ - --capabilities CAPABILITY_IAM CAPABILITY_NAMED_IAM \ - --profile claude --region us-west-2 -``` - -### Step 3 — verify the bake landed and a container job runs - -On any node from the new cluster: - -```bash -which enroot # /usr/bin/enroot (pre-baked) -ls /opt/aws/pcs/scheduler/slurm-${SLURM_VERSION}/lib/slurm/spank_pyxis.so # built for matching Slurm -cat /etc/aws/pcs/scheduler/slurm-${SLURM_VERSION}/plugstack.conf.d/pyxis.conf # references the .so -test ! -s /var/log/pcs-post-install.log && echo "post-install did not run (PostInstallScriptUrl='')" -``` - -Then a container job through the login node (same form as Test 2): - -```bash -srun --partition=cpu1 --nodes=1 --ntasks=1 \ - --container-image=ubuntu:22.04 bash -c "echo PYXIS_FROM_AMI_OK" -``` - -**Expected:** `enroot` and the per-version Pyxis files exist **without** the -post-install hook running (because `PostInstallScriptUrl=""`); the container job -prints `PYXIS_FROM_AMI_OK`. Slurmd starts cleanly (no `Incompatible Slurm plugin -version` in `journalctl -u slurmd`). - -### Step 4 — clean up - -```bash -aws cloudformation delete-stack --stack-name pcs-amitest-${SLURM_VERSION/./} --profile claude --region us-west-2 -aws cloudformation delete-stack --stack-name pcs-dlami-${SLURM_VERSION/./} --profile claude --region us-west-2 -``` - -The DLAMI stack's AMI itself is **not** automatically deregistered when the stack is -deleted — if you need to free its EBS snapshots, deregister the AMI manually -(`aws ec2 deregister-image --image-id $AMI_ID`) and delete the snapshot. - ---- - -## Test 9: EFA on CPU HPC instances (hpc6a / hpc7a / hpc8a) - -**When to run:** when `add-cng.yaml`'s EFA wiring (`EnableEfa`, `EfaInterfaceCount`, -`PlacementGroupName`) or the deploy-all forwarding params (`OnDemandEnableEfa`, -`OnDemandEfaInterfaceCount`, `OnDemandPlacementGroupName`) change. Skip if only -GPU/monitoring/AMI paths were touched. - -This validates that the on-demand CPU CNG actually launches with EFA NICs in a -cluster placement group, and that MPI / libfabric over EFA works end-to-end. The -verified-configurations table above documents the bandwidth numbers; this section -documents the **how-to** so a contributor can reproduce. - -### Step 1 — deploy with EFA on the CPU CNG - -```bash -# hpc7a / hpc8a have 2 EFA NICs; hpc6a has 1. -INSTANCE_TYPE=hpc7a.96xlarge -EFA_NICS=2 - -# AZ availability is region-specific — confirm with describe-instance-type-offerings: -# hpc7a is in us-east-2b (and others); hpc8a is in us-east-2b / eu-north-1 / ap-northeast-1; -# hpc6a is in us-east-2 (b) / us-west-2 / eu-west-1 etc. AZ MUST contain the type. -AWS_AZ=us-east-2b - -aws cloudformation create-stack \ - --stack-name pcs-hpc-efa \ - --region us-east-2 \ - --template-url https://awsome-distributed-ai.s3.amazonaws.com/templates/pcs-ml-cluster-deploy-all.yaml \ - --parameters \ - ParameterKey=PrimarySubnetAZ,ParameterValue=$AWS_AZ \ - ParameterKey=DeployOnDemandCNG,ParameterValue=true \ - ParameterKey=OnDemandInstanceType,ParameterValue=$INSTANCE_TYPE \ - ParameterKey=OnDemandCngName,ParameterValue=hpc \ - ParameterKey=OnDemandQueueName,ParameterValue=hpc \ - ParameterKey=OnDemandMinCount,ParameterValue=0 \ - ParameterKey=OnDemandMaxCount,ParameterValue=2 \ - ParameterKey=OnDemandEnableEfa,ParameterValue=true \ - ParameterKey=OnDemandEfaInterfaceCount,ParameterValue=$EFA_NICS \ - --capabilities CAPABILITY_IAM CAPABILITY_NAMED_IAM -``` - -The `OnDemandCNGStack` nested stack auto-creates a cluster placement group and -exposes the name as a stack output (`PlacementGroupName`). To share the PG across -multiple CNGs (heterogeneous tightly-coupled jobs), pass -`OnDemandPlacementGroupName=` instead. - -### Step 2 — verify EFA visibility on a compute node - -After the stack reaches CREATE_COMPLETE and an `srun` has woken up a node: - -```bash -# On a compute node (via Slurm srun from login): -srun -p hpc -N 1 -n 1 bash -c ' - /opt/amazon/efa/bin/fi_info -p efa | head -20 # provider=efa, FI_EP_RDM - lspci | grep -iE "EFA|Elastic" # 2 EFA + 2 ENA on hpc7a/hpc8a -' -``` - -Expected (hpc7a / hpc8a): -- `fi_info -p efa`: shows `efa-direct` and `efa` fabrics on `rdmap36s0` and - `rdmap42s0` (or matching device names). -- `lspci`: 2 lines `Elastic Fabric Adapter` (or `Device efa3` on hpc8a) + 2 lines - `Elastic Network Adapter`. - -### Step 3 — OSU MPI micro-benchmarks - -Build OSU 7.4 on `/fsx` once (shared across compute nodes) using the -PCS-Ready DLAMI's `/opt/amazon/openmpi`: - -```bash -mkdir -p /fsx/osu && cd /fsx/osu -curl -fL -o osu.tgz https://mvapich.cse.ohio-state.edu/download/mvapich/osu-micro-benchmarks-7.4.tar.gz -tar xf osu.tgz && cd osu-micro-benchmarks-7.4 -PATH=/opt/amazon/openmpi/bin:$PATH ./configure CC=mpicc CXX=mpicxx --prefix=/fsx/osu -PATH=/opt/amazon/openmpi/bin:$PATH make -j8 -``` - -Submit a 2-node sbatch with the AWS-tuned EFA env (the canonical reference -parameters, see "Tuning notes" in [Verified configurations](#hpc-efa-on-cpu-instances-ondemandenableefatrue)): - -```bash -cat > /fsx/osu/osu-bench.sbatch <<'EOF' -#!/bin/bash -#SBATCH --job-name=osu-efa -#SBATCH --partition=hpc -#SBATCH --nodes=2 -#SBATCH --exclusive -#SBATCH --output=/fsx/osu/logs/%x_%j.out - -set -ex -OSU=/fsx/osu/osu-micro-benchmarks-7.4 -export PATH=/opt/amazon/openmpi/bin:$PATH - -export FI_PROVIDER=efa -export FI_EFA_FORK_SAFE=1 # huge page stays at the default (=1, on) -export OMPI_MCA_pml=cm # AWS Open MPI 4.1.7 has no pml=ofi -export OMPI_MCA_mtl=ofi -export OMPI_MCA_mtl_ofi_provider_include=efa -export OMPI_MCA_btl=^openib,tcp - -MPI_X="-x FI_PROVIDER -x FI_EFA_FORK_SAFE \ - -x OMPI_MCA_pml -x OMPI_MCA_mtl -x OMPI_MCA_mtl_ofi_provider_include \ - -x OMPI_MCA_btl" - -mpirun -np 2 -N 1 $MPI_X $OSU/c/mpi/pt2pt/standard/osu_latency -mpirun -np 2 -N 1 $MPI_X $OSU/c/mpi/pt2pt/standard/osu_bw -mpirun -np 2 -N 1 $MPI_X $OSU/c/mpi/pt2pt/standard/osu_bibw -mpirun -np 32 -N 16 $MPI_X $OSU/c/mpi/pt2pt/standard/osu_mbw_mr -mpirun -np 32 -N 16 $MPI_X $OSU/c/mpi/collective/blocking/osu_allreduce -EOF - -mkdir -p /fsx/osu/logs -sbatch -p hpc /fsx/osu/osu-bench.sbatch -``` - -The reference numbers per instance type are in -[Verified configurations](#hpc-efa-on-cpu-instances-ondemandenableefatrue) above. - -### Step 4 — observe NIC-level traffic in Grafana (optional) - -The monitoring stack's Compute Node Details dashboard has dedicated EFA panels -sourced from the `node_amazonefa_*` metrics produced by the v2.7+ -`efa-metrics.sh` textfile collector: - -- RDMA Read / Write Throughput -- SRD Retransmitted Packets -- Work-Request Errors - -For per-NIC `tx_bytes` / `rx_bytes` rate during a benchmark, query Prometheus -directly: - -```promql -rate(node_amazonefa_tx_bytes[30s]) * 8 / 1e9 # Gbps per (instance, device) -sum by (instance) (rate(node_amazonefa_tx_bytes[30s])) * 8 / 1e9 # both NICs -``` - -(The textfile collector cadence is 30s; rates over windows shorter than that are -zeros most of the time. OSU sub-tests are also short — 10–30s each — so wall-clock -peak in Prometheus typically reads below the OSU-reported peak.) - -### Step 5 — clean up - -When done with EFA testing, just delete the CNG stack (or the whole deploy-all -stack). The auto-created cluster placement group is owned by the CNG stack and -is removed automatically. Slurm puts the EFA compute nodes to sleep on idle -(`SuspendTime`), so leaving the cluster up between benchmarks does not keep the -hpc7a/hpc8a instances running. - ---- - -## Test 10: FSx storage health and performance - -Run this test whenever Lustre-related template changes are made (mount options, -`lctl` tunables, stripe configuration, `FSxLustreEnableEfa`, Lustre version, -or any UserData change that touches `/fsx`). The test has two parts: - -- **Part A — Health check**: confirms the filesystem mounts, is usable, and - matches what CFN asked for. Run after every deploy; fast (~2 min). -- **Part B — Performance baseline + regression/improvement test**: measures - storage throughput under controlled conditions. Run before and after any - Lustre performance-related change to detect regressions or confirm - improvements. - ---- - -### Part A — Health check - -#### A1. Filesystems mounted on every node - -On the login node and at least one compute node (via `srun`): - -```bash -mount | grep -E ' /home | /fsx ' -df -h /home /fsx -cat /proc/mounts | grep lustre # confirm mount options (noatime, flock, lazystatfs) -``` - -Expected: -- `/fsx` mounted as type `lustre`, options include `noatime`, `flock`, `lazystatfs` -- `/home` mounted as type `nfs`, options include `nconnect=16,rsize=1048576,wsize=1048576` -- Both `df -h` show `Avail` > 0 - -Troubleshooting: if a mount is missing, check `/var/log/cloud-init-output.log`. - -#### A2. Read/write sanity - -```bash -# /fsx (Lustre) -dd if=/dev/zero of=/fsx/.healthcheck bs=1M count=1024 conv=fsync 2>&1 | tail -1 -dd if=/fsx/.healthcheck of=/dev/null bs=1M 2>&1 | tail -1 -rm /fsx/.healthcheck - -# /home (OpenZFS) -dd if=/dev/zero of=/home/ubuntu/.healthcheck bs=1M count=100 conv=fsync 2>&1 | tail -1 -dd if=/home/ubuntu/.healthcheck of=/dev/null bs=1M 2>&1 | tail -1 -rm /home/ubuntu/.healthcheck -``` - -#### A3. FSx-side parameters match CFN inputs - -```bash -FSX_ID=$(aws cloudformation describe-stacks --stack-name \ - --query 'Stacks[0].Outputs[?OutputKey==`FSxLustreFilesystemId`].OutputValue' \ - --region --output text) -aws fsx describe-file-systems --file-system-ids "$FSX_ID" --region \ - --query 'FileSystems[0].[StorageCapacity,StorageType,LustreConfiguration.[DeploymentType,PerUnitStorageThroughput,DataCompressionType,EfaEnabled,MetadataConfiguration.Mode]]' \ - --output text -``` - -Expected defaults: `1200 | SSD | PERSISTENT_2 | 250 | LZ4 | False | AUTOMATIC` - -When `FSxLustreEnableEfa=true`: `EfaEnabled = True`, `Capacity >= 19200` (at -PerUnitStorageThroughput=250). A CFN Rule on both the prerequisites and -deploy-all templates fails the stack at create time when combined with -PERSISTENT_1. - -#### A4. Storage dashboard in Grafana - -Open Grafana → Storage dashboard. Verify `/fsx` panels (Throughput, IOPS, -Free Capacity) populate within ~5 min of a workload. The `dd` from A2 is -enough to seed values. - ---- - -### Part B — Performance baseline and regression test - -Use this procedure to: -1. Record a **baseline** (before a change) -2. Apply the change (mount option, lctl tunable, stripe config, etc.) -3. Record the **after** measurement -4. Compare — confirm no regression and quantify improvement - -#### B1. Preparation (run once per test cluster) - -```bash -# 10K small files for metadata testing (stat / readdir) -STAT_DIR=/fsx/perf-bench/stat-10k -mkdir -p $STAT_DIR -seq 1 10000 | xargs -P 64 -I{} touch $STAT_DIR/file_{} - -# 10K × 4KB files for smallfile read testing -SF_DIR=/fsx/perf-bench/smallfile-10k -mkdir -p $SF_DIR -seq 1 10000 | xargs -P 64 -I{} dd if=/dev/urandom of=$SF_DIR/file_{} bs=4096 count=1 2>/dev/null - -# 4GB sequential file for throughput testing -dd if=/dev/zero of=/fsx/perf-bench/seq-4g bs=1M count=4096 oflag=direct -``` - -#### B2. Single-node benchmarks (login node) - -```bash -RESULTS=/fsx/perf-bench/results-$(date +%Y%m%d-%H%M)-${LABEL:-baseline} -mkdir -p $RESULTS - -# (1) df latency — measures statfs() performance -for i in $(seq 1 100); do - ts_s=$(date +%s%N); df /fsx >/dev/null; ts_e=$(date +%s%N) - echo $(( (ts_e - ts_s) / 1000000 )) -done > $RESULTS/df_latency_ms.txt -echo "df median: $(sort -n $RESULTS/df_latency_ms.txt | awk 'NR==50{print}') ms" - -# (2) stat 10K files — measures metadata read (MDS) throughput -echo 3 | sudo tee /proc/sys/vm/drop_caches >/dev/null; sleep 2 -ts_s=$(date +%s%N) -find /fsx/perf-bench/stat-10k -type f | xargs stat -c '%s' >/dev/null -ts_e=$(date +%s%N) -echo "stat 10K: $(( (ts_e - ts_s) / 1000000 )) ms" | tee $RESULTS/stat_10k.txt - -# (3) smallfile 10K × 4KB read — measures many-file open+read (Python imports, HF cache pattern) -echo 3 | sudo tee /proc/sys/vm/drop_caches >/dev/null; sleep 2 -ts_s=$(date +%s%N) -find /fsx/perf-bench/smallfile-10k -type f -exec cat {} + >/dev/null -ts_e=$(date +%s%N) -echo "smallfile 10K read: $(( (ts_e - ts_s) / 1000000 )) ms" | tee $RESULTS/smallfile_read.txt - -# (4) Sequential read 4GB — measures bulk data throughput (OST bandwidth) -echo 3 | sudo tee /proc/sys/vm/drop_caches >/dev/null; sleep 2 -dd if=/fsx/perf-bench/seq-4g of=/dev/null bs=1M 2>&1 | tee $RESULTS/dd_read.txt - -# (5) Sequential write 4GB -dd if=/dev/zero of=/fsx/perf-bench/seq-write-4g bs=1M count=4096 conv=fsync 2>&1 | tee $RESULTS/dd_write.txt -rm -f /fsx/perf-bench/seq-write-4g -``` - -#### B3. Multi-node benchmarks (compute nodes, via srun) - -These are the most sensitive tests for detecting regressions under concurrent -load — the typical ML training scenario. - -```bash -# Prerequisite: at least 4 compute nodes available -# sinfo -N should show 4 nodes idle or idle~ - -# (6) Multi-node stat — N nodes × M procs concurrent stat of 10K files -# This is the primary regression indicator for metadata changes. -srun -N 4 --ntasks-per-node=1 -p cpu1 bash -c \ - 'echo 3 | sudo tee /proc/sys/vm/drop_caches >/dev/null' -sleep 3 -ts_s=$(date +%s%N) -srun -N 4 --ntasks-per-node=4 -p cpu1 bash -c \ - "find /fsx/perf-bench/stat-10k -type f | xargs stat -c '%s' >/dev/null" -ts_e=$(date +%s%N) -echo "multi-node stat 16p: $(( (ts_e - ts_s) / 1000000 )) ms" | tee $RESULTS/multi_stat.txt - -# (7) Multi-node sequential read — all nodes read the same 4GB file -srun -N 4 --ntasks-per-node=1 -p cpu1 bash -c \ - 'echo 3 | sudo tee /proc/sys/vm/drop_caches >/dev/null' -sleep 3 -ts_s=$(date +%s%N) -srun -N 4 --ntasks-per-node=1 -p cpu1 bash -c \ - "dd if=/fsx/perf-bench/seq-4g of=/dev/null bs=1M 2>/dev/null" -ts_e=$(date +%s%N) -echo "multi-node read 4N: $(( (ts_e - ts_s) / 1000000 )) ms" | tee $RESULTS/multi_read.txt - -# (8) flock correctness — concurrent lock serialization -LOCK=/fsx/perf-bench/locktest -rm -f $LOCK -(flock -x 200; echo "lock1 $(date +%s%N)"; sleep 0.1; echo "unlock1 $(date +%s%N)") 200>$LOCK & -(sleep 0.01; flock -x 200; echo "lock2 $(date +%s%N)"; sleep 0.1) 200>$LOCK & -wait -echo "flock: PASS" | tee $RESULTS/flock.txt -``` - -#### B4. A/B comparison for mount options or lctl changes - -To compare two configurations (e.g. `noatime` vs `relatime`), use -`mount -o remount` to switch without a full redeploy: - -```bash -# Switch ALL nodes (login + compute) to config A -sudo mount -o remount,relatime /fsx -srun -N 4 --ntasks-per-node=1 -p cpu1 bash -c 'sudo mount -o remount,relatime /fsx' -# Run B2 + B3 with LABEL=before - -# Switch ALL nodes to config B -sudo mount -o remount,noatime /fsx -srun -N 4 --ntasks-per-node=1 -p cpu1 bash -c 'sudo mount -o remount,noatime /fsx' -# Run B2 + B3 with LABEL=after -``` - -For `lctl` changes: apply the tunable, drop caches, re-run. No remount needed. - -#### B5. Interpreting results - -| Metric | What it measures | Sensitive to | +| Feature | Verified | Result | |---|---|---| -| df latency | `statfs()` → all OSTs | `lazystatfs` mount option; OST count | -| stat 10K | MDS metadata read (stat RPCs) | `noatime`; `mdc.*.max_rpcs_in_flight`; `statahead_max` | -| smallfile 10K read | open + read + close × many files | `noatime`; `mdc.*.max_rpcs_in_flight`; read-ahead | -| dd seq read | Single-stream bulk data throughput | `osc.*.max_rpcs_in_flight`; `max_pages_per_rpc`; stripe; provisioned throughput | -| dd seq write | Single-stream write throughput | `osc.*.max_dirty_mb`; `max_rpcs_in_flight`; stripe | -| multi-node stat | Concurrent MDS load under contention | `noatime` (scales with node count); `mdc.*` tunables | -| multi-node read | Aggregate read bandwidth | Provisioned throughput; node count × per-client BW | -| flock | Locking correctness | `flock` mount option | - -**Regression criteria**: a >10% degradation on any metric (excluding `df` -latency which is <5 ms and subject to jitter) should block the change until -investigated. - -**Expected magnitudes for common changes**: -- `noatime` addition: stat -0% to -5% single-node; -4% to -30% multi-node - (scales with concurrency) -- `mdc.*.max_rpcs_in_flight` 8→64: stat -30% to -60%; smallfile -20% to -50% -- `osc.*.max_rpcs_in_flight` 8→64: dd read +50% to +200% (if provisioned - throughput allows) -- `lfs setstripe -c -1 -S 16M`: dd read/write +100% to +400% (if OSTs > 2) - ---- - -### Baseline results (2026-06-11, us-east-2, 1.2 TiB PERSISTENT_2 / 2 OST, c6i.4xlarge) - -#### Single-node - -| Metric | Value | Notes | +| **Multi-AZ subnets** | 3 private subnets, 3 distinct AZs, non-overlapping `/18` CIDRs (`10.1.0.0/18`, `10.1.64.0/18`, `10.1.128.0/18`); compute nodes actually scheduled into the additional-AZ subnets (`10.1.x`) | ✅ | +| **SSHAccessCidr** | `LoginAccessSecurityGroup` opens **port 22 only** (443 absent since `GrafanaAccessCidr` empty) from the given CIDR, attached to the login node only; direct `ssh` to the login public IP succeeds | ✅ | +| **MonitoringStack=Prometheus-LoginNode** | 6 login containers up (prometheus/grafana/nginx/cloudwatch-exporter/node-exporter/pushgateway) | ✅ | +| **OpenLDAP server (boot-time, automatic)** | `slapd` active, DB on `/home/ldap-db`, OUs + `clusterusers` gid 3000 created, admin password in SSM; `ldap-add-user` helper auto-installed to `/usr/local/bin`; apt/dpkg-lock wait survived first-boot unattended-upgrades | ✅ | +| **`directory-role` tag discovery** | login tagged `directory-role=server` (independent of `monitoring-role`); compute clients discover the server by that tag, scoped to `pcs-cluster-id` | ✅ | +| **Compute SSSD client (boot-time, automatic)** | fresh compute nodes auto-install SSSD + sssd-tools, resolve `testuser1` via the login LDAP — no manual step | ✅ | +| **Slurm job as LDAP user** | `srun` as `testuser1` runs with `uid=10001` on the compute node | ✅ | +| **Multi-node UID consistency** | two different compute nodes both resolve `testuser1`→`uid=10001` | ✅ | +| **User delete propagation** | `ldapdelete` + `sss_cache -E` removes the user from `getent` immediately (sssd-tools now installed) | ✅ | +| **Template lint** | `validate-template` passes on all 8 edited `assets/*.yaml` (incl. 3 GPU templates + 2 IAM stacks) | ✅ | +| **`OnDemandEfaInterfaceCount` (0/1/2 collapse)** | `=1` → hpc6a launch template has 1 EFA NIC; `=2` → hpc8a has 2 EFA NICs (`DeviceIndex 0/1`, `InterfaceType: efa`). Real EFA traffic on hpc8a (2 nodes, `FI_PROVIDER=efa`): `osu_bw` peak **26.3 GB/s** (~210 Gbps, 1 pair); `osu_mbw_mr` peak **42.9 GB/s** (~343 Gbps, 16 pairs/node multi-rail) | ✅ | +| **Slurm managed accounting + multi-user (Test 12)** | on a `ManagedAccounting=enabled` cluster: LDAP users alice/bob registered in `sacctmgr` (account=ml-team) as root admin; jobs submitted as each LDAP user complete and `sacct -a` records them under the correct `User`+`Account` (`scontrol show job`: `UserId=alice(10001) Account=ml-team`) | ✅ | + +--- + +## Region coverage + +The templates fetch nested stacks + boot scripts from a single S3 bucket +(`S3BucketName`) and resolve the PCS-Ready DLAMI from SSM per region. + +Columns: **Deploy** (deploy-all → CREATE_COMPLETE), **Mon** (6 monitoring +containers up on the login node), **Storage** (FSx Lustre `/fsx` + OpenZFS +`/home` created & mounted — with the OpenZFS deployment type that worked), +**Pyxis** (Enroot/Pyxis container job runs), **GPU** (a GPU CNG launched and ran +on reserved capacity — Capacity Block or ODCR; ✅ marks regions where this was +exercised in **earlier** GPU validation rounds, not necessarily on the row's +deploy/monitoring/storage verification date), **Verified** (date, UTC). + +| Region | Deploy | Mon | Storage (OpenZFS type) | Pyxis | GPU | Verified | +|---|---|---|---|---|---|---| +| **us-east-1** (N. Virginia) | ✅ | ✅ | ✅ `SINGLE_AZ_HA_2` | ✅ | ✅ | 2026-06-17 | +| **us-east-2** (Ohio) | ✅ | ✅ | ✅ `SINGLE_AZ_HA_2` | ✅ | ✅ | 2026-06-17 | +| **us-west-2** (Oregon) | ✅ | ✅ | ✅ `SINGLE_AZ_HA_2` | ✅ | ✅ | 2026-06-17 | +| **ap-northeast-1** (Tokyo) | ✅ | ✅ | ✅ `SINGLE_AZ_HA_2` | ✅ | — | 2026-06-17 | +| **ap-south-1** (Mumbai) | ✅ | ✅ | ✅ `SINGLE_AZ_HA_2` | ✅ | ✅ | 2026-06-17 | +| **ap-northeast-3** (Osaka) | ✅ | ✅ | ✅ `SINGLE_AZ_1` ¹ | ✅ | — | 2026-06-17 | +| **ap-southeast-1** (Singapore) | ✅ | ✅ | ✅ `SINGLE_AZ_HA_2` | ✅ | — | 2026-06-17 | +| **ap-southeast-2** (Sydney) | ✅ | ✅ | ✅ `SINGLE_AZ_HA_2` | ✅ | — | 2026-06-17 | +| **eu-central-1** (Frankfurt) | ✅ | ✅ | ✅ `SINGLE_AZ_HA_2` | ✅ | — | 2026-06-17 | +| **eu-north-1** (Stockholm) | ✅ | ✅ | ✅ `SINGLE_AZ_HA_2` | ✅ | — | 2026-06-17 | + +¹ **Osaka (ap-northeast-3) does not support the default `OpenZFSDeploymentType=SINGLE_AZ_HA_2`** +— the default deploy fails at the OpenZFS filesystem with +`Invalid deploymentType (BadRequest)`. Deploy there with +`OpenZFSDeploymentType=SINGLE_AZ_1` (+ a valid `HomeThroughput`, e.g. 256). This +is the documented "not available in every Region" case (see the parameter's +description in `ml-cluster-prerequisites.yaml`); `SINGLE_AZ_1` is available in +all regions and is the safe fallback. + +### Remaining PCS launch regions (not yet run) + +AWS PCS is available in 18 regions total (per +`/aws/service/global-infrastructure/services/pcs/regions`). Beyond the 10 tested +above, these are expected to work but have not been run — confirm +`LustreDeploymentType` / `OpenZFSDeploymentType` availability before relying on +the defaults (`PERSISTENT_2` / `SINGLE_AZ_HA_2`), as Osaka above shows: +`eu-south-1` (Milan), `eu-south-2` (Spain), `eu-west-1` (Ireland), `eu-west-2` +(London), `eu-west-3` (Paris), `sa-east-1` (São Paulo), `us-gov-east-1`, +`us-gov-west-1` (GovCloud). + +Cross-region note: nested-stack `TemplateURL` and the in-instance +`aws s3 cp` of boot scripts both work against an S3 bucket in a **different** +region (S3 global namespace; no `--region` needed) — verified ap-south-1 → +us-east-1 bucket. The PCS-Ready DLAMI SSM parameter +(`/aws/service/pcs/ami/dlami-base-ubuntu2404/x86_64/latest/ami-id`) resolves in +every region tested. `OpenZFSDeploymentType=SINGLE_AZ_HA_2` and +`LustreDeploymentType=PERSISTENT_2` (the defaults) were available in all five. + +> **Testing unpublished templates:** an empty `PostInstallScriptUrl` (the default) +> resolves to `s3:///scripts/install-enroot-pyxis.sh` — i.e. +> the boot script comes from the **same bucket as the nested templates**. So when +> testing changes that aren't in the public bucket yet, point `S3BucketName` at your own +> bucket **and** sync the scripts there (the `aws s3 sync assets/ …` step covers both); +> otherwise the first-boot fetch fails and nodes come up without Enroot/Pyxis +> (`/var/log/pcs-post-install.log` shows the failed `aws s3 cp`). The cluster still +> reaches CREATE_COMPLETE; only the container runtime is missing. See +> [docs/DEPLOY-TESTING.md](../docs/DEPLOY-TESTING.md). + +--- + +## Canonical assets reused (not duplicated) + +| Workload | Source in this repo | PCS-specific delta | |---|---|---| -| df latency (median) | 1 ms | lazystatfs server-default | -| stat 10K files | 1851 ms | 5,400 stat ops/sec | -| smallfile 10K × 4KB read | 9374 ms | 1,067 files/sec | -| dd sequential read 4GB | 620 MB/s | Near provisioned limit (1.2 TiB × 250 MB/s/TiB ÷ 1024 ≈ 293 MB/s theoretical baseline; burst to 620 due to FSx credit system) | -| dd sequential write 4GB | 613 MB/s | | - -#### Multi-node (4 × c6i.4xlarge) - -| Metric | relatime | noatime | Delta | -|---|---|---|---| -| 16-stream stat 10K files | 5033 ms | 4812 ms | **-4.4%** | +| NCCL `all_reduce` | [`micro-benchmarks/nccl-tests/slurm/nccl-tests-container.sbatch`](../../../micro-benchmarks/nccl-tests) | Partition name, Enroot import on login node | +| FSDP Llama-2 7B | [`3.test_cases/pytorch/FSDP`](../../../3.test_cases/pytorch/FSDP) | Cache on `/fsx`, 2 nodes | +| Megatron-LM GPT-3 (TP/PP/DP) | [`3.test_cases/megatron/megatron-lm`](../../../3.test_cases/megatron/megatron-lm) | Import `.sqsh` to `/fsx`, data under `/fsx/gpt2/`, 4 nodes | +| GPU Health Check | [`4.validation_and_observability/2.gpu-cluster-healthcheck`](../../../4.validation_and_observability/2.gpu-cluster-healthcheck) | sbatch wrapper, partition name | -#### Interpretation +--- -On a small filesystem (2 OSTs, MDS far from saturation), single-node deltas -are in the noise. Multi-node shows the beginning of MDS contention relief from -`noatime`. At production scale (64+ nodes, 10+ OSTs), improvements from -`noatime` + `mdc` tunables are expected to be 10–30× larger. +## Notes -These numbers serve as the **regression baseline** for this filesystem size. -When running the same tests on a larger filesystem (e.g. 19200 GiB / ~16 OSTs -for EFA testing), record a new baseline — absolute numbers will differ but the -relative before/after comparison remains valid. +- All Slurm commands run as `ubuntu` from the login node (SSM or SSH) +- Slurm binaries: `export PATH=/opt/aws/pcs/scheduler/slurm-25.11/bin:$PATH` +- Template lint: `aws cloudformation validate-template --template-body file://assets/.yaml` +- Regression criteria (storage): >10% degradation blocks the change --- @@ -924,13 +158,14 @@ aws cloudformation delete-stack --stack-name aws cloudformation wait stack-delete-complete --stack-name ``` -Nested stacks (and FSx) are deleted automatically — back up FSx data first. A Capacity -Block keeps billing for its full reserved window regardless and is not released by stack -deletion. - -## References +Nested stacks (and FSx) are deleted automatically — back up FSx data first. -- [AWS PCS Documentation](https://docs.aws.amazon.com/pcs/) -- [aws-parallelcluster-monitoring](https://github.com/aws-samples/aws-parallelcluster-monitoring) -- [Slurm OpenMetrics](https://slurm.schedmd.com/rest.html#openmetrics) · [DCGM Exporter](https://github.com/NVIDIA/dcgm-exporter) -- [AI/ML for AWS PCS Workshop](https://catalog.workshops.aws/ml-on-pcs/) +If DELETE_FAILED on CNG stacks (PCS timing dependency), delete PCS CNGs first: +```bash +CLUSTER_ID= +for cng in $(aws pcs list-compute-node-groups --cluster-identifier $CLUSTER_ID --query 'computeNodeGroups[].id' --output text); do + aws pcs delete-compute-node-group --cluster-identifier $CLUSTER_ID --compute-node-group-identifier $cng +done +sleep 60 +aws cloudformation delete-stack --stack-name +``` diff --git a/architectures/aws-pcs/tests/compute-test.md b/architectures/aws-pcs/tests/compute-test.md new file mode 100644 index 000000000..ab39ee9ec --- /dev/null +++ b/architectures/aws-pcs/tests/compute-test.md @@ -0,0 +1,93 @@ +# Compute Tests (Tests 4-6) + +Validates CPU queue, GPU families (single-NIC G-series + multi-NIC P-series), +and NCCL multi-node communication over EFA. + +--- + +## Test 4: G-series GPU (single NIC) + +Single-NIC GPU instances (`g5`/`g6`) use `add-cng.yaml` — deploy them as the On-Demand +CNG, e.g. `OnDemandInstanceType=g6.12xlarge`, `OnDemandQueueName=gpu-g6` (see +[README Example 2](../README.md#5-usage-examples)). + +```bash +srun --partition=gpu-g6 --nodes=1 --gres=gpu:1 \ + --container-image=docker://nvidia/cuda:12.4.1-base-ubuntu22.04 nvidia-smi +``` + +**Expected:** the container runs and `nvidia-smi` lists the node's GPU(s) — confirms +single-NIC GPU + Pyxis on a G-series queue. + +--- + +## Test 5: P5/P6 GPU (multi-NIC) + +Multi-NIC GPU node groups, selected automatically by `PseriesInstanceType` +(`p5.48xlarge`/`p5e`/`p5en` → `add-cng-p5.yaml`; `p6-b200.48xlarge` → `add-cng-p6-b200`; +`p6-b300.48xlarge` → `add-cng-p6-b300`). The EFA interface count is derived from the +type — no parameter to set. A one-line interactive `srun` is enough for a GPU sanity +check (no batch script needed); set `--partition` to your GPU queue: + +```bash +srun --partition=gpu-p6b200 --nodes=1 --gres=gpu:8 \ + --container-image=docker://nvidia/cuda:12.4.1-base-ubuntu22.04 nvidia-smi +``` + +**Expected:** `nvidia-smi` lists 8 GPUs of the expected model (H100 / H200 / B200 / +B300). This confirms the multi-NIC launch template booted and Pyxis works on the GPU +node. EFA itself is exercised by Test 6. + +--- + +## Test 6: NCCL multi-node (EFA) + +2-node × 8-GPU `all_reduce_perf` over EFA, using the repo's canonical launcher +[`micro-benchmarks/nccl-tests/slurm/nccl-tests-container.sbatch`](../../../micro-benchmarks/nccl-tests/slurm/nccl-tests-container.sbatch) +(it reads `$IMAGE`, default `/fsx/nccl-tests.sqsh`). Only two PCS-specific deltas: + +1. **Import the image on the login node** — `enroot import` builds its overlayfs on the + node-local root disk (the login node has a 300 GiB root via `RootVolumeSize`); FSx + Lustre can't host that overlay, so only the resulting `.sqsh` goes to shared `/fsx`. + Pin a specific image tag for reproducible numbers (don't use `latest`): + + ```bash + # On the login node (direct, not a batch job). enroot URI form is + # docker://[REGISTRY#]REPO:TAG — the registry needs a '#', or it 401s on Docker Hub. + TAG=cuda12.8.1-efa1.43.2-ofiv1.16.3-ncclv2.27.7-1-testsv2.16.9 + enroot import -o /fsx/nccl-tests.sqsh "docker://public.ecr.aws#hpc-cloud/nccl-tests:${TAG}" + ``` + +2. **Submit with your GPU queue** — set the partition to the queue you deployed + (`gpu-p5` / `gpu-p6b200` / `gpu-p6b300`); the canonical script defaults to 2 nodes, + 8 tasks/node: + + ```bash + cd /fsx && sbatch --partition=gpu-p6b200 \ + /fsx/awsome-distributed-ai/micro-benchmarks/nccl-tests/slurm/nccl-tests-container.sbatch + ``` + +**Expected** (in `nccl-all_reduce_perf_.out`): +- EFA is the provider: `NET/OFI Selected provider is efa, fabric is efa-direct (found N nics)` + (N = EFA interface count: 32 for p5/p5e, 16 for p5en, 8 for p6-b200, 16 for p6-b300). +- Correctness: `# Out of bounds values : 0 OK`, every size `#wrong: 0`. +- `busbw` rises with message size — on 2× p6-b300, ~654 GB/s at 16 GiB but **~751 GB/s at + 64 GiB** (the 16 EFA NICs aren't saturated at 16 GiB); ~480 GB/s on 2× p5. + +> **Sizing the sweep on B300.** The canonical script sweeps to `-e 16G`. On p6-b300 that +> under-reports peak bandwidth — raise it to `-e 64G` (edit the `all_reduce_perf` line) to +> see the cards saturate. **Don't go to 128 GiB/256 GiB**: the all_reduce buffer exceeds +> B300 GPU memory and the job is OOM-killed. For a higher peak, add nodes, not buffer size. + +### Verified — p6-b200 ×4 (32 GPU, ap-south-1) + +`all_reduce_perf` over EFA (8 EFA NICs/node, `NET/OFI provider efa, fabric +efa-direct`, NICs auto-selected). busbw by message size: 128 MiB 295, 256 MiB +325, 1 GiB 367, 4 GiB 375, **16 GiB 377 GB/s** (peak); 0 wrong / 0 out-of-bounds. +The curve is healthy (latency-bound small → bandwidth-bound large, asymptotes at +377), a reasonable effective busbw for p6-b200's ≈800 Gbps/node node-to-node EFA. +This is **not a peak-tuned number** — NICs auto-selected (no `NCCL_ALGO` / +topology tuning), general-purpose nccl-tests image; for a higher peak use a +B200-tuned NCCL build + the `topology-aware-nccl-tests` sbatch. + +--- diff --git a/architectures/aws-pcs/tests/gpu-healthcheck-test.md b/architectures/aws-pcs/tests/gpu-healthcheck-test.md new file mode 100644 index 000000000..ffe4e628f --- /dev/null +++ b/architectures/aws-pcs/tests/gpu-healthcheck-test.md @@ -0,0 +1,79 @@ +# Test 13: GPU Cluster Health Check + +Validates GPU hardware, EFA, and NVLink on a PCS GPU node group using the +repository's [GPU Cluster Health Check Suite](../../../4.validation_and_observability/2.gpu-cluster-healthcheck). + +**This page documents only the PCS-specific deltas.** For the suite itself — what +each check does, the severity classification (PASS / MONITOR / REBOOT / ISOLATE), +the lightweight vs intensive check levels, instance profiles, and the per-check +reference — see the suite's own +[README](../../../4.validation_and_observability/2.gpu-cluster-healthcheck/README.md). +Don't duplicate that here. + +**Prerequisites:** a deployed GPU CNG (P5/P5e/P5en/P6-B200/P6-B300) and the suite +available on shared `/fsx` (so every compute node sees the same scripts). + +--- + +## PCS-specific deltas + +1. **Stage the suite on `/fsx`** (shared) so all GPU nodes run the same copy: + ```bash + cd /fsx && git clone https://github.com/awslabs/awsome-distributed-ai.git --depth 1 + HC=/fsx/awsome-distributed-ai/4.validation_and_observability/2.gpu-cluster-healthcheck + ``` + +2. **Drive it through the PCS Slurm queue** — use `srun`/`sbatch` against your GPU + partition name (e.g. `gpu-p5`, `gpu-b200`), not the suite's bare host invocation: + ```bash + export PATH=/opt/aws/pcs/scheduler/slurm-25.11/bin:$PATH + # lightweight suite (checks 0-3) on one GPU node + srun -p -N1 -n1 --exclusive bash $HC/gpu-healthcheck.sh --suite lightweight + # multi-node NCCL (check 5) — runs inside the Slurm allocation + srun -p -N2 -n2 --exclusive bash $HC/gpu-healthcheck.sh --check 5 + ``` + The suite's `slurm/` directory also ships `sbatch` wrappers and a + `prolog-gpu-healthcheck.sh` you can wire into PCS as a Slurm prolog (see the + suite README's Slurm section) — no PCS-specific change needed there. + +3. **When to run on a PCS cluster:** after any GPU CNG deploy or instance + replacement (lightweight), and before long multi-day training runs + (lightweight + check 5). Check 5 overlaps with [Test 6 NCCL](./compute-test.md) + but adds the suite's per-instance bandwidth thresholds. + +--- + +## Verified on real hardware + +Lightweight suite (checks 0-3) on **p6-b200** (ap-south-1, via `srun -p gpu`): +**4/4 PASS** — + +| Check | Result | +|---|---| +| 0 nvidia-smi | PASS — 8 GPUs, no Xid/SXid, no ECC/retired-page errors | +| 1 DCGM L2 | PASS — all Level-2 diagnostics | +| 2 EFA enumeration | PASS — 8 EFA PCI / 10 RDMA devices (`fi_info` warning is non-blocking: libfabric lives in the container, not on the host) | +| 3 Topology | PASS — 8 GPUs, connectivity validated (a B200 "unsupported P2P path" warning is expected and non-blocking) | + +EFA device count matching the instance profile and the multi-node NCCL bandwidth +threshold are also covered by [compute-test.md](./compute-test.md) (NCCL +all_reduce hit 377 GB/s on 4× p6-b200). + +### Intensive suite (checks 4-6) on p6-b200 ×2 (ap-south-1) + +The intensive suite (`--suite intensive`) needs **exclusive nodes for 1-3 hr** (DCGM L4 +alone is 45 min – 2.25 hr/node), so it's a maintenance-window / pre-long-run check, not a +per-deploy gate. + +- **Check 6 — EFA loopback: PASS** on 2× p6-b200 (8 EFA domains tested per node, both nodes). +- **Check 4 — DCGM L4:** budget the full 45 min+/node; it disables MIG, stops concurrent + GPU telemetry (`dcgm-exporter`), and runs the EUD + pulse-power stress, so it cannot share + a node with other GPU work. + +> **⚠️ Validate multi-node NCCL via [compute-test.md Test 6](./compute-test.md#test-6-nccl-multi-node-efa), +> not intensive check 5.** The suite's `checks/5-nccl-allreduce.sh` defaults to an ECR image +> URI that enroot rejects (needs the `#` registry separator, +> `docker://public.ecr.aws#hpc-cloud/nccl-tests:`), and invokes `all_reduce_perf` by +> bare name when that image keeps the binaries under `/opt/nccl-tests/build/` (not on +> `PATH`). The canonical `nccl-tests-container.sbatch` in Test 6 handles both correctly and +> ran to 377 GB/s on this cluster. diff --git a/architectures/aws-pcs/tests/hpc-efa-test.md b/architectures/aws-pcs/tests/hpc-efa-test.md new file mode 100644 index 000000000..5ed3fe1c7 --- /dev/null +++ b/architectures/aws-pcs/tests/hpc-efa-test.md @@ -0,0 +1,159 @@ +# HPC EFA Tests (Test 9) + +Validates EFA on CPU HPC instances (hpc6a/hpc7a/hpc8a) including placement +group auto-creation, multi-NIC wiring, and OSU MPI benchmark results. + +--- + +## Test 9: EFA on CPU HPC instances (hpc6a / hpc7a / hpc8a) + +**When to run:** when `add-cng.yaml`'s EFA wiring (`EfaInterfaceCount`, +`PlacementGroupName`) or the deploy-all forwarding params +(`OnDemandEfaInterfaceCount`, `OnDemandPlacementGroupName`) change. Skip if only +GPU/monitoring/AMI paths were touched. + +> **EFA is enabled by the interface count.** `OnDemandEfaInterfaceCount=0` +> (default) = no EFA; `1` or `2` enables EFA with that many NICs. There is no +> separate `OnDemandEnableEfa` flag (removed — the count alone drives it). + +This validates that the on-demand CPU CNG actually launches with EFA NICs in a +cluster placement group, and that MPI / libfabric over EFA works end-to-end. The +verified-configurations table above documents the bandwidth numbers; this section +documents the **how-to** so a contributor can reproduce. + +### Step 1 — deploy with EFA on the CPU CNG + +```bash +# hpc7a / hpc8a have 2 EFA NICs; hpc6a has 1. +INSTANCE_TYPE=hpc7a.96xlarge +EFA_NICS=2 + +# AZ availability is region-specific — confirm with describe-instance-type-offerings: +# hpc7a is in us-east-2b (and others); hpc8a is in us-east-2b / eu-north-1 / ap-northeast-1; +# hpc6a is in us-east-2 (b) / us-west-2 / eu-west-1 etc. AZ MUST contain the type. +AWS_AZ=us-east-2b + +aws cloudformation create-stack \ + --stack-name pcs-hpc-efa \ + --region us-east-2 \ + --template-url https://awsome-distributed-ai.s3.amazonaws.com/templates/pcs-ml-cluster-deploy-all.yaml \ + --parameters \ + ParameterKey=PrimarySubnetAZ,ParameterValue=$AWS_AZ \ + ParameterKey=DeployOnDemandCNG,ParameterValue=true \ + ParameterKey=OnDemandInstanceType,ParameterValue=$INSTANCE_TYPE \ + ParameterKey=OnDemandCngName,ParameterValue=hpc \ + ParameterKey=OnDemandQueueName,ParameterValue=hpc \ + ParameterKey=OnDemandMinCount,ParameterValue=0 \ + ParameterKey=OnDemandMaxCount,ParameterValue=2 \ + ParameterKey=OnDemandEfaInterfaceCount,ParameterValue=$EFA_NICS \ + --capabilities CAPABILITY_IAM CAPABILITY_NAMED_IAM +``` + +The `OnDemandCNGStack` nested stack auto-creates a cluster placement group and +exposes the name as a stack output (`PlacementGroupName`). To share the PG across +multiple CNGs (heterogeneous tightly-coupled jobs), pass +`OnDemandPlacementGroupName=` instead. + +### Step 2 — verify EFA visibility on a compute node + +After the stack reaches CREATE_COMPLETE and an `srun` has woken up a node: + +```bash +# On a compute node (via Slurm srun from login): +srun -p hpc -N 1 -n 1 bash -c ' + /opt/amazon/efa/bin/fi_info -p efa | head -20 # provider=efa, FI_EP_RDM + lspci | grep -iE "EFA|Elastic" # 2 EFA + 2 ENA on hpc7a/hpc8a +' +``` + +Expected (hpc7a / hpc8a): +- `fi_info -p efa`: shows `efa-direct` and `efa` fabrics on `rdmap36s0` and + `rdmap42s0` (or matching device names). +- `lspci`: 2 lines `Elastic Fabric Adapter` (or `Device efa3` on hpc8a) + 2 lines + `Elastic Network Adapter`. + +### Step 3 — OSU MPI micro-benchmarks + +Build OSU 7.4 on `/fsx` once (shared across compute nodes) using the +PCS-Ready DLAMI's `/opt/amazon/openmpi`: + +```bash +mkdir -p /fsx/osu && cd /fsx/osu +curl -fL -o osu.tgz https://mvapich.cse.ohio-state.edu/download/mvapich/osu-micro-benchmarks-7.4.tar.gz +tar xf osu.tgz && cd osu-micro-benchmarks-7.4 +PATH=/opt/amazon/openmpi/bin:$PATH ./configure CC=mpicc CXX=mpicxx --prefix=/fsx/osu +PATH=/opt/amazon/openmpi/bin:$PATH make -j8 +``` + +Submit a 2-node sbatch with the AWS-tuned EFA env (the canonical reference +parameters, see "Tuning notes" in the verified-configuration numbers in [tests/README.md](./README.md)): + +```bash +cat > /fsx/osu/osu-bench.sbatch <<'EOF' +#!/bin/bash +#SBATCH --job-name=osu-efa +#SBATCH --partition=hpc +#SBATCH --nodes=2 +#SBATCH --exclusive +#SBATCH --output=/fsx/osu/logs/%x_%j.out + +set -ex +OSU=/fsx/osu/osu-micro-benchmarks-7.4 +export PATH=/opt/amazon/openmpi/bin:$PATH + +export FI_PROVIDER=efa +export FI_EFA_FORK_SAFE=1 # huge page stays at the default (=1, on) +export OMPI_MCA_pml=cm # AWS Open MPI 4.1.7 has no pml=ofi +export OMPI_MCA_mtl=ofi +export OMPI_MCA_mtl_ofi_provider_include=efa +export OMPI_MCA_btl=^openib,tcp + +MPI_X="-x FI_PROVIDER -x FI_EFA_FORK_SAFE \ + -x OMPI_MCA_pml -x OMPI_MCA_mtl -x OMPI_MCA_mtl_ofi_provider_include \ + -x OMPI_MCA_btl" + +mpirun -np 2 -N 1 $MPI_X $OSU/c/mpi/pt2pt/standard/osu_latency +mpirun -np 2 -N 1 $MPI_X $OSU/c/mpi/pt2pt/standard/osu_bw +mpirun -np 2 -N 1 $MPI_X $OSU/c/mpi/pt2pt/standard/osu_bibw +mpirun -np 32 -N 16 $MPI_X $OSU/c/mpi/pt2pt/standard/osu_mbw_mr +mpirun -np 32 -N 16 $MPI_X $OSU/c/mpi/collective/blocking/osu_allreduce +EOF + +mkdir -p /fsx/osu/logs +sbatch -p hpc /fsx/osu/osu-bench.sbatch +``` + +The reference numbers per instance type are in +the verified-configuration numbers in [tests/README.md](./README.md) above. + +### Step 4 — observe NIC-level traffic in Grafana (optional) + +The monitoring stack's Compute Node Details dashboard has dedicated EFA panels +sourced from the `node_amazonefa_*` metrics produced by the v2.7+ +`efa-metrics.sh` textfile collector: + +- RDMA Read / Write Throughput +- SRD Retransmitted Packets +- Work-Request Errors + +For per-NIC `tx_bytes` / `rx_bytes` rate during a benchmark, query Prometheus +directly: + +```promql +rate(node_amazonefa_tx_bytes[30s]) * 8 / 1e9 # Gbps per (instance, device) +sum by (instance) (rate(node_amazonefa_tx_bytes[30s])) * 8 / 1e9 # both NICs +``` + +(The textfile collector cadence is 30s; rates over windows shorter than that are +zeros most of the time. OSU sub-tests are also short — 10–30s each — so wall-clock +peak in Prometheus typically reads below the OSU-reported peak.) + +### Step 5 — clean up + +When done with EFA testing, just delete the CNG stack (or the whole deploy-all +stack). The auto-created cluster placement group is owned by the CNG stack and +is removed automatically. Slurm puts the EFA compute nodes to sleep on idle +(`SuspendTime`), so leaving the cluster up between benchmarks does not keep the +hpc7a/hpc8a instances running. + +--- diff --git a/architectures/aws-pcs/tests/iam-test.md b/architectures/aws-pcs/tests/iam-test.md new file mode 100644 index 000000000..2b5a0691a --- /dev/null +++ b/architectures/aws-pcs/tests/iam-test.md @@ -0,0 +1,106 @@ +# Test 14: IAM Policies (cluster-admin / cluster-user) + +Validates the two IAM policy stacks ([`cluster-admin-iam.yaml`](../assets/cluster-admin-iam.yaml) +and [`cluster-user-iam.yaml`](../assets/cluster-user-iam.yaml), documented in +[docs/IAM.md](../docs/IAM.md)): + +1. A principal with **only** the cluster-admin policy can deploy and tear down a full + cluster — without `iam:CreatePolicy`. +2. A principal with **only** the cluster-user policy is constrained to SSM access on the + login node and cannot read the OpenLDAP admin password. + +This is the representative "two-role" use case: an admin who owns cluster lifecycle, and a +user who only runs jobs / views dashboards. + +--- + +## Setup + +Deploy both policy stacks, then create a throwaway role attached to **only** the admin +(or only the user) policy — nothing else — so the test reflects exactly what each policy +grants. + +```bash +# Deploy the policy + group stacks (see docs/IAM.md for the parameters) +aws cloudformation create-stack --stack-name pcs-iam-admin \ + --template-url https://.s3.amazonaws.com/cluster-admin-iam.yaml \ + --capabilities CAPABILITY_IAM CAPABILITY_NAMED_IAM --region +aws cloudformation create-stack --stack-name pcs-iam-user \ + --template-url https://.s3.amazonaws.com/cluster-user-iam.yaml \ + --capabilities CAPABILITY_IAM CAPABILITY_NAMED_IAM --region +``` + +Attach the resulting managed policy (admin or user) to a dedicated test role you can +`assume-role` into, or use `aws iam simulate-principal-policy` against the policy ARN for +a no-deploy check. + +--- + +## Test A — admin policy deploys and deletes a cluster + +Assume the admin-only role, then run a representative deploy (multi-user + SSH + monitoring +all on, to exercise the IAM, SSM, and KMS permissions): + +```bash +aws cloudformation create-stack --stack-name pcs-iam-test \ + --template-url https://.s3.amazonaws.com/pcs-ml-cluster-deploy-all.yaml \ + --parameters \ + ParameterKey=PrimarySubnetAZ,ParameterValue= \ + ParameterKey=DirectoryService,ParameterValue=OpenLDAP-LoginNode \ + ParameterKey=SSHAccessCidr,ParameterValue=/32 \ + ParameterKey=MonitoringStack,ParameterValue=Prometheus-LoginNode \ + --capabilities CAPABILITY_IAM CAPABILITY_NAMED_IAM --region +# ... wait for CREATE_COMPLETE, then: +aws cloudformation delete-stack --stack-name pcs-iam-test --region +``` + +**Expected:** +- Every nested stack reaches `CREATE_COMPLETE` with **no `AccessDenied`**. +- The same admin-only role tears the stack down to `DELETE_COMPLETE` with no `AccessDenied`. +- It works **without `iam:CreatePolicy`** — the instance role's permissions are attached + inline (`PutRolePolicy`), not as a managed policy, so deploy-all needs no policy-creation + right. + +No-deploy equivalent with `simulate-principal-policy` (pass proper resource ARNs and the +`iam:PassedToService` context): `cloudformation:CreateStack`, +`ec2:RunInstances`/`CreateSubnet`/`CreateSecurityGroup`, `fsx:CreateFileSystem`, +`pcs:CreateCluster`, `iam:CreateRole`/`PutRolePolicy`/`CreateInstanceProfile`/`PassRole` +(to `ec2` and `pcs`), `ssm:PutParameter`/`GetParameter` on `/pcs/*/ldap/*`, `kms:Decrypt` +→ all `allowed`; `iam:CreatePolicy` → `implicitDeny` (expected, and not needed). + +--- + +## Test B — user policy is constrained + +With the user-only role / policy: + +**Expected (allowed):** `pcs:GetCluster`/`ListComputeNodeGroups`, +`cloudformation:DescribeStacks`, `ec2:DescribeInstances`, `ssm:StartSession` **on the login +node**, and reading `/pcs/*/grafana/*` (the Grafana password). + +**Expected (denied — `implicitDeny`):** `cloudformation:CreateStack`, `pcs:CreateCluster`, +`ec2:RunInstances`, `fsx:CreateFileSystem`, `iam:CreateRole`; `ssm:StartSession` on a +**compute** node; and — critically for security — reading `/pcs/*/ldap/*` (the OpenLDAP +admin password). + +```bash +# Example simulate checks (repeat per action/resource): +aws iam simulate-principal-policy --policy-source-arn \ + --action-names ssm:StartSession \ + --resource-arns arn:aws:ec2:::instance/ +# → allowed for the login node, implicitDeny for a compute node +``` + +--- + +## Verified + +Run end-to-end on a real account (us-east-2) against the major-update templates, with a +dedicated role attached to only the admin (or user) policy: + +- Both IAM stacks `CREATE_COMPLETE`. +- Admin-only role: deploy-all (DirectoryService + SSHAccessCidr + monitoring) → + `CREATE_COMPLETE`, `delete-stack` → `DELETE_COMPLETE`, both with no `AccessDenied`, and + with `iam:CreatePolicy` simulating `implicitDeny` (confirming it isn't needed). +- User-only role: status/describe/login-SSM allowed; create/compute-SSM denied; can read + `/pcs/*/grafana/*` but **not** `/pcs/*/ldap/*`. diff --git a/architectures/aws-pcs/tests/infra-test.md b/architectures/aws-pcs/tests/infra-test.md new file mode 100644 index 000000000..ba1d9c6c6 --- /dev/null +++ b/architectures/aws-pcs/tests/infra-test.md @@ -0,0 +1,181 @@ +# Infrastructure Tests (Tests 1-3, 8) + +Validates cluster infrastructure: monitoring stack, container runtime (Enroot/Pyxis) +first-boot install, and the pre-baked AMI build path. + +--- + +## Test 1: Monitoring stack + +With `MonitoringStack=Prometheus-LoginNode` (default), Prometheus/Grafana/exporters install on the +login node and DCGM/node exporters on the compute nodes. + +```bash +# On the login node: +docker ps --format "table {{.Names}}\t{{.Status}}" # login: prometheus, grafana, nginx, cloudwatch-exporter, node-exporter, pushgateway +tail -5 /var/log/monitoring-install.log # ends "...complete (exit 0)" +ls -la /opt/aws-parallelcluster-monitoring # installed on node-local /opt (NOT /home) +curl -s http://localhost:9090/api/v1/targets | \ + python3 -c 'import sys,json;[print(t["labels"].get("instance"),t["health"]) for t in json.load(sys.stdin)["data"]["activeTargets"]]' +curl -s http://localhost:6817/metrics | head # Slurm OpenMetrics +``` + +**Expected:** the six login containers are `Up`; install log exits 0; the tree is under +`/opt`; all Prometheus targets `up`; Slurm OpenMetrics returns Prometheus-format text. +For dashboard access (SSM port-forward or public CIDR) see +[README §8 Monitoring](../README.md#8-monitoring). Use `MonitoringVersion=v2.9.1`+ on PCS. + +--- + +## Test 2: Enroot/Pyxis container runtime (first-boot install) + +This is the default path used by every cluster that doesn't override `AmiId`: +`PostInstallScriptUrl` runs `install-enroot-pyxis.sh` once on each node at first boot. +The pre-baked-AMI path is validated separately as [Test 8](#test-8-pre-baked-ami-build-standalone-dlami-template). + +Deploy `pcs-ml-cluster-deploy-all.yaml` with both `AmiId` and `PostInstallScriptUrl` +left at their defaults (so SSM auto-resolves the latest PCS-Ready DLAMI and +post-install runs the Enroot/Pyxis installer), then on any node: + +> **The default `PostInstallScriptUrl` is now an `s3://` URL** (empty → +> `s3:///scripts/install-enroot-pyxis.sh`, fetched with +> the instance role). Verified end-to-end against a **private** test bucket: the +> node's post-install log shows `Downloading post-install script from s3://…` and +> `pyxis.conf` is installed — so no public S3 is required (works in dev accounts). An +> `http(s)://` value still works too (curl), for public/GitHub-raw scripts. + +```bash +which enroot # /usr/bin/enroot +ls /opt/aws/pcs/scheduler/slurm-*/lib/slurm/spank_pyxis.so # per-version Pyxis SPANK plugin +cat /etc/aws/pcs/scheduler/slurm-*/plugstack.conf.d/pyxis.conf # points at the matching .so +tail -1 /var/log/pcs-post-install.log # "...completed (exit 0)" +``` + +**Expected:** `enroot` on `PATH`; a `spank_pyxis.so` under the **cluster's** Slurm version +dir, and the plugstack `pyxis.conf` referencing that exact path; post-install log exits 0. +The Test 1/6/7 container jobs are the functional proof that Pyxis works. + +> **⚠️ Regression-test rule for `assets/scripts/install-enroot-pyxis.sh`.** This script has bitten +> us repeatedly in ways a single 25.11 GPU run does not catch. **Any change to it MUST be +> retested across the full matrix at the top of this guide**, specifically: +> - **All supported Slurm versions** (25.05 **and** 25.11). The Pyxis SPANK plugin is +> ABI-locked to its Slurm version — a plugin built for the wrong version stops slurmd from +> starting (`Incompatible Slurm plugin version`). The script builds Pyxis for the version +> passed in `PCS_SLURM_VERSION` and installs the `.so` to a per-version path; a regression +> here only shows on the *other* version. +> - **The pre-baked AMI path too** ([Test 8](#test-8-pre-baked-ami-build-standalone-dlami-template)). +> `pcs-ready-dlami-with-enroot-pyxis.yaml` carries its **own copy** of the Enroot/Pyxis +> steps in its Image Builder UserData — editing `install-enroot-pyxis.sh` does **not** +> change the AMI path until you rebuild. Build an AMI per supported `SlurmVersion`, +> deploy a cluster pinned to it (`AmiId=` + `PostInstallScriptUrl=' '`), and +> run a container job. +> - **On a clean first boot**, not a hand-patched node — post-install runs before +> slurmd/profile.d/controller exist, and several bugs only appear there. + +--- + +## Test 3: CPU queue + +Deployed by default as `cpu1` (`DeployOnDemandCNG=true`, `c6i.4xlarge`, 0–4 dynamic). + +```bash +sinfo # cpu1 partition present, nodes idle~ +srun --partition=cpu1 --nodes=1 hostname # a node powers up and runs +``` + +**Expected:** `cpu1` shows in `sinfo`; a dynamically-scaled node launches and the job +returns its hostname. + +--- + +--- + +## Test 8: Pre-baked AMI build (standalone DLAMI template) + +**When to run:** when `pcs-ready-dlami-with-enroot-pyxis.yaml` or any code it bakes +in (`assets/scripts/install-enroot-pyxis.sh`) changes — the cluster stack does NOT run +Image Builder, so a fix to the install script is only in the AMI after a rebuild. +Skip this test if you only touched the cluster templates. + +This is an **independent flow**, not a deploy-all parameter: build an AMI with the +standalone template, then deploy a cluster pinned to that AMI ID with +`PostInstallScriptUrl=' '` so nothing else runs at boot. + +The AMI is **single-Slurm-version by design** (Pyxis SPANK plugin ABI is +version-locked) — so when you run this test, run it for **every supported +`SlurmVersion`** that the install-script change could affect (typically both 25.05 +and 25.11). + +### Step 1 — build the AMI (~30 min one-time per Slurm version) + +```bash +SLURM_VERSION=25.11 # repeat for 25.05 if relevant + +aws cloudformation create-stack \ + --stack-name pcs-dlami-${SLURM_VERSION/./} \ + --template-url https://.s3.amazonaws.com/pcs-ready-dlami-with-enroot-pyxis.yaml \ + --parameters ParameterKey=SlurmVersion,ParameterValue=${SLURM_VERSION} \ + --capabilities CAPABILITY_IAM CAPABILITY_NAMED_IAM \ + --region us-west-2 + +aws cloudformation wait stack-create-complete \ + --stack-name pcs-dlami-${SLURM_VERSION/./} \ + --region us-west-2 + +AMI_ID=$(aws cloudformation describe-stacks \ + --stack-name pcs-dlami-${SLURM_VERSION/./} \ + --query 'Stacks[0].Outputs[?OutputKey==`DLAMIforPCSAmiId`].OutputValue' \ + --output text --region us-west-2) +echo "$AMI_ID" # ami-0xxxxxxxxxxxxxxxx +``` + +### Step 2 — deploy a cluster pinned to that AMI + +```bash +aws cloudformation create-stack \ + --stack-name pcs-amitest-${SLURM_VERSION/./} \ + --template-url https://.s3.amazonaws.com/pcs-ml-cluster-deploy-all.yaml \ + --parameters \ + ParameterKey=PrimarySubnetAZ,ParameterValue=us-west-2a \ + ParameterKey=SlurmVersion,ParameterValue=${SLURM_VERSION} \ + ParameterKey=AmiId,ParameterValue=${AMI_ID} \ + ParameterKey=PostInstallScriptUrl,ParameterValue=' ' \ + --capabilities CAPABILITY_IAM CAPABILITY_NAMED_IAM \ + --region us-west-2 +``` + +### Step 3 — verify the bake landed and a container job runs + +On any node from the new cluster: + +```bash +which enroot # /usr/bin/enroot (pre-baked) +ls /opt/aws/pcs/scheduler/slurm-${SLURM_VERSION}/lib/slurm/spank_pyxis.so # built for matching Slurm +cat /etc/aws/pcs/scheduler/slurm-${SLURM_VERSION}/plugstack.conf.d/pyxis.conf # references the .so +test ! -s /var/log/pcs-post-install.log && echo "post-install did not run (PostInstallScriptUrl=' ')" +``` + +Then a container job through the login node (same form as Test 2): + +```bash +srun --partition=cpu1 --nodes=1 --ntasks=1 \ + --container-image=ubuntu:22.04 bash -c "echo PYXIS_FROM_AMI_OK" +``` + +**Expected:** `enroot` and the per-version Pyxis files exist **without** the +post-install hook running (because `PostInstallScriptUrl=' '`); the container job +prints `PYXIS_FROM_AMI_OK`. Slurmd starts cleanly (no `Incompatible Slurm plugin +version` in `journalctl -u slurmd`). + +### Step 4 — clean up + +```bash +aws cloudformation delete-stack --stack-name pcs-amitest-${SLURM_VERSION/./} --region us-west-2 +aws cloudformation delete-stack --stack-name pcs-dlami-${SLURM_VERSION/./} --region us-west-2 +``` + +The DLAMI stack's AMI itself is **not** automatically deregistered when the stack is +deleted — if you need to free its EBS snapshots, deregister the AMI manually +(`aws ec2 deregister-image --image-id $AMI_ID`) and delete the snapshot. + +--- diff --git a/architectures/aws-pcs/tests/lint-docs.sh b/architectures/aws-pcs/tests/lint-docs.sh new file mode 100755 index 000000000..1adf2b74f --- /dev/null +++ b/architectures/aws-pcs/tests/lint-docs.sh @@ -0,0 +1,79 @@ +#!/usr/bin/env bash +# Docs ⇄ template consistency lint for the AWS PCS reference architecture. +# +# Catches the most common drift after a parameter rename/removal: docs (and +# other templates) still referencing an old parameter name, a removed parameter +# presented as current, or a known-stale phrase. Run from anywhere; paths are +# resolved relative to this script's location (architectures/aws-pcs/). +# +# bash architectures/aws-pcs/tests/lint-docs.sh +# +# Exit code is non-zero if any check fails, so it can gate a PR in CI. +# +# When you intentionally rename/remove a parameter, update the BANNED list below +# in the same change — that is the point: the lint forces docs to keep up. + +set -uo pipefail + +ROOT="$(cd "$(dirname "${BASH_SOURCE[0]}")/.." && pwd)" # architectures/aws-pcs +cd "$ROOT" + +# Files that describe the user-facing interface (must not mention removed params). +DOC_GLOBS=(README.md docs/*.md tests/*.md) + +fail=0 +report() { echo "FAIL: $1"; fail=1; } + +# 1. Removed / renamed parameters must not appear in docs as if current. +# Each entry is TAB-separated: \t\t. +# A hit is a failure UNLESS the line also matches the allowed-context regex +# (used to permit explicit "(renamed from X)" / "→" / "internally" history +# notes). Use NEVERMATCH as the allow-regex when no exception applies. +BANNED=( + $'OnDemandEnableEfa\tremoved\tOnDemandEnableEfa was replaced by OnDemandEfaInterfaceCount (0/1/2)' + $'GrafanaPublicAccessCidr\t[Rr]enamed|→ ?`?GrafanaAccessCidr\tGrafanaPublicAccessCidr was renamed to GrafanaAccessCidr' + $'DeployMonitoring=true\tMonitoringStack|internally|nested\tDeployMonitoring (bool) was replaced by MonitoringStack at the deploy-all layer' + $'S3 (public )?hosting is not allowed\tNEVERMATCH\tPostInstallScriptUrl now accepts s3:// URLs' + $'architectures/aws-pcs/iam/\tNEVERMATCH\tthe iam/ directory was removed; use docs/IAM.md + assets/cluster-*-iam.yaml' +) +for entry in "${BANNED[@]}"; do + IFS=$'\t' read -r pat allow msg <<<"$entry" + hits=$(grep -rnE "$pat" "${DOC_GLOBS[@]}" 2>/dev/null | grep -vE "$allow" || true) + if [ -n "$hits" ]; then + report "stale reference ($msg):" + echo "$hits" | sed 's/^/ /' + fi +done + +# 2. PostInstallScriptUrl: docs must not say empty = skip (empty now = auto-install; +# a single space is the skip sentinel). The literal `PostInstallScriptUrl=""` +# (empty string) shown as a skip is the wrong pattern; the correct "single +# space to skip" wording is fine. +hits=$(grep -rnE 'PostInstallScriptUrl=""' "${DOC_GLOBS[@]}" 2>/dev/null || true) +if [ -n "$hits" ]; then + report 'PostInstallScriptUrl="" (empty) shown as skip — empty now auto-installs; use a single space to skip:' + echo "$hits" | sed 's/^/ /' +fi + +# 3. Every parameter the deploy-all template declares should be documented in +# PARAMETERS.md (catches a new param added without a docs row). +params=$(awk '/^Parameters:/{p=1;next} /^[A-Za-z]/{p=0} p&&/^ [A-Z][A-Za-z0-9]+:/{gsub(/[: ]/,"");print}' assets/pcs-ml-cluster-deploy-all.yaml | sort -u) +for prm in $params; do + grep -q "\`$prm\`" docs/PARAMETERS.md || report "deploy-all parameter '$prm' is not documented in docs/PARAMETERS.md" +done + +# 4. Same-file Markdown anchor links in README.md resolve to a real heading. +# (Cross-file and external links are out of scope — kept simple on purpose.) +while IFS= read -r anchor; do + # build the set of heading slugs in README + slugs=$(grep -E '^#{1,6} ' README.md \ + | sed -E 's/^#{1,6} //; s/`//g' \ + | tr '[:upper:]' '[:lower:]' \ + | sed -E 's/[^a-z0-9 -]//g; s/ /-/g') + grep -qx "$anchor" <<<"$slugs" || report "README.md internal anchor '#$anchor' has no matching heading" +done < <(grep -oE '\]\(#[a-z0-9-]+\)' README.md | sed -E 's/\]\(#//; s/\)//' | sort -u) + +if [ "$fail" -eq 0 ]; then + echo "docs lint: PASS (no stale parameter references, no empty=skip wording, all deploy-all params documented, README anchors resolve)" +fi +exit $fail diff --git a/architectures/aws-pcs/tests/multi-user-test.md b/architectures/aws-pcs/tests/multi-user-test.md new file mode 100644 index 000000000..3e5d1302c --- /dev/null +++ b/architectures/aws-pcs/tests/multi-user-test.md @@ -0,0 +1,481 @@ +# Tests 11-12: Multi-User Directory (OpenLDAP) + Slurm Accounting + +Run this test when `DirectoryService=OpenLDAP-LoginNode` is enabled. Validates that the +OpenLDAP directory on the login node is functional, that LDAP users are +visible on compute nodes via SSSD, and that Slurm can run jobs as those users. + +**Prerequisites**: a deployed cluster with `DirectoryService=OpenLDAP-LoginNode` on the login +CNG and at least one compute CNG. All commands run from the login node unless +noted otherwise. + +--- + +## Part A — LDAP server health (login node) + +### A1. slapd service running + +```bash +systemctl status slapd +``` + +Expected: `Active: active (running)` + +### A2. LDAP database on shared storage + +```bash +ls -la /home/ldap-db/ +``` + +Expected: MDB data files (`data.mdb`, `lock.mdb`) owned by `openldap:openldap`. +This location (shared OpenZFS `/home`) means the DB survives login node +replacement. + +### A3. Base DN searchable + +```bash +ldapsearch -x -H ldap://localhost -b "dc=cluster,dc=internal" -s base +``` + +Expected: returns the base entry without errors. + +### A4. Organizational units exist + +```bash +ldapsearch -x -H ldap://localhost -b "dc=cluster,dc=internal" "(objectClass=organizationalUnit)" dn +``` + +Expected: `ou=People,dc=cluster,dc=internal` and `ou=Groups,dc=cluster,dc=internal` + +### A5. Admin password in SSM + +```bash +CLUSTER_ID= +aws ssm get-parameter --name "/pcs/${CLUSTER_ID}/ldap/admin-password" \ + --with-decryption --query 'Parameter.Value' --output text +``` + +Expected: returns a non-empty string (the auto-generated password). + +--- + +## Part B — User lifecycle + +### B1. Create a test user + +```bash +export LDAP_ADMIN_PASSWORD=$(aws ssm get-parameter \ + --name "/pcs/${CLUSTER_ID}/ldap/admin-password" \ + --with-decryption --query 'Parameter.Value' --output text) +export LDAP_DOMAIN_SUFFIX="dc=cluster,dc=internal" + +# Using the helper script +sudo -E bash /usr/local/bin/ldap-add-user.sh testuser1 10001 3000 +``` + +Or manually: +```bash +ldapadd -x -H ldap://localhost -D "cn=admin,dc=cluster,dc=internal" -w "$LDAP_ADMIN_PASSWORD" < /home/testuser1/multinode-test.txt' +srun -N 1 -n 1 -p cpu1 --exclude=$(srun -N 1 -n 1 -p cpu1 hostname) bash -c 'cat /home/testuser1/multinode-test.txt' +``` + +Expected: the second node reads what the first wrote (shared OpenZFS `/home`). + +### D3. New compute node picks up existing users + +If a new compute node scales up after users were created: +```bash +# Force a new node to spin up +srun -N 2 -n 2 -p cpu1 bash -c 'getent passwd testuser1' +``` + +Expected: the freshly-started node resolves `testuser1` via SSSD (it queries +the login node's LDAP at boot via the client setup script). + +--- + +## Part E — LDAP server resilience + +### E1. SSSD cache survives brief LDAP outage + +```bash +# Stop slapd briefly on login node +sudo systemctl stop slapd + +# Compute node should still resolve from cache +srun -N 1 -n 1 -p cpu1 bash -c 'getent passwd testuser1' + +# Restart +sudo systemctl start slapd +``` + +Expected: cached user resolves even with slapd down (SSSD `cache_credentials=true`). + +### E2. LDAP DB survives login node replacement + +After login node terminate + PCS replacement: +```bash +# On new login node +ls /home/ldap-db/data.mdb +ldapsearch -x -H ldap://localhost -b "ou=People,dc=cluster,dc=internal" uid +``` + +Expected: previously-created users still in the directory (DB on shared OpenZFS +survived the instance replacement). + +--- + +## Verdict checklist + +| Check | Expected | +|---|---| +| slapd running on login node | ✅ | +| LDAP DB on /home/ldap-db (shared OpenZFS) | ✅ | +| Admin password in SSM Parameter Store | ✅ | +| User created via ldapadd/helper script | ✅ | +| User visible on login node (`getent passwd`) | ✅ | +| User visible on compute node (`srun getent passwd`) | ✅ | +| Home dir auto-created on first login | ✅ | +| User deleted, no longer resolvable | ✅ | +| Slurm job runs as LDAP user | ✅ | +| Multiple nodes resolve same UID | ✅ | +| Home dir accessible from all nodes | ✅ | +| New compute node resolves existing users | ✅ | +| SSSD cache works during brief LDAP outage | ✅ | +| LDAP DB survives login node replacement | ✅ | + +--- + +# Test: Slurm Managed Accounting + Multi-User + +Validates that Slurm managed accounting works with LDAP multi-user setup. +Covers user/account creation, resource limit enforcement, job tracking, and +reporting. + +**Prerequisites**: +- Cluster with `ManagedAccounting=enabled` and `DirectoryService=OpenLDAP-LoginNode` +- Slurm 25.11 (accounting is available on 24.11+; templates default to 25.11) +- At least one compute node available + +All commands run on the login node as root unless noted. + +--- + +## Part A — Accounting infrastructure + +### A1. Verify accounting is active + +```bash +export PATH=/opt/aws/pcs/scheduler/slurm-25.11/bin:$PATH +sacctmgr show cluster +``` + +Expected: cluster name listed with a valid `ControlHost`. + +### A2. Default account exists + +```bash +sacctmgr show account +``` + +If no accounts exist yet, create the default: +```bash +sacctmgr -i add account default Description="Default account" +``` + +--- + +## Part B — User + account registration + +### B1. Create LDAP users (if not done already) + +```bash +ADMIN_PW=$(aws ssm get-parameter --name "/pcs/${CLUSTER_ID}/ldap/admin-password" \ + --with-decryption --query 'Parameter.Value' --output text --region us-east-2) + +sudo LDAP_ADMIN_PASSWORD="$ADMIN_PW" ldap-add-user alice 10001 3000 +sudo LDAP_ADMIN_PASSWORD="$ADMIN_PW" ldap-add-user bob 10002 3000 +``` + +### B2. Register users in Slurm accounting + +```bash +sacctmgr -i add account ml-team Description="ML Research Team" +sacctmgr -i add user alice Account=ml-team +sacctmgr -i add user bob Account=ml-team + +# Verify +sacctmgr show user alice bob format=User,Account,DefaultAccount +``` + +Expected: +``` + User Account DefaultAccount +--------- ---------- -------------- + alice ml-team ml-team + bob ml-team ml-team +``` + +### B3. Set resource limits (optional) + +```bash +# Cap alice at 100 CPU-hours +sacctmgr -i modify user alice set GrpTRESRunMins=cpu=6000 + +# Verify +sacctmgr show user alice format=User,Account,GrpTRESRunMins +``` + +--- + +## Part C — Job submission and tracking + +### C1. Submit jobs as LDAP users + +```bash +# As alice +sudo su - alice -c 'export PATH=/opt/aws/pcs/scheduler/slurm-25.11/bin:$PATH; \ + sbatch --partition=cpu1 --wrap="sleep 10; hostname; id" -J alice-test' + +# As bob +sudo su - bob -c 'export PATH=/opt/aws/pcs/scheduler/slurm-25.11/bin:$PATH; \ + sbatch --partition=cpu1 --wrap="sleep 10; hostname; id" -J bob-test' +``` + +Wait for jobs to complete: +```bash +squeue # should show jobs running then empty +``` + +### C2. Verify job history with sacct + +```bash +sacct --starttime=$(date -d "1 hour ago" +%Y-%m-%dT%H:%M) \ + --format="JobID,User,JobName,Partition,Account,AllocCPUS,State,ExitCode,Elapsed" +``` + +Expected: both jobs listed with correct User, Account, State=COMPLETED. + +### C3. Per-user reporting + +```bash +# Jobs by alice +sacct -u alice --format="JobID,JobName,State,Start,End,Elapsed" + +# Utilization by account +sreport cluster AccountUtilizationByUser \ + start=$(date -d "1 hour ago" +%Y-%m-%dT%H:%M) \ + format="Account,Login,Used" +``` + +### C4. Resource limit enforcement (if B3 was set) + +```bash +# Submit a job that would exceed alice's limit +sudo su - alice -c 'export PATH=/opt/aws/pcs/scheduler/slurm-25.11/bin:$PATH; \ + sbatch --partition=cpu1 --ntasks=96 --time=120:00 --wrap="sleep 7200" -J limit-test' +``` + +Check if job is pending with reason `AssocGrpCPURunMinutesLimit`: +```bash +squeue -u alice --format="%i %j %T %r" +``` + +--- + +## Part D — Fairshare (optional) + +### D1. Set different fairshare weights + +```bash +sacctmgr -i modify account ml-team set FairShare=100 +sacctmgr -i add account ops-team Description="Operations" +sacctmgr -i modify account ops-team set FairShare=50 +``` + +### D2. Check fairshare values + +```bash +sshare -a --format=Account,User,RawShares,NormShares,RawUsage,FairShare +``` + +--- + +## Part E — AccountingPolicyEnforcement + +### E1. Test with enforcement=none (default) + +An unregistered user can still submit jobs: +```bash +# Create a user NOT in sacctmgr +sudo LDAP_ADMIN_PASSWORD="$ADMIN_PW" ldap-add-user charlie 10003 3000 +sudo su - charlie -c 'export PATH=/opt/aws/pcs/scheduler/slurm-25.11/bin:$PATH; \ + srun -p cpu1 -N1 -n1 hostname' +``` + +Expected: job runs (no accounting enforcement). + +### E2. Test with enforcement=associations,limits,safe + +If `AccountingPolicyEnforcement=associations,limits,safe` was set at cluster +creation: +```bash +# charlie (not in sacctmgr) should be rejected +sudo su - charlie -c 'export PATH=/opt/aws/pcs/scheduler/slurm-25.11/bin:$PATH; \ + srun -p cpu1 -N1 -n1 hostname' +``` + +Expected: `error: Unable to allocate resources: Invalid account or account/partition combination specified` + +--- + +## Verdict checklist + +| Check | Expected | +|---|---| +| sacctmgr show cluster → cluster listed | ✅ | +| Users registered in accounting (alice, bob) | ✅ | +| Jobs run as LDAP users complete successfully | ✅ | +| sacct shows correct user/account/state per job | ✅ | +| sreport shows per-user utilization | ✅ | +| Resource limit enforcement (if set) blocks over-limit jobs | ✅ | +| Unregistered user behavior matches AccountingPolicyEnforcement setting | ✅ | diff --git a/architectures/aws-pcs/tests/storage-test.md b/architectures/aws-pcs/tests/storage-test.md new file mode 100644 index 000000000..5c8de69a2 --- /dev/null +++ b/architectures/aws-pcs/tests/storage-test.md @@ -0,0 +1,131 @@ +# Storage Tests (Test 10) + +Validates FSx for Lustre + OpenZFS health, mount options, and performance +regression/improvement testing. + +--- + +## Test 10: FSx storage health + +Validates that both shared filesystems (Lustre on `/fsx`, OpenZFS on `/home`) +mount cleanly on every node, are usable, and that the FSx-side configuration +matches what the template asked for. Most of this is exercised implicitly by +Tests 1–9 (the post-install script, monitoring stack, OSU / FSDP all touch +`/fsx`); this test is the explicit health check to run after a fresh deploy +or after touching `ml-cluster-prerequisites.yaml` / FSx-related parameters. + +### Step 1 — both filesystems mounted on every node + +On the login node and at least one compute node (via `srun`): + +```bash +mount | grep -E ' /home | /fsx ' +df -h /home /fsx +``` + +Expected: +- `/fsx` mounted as type `lustre`, source `.fsx..amazonaws.com@tcp:/`, + size matches the `Capacity` parameter (1200 GiB default; 19200 GiB or larger + when `FSxLustreEnableEfa=true`). +- `/home` mounted as type `nfs` over OpenZFS, NFS options include + `nconnect=16,rsize=1048576,wsize=1048576` (the deploy-all UserData mount + string). +- Both `df -h` reports show `Avail` greater than zero. + +If a mount is missing on a freshly booted node, check +`/var/log/cloud-init-output.log` and `/var/log/pcs-post-install.log` for the +`mount` line. The most common first-boot failure is the OpenZFS DNS name not +being resolvable yet (NFS settle race); the post-install log will show +`mount.nfs: Failed to resolve server`. + +### Step 2 — read/write sanity + +```bash +# /fsx (Lustre): write 1 GiB, read it back +dd if=/dev/zero of=/fsx/.healthcheck bs=1M count=1024 conv=fsync 2>&1 | tail -1 +dd if=/fsx/.healthcheck of=/dev/null bs=1M 2>&1 | tail -1 +rm /fsx/.healthcheck + +# /home (OpenZFS / NFS): same, 100 MiB (it's a small home filesystem) +dd if=/dev/zero of=/home/ubuntu/.healthcheck bs=1M count=100 conv=fsync 2>&1 | tail -1 +dd if=/home/ubuntu/.healthcheck of=/dev/null bs=1M 2>&1 | tail -1 +rm /home/ubuntu/.healthcheck +``` + +Expected: both write+read complete without error. Throughput is bounded by the +single-stream limits of NFS / Lustre on a single node — this is a sanity check, +not a benchmark. For Lustre throughput numbers, see the FSx Lustre User Guide +(provisioned throughput = `Capacity * PerUnitStorageThroughput / 1024` MB/s). + +### Step 3 — FSx-side parameters match the CFN inputs + +```bash +FSX_ID=$(aws cloudformation describe-stacks \ + --stack-name \ + --query 'Stacks[0].Outputs[?OutputKey==`FSxLustreFilesystemId`].OutputValue' \ + --region --output text) + +aws fsx describe-file-systems --file-system-ids "$FSX_ID" --region \ + --query 'FileSystems[0].[StorageCapacity,StorageType,LustreConfiguration.[DeploymentType,PerUnitStorageThroughput,DataCompressionType,EfaEnabled,MetadataConfiguration.Mode]]' \ + --output text +``` + +Expected (default deploy): +- `StorageCapacity` = your `Capacity` parameter (1200 by default) +- `StorageType` = `SSD` +- `DeploymentType` = `PERSISTENT_2` (default; or `PERSISTENT_1` if you set it) +- `PerUnitStorageThroughput` = 250 (default) +- `DataCompressionType` = `LZ4` (default) +- `EfaEnabled` = `False` (default; `True` when `FSxLustreEnableEfa=true`. EFA on + FSx is a PERSISTENT_2-only feature — the prerequisites and deploy-all templates + enforce this with a CFN Rule that fails the stack at create time when + `FSxLustreEnableEfa=true` is combined with `LustreDeploymentType=PERSISTENT_1`) +- `MetadataConfiguration.Mode` = `AUTOMATIC` on PERSISTENT_2 + +For the OpenZFS `/home` filesystem: + +```bash +FSXO_ID=$(aws cloudformation describe-stacks \ + --stack-name \ + --query 'Stacks[0].Outputs[?OutputKey==`FSxOFilesystemId`].OutputValue' \ + --region --output text) + +aws fsx describe-file-systems --file-system-ids "$FSXO_ID" --region \ + --query 'FileSystems[0].[StorageCapacity,OpenZFSConfiguration.[DeploymentType,ThroughputCapacity]]' \ + --output text +``` + +### Step 4 — Storage dashboard in Grafana + +The `Compute Node Details` and `HPC Cluster Monitoring → Storage` Grafana +dashboards are populated by the **CloudWatch Exporter** (FSx CloudWatch +metrics, scraped by the monitoring stack on the login node — see the +[aws-parallelcluster-monitoring v2.6 release notes](https://github.com/aws-samples/aws-parallelcluster-monitoring/releases/tag/v2.6)). + +Open Grafana (Test 1's port-forward / public CIDR), go to the Storage +dashboard, and verify the `/fsx` panels (Throughput, IOPS, Free Capacity, +Client Connections) populate within ~5 minutes of a workload starting. CW +metrics have a ~5 min publishing delay, so a brand-new filesystem with no I/O +shows blank panels for a while; the `dd` from Step 2 is enough to seed values. + +### `FSxLustreEnableEfa=true` specifics + +When `FSxLustreEnableEfa=true` is set on a PERSISTENT_2 SSD filesystem, the +extra checks beyond the above are: + +- `aws fsx describe-file-systems` `LustreConfiguration.EfaEnabled = true` +- `Capacity` is at-or-above the EFA minimum for the chosen + `PerUnitStorageThroughput` tier (19200 GiB for tier 250; the FSx for + Lustre User Guide has the full matrix). Below the minimum, the FSx side + rejects the `CreateFileSystem` with `Invalid storage capacity provided: + N GiB. Minimum storage capacity for an EFA enabled LUSTRE file systems + with deployment type PERSISTENT_2, per unit storage throughput X and + storage type SSD is M`. The Lustre nested stack fails first, then the + whole stack rolls back; that's the expected behavior for an undersized + Capacity. + +The FSx-side EFA endpoints are usable from EFA-capable clients (CPU CNGs +deployed with `OnDemandEfaInterfaceCount > 0`, P5/P6 GPU CNGs). Plain Lustre client +mounts continue to work over TCP for non-EFA nodes; EFA support is additive. + +--- diff --git a/architectures/aws-pcs/tests/training-test.md b/architectures/aws-pcs/tests/training-test.md new file mode 100644 index 000000000..1d8867d6e --- /dev/null +++ b/architectures/aws-pcs/tests/training-test.md @@ -0,0 +1,129 @@ +# Training Tests (Test 7) + +Validates end-to-end distributed training using the repository's canonical +FSDP Llama-2 7B test case. + +--- + +## Test 7: FSDP sample training + +A short FSDP Llama-2 7B run, using the repo's canonical training case +[`3.test_cases/pytorch/FSDP`](../../../3.test_cases/pytorch/FSDP) and its +[`slurm/llama2_7b-training.sbatch`](../../../3.test_cases/pytorch/FSDP/slurm/llama2_7b-training.sbatch). +Follow that case's README; the only PCS-specific deltas are where things live on the +shared filesystems and the node count: + +1. **Build the venv on shared `/fsx`** so every compute node sees it (the canonical + `slurm/create_venv.sh` creates `./env` in place — run it from a `/fsx` checkout): + + ```bash + cd /fsx && git clone --depth 1 https://github.com/awslabs/awsome-distributed-ai.git + cd /fsx/awsome-distributed-ai/3.test_cases/pytorch/FSDP/slurm && bash create_venv.sh + export HF_TOKEN=hf_xxx # gated Llama-2 tokenizer + ``` + +2. **Keep the HuggingFace cache on `/fsx` (Lustre), not `/home` (NFS).** Concurrent rank + file-locking on NFS throws `OSError: [Errno 116] Stale file handle`; export + `HF_HOME=/fsx/.hf-cache` before submitting. + + > **HuggingFace rate limits.** The test case streams its dataset from the Hub, so + > every rank pulls from HF at startup. Large datasets (e.g. `allenai/c4`, 1024 + > shards) can return `429 Too Many Requests` under many concurrent ranks even with + > an `HF_TOKEN` — the dataset is fetched, not the cluster, so this is an HF-account + > limit, not a cluster issue. For large multi-node runs, use an HF account with a + > higher rate limit, an HF mirror, or pre-tokenized data staged on `/fsx` (no + > streaming). + +3. **Submit 2 nodes** (the canonical sbatch defaults to 4) on your GPU queue. The venv must + be on `PATH` for `torchrun` to resolve on every node — point `PATH` at the shared venv via + `--export` (the canonical sbatch doesn't `activate` it): + + ```bash + sbatch --nodes=2 --partition=gpu-p6b200 \ + --export=ALL,PATH=/fsx/awsome-distributed-ai/3.test_cases/pytorch/FSDP/slurm/env/bin:$PATH,HF_HOME=/fsx/.hf-cache \ + llama2_7b-training.sbatch + ``` + +### Option B — run it in an Enroot/Pyxis container instead of the venv + +The same canonical sbatch switches to container mode when `CONTAINER_IMAGE` is set (it adds +`--container-image`/`--container-mounts` and runs `./train.py` inside). Build the image once +on the login node and submit with `CONTAINER_IMAGE` — no venv needed: + +```bash +# On the login node (300 GiB root disk + Docker), build + import to /fsx: +cd /fsx/awsome-distributed-ai/3.test_cases/pytorch/FSDP +sudo docker build -t fsdp:pytorch -f Dockerfile . +enroot import -o /fsx/pytorch-fsdp.sqsh dockerd://fsdp:pytorch + +# Submit (container mode; mounts $(pwd) into /fsx inside the container): +cd slurm && sbatch --nodes=2 --partition=gpu-p6b300 \ + --export=ALL,CONTAINER_IMAGE=/fsx/pytorch-fsdp.sqsh,HF_HOME=/fsx/.hf-cache,HF_TOKEN=hf_xxx \ + llama2_7b-training.sbatch +``` + +> If the Dockerfile's `FROM` tag (a `public.ecr.aws/hpc-cloud/nccl-tests` tag) has been +> rotated out of the registry, substitute a current tag from that repo before building. + +**Expected** (in `logs/llama2_7b-FSDP_.out`), either path: +- NCCL initializes over EFA (`found N nics`) and training logs ~100 steps + a validation + step, saving checkpoints under `./checkpoints`. +- Throughput per step, e.g. on 2× p6-b300 **~200 TFLOPS/GPU, ~77k tokens/s** (venv) / + **~193 TFLOPS/GPU** (container); ~60 TFLOPS on 2× p5/H100. (Loss is constant at ln(vocab) + in this smoke test — a known dataloader/vocab quirk of the test case, not a cluster problem.) + +> **Multi-NIC tip:** the canonical sbatch already sets `NCCL_SOCKET_IFNAME=^docker,lo,veth,eth` +> (NCCL auto-selects). Do **not** pin a single interface on P5/P6 — all NICs share one +> subnet and pinning one breaks the cross-node NCCL bootstrap ring. + +--- + +## Test 7b: Megatron-LM GPT-3 (tensor + pipeline parallel) + +A second distributed-training path that exercises 3D parallelism (TP/PP/DP) rather than +FSDP, using the repo's canonical Megatron-LM case +[`3.test_cases/megatron/megatron-lm`](../../../3.test_cases/megatron/megatron-lm). Unlike +the FSDP case it does **not** stream from HuggingFace — data is tokenized once to local +`.bin`/`.idx` on `/fsx`, so it has no HF rate-limit exposure. Two PCS-specific deltas: + +1. **Build + import the image on the login node** (300 GiB root; the overlay can't live on + Lustre), writing the `.sqsh` to `/fsx`: + + ```bash + cd /fsx/awsome-distributed-ai/3.test_cases/megatron/megatron-lm + sudo docker build -t aws-megatron-lm -f 0.distributed-training.Dockerfile . + enroot import -o /fsx/aws-megatron-lm.sqsh dockerd://aws-megatron-lm:latest + ``` + +2. **Preprocess once, then train**, passing `IMAGE`/`DATA_PATH`/`FSX_MOUNT` as env (the + sbatch scripts read them). Preprocessing is CPU work but still needs a Pyxis-capable + node and the login `srun` client: + + ```bash + cd slurm/gpt3 + IMAGE=/fsx/aws-megatron-lm.sqsh DATA_PATH=/fsx FSX_MOUNT=/fsx:/fsx \ + sbatch -p gpu 1.data-preprocessing.sbatch # → /fsx/my-gpt2_text_document.{bin,idx} + # the training sbatch expects data under ${DATA_PATH}/gpt2/ — symlink it there + mkdir -p /fsx/gpt2 && ln -sf /fsx/my-gpt2_text_document.* /fsx/gpt2/ \ + && ln -sf /fsx/gpt2-vocab.json /fsx/gpt2-merges.txt /fsx/gpt2/ + IMAGE=/fsx/aws-megatron-lm.sqsh DATA_PATH=/fsx FSX_MOUNT=/fsx:/fsx \ + sbatch --nodes=4 -p gpu 2.distributed-training.sbatch + ``` + + For a short smoke run, add `--exit-interval 20` and `--log-throughput` to the megatron + args in `2.distributed-training.sbatch`. At `NODES≤4` the script picks `TP=4, PP=2, + GBS=288` automatically. + +**Expected:** NCCL over EFA (`NET/OFI Selected provider is efa, fabric is efa-direct (found +8 nics)` on p6-b200); `0` nan/skipped after warmup; `lm loss` begins descending once the +LR warmup kicks in (loss-scale auto-tuning skips the first ~16 steps with `lr=0` by design). + +### Verified — p6-b200 ×4 (32 GPU, ap-south-1) + +GPT-3 36-layer / hidden 4096 / 32 heads, seq 2048, TP=4 PP=2 (DP=4), GBS=288, fp16 + +activation recompute. Steady **~134 TFLOP/s/GPU** (iters 3–20, ~6.4 s/iter); `lm loss` +**10.91 → 10.46** over iters 17–20 as the cosine LR warmup ramps (grad norm 57→45); 0 nan, +0 skipped post-warmup. This is a general-purpose (un-tuned) container — not a peak MFU +number — but confirms TP+PP+DP 3D parallelism trains correctly over EFA on B200. + +---