Enable simple, text-based Slurm accounting - #7
Merged
Merged
Conversation
sean-smith
reviewed
Oct 3, 2023
| # NOTE: JobCompType entry will be duplicated, hence will cause a harmless | ||
| # warning message in `systemctl status --no-pager slurmctld`. | ||
| - JobCompType: jobcomp/filetxt | ||
| - JobCompLoc: /home/slurm/slurm-job-completions.txt |
There was a problem hiding this comment.
does /home/slurm exist on the cluster?
Contributor
Author
|
yes, Pcluster creates this dir on head node.
…On Wed, 4 Oct 2023, 07:52 Sean Smith, ***@***.***> wrote:
***@***.**** commented on this pull request.
------------------------------
In
1.architectures/2.aws-parallelcluster/distributed-training-p4de-base.yaml
<#7 (comment)>
:
> @@ -24,6 +24,17 @@ Scheduling:
Scheduler: slurm
SlurmSettings:
ScaledownIdletime: 60
+ CustomSlurmSettings:
+ # Simple accounting to text file /home/slurm/slurm-job-completions.txt.
+ #
+ # Must be disabled should you prefer to setup Slurm accounting to database
+ # (https://docs.aws.amazon.com/parallelcluster/latest/ug/slurm-accounting-v3.html).
+ #
+ # NOTE: JobCompType entry will be duplicated, hence will cause a harmless
+ # warning message in `systemctl status --no-pager slurmctld`.
+ - JobCompType: jobcomp/filetxt
+ - JobCompLoc: /home/slurm/slurm-job-completions.txt
does /home/slurm exist on the cluster?
—
Reply to this email directly, view it on GitHub
<#7 (review)>,
or unsubscribe
<https://github.com/notifications/unsubscribe-auth/AAR3PLMFYIPVYTJDAICEBK3X5SQMLAVCNFSM6AAAAAA5RDSO3GVHI2DSMVQWIX3LMV43YUDVNRWFEZLROVSXG5CSMV3GSZLXHMYTMNJWGMYTAMBSG4>
.
You are receiving this because you authored the thread.Message ID:
***@***.***
com>
|
|
LGTM |
sean-smith
approved these changes
Oct 6, 2023
nkumaraws
added a commit
that referenced
this pull request
Mar 6, 2026
- Add license headers (MIT-0 for .sh/.py, Apache-2.0 for .yaml) - Remove hardcoded nsys version 2025.6.1; use glob-based auto-discovery - Replace shell variable placeholders with concrete values in manifest - Add envsubst instructions for customization - Fix kubectl apply paths to include EKS/ prefix in README - Add EFA production profiling note (gotcha #7) - Update EFA env vars: use NCCL_TUNER_PLUGIN, remove NCCL_PROTO/ALGO/FORK_SAFE - Update NCCL_SOCKET_IFNAME to exclude docker/veth interfaces - Fix missing newline at end of README
KeitaW
pushed a commit
that referenced
this pull request
Mar 11, 2026
…1008) * Add Nsight Systems host-mount profiling for EKS (HyperPod / DLAMI) Add a simpler alternative to the sidecar injector approach for profiling distributed PyTorch training on EKS clusters where nsys is pre-installed (SageMaker HyperPod EKS, DLAMI-based nodes). New files: - EKS/nsys-profile.sh: Profiling wrapper with auto-detection, PyTorch NVTX annotations, Python sampling, CUDA memory tracking, selective rank profiling - EKS/nsys_analyze.py: Automated bottleneck analysis with kernel classification, cross-worker comparison, and actionable recommendations - kubernetes/llama3_2_1b-fsdp-nsight.yaml: Reference PyTorchJob manifest for profiled Llama 3.2 1B FSDP training Updated: - README.md: Added Section 8 documenting the host-mount approach with quick start guide, prerequisites, configuration, report retrieval, and gotchas * refactor: Move nsight profiling manifest into EKS/ for self-contained packaging Move llama3_2_1b-fsdp-nsight.yaml from 3.test_cases/pytorch/FSDP/kubernetes/ into 4.validation_and_observability/5.nsight/EKS/ so all nsight EKS assets (wrapper script, analysis script, and manifest) live in one directory. Update README reference accordingly. * Address PR #1008 review comments from KeitaW and paragao - Add license headers (MIT-0 for .sh/.py, Apache-2.0 for .yaml) - Remove hardcoded nsys version 2025.6.1; use glob-based auto-discovery - Replace shell variable placeholders with concrete values in manifest - Add envsubst instructions for customization - Fix kubectl apply paths to include EKS/ prefix in README - Add EFA production profiling note (gotcha #7) - Update EFA env vars: use NCCL_TUNER_PLUGIN, remove NCCL_PROTO/ALGO/FORK_SAFE - Update NCCL_SOCKET_IFNAME to exclude docker/veth interfaces - Fix missing newline at end of README * fix: Address remaining PR #1008 review comments (license, EFA flags, placeholders) - Fix license header: Apache-2.0 -> MIT-0 per repo standard (KeitaW) - Replace :latest image tag with :<IMAGE_TAG> placeholder (KeitaW) - Remove broken envsubst instructions; use <PLACEHOLDER> style consistently (KeitaW) - Simplify EFA env vars per paragao: keep only FI_PROVIDER, NCCL_BUFFSIZE, NCCL_P2P_NET_CHUNKSIZE, NCCL_TUNER_PLUGIN; remove FI_EFA_USE_DEVICE_RDMA, NCCL_NET, duplicate NCCL_SOCKET_IFNAME, and LD_LIBRARY_PATH - Remove envsubst references from README to match YAML changes
7 tasks
DaisukeMiyamoto
added a commit
to DaisukeMiyamoto/awsome-distributed-ai
that referenced
this pull request
Jun 4, 2026
…gateway too (review awslabs#7/awslabs#8) Opening 443 to a CIDR exposes more than the password-gated Grafana: the login node's nginx also reverse-proxies the unauthenticated /prometheus/, /pushgateway/ and /slurmexporter/ paths, so anyone in the CIDR can read all cluster metrics (and push) without credentials. Document this real exposure in the parameter description, the README Option B security notes, and docs/PARAMETERS.md. Keep 0.0.0.0/0 accepted (no hard block) — it's useful for short-lived PoC / workshop access where per-user local SSM permissions are impractical — but warn to narrow or clear it when done. (awslabs#7 documented here; the upstream fix to gate those nginx locations is tracked separately for aws-parallelcluster-monitoring.)
KeitaW
added a commit
that referenced
this pull request
Jun 6, 2026
…tainer runtime to the PCS reference cluster (#1120) * update README for PCS * add reference cluster with PCS * Update PCS README: add deployment options, 1-click deploy, and fix S3 bucket name * Add unified multi-NIC template for P5/P6 instances and update documentation * Improve PCS deployment: rename to Pseries, update docs and examples * Add placement group for on-demand P5/P6 instances in add-cng-p5 * Add launch button and refine PCS documentation * add launch button image * fix typos * fix typos * Upgrade Slurm support to 25.05/25.11, default to 25.11 * Switch default AMI to PCS-specific DLAMI image * Change mount path of FSxL from /shared to /fsx * Simplify DLAMI build to use PCS-ready base image with Enroot/Pyxis only * fix typos * fix typos * update readme in pcs * add BaseAMI parameter for ImageBuilder * use awsome-distributed-ai bucket * fix typos * update architecture * update architecture image * Address PR review feedback: fix permissions, add documentation - Change FSx Lustre mount permissions from 0777 to 1777 (sticky bit) - Add upstream repository reference to README - Add Testing and Validation section documenting tested configurations * Refactor cluster deployment to use modular add-cng templates - Separate cluster core (Slurm scheduler) from compute nodes - Use add-cng.yaml for both login and compute node groups - Make queue creation optional in add-cng templates (empty QueueName for login nodes) - Remove AmiId parameter from cluster.yaml (only needed for compute nodes) - Update pcs-ml-cluster-deploy-all.yaml to deploy login and cpu1 as separate nested stacks - Change default DeployOnDemandCNG to true (cpu1 queue enabled by default) - Update README with new deployment architecture and examples * Use variables for AZ ID in deployment examples - Add AZ_ID variable at the start of each example - Makes it easier to change AZ consistently across examples - Aligns with CAPACITY_RESERVATION_ID variable pattern in Example 4 * Add AWS ParallelCluster Monitoring reference - Add link to aws-parallelcluster-monitoring GitHub repository - Provides comprehensive monitoring solution for HPC clusters with Prometheus and Grafana * Add AWS Parallel Computing Service description - Clarify what AWS PCS is in the introduction - Mention it's a fully managed service for HPC with Slurm scheduler * Add AWS Parallel Computing Service to top-level README - Include AWS PCS in the list of supported compute platforms - Update 9.aws-pcs description to clarify it uses Slurm scheduler * Add AWS ParallelCluster Monitoring integration to PCS templates Integrates aws-parallelcluster-monitoring stack (Prometheus, Grafana, DCGM) into PCS cluster deployment with optional, backward-compatible configuration. Changes: - Add monitoring-iam-policy.yaml: IAM permissions for CloudWatch, SSM, Pricing API - Update cluster.yaml: Add RoleName output for IAM policy attachment - Update add-cng.yaml: Add DeployMonitoring/MonitoringVersion parameters and UserData script - Update add-cng-p5.yaml: Add monitoring parameters and UserData for P5/P6 nodes - Update pcs-ml-cluster-deploy-all.yaml: Add MonitoringIAMPolicyStack and pass parameters to all CNGs - Update README.md: Add monitoring access instructions via Session Manager Monitoring deployment: - Login node: Prometheus, Grafana, custom metrics exporters - Compute nodes with GPU: DCGM exporter (auto-detected) - Compute nodes without GPU: Node exporter Key features: - Backward compatible: DeployMonitoring defaults to 'false' in add-cng.yaml - Enabled by default in pcs-ml-cluster-deploy-all.yaml (DeployMonitoring='true') - Installs appropriate components based on node type via post-install.sh from upstream - Access via Session Manager port forwarding (local:8443 -> remote:443) - Grafana password stored in SSM Parameter Store * Add Ubuntu 24.04 compatibility workaround for monitoring stack The upstream aws-parallelcluster-monitoring installer hardcodes PLATFORM_USER=ec2-user for PCS, but Ubuntu 24.04 uses ubuntu as the default user. Changes: - Add symlink workaround in UserData: /home/ec2-user -> /home/ubuntu - Apply to both add-cng.yaml and add-cng-p5.yaml - Document known limitation in README.md The workaround creates a symlink before running post-install.sh, ensuring the monitoring stack installs to the correct home directory on Ubuntu-based AMIs. * Fix RoleName output in cluster.yaml The upstream pcs-iip-minimal.yaml does not export RoleName. Construct RoleName directly using the same pattern as the nested stack parameter: AWSPCS-pcs-<hash>-<region>-role * Add monitoring stack integration for AWS PCS - Add aws-parallelcluster-monitoring integration with IAM policies - Add Slurm OpenMetrics configuration (MetricsType, CommunicationParameters) - Add MonitoringRole parameter (login/compute/none) for flexible deployment - Add monitoring-role and pcs-cluster-id tags for login nodes - Add InstanceMetadataTags support for IMDS tag access - Add PCS API permissions (pcs:GetCluster, pcs:ListComputeNodeGroups) - Fix shell compatibility (== to = for POSIX sh) - Add test guide for monitoring stack validation - Add architecture diagram for monitoring components * Address PR #1109 review feedback (Round 2) - Fix stray blank line in root README workshop table - Change ml-cluster-prerequisites.yaml license from Apache-2.0 to MIT-0 - Fix Pyxis Slurm version compatibility: install for both 25.05 and 25.11 - Update PseriesMinCount description to clarify static vs dynamic scaling - Add AMI build time note (~30 minutes) to README * Fix monitoring installation bug with robust pattern matching Improved sed pattern to fix aws-parallelcluster-monitoring upstream bug: - Old: Line number based (83s/...) - fragile if file changes - New: Pattern based with whitespace preservation - more robust - Pattern: s/^\([[:space:]]*\)local login_id$/\1login_id=/g Also fixed ClusterName parameter in MonitoringIAMPolicyStack to use ClusterId instead of stack name for correct SSM parameter path. Tested with successful automated deployment on pcs-monitoring-test-v4. * Add architectures/aws-pcs directory for new structure Created new architectures directory structure alongside existing 1.architectures to prepare for repository reorganization. Contents: - architectures/aws-pcs/: Complete AWS PCS architecture templates - All CloudFormation templates with monitoring integration - README and architecture diagrams This maintains backward compatibility with 1.architectures/9.aws-pcs while establishing the new standardized structure. * Add tests directory to architectures/aws-pcs Migrated test documentation from 1.architectures/9.aws-pcs/tests/ to new standardized structure. Contents: - README.md: Test directory overview - monitoring-stack-test.md: Monitoring stack validation guide * Remove old 1.architectures/9.aws-pcs directory Removed legacy directory structure as content has been migrated to architectures/aws-pcs/ in the new standardized structure. All templates, documentation, and tests are now available in: architectures/aws-pcs/ * Update README.md to reference new architectures/aws-pcs path Changed link from 1.architectures/9.aws-pcs to architectures/aws-pcs to reflect the new directory structure. * Add ClusterName tag to launch templates Added ClusterName tag to instance TagSpecifications in both add-cng.yaml and add-cng-p5.yaml for better instance identification and filtering. Tag value: ClusterId parameter * Apply monitoring integration to architectures/aws-pcs Synced monitoring integration changes to new directory structure: - Add DeployMonitoring and MonitoringVersion parameters - Add MonitoringRole parameter (login/compute/none) - Add ClusterName tag to launch templates - Add monitoring stack installation with bug fix for upstream issue - Add Slurm OpenMetrics configuration - Add monitoring IAM policy and IMDS tag support * Add monitoring usage documentation to README * Fix MonitoringIAMPolicy ClusterName parameter and add missing template * Integrate IAM resources into cluster.yaml to eliminate external dependencies - Remove nested stack dependency on pcs-iip-minimal.yaml - Move IAM Role and Instance Profile directly into cluster.yaml - Integrate MonitoringPolicy as conditional ManagedPolicy in cluster.yaml - Remove HpcRecipesS3Bucket and HpcRecipesBranch parameters - Remove MonitoringIAMPolicyStack nested stack from pcs-ml-cluster-deploy-all.yaml - Simplify deployment by reducing nested stacks from 6 to 4 * Add Prerequisites stack reuse and UserData Enroot/Pyxis installation This commit implements two major features for faster testing iterations: 1. Prerequisites Stack Reuse - Replace 7 Existing* parameters with single PrerequisitesStackName - Use CloudFormation Exports/Imports for automatic resource resolution - Reduce subsequent cluster deployment time from 20-30 min to 5-10 min - CloudFormation enforces deletion protection for shared infrastructure 2. UserData Enroot/Pyxis Installation - Add InstallEnrootPyxis parameter for optional container support - Install Enroot 3.5.0 and Pyxis v0.20.0 on first boot - Support both Slurm 25.05 and 25.11 versions - Enable rapid testing without ImageBuilder (8-12 min boot vs 3 min) Changes: - pcs-ml-cluster-deploy-all.yaml: Add PrerequisitesStackName parameter and reorganize metadata - add-cng.yaml, add-cng-p5.yaml: Add UserData installation logic with proper shell escaping - ml-cluster-prerequisites.yaml: Add Export aliases for compatibility - architecture-components.md: Document main components for diagram generation Deployment modes: - Mode 1 (Production): BuildAMI=true, InstallEnrootPyxis=false - Mode 2 (Testing): BuildAMI=false, InstallEnrootPyxis=true - Mode 3 (Minimal): BuildAMI=false, InstallEnrootPyxis=false * Refactor PCS deployment templates and update documentation - Refactor pcs-ml-cluster-deploy-all.yaml to call cluster.yaml + add-cng.yaml directly - Remove cluster-with-compute-nodes.yaml intermediate layer - Add separate stacks: ClusterStack, LoginNodeGroupStack, OnDemandCNGStack, PseriesCNGStack - Maintain README compatibility: cluster.yaml can be deployed standalone - Delete unused monitoring-iam-policy.yaml - MonitoringPolicy is now integrated into cluster.yaml - Update CLAUDE.md with modular deployment guide - Add standalone Prerequisites deployment instructions - Document environment variable usage for rapid iteration - Add Enroot/Pyxis testing procedures - Include verification and troubleshooting steps - Backup cluster-with-compute-nodes.yaml for reference * Fix shell variable escaping in UserData for Enroot/Pyxis installation - Change all shell variables from $ to 94274 for proper CloudFormation !Sub escaping - Fixes 404 error when downloading Enroot packages (variables were not expanded) - Affects: ENROOT_RELEASE, PYXIS_RELEASE, SLURM_VERSION, arch, distribution, etc. - Now curl downloads correct URLs like v3.5.0 instead of v$ENROOT_RELEASE * Externalize Enroot/Pyxis installation to separate script - Create install-enroot-pyxis.sh script with proper variable handling - Upload script to S3 (midaisuk-llm-dev/scripts/) - Simplify UserData to download and execute external script - Fixes CloudFormation !Sub variable escaping issues - Benefits: maintainability, testability, no variable escaping problems * Fix Enroot/Pyxis installation script URL to use GitHub raw URL - Change S3 URL to GitHub raw URL in add-cng.yaml - Add security restrictions to CLAUDE.md (no S3 public access) - Verified installation works on login node i-0738c0e1f93dd452c - Enroot 3.5.0 and Pyxis v0.20.0 successfully installed * Generalize Enroot/Pyxis UserData into a post-install script hook Replace the InstallEnrootPyxis boolean with a generic PostInstallScriptUrl / PostInstallScriptArgs pair (PCS equivalent of ParallelCluster OnNodeConfigured) across add-cng.yaml, add-cng-p5.yaml and pcs-ml-cluster-deploy-all.yaml. - add-cng-p5.yaml: drop the duplicated inline Enroot/Pyxis install block; both CNG templates now fetch and run the same external script via the hook. - deploy-all defaults reflect the most common setup: BuildAMI=false, DeployMonitoring=true, PostInstallScriptUrl set to the upstream installer URL. - DeployMonitoring keeps its platform-specific logic unchanged. - Document runtime dependencies (Slurm headers, OS family, egress, apt lock, GPU toolkit) in the script header and README; post-install log is now /var/log/pcs-post-install.log. - Remove stale cluster-with-compute-nodes.yaml.backup and refresh architecture diagram. * Add monitoring stack test cluster for Ubuntu 24.04 verification Minimal single-login-node PCS template (pcluster-monitoring-test-cluster.yaml) plus a deploy helper and notes, used to verify the aws-parallelcluster-monitoring stack installs cleanly on the PCS Ubuntu 24.04 DLAMI before the upstream fixes land. Points at the DaisukeMiyamoto fork branch for testing. * Pin aws-parallelcluster-monitoring to release tag v2.6.2 Replace the 'latest'/'main' monitoring references with a pinned release tag for upstream stability. MonitoringVersion now defaults to v2.6.2 and is used both to fetch post-install.sh from the matching tag and as the version argument passed to it, across add-cng.yaml, add-cng-p5.yaml and pcs-ml-cluster-deploy-all.yaml. Note: v2.6.2 still lacks the Ubuntu fixes (upstream PR #44), so the ec2-user and local-login_id UserData workarounds remain until a release with those fixes can be pinned. * Update README to pin MonitoringVersion to v2.6.2 Reflect the version-pinning policy in the monitoring section: default tag v2.6.2, explain it is used for both the post-install.sh fetch and the version argument, and note why latest/main are avoided. * Fix Enroot/Pyxis install on GPU nodes: gpg no-tty and stable nvidia repo path * Add P5 GPU training validation tests and 300 GiB root volume option - Add tests/gpu-training-validation/ (nvidia-smi, 2-node NCCL all_reduce both AMI-native and via the prebuilt public.ecr.aws/hpc-cloud/nccl-tests image, and a Megatron-LM Llama-2 smoke test), plus RESULTS.md documenting the run. - Add RootVolumeSize parameter (default 300 GiB) + BlockDeviceMappings to add-cng.yaml and add-cng-p5.yaml; the ~75 GiB DLAMI default overflows when pulling/importing large container images. * Bump monitoring to v2.6.3 and drop Ubuntu ec2-user workarounds aws-parallelcluster-monitoring v2.6.3 (PR #44) natively detects the Ubuntu 'ubuntu' user, so the templates no longer need the ec2-user shim or the 'local login_id' sed patch. Pin MonitoringVersion default to v2.6.3 across add-cng.yaml, add-cng-p5.yaml and pcs-ml-cluster-deploy-all.yaml, simplify the UserData monitoring block, and update the README and test docs accordingly. * Tag nodes Name=HeadNode/Compute for monitoring dashboards The monitoring Grafana dashboards filter on the EC2 Name tag relabeled into the instance_name label, expecting the literals 'HeadNode' (login/head) and 'Compute' (compute). Derive Name from whether the node group has a Slurm queue (CreateQueue), and preserve the per-node-group name on a new CngName tag. * Make RootVolumeSize consistent across CNG templates and deploy-all - add-cng.yaml: add MinValue: 75 to match add-cng-p5.yaml. - pcs-ml-cluster-deploy-all.yaml: add a RootVolumeSize parameter (default 300 GiB) and pass it down to the login, on-demand, and P-series nested stacks so the whole-cluster deploy can control root volume size in one place. * Remove duplicate MinValue in add-cng.yaml RootVolumeSize * PCS monitoring: free the Name tag, identify node type via monitoring-role, add MonitoringRepo - Tag login and compute nodes with monitoring-role (login|compute) so the monitoring stack identifies node type by that tag instead of Name. Name now defaults to PCS-<cngname> and is free for arbitrary use. Replaces the IsMonitoringLogin condition with IsMonitoringRoleSet so the tag is emitted for both roles; deploy-all passes MonitoringRole=compute to the compute CNGs. - Add a MonitoringRepo parameter (default aws-samples/aws-parallelcluster-monitoring) and pass both the ref and repo to post-install.sh, so a fork + branch can be used to test unreleased monitoring changes. MonitoringVersion may now be a branch. - README: document MonitoringRepo, the monitoring-role node-type convention, and fork/branch testing. * Add P6-B200 and P6-B300 GPU compute node group templates Add per-family multi-NIC EFA templates for the Blackwell P6 instances, since their network-card counts differ from P5 (P6-B200 has 8 network cards, P6-B300 has 17, vs 16/32 for P5). Each template hardcodes the right number of EFA interfaces using the proven P5 pattern (primary NCI0 on DeviceIndex 0, the rest on DeviceIndex 1, all InterfaceType: efa), and drops the NetworkInterfaceCount parameter / Use32Interfaces condition since the count is fixed per family. All other parameters and resources (PCS CNG/Queue, UserData, RootVolumeSize, monitoring, Capacity Block) are identical to add-cng-p5.yaml. Not yet validated on real p6 hardware and not yet wired into the deploy-all orchestration template. * deploy-all: select P-series NIC template by instance type PseriesInstanceType now accepts p6-b200.48xlarge and p6-b300.48xlarge in addition to the p5 family. Conditions (PseriesIsP5/PseriesIsP6B200/ PseriesIsP6B300) derive the family from the instance type and create the matching nested stack: add-cng-p5.yaml (NetworkInterfaceCount), add-cng-p6-b200 .yaml (8 NICs), or add-cng-p6-b300.yaml (17 NICs). NetworkInterfaceCount is passed only to the P5 stack; the P6 stacks omit it since their NIC count is fixed. Note NetworkInterfaceCount as P5-only in its description. * P6 templates: use EFA 'Use case 2' NIC layout (ENA primary + efa-only/ENA pairs) Align the P6-B200/P6-B300 network interface config with the EC2 EFA docs 'Use case 2' (maximum EFA + ENA bandwidth), matching the ParallelCluster B300 LaunchTemplateOverrides tutorial: - Primary NCI 0, DeviceIndex 0: ENA (the primary cannot be EFA-only). - NCI 1..N-1: an efa-only interface on DeviceIndex 0 and an ENA interface on DeviceIndex 1. This replaces the earlier P5-style layout (all interfaces InterfaceType: efa, primary on a single card) which is not the AWS-recommended P6 configuration. P6-B200 = 8 network cards (15 interfaces); P6-B300 = 17 network cards (33). * P6 templates: use InterfaceType efa (not efa-only) for PCS compatibility The earlier 'Use case 2' layout (efa-only on dev0 + ENA on dev1) failed to launch through PCS: PCS does not propagate efa-only network interfaces from the launch template to the RunInstances call, and instead requested EFA on network card 0 (EFA limit 0 on P6), causing Client.AttachmentLimitExceeded. Switch to the layout that actually launches via PCS, validated on real p6-b300.48xlarge hardware (Capacity Block, us-west-2): - Network card 0, device index 0: plain ENA (P6 NIC 0 does not support EFA). - Network cards 1..N-1, device index 0: InterfaceType: efa. P6-B200 = 8 network cards, P6-B300 = 17. The bare-EC2 RunInstances accepts both efa-only and efa layouts, but PCS only handles 'efa', so we use 'efa'. * CNG templates: track latest launch template version CustomLaunchTemplate.Version was hardcoded to 1, so when a template change produced a new launch template version (e.g. NIC or UserData edits), the compute node group kept using version 1 and silently ignored the update — observed on p6-b300 where a fixed NIC layout (v2) never reached the node group. Use !GetAtt PCSLaunchTemplate.LatestVersionNumber in add-cng.yaml, add-cng-p5.yaml, add-cng-p6-b200.yaml, and add-cng-p6-b300.yaml so node groups follow template updates. Verified on real p6-b300: the node group moved to v2 and launched. * add-cng-p6-b200: use EFA on all 8 network cards Unlike p6-b300 (network card 0 cannot do EFA; MaxEfaInterfaces 16 of 17), p6-b200 supports EFA on all 8 network cards (MaxEfaInterfaces=8). Configure every card with InterfaceType: efa (primary NCI0 on DeviceIndex 0, NCI1-7 on DeviceIndex 1) to maximize EFA bandwidth, matching the proven p5 all-efa layout. Not yet validated on real p6-b200 hardware (Capacity Block opens 2026-06-03). * P6 templates: set family-specific CngName defaults CngName defaulted to 'gpu-p5' (leftover from copying add-cng-p5.yaml). Set 'gpu-p6-b200' and 'gpu-p6-b300' so standalone deploys get a sensible node group name. (deploy-all overrides this via PseriesCngName, so it only affects direct use of these templates.) * PCS monitoring: retry post-install on transient failure; docs to /opt path The monitoring tree now installs on node-local /opt (fixed in the aws-parallelcluster-monitoring fork) instead of the shared NFS /home, which was racing across nodes at first boot (Stale file handle). As a belt-and-suspenders safety net, the UserData monitoring step in add-cng.yaml / add-cng-p6-b200.yaml / add-cng-p6-b300.yaml now retries post-install.sh up to 3x (the installer is idempotent) so a one-off transient first-boot failure self-heals. Also update test docs / sample template references from /home/ubuntu/aws-parallelcluster-monitoring to /opt/aws-parallelcluster-monitoring. * PCS docs/params: human-friendly 1-click, FSx Region deployment types, README restructure - deploy-all: add ParameterLabels (friendly console names) + group RootVolumeSize; add LustreDeploymentType (PERSISTENT_1/2) and OpenZFSDeploymentType (SINGLE_AZ*) so Region-dependent FSx tiers are selectable (defaults unchanged); relax PerUnitStorageThroughput to a free-form number. - ml-cluster-prerequisites: wire the two FSx deployment-type params; emit Lustre MetadataConfiguration only for PERSISTENT_2. - cluster.yaml: remove 5 unused params (VpcId/PublicSubnetId/FSx*); only the private subnet + security group are needed by the PCS cluster. - README: update Key Features/Architecture to the current default (PCS-ready DLAMI, no AMI build, built-in monitoring, Enroot/Pyxis first-boot-vs-AMI choice, P5/P6 multi-NIC auto-select); restructure (consolidate template selection into one reference table, trim examples, slim Monitoring); add GPU Node List screenshot. * PCS GPU CNG: auto-derive EFA NIC count from instance type; clarify CapacityReservationId is Capacity-Block-only - add-cng-p5: drop the NetworkInterfaceCount parameter; derive the EFA interface count from InstanceType (p5en.48xlarge = 16, p5/p5e = 32) via the Use32Interfaces condition; constrain InstanceType to the P5 family. deploy-all: remove the NetworkInterfaceCount parameter/label/group/passthrough; PseriesInstanceType now drives both the template choice and the NIC count. - CapacityReservationId (all P-series templates + deploy-all): document that SETTING it forces MarketType=capacity-block (Capacity Blocks for ML), and that an ODCR ID must NOT go here — an open ODCR is consumed automatically by On-Demand launches, so leave it empty for On-Demand/ODCR. Relabel and update README capacity guidance. * deploy-all params: generic PostInstallScriptUrl override note; move MonitoringRepo/Version to developer group - PostInstallScriptUrl: drop the fork-specific (DaisukeMiyamoto) override example; keep the generic 'override with another HTTP(S) URL' note. - Move MonitoringVersion/MonitoringRepo out of '2. PCS Cluster Configuration' into the final developer/advanced group (alongside the nested-template S3 settings), since they are developer-facing source overrides. * README: lead with one-click ML-training-ready; align params to console groups; clarify GPU/EFA, monitoring - Key Features: lead with 'one click to an ML-training-ready cluster' (the main value); demote no-AMI-build/PCS-ready-DLAMI to a supporting note. - Quick Start + all examples: set AZ_ID in the first line so the one required choice (Availability Zone) is explicit. - Configuration: restructure the parameter tables into the 7 console parameter groups (matches deploy-all's AWS::CloudFormation::Interface). - GPU compute: unify NIC/network-card wording to 'EFA interfaces'; drop the misleading p5en-only NVSwitch note (all P5/P6 use NVLink/NVSwitch intra-node). - Monitoring: reference the GPU dashboard screenshot from the section intro; add a pointer to the AWS-managed Prometheus/Grafana option (4.prometheus-grafana). * README: add a quick NCCL multi-node GPU test (enroot import + Pyxis sbatch) Adds a 'Running a multi-node GPU job (NCCL test)' section: import nccl-tests to /fsx via Enroot, submit a 2-node 16-GPU all_reduce_perf via Pyxis, and interpret the result (EFA provider line, '# Out of bounds values : 0 OK', busbw). Points to the FSDP test case for full training. * Bump default MonitoringVersion to v2.6.5 (PRs #48/#49 merged upstream) aws-parallelcluster-monitoring v2.6.4 (PR #48, /opt install — fixes the shared-/home Stale-file-handle race) and v2.6.5 (PR #49, dcgm-exporter Docker-29.x tag + throttle tiles) are now released upstream. Bump the templates' default MonitoringVersion from v2.6.3 -> v2.6.5 so the default deploy uses the fixed upstream release directly — no fork/branch override needed. MonitoringRepo already defaults to aws-samples upstream. * README: reference monitoring v2.6.5 (PCS /opt + DCGM Docker-29.x fixes now released upstream) * docs: split full parameter reference into PARAMETERS.md; slim README Configuration Move the 7-group parameter detail tables to a dedicated PARAMETERS.md and leave the README Configuration section as a short overview + most-used-parameters table + link. The concept guides (container runtime, FSx Region availability, GPU compute) stay in the README since they inform choices. README ~460 -> ~412 lines. * README: run NCCL enroot import on the login node; drop PR references from Monitoring - NCCL test: import the image directly on the login node (it has Enroot and writes to shared /fsx) instead of via srun on a GPU node. - Monitoring: remove the upstream PR link; describe the v2.6.5 fixes without referencing PR numbers. * README: move Templates after Cleanup; order Configuration as container/GPU/storage; number sections - Reorder Configuration subsections to Container runtime -> GPU compute -> Storage (user-facing choices first). - Move the Templates reference section to after Cleanup (reference material near the end). - Number all top-level sections 1-13; fix the affected internal anchor links. * deploy-all: optional Grafana public access via GrafanaPublicAccessCidr (login-only SG) Add GrafanaPublicAccessCidr (default empty = SSM-only, unchanged behavior). When set to a CIDR, deploy-all creates a login-only security group opening HTTPS/443 to that CIDR and attaches it to the login node via a new add-cng.yaml ExtraSecurityGroupId parameter — so Grafana is reachable at https://<login-public-ip>/grafana/ without SSM port forwarding. Crucially the SG is attached ONLY to the login node, not the shared cluster SG, so compute nodes and FSx are not exposed. Validated; AllowedPattern guards the CIDR; README + PARAMETERS.md updated. * README: document Grafana access as two options (SSM port-forward vs GrafanaPublicAccessCidr) Restructure 'Accessing Grafana' into Option A (SSM port forwarding, default/private) and Option B (direct public access via GrafanaPublicAccessCidr — login-only SG on 443), with the shared admin-password step up top and security notes (login-only exposure, tight CIDR over 0.0.0.0/0, self-signed cert). Update the Key Features monitoring bullet. * README: link Quick Start to the Grafana public-access option (GrafanaPublicAccessCidr) * README: fold Cleanup into Quick Start (Console or CLI); merge User Management into Additional Resources; renumber - Move cleanup to the end of Quick Start; note deletion via CloudFormation Management Console (select stack -> Delete) or CLI. - Merge the User Management LDAP link into Additional Resources. - Renumber sections (now 1-11). * docs: move architecture-components.md and PARAMETERS.md into docs/; fix links Create architectures/aws-pcs/docs/ and relocate the two reference docs there. Update README links to docs/PARAMETERS.md and PARAMETERS.md's back-links to ../README.md. * tests: drop obsolete monitoring test assets; refresh tests/README; bump monitoring-stack-test to v2.6.5 - Remove pcluster-monitoring-test-cluster.yaml + deploy-monitoring-test.sh + monitoring-ubuntu24-fixes.md (the Ubuntu/PCS fixes are resolved upstream in v2.6.x, and deploy-all covers the monitoring test path). Also removed the stale template from the test bucket. - tests/README: drop the Ubuntu-workaround use case, add the GPU Training Validation entry, note Grafana public access. - monitoring-stack-test.md: v2.6.3 -> v2.6.5 references. * tests: consolidate into a single guide; flatten scripts; cover full matrix - Merge monitoring-stack-test.md + gpu-training-validation/ into one tests/README.md with a coverage matrix (monitoring; Enroot/Pyxis via UserData AND ImageBuilder; CPU; G-series single-NIC GPU; P5/P6-b200/p6-b300 multi-NIC; NCCL; FSDP), each test stating what to run and the expected result inline. - Flatten gpu-training-validation/ to tests/ root; replace the Megatron stage with a working FSDP smoke test (03-fsdp-llama2.sbatch). - enroot import now runs directly on the login node (300 GiB root), not a batch job; drop the /dev/shm staging and the AMI-native NCCL variant. - Fix the README link to the new tests guide. * README: add OpenZFS (/home) throughput row to the FSx storage table The Storage table listed Lustre PerUnitStorageThroughput but omitted the OpenZFS HomeThroughput. Add it with the per-deployment-type valid values (SINGLE_AZ_2/HA_2 = 160..10240, SINGLE_AZ_HA_1 = 128..4096, SINGLE_AZ_1 = 64..4096), per the FSx docs. * tests/README: add a Verified configurations section Summarize the configs validated on real hardware (2x p6-b300 ~760 GB/s NCCL / ~195 TFLOPS FSDP; 2x p5 ~480 GB/s / ~60 TFLOPS; CPU+monitoring; Grafana public access), the deploy path used, and the EFA-NIC-count / FSDP-loss notes. * templates: fix inaccurate top-level Descriptions - cluster.yaml: was described as a full HPC cluster with login+compute nodes, EFS, and a new VPC — none of which it does. Corrected to: PCS scheduler core only (no nodes), attaches to an existing subnet/SG, optional monitoring IAM + Slurm accounting. - add-cng.yaml: expand the terse one-liner to describe single-NIC CNGs (login/CPU/ single-NIC GPU), FSx mounts, Enroot/Pyxis + monitoring, queue-only-when-QueueName. - ml-cluster-prerequisites.yaml: mention the FSx for OpenZFS (/home) filesystem it also creates (was Lustre-only), and that FSx deployment types/throughput are parameters. * tests/README: add the validated 2x p6-b200 row (NCCL ~654 GB/s/8 nics, FSDP ~223 TFLOPS/GPU) * install-enroot-pyxis.sh: wait for apt/dpkg lock + retry apt (fix cold-boot exit 100) At first boot the script raced Ubuntu's unattended-upgrades for the dpkg lock; the very first apt-get aborted with 'Could not get lock /var/lib/dpkg/lock-frontend ... held by (unattended-upgr)' under set -e (observed as post-install exit 100, leaving the node without Enroot/Pyxis). Add wait_for_apt_lock() (polls the dpkg/apt locks via fuser, up to 300s) and an apt_get() wrapper (DPkg::Lock::Timeout=300 + up to 5 retries); route all apt calls through it. Verified: a manual rerun after the lock cleared installs cleanly. * tests/README: flag B300 NCCL bandwidth as needs-larger-message/more-nodes follow-up Mark the ~760 GB/s as busbw and add a footnote: B300 has 2x the EFA cards of B200 but the 2-node/16 GiB busbw was only ~1.16x higher, so the 2-node test likely doesn't saturate all 16 cards — re-test with larger message sizes (~64 GiB) and more nodes before treating it as B300 peak network bandwidth. * docs: add ROADMAP.md (future implementation TODO) + link from README Add architectures/aws-pcs/docs/ROADMAP.md as a markdown checklist of implementation items under consideration (P6-B300 at-scale NCCL re-test, targeted ODCR, Trn template, Multi-AZ FSx, P6 in the AMI recipe, DCGM throttle-reason signal, test-matrix automation). GitHub Issues are disabled on the fork, so an in-repo, version-controlled checklist is the SSOT. Link it (and PARAMETERS.md) from the README Additional Resources. * docs/ROADMAP: reset to the agreed open items (Multi-AZ prereq, managed monitoring, Trainium validation, test automation) * add-cng templates: make AmiId optional (auto-resolve PCS-ready DLAMI from SSM when empty) Change AmiId in all four add-cng templates from a required AWS::EC2::Image::Id to an optional String (default ''): when empty, the node group's AmiId resolves {{resolve:ssm:/aws/service/pcs/ami/dlami-base-ubuntu2404/x86_64/latest/ami-id}} via a HasAmiId condition; a non-empty ami-xxxx still overrides (e.g. a custom Enroot/Pyxis AMI). deploy-all is unaffected (it always passes a resolved/built AmiId). Users no longer need to look up the AMI ID. * docs/ROADMAP: add user-management backend (LDAP/AD) integration item * docs/ROADMAP: add P6e-GB200/GB300 (Grace-Blackwell, arm64) support item * README: warn against combining BuildAMI=true with the default PostInstallScriptUrl (double install) PostInstallScriptUrl defaults to the Enroot/Pyxis installer and deploy-all passes it to the CNGs regardless of BuildAMI. So BuildAMI=true + the default URL makes each node both boot the pre-baked AMI AND re-run the installer at first boot. Document that BuildAMI=true should be paired with PostInstallScriptUrl=""; reword 'two combinable options' to 'choose one'. * tests/README: note BuildAMI=true must pair with PostInstallScriptUrl="" (avoid double install) * tests: record BuildAMI=true deploy-all validation (no double-install) * aws-pcs: fix PostInstallScriptUrl org (awslabs) and make enroot cache per-user - PostInstallScriptUrl default + tests README clone URL pointed at aws-samples; this repo lives under awslabs, so the raw URLs 404 at first boot. (review #1) - enroot ENROOT_CACHE_PATH was shared (/tmp/enroot) while data/runtime are per-user, so a second user importing an image sharing a cached layer hit an unreadable 0640 blob and the job died. Make the cache per-user too. (review #2) * aws-pcs: validate FSx throughput against deployment type (review #3/#4) FSx throughput is a discrete grid that depends on the deployment type, but the two were uncoupled: the default PerUnitStorageThroughput=250 is invalid for the PERSISTENT_1 fallback the parameter recommends, and HomeThroughput=320 is invalid for the SINGLE_AZ_1 fallback the description steers users toward -- both surfaced as an opaque CreateFileSystem failure deep in the nested prerequisites stack. - Make both throughput params String + AllowedValues (full valid grid -> console dropdown, no typos), keeping the existing defaults (250 / 320) unchanged. - Add a Rules section to both ml-cluster-prerequisites.yaml and pcs-ml-cluster-deploy-all.yaml asserting (DeploymentType, throughput) is a valid pair, so a mismatch fails at stack-create with a clear message. - Cross-reference the coupling in both parameter descriptions. * deploy-all: force PostInstallScriptUrl empty when BuildAMI=true (review #6) BuildAMI=true bakes Enroot/Pyxis into the custom AMI, but PostInstallScriptUrl was still passed (defaulted) to every CNG stack, so the installer ran a second time at first boot. The 'set it empty when BuildAMI=true' rule lived only in the PR body, not in the console parameter descriptions, and nothing enforced it. Guard all five CNG passthroughs with !If [BuildAMIEnabled, '', !Ref ...] so the template enforces it, and cross-reference the interaction in both the BuildAMI and PostInstallScriptUrl parameter descriptions. * install-enroot-pyxis.sh: make idempotent + fix slurmd PATH (review #6/#18) PostInstallScriptUrl is a generic first-boot hook, so the template no longer forces it empty under BuildAMI=true (reverted). Instead the installer itself is now safe to re-run on an AMI where Enroot/Pyxis is already baked in: - Skip the Enroot .deb install when 'enroot version' already equals the target (reinstall only on a version mismatch). - Skip Pyxis per Slurm version when spank_pyxis.so + the plugstack symlink already exist, instead of rebuilding from source every boot. - Derive the slurmd PATH from the slurm-*/bin dirs actually present (a hardcoded slurm-25.11 silently mismatched SlurmVersion=25.05 clusters), and guard the append with grep -qxF so a re-run does not add a duplicate PATH line. - Document the execution context (runs before slurmd starts; does not restart it) and the idempotency contract in the script header. * deploy-all: keep PostInstallScriptUrl as a generic hook, warn (not block) on BuildAMI PostInstallScriptUrl is intended as a generic first-boot hook, not Enroot/Pyxis- specific, so do NOT force it empty under BuildAMI=true (that would block other custom first-boot setup). Revert the !If guard to a plain !Ref passthrough and reword the BuildAMI / PostInstallScriptUrl descriptions to warn about the default installer reinstalling on a pre-baked AMI. The installer is now idempotent (separate commit), so a redundant default re-run is a fast no-op anyway. * docs: GrafanaPublicAccessCidr exposes unauthenticated Prometheus/Pushgateway too (review #7/#8) Opening 443 to a CIDR exposes more than the password-gated Grafana: the login node's nginx also reverse-proxies the unauthenticated /prometheus/, /pushgateway/ and /slurmexporter/ paths, so anyone in the CIDR can read all cluster metrics (and push) without credentials. Document this real exposure in the parameter description, the README Option B security notes, and docs/PARAMETERS.md. Keep 0.0.0.0/0 accepted (no hard block) — it's useful for short-lived PoC / workshop access where per-user local SSM permissions are impractical — but warn to narrow or clear it when done. (#7 documented here; the upstream fix to gate those nginx locations is tracked separately for aws-parallelcluster-monitoring.) * docs: note p6-b300 GPU metrics gap (DCGM 4.2.0 pin); track upstream (review #9) Monitoring v2.6.5 pins dcgm-exporter to DCGM 4.2.0, which doesn't support B300 (needs >=4.4.0); the pin is forced by the Docker 29.x OCI-index pull issue (#47). So p6-b300 GPU dashboards stay empty while every other GPU type is fine. Rather than patch the monitoring internals from the PCS template layer (fragile, and the fix belongs upstream), filed aws-parallelcluster-monitoring#50 proposing a configurable dcgm-exporter image (env var, current pin as default) so a B300-capable build can be supplied by digest without a fork. Document the known limitation in the README monitoring section and the tests verified-configurations matrix (the B300 row no longer overclaims '16 GPUs in Grafana') with a link to the upstream issue. * GPU add-cng: lock InstanceType with AllowedValues + document the per-family split (review #5/#10) Decided to keep one add-cng template per GPU family rather than consolidate: the templates are ~85% identical but their NetworkInterfaces EFA layouts genuinely differ (p5/p6-b200 put EFA on card 0 with DeviceIndex 1; p6-b300 has an ENA-only card 0 with EFA on cards 1-16 at DeviceIndex 0, 17 cards total). A single template would need ~32 per-card !If wrappers, or an Fn::ForEach transform whose required CAPABILITY_AUTO_EXPAND breaks the one-click quick-create links -- and would mean editing the already-validated p5/b200 blocks. - #5: add AllowedValues to InstanceType on the p6-b200 and p6-b300 templates (p5 already had it), so a type whose NIC layout doesn't match the template's block fails at stack validation instead of with an opaque launch error. - #10: document the split rationale in a header comment on all three templates and add an Fn::ForEach consolidation item to docs/ROADMAP.md as future work. * tests: reuse canonical NCCL/FSDP assets instead of shipping copies (review #11-#14) The tests/ dir duplicated assets the repo already maintains, which would drift and violated the 'reuse existing assets' contribution rule. Drop the four local scripts and point the guide at the canonical sources, documenting only the PCS-specific deltas: - Remove 01-nvidia-smi.sbatch (#13) -> a one-line interactive srun in Test 5. - Remove 02-nccl-tests.sbatch + 02-import-nccl-image.sh (#11/#14) -> Test 6 uses micro-benchmarks/nccl-tests/slurm/nccl-tests-container.sbatch; PCS deltas = import on the login node (Lustre can't host the enroot overlay) with a PINNED image tag (no more 'latest'), and the GPU queue partition name. - Remove 03-fsdp-llama2.sbatch (#12) -> Test 7 uses 3.test_cases/pytorch/FSDP (create_venv.sh + slurm/llama2_7b-training.sbatch); PCS deltas = venv + HF_HOME on /fsx (not NFS /home), and --nodes=2. - Replace the intro script table with a canonical-asset reference table. Verified-configurations matrix and footnotes are kept (real-HW results). * install-enroot-pyxis.sh: document that Slurm 24.11 is out of scope (review #17) The reviewer noted the script only builds Pyxis for 25.05/25.11 while PCS also supports 24.11. That scoping is deliberate and already consistent: cluster.yaml's SlurmVersion AllowedValues are 25.05/25.11 only, so the templates can't create a 24.11 cluster. Rather than add an unvalidated 24.11 build, document in the script header that 24.11 is intentionally out of scope and to keep SLURM_VERSIONS in sync with cluster.yaml's AllowedValues. * docs: AMI-pinning tip for production + align double-install notes with idempotent installer (review #16) #16: AmiId defaults to the SSM /latest/ PCS-ready DLAMI path (no version-pinned parameter exists) and CloudFormation re-resolves it on stack updates, so scale-out nodes can drift to a newer AMI. Add a production tip to README: resolve it once and pass the result as AmiId so every node is identical. Also align the BuildAMI + PostInstallScriptUrl notes (README + tests/README) with the #6 resolution: the installer is now idempotent and PostInstallScriptUrl stays a generic hook, so leaving the default under BuildAMI=true is a fast no-op (not a conflict); PostInstallScriptUrl="" is recommended only for the cleanest boot. * deploy-all + CNG: add DcgmExporterImage param to override dcgm-exporter (review #9, monitoring #50) Wire an optional DcgmExporterImage through deploy-all -> every CNG stack -> the monitoring UserData, which exports DCGM_EXPORTER_IMAGE before running post-install.sh. Empty (default) keeps the monitoring stack's own default dcgm pin (DCGM 4.2.0, up to B200), so existing behaviour is unchanged. For p6-b300 (needs DCGM >= 4.4.0) set a newer build, ideally by digest to bypass the Docker 29.x OCI-index pull failure. Pairs with aws-parallelcluster-monitoring #50 (DCGM_EXPORTER_IMAGE support); lets a one-click deploy enable B300 GPU dashboards without a monitoring fork. * tests/README + README: reflect B300 validation findings (review #9 follow-up) From the live 2x p6-b300 run (Slurm 25.11, Docker 29.5.3, DCGM 4.5.2 by digest): - B300 GPU metrics DO populate once DcgmExporterImage is set to a B300-capable build by digest (bypasses the Docker-29.x OCI-index pull failure). Replaces the earlier 'GPU metrics empty' note; verified all 16 B300 GPUs in Grafana. Needs monitoring DCGM_EXPORTER_IMAGE support (aws-parallelcluster-monitoring#51, for #50). - NCCL all_reduce scales past the canonical 16 GiB sweep: ~654 GB/s @16 GiB but ~751 GB/s @64 GiB on 2x p6-b300 (16 nics) -- so 16 GiB was unsaturated, answering the review's 'larger message size' ask. 128/256 GiB OOM (buffer > B300 memory); scale nodes not buffer for a higher peak. - FSDP validated two ways on p6-b300: shared-/fsx venv (~200 TFLOPS/GPU) and Enroot/Pyxis container via CONTAINER_IMAGE (~193 TFLOPS/GPU). Test 7 now documents both, plus the venv PATH --export needed for torchrun to resolve on every node. - Updated the verified-configurations matrix and the README NCCL/monitoring notes to match. * install-enroot-pyxis.sh: pin slurmd PATH to the cluster's Slurm version (fix Pyxis break on multi-version DLAMI) Regression found on a live p6-b300 cluster: every --container-image job died with 'spank_pyxis.so: task_init() failed'. Root cause: the slurmd PATH line put every /opt/aws/pcs/scheduler/slurm-*/bin on PATH, and glob order puts the OLDEST (24.11) first. The Pyxis/Enroot PMI hook runs 'scontrol show config'; the 24.11 scontrol fatally fails to parse the 25.11 cluster's slurm.conf ('unrecognized key: MetricsType' -- set by cluster.yaml for the monitoring OpenMetrics endpoint), aborting the hook. Fix: resolve the cluster's actual Slurm version from the live controller (slurmd has --conf-server) via a probe 'scontrol show config', then put only that version's bin on PATH (matched on major.minor). Falls back to the newest installed version if the controller isn't reachable (e.g. during an AMI build). Idempotent: strips any prior PATH= line before appending. Verified on the live p6-b300 cluster (Slurm 25.11): picks slurm-25.11/bin, and NCCL + FSDP --container-image jobs run cleanly. This supersedes the earlier glob approach that itself regressed the previously-working hardcoded single-version PATH. * Default MonitoringVersion to v2.9.1 (DCGM_EXPORTER_IMAGE support for B300) aws-parallelcluster-monitoring v2.9.1 shipped the DCGM_EXPORTER_IMAGE override (PR #51, for #50) that lets DcgmExporterImage enable B300 GPU metrics without a fork, and carries the PCS /opt install + Docker-29.x DCGM fixes from v2.6.4+. Bump the MonitoringVersion default v2.6.5 -> v2.9.1 across deploy-all + the four add-cng templates, and update the README / PARAMETERS / tests prose to match. tests/README also records the v2.9.1 B300 validation (GPU metrics via digest, NCCL 64 GiB ~751 GB/s, FSDP venv ~205 TFLOPS / container ~193 TFLOPS). * install-enroot-pyxis.sh: derive slurmd PATH from the slurmd unit, not scontrol (boot-safe) The previous version detected the cluster Slurm version by running scontrol show config -- but this script runs at first boot BEFORE slurmd starts, when the config source isn't established yet, so scontrol prints nothing and (under set -eo pipefail) the failing command substitution aborted the whole post-install with exit 1. The PATH line was never written (slurmd still ran via its absolute ExecStart path, but the post-install reported failure and the Pyxis PMI hook had no usable scontrol on container nodes). Derive the version from the slurmd systemd unit's ExecStart instead (/opt/aws/pcs/scheduler/slurm-<ver>/sbin/slurmd), which PCS sets to the chosen version and is readable at first boot. Pins PATH to exactly that version's bin; falls back to the newest installed version if the unit can't be read (AMI build). Verified on a live p6-b300 (Slurm 25.11): extracts slurm-25.11/bin correctly. * cluster.yaml: gate MetricsType on Slurm 25.11+ (fix 25.05 cluster create failure) PCS rejects cluster creation when SlurmCustomSettings includes MetricsType on Slurm < 25.11 ('MetricsType is only supported for Slurm version 25.11 or later'), but the template emitted MetricsType=metrics/openmetrics + CommunicationParameters= enable_http whenever DeployMonitoring=true, regardless of version. So a 25.05 cluster with monitoring (a valid AllowedValues combo) always failed to create. Gate those two settings on a new IncludeMetricsConfig condition (monitoring AND SlurmVersion=25.11). Also redefine IncludeSlurmConfig so monitoring-on-25.05 doesn't force an otherwise-empty SlurmCustomSettings list (which PCS also rejects): emit SlurmConfiguration only when accounting, an accounting policy, or the metrics config actually applies. The Slurm OpenMetrics dashboards simply won't populate on 25.05; all other monitoring (node/GPU/CloudWatch) is unaffected. * install-enroot-pyxis.sh: resolve slurmd PATH from PCS profile.d/unit (robust, boot-safe) slurmd needs the cluster's Slurm bin on its PATH for the Pyxis PMI hook (50-slurm-pmi.sh runs 'scontrol show config' in the job env; PCS's /etc/profile.d/slurm.sh only affects login shells, and slurmd's own PATH omits scontrol -- so without this, --container-image jobs fail with 'Command not found: scontrol'). Resolve the cluster's version from the first source present at post-install time: (1) the /etc/profile.d/slurm.sh symlink PCS points at the chosen version, then (2) the slurmd systemd unit ExecStart, then (3) newest installed version. Each step is guarded with an if + '|| true' so a missing source can't trip set -e (the prior '[ -n VER ] && ...' form returned non-zero and aborted post-install with exit 1 when the version source wasn't ready yet at first boot). Pins exactly one version's bin, so an older scontrol never shadows and breaks on a newer slurm.conf key (MetricsType). Verified resolution under set -e on a live 25.11 node. * Pass cluster Slurm version to post-install via PCS_SLURM_VERSION env (fix Pyxis PATH) The node can't discover the cluster's Slurm version at post-install time (cloud-init, before slurmd/profile.d/controller exist), and the PCS DLAMI ships multiple versions where an older scontrol breaks the Pyxis PMI hook on a newer slurm.conf. Pass the version in explicitly instead of guessing: - install-enroot-pyxis.sh: read PCS_SLURM_VERSION and pin slurmd's PATH to that version's bin; fall back to newest installed if unset (manual run). This is also the natural way to get the version once PCS ships a native post-install hook. - add-cng*.yaml (x4): add a SlurmVersion parameter (Default 25.11, AllowedValues 25.05/25.11 -- so direct add-cng use or an older deploy-all that omits it still works) and export PCS_SLURM_VERSION before running the post-install script. - deploy-all: pass SlurmVersion through to every CNG stack. - pcs-ready-dlami (AMI build): the AMI is cluster-agnostic, so list all slurm-*/bin newest-first on slurmd's PATH (no older-version shadowing) instead of hardcoding 25.11; a PostInstallScriptUrl boot then rewrites it to the exact cluster version. * install-enroot-pyxis.sh: build Pyxis only for the cluster's Slurm version + per-version plugin path Two linked fixes for the SPANK plugin version mismatch that crashed slurmd on a 25.05 cluster ('Incompatible Slurm plugin version (25.11.2)'): 1. The Pyxis Makefile installs spank_pyxis.so to $libdir/slurm and bakes that absolute path into the generated pyxis.conf. The old code used the single shared /usr/local/lib for every version, so building 25.05 then 25.11 left only the 25.11 .so, and the 25.05 plugstack pointed slurmd at it -> slurmd refused to start. Install each version to a per-version libdir (/opt/aws/pcs/scheduler/slurm-<ver>/lib) and write that exact path into the version's plugstack pyxis.conf. 2. Build ONLY the cluster's version when PCS_SLURM_VERSION is provided (the common path), instead of all of them -- no wasted build, no chance of a wrong-version .so. Falls back to all supported versions when unset (manual run / cluster-agnostic AMI). Idempotency check now keys on the per-version .so, not the shared path. * tests/README: add pre-merge full test matrix + Enroot/Pyxis regression-test rule Several real bugs this round only appeared in specific combinations (25.05-only MetricsType rejection, a Pyxis SPANK plugin built for the wrong Slurm version stopping slurmd, a first-boot post-install failure). Document the bar for merging: - A 'Pre-merge full test matrix' up top: every Slurm version (25.05 AND 25.11), both container-runtime paths (first-boot install AND BuildAMI=true), clean first boot, CPU + each GPU family, monitoring, and template lint. - A regression-test rule in Test 2: ANY change to install-enroot-pyxis.sh must be retested on all supported Slurm versions and via a BuildAMI=true run (the AMI path carries its own copy of the logic), on a clean first boot. Also corrected the Pyxis plugin path to the per-version location. * AMI build: single-version Pyxis (add SlurmVersion param, plumb from deploy-all) The pcs-ready-dlami AMI used to bake Pyxis for every supported Slurm version into a shared /usr/local/lib/slurm path. Each make install clobbered the previous version's spank_pyxis.so, leaving only the last version (25.11) -- so an AMI used by a 25.05 cluster crashed slurmd at start with 'Incompatible Slurm plugin version'. Mirror the post-install fix: build Pyxis only for the cluster's Slurm version, passed in as a new SlurmVersion CFN parameter (Default 25.11, AllowedValues 25.05/25.11). The slurmd PATH is then a single, unambiguous bin dir for that version -- no multi-version glob, no version-detection, no .so collision. - pcs-ready-dlami-with-enroot-pyxis.yaml: add SlurmVersion parameter; switch the EnrootPyxisComponent Data to !Sub so ${SlurmVersion} is injected; bash vars in the block are escaped to ${!var}; SLURM_VERSIONS pinned to the parameter; slurmd PATH set directly to /opt/aws/pcs/scheduler/slurm-${SlurmVersion}/bin. - pcs-ml-cluster-deploy-all.yaml: pass SlurmVersion through to DLAMIStack so the AMI always matches the cluster's Slurm version. Backward compatible: the AMI build template still has a SlurmVersion default (25.11), matching the cluster.yaml default. * docs: add OPERATIONS.md and slim README/tests caveats to one-line pointers Move the recurring operational caveats (Slurm 25.05 vs 25.11 capability differences, AMI single-version rule, MonitoringVersion migration notes, B300 DcgmExporterImage setup, AMI pinning for production, FSx deployment-type/throughput coupling, P6-B300 NIC topology lock-in) into a dedicated docs/OPERATIONS.md so they live in one place and stop expanding the README. Also document things this round's testing made explicit: - AMI build is single-Slurm-version by design (the SPANK plugin ABI is version-locked, so SlurmVersion on the DLAMI stack must match the cluster). - The PostInstall path passes the cluster's Slurm version via the PCS_SLURM_VERSION env var (UserData -> install-enroot-pyxis.sh) -- the node can't discover it at first boot, before slurmd / profile.d / the controller config exist. - v2.6.5 -> v2.9.1 brings Grafana 13 and DCGM_EXPORTER_IMAGE support. README and tests/README keep the structural content; the long caveat blocks become one-line pointers into OPERATIONS.md sections. tests/README adds measured stack creation times for the four CPU patterns (PostInstall and BuildAMI x 25.05/25.11) and the GPU+CB run, and the Pre-merge matrix's BuildAMI row now requires verifying the SlurmVersion=cluster.SlurmVersion match. * docs/PARAMETERS.md: add DcgmExporterImage; cross-reference OPERATIONS.md from SlurmVersion + BuildAMI + MonitoringVersion DcgmExporterImage was missing from the parameter reference (it landed in the deploy-all template earlier in this branch). Add it under Developer/Advanced with a pointer to the validated B300 digest in OPERATIONS.md §3.1. Tighten three existing rows that have non-trivial operational implications now in OPERATIONS.md: - SlurmVersion -> §1 (Slurm OpenMetrics is 25.11+ only; the value also drives Pyxis build version and the AMI it gets baked into). - BuildAMI -> §2 (the AMI is single-Slurm-version; SlurmVersion on the AMI build must match the cluster's). - MonitoringVersion -> §3 (v2.6.5 -> v2.9.1 migration notes, Grafana 13). * docs/OPERATIONS.md: explain why single-version Pyxis is fine across cluster Slurm upgrades Add §2.3 covering the upgrade-compatibility argument for the single-version pin: per the Slurm upgrade policy, scontrol/srun from version N interoperate with a slurmctld of N+1 (or N+2/N+3 on 24.11+), so a cluster upgrade does not break nodes built for the prior version; advancing means setting the new SlurmVersion and redeploying so Pyxis is rebuilt for it. Pairs with Reply E to KeitaW's #844 review thread. * docs/OPERATIONS.md: record the cgroup-v2 prolog race as a known issue Hit reproducibly during the 4-pattern CPU validation: the very first srun against a freshly-launched cpu1 node fails with 'Job prolog failed' and drains the node, while subsequent srun's land on sibling nodes and run cleanly. Cause is below the templates: PCS bootstrap starts slurmd, then ~9s later systemd forces a slurmd restart because cgroup-v2 cpuset.cpus/cpuset.mems setup fails ('No space left on device' is the Linux misleading-error for an empty/invalid cpuset value). The user srun arrives in that gap, the prolog handshake never reaches the re-started slurmd, and the controller times out after 420s. Nothing this PR's scripts modify (cgroup, slurmd unit, post-install timing) is in the path. Recorded under §7 Known issues with the journal trace and the workaround ('scontrol update nodename=cpu1-N state=resume', or just resubmit and let it land on a sibling). Recommendations recap renumbered to §8; existing OPERATIONS.md anchors referenced from README/tests/PARAMETERS are unaffected (#1-#4). * docs/ROADMAP.md: add Software stack section (Spack, Intel oneAPI, NVIDIA HPC SDK, modules) The cluster currently ships only the PCS DLAMI's pre-installed CUDA/NCCL/EFA stack plus Enroot/Pyxis for containers — no native HPC package manager and no first-class hook for traditional toolchains (Intel oneAPI, NVIDIA HPC SDK). Track these as future work so users running source-built MPI/BLAS/scientific apps or Fortran workloads have a documented path; once Spack lands, an Lmod-style module system rounds it out. * README: surface SlurmVersion in §4 + document the AMI build's pipeline rationale Two pending edits: - §4 'Most-used parameters' table didn't list SlurmVersion. It became a structural parameter this round (drives Slurm OpenMetrics availability, the Pyxis build version, and the AMI's baked Slurm) and deserves a row alongside BuildAMI etc. Cross-references OPERATIONS.md §1 for the full version trade-off. - §4 Container runtime: spell out what the in-stack BuildAMI=true template earns beyond a one-shot Image Builder run (managed pipeline: scheduled rebuilds, AMI lifecycle deprecation, SSM parameter publishing of the latest AMI ID), so the rationale matches Reply F to KeitaW's follow-up review and the next reviewer doesn't re-litigate it. * README: tighten Key Features (capacity scope, drop one-click duplication) Two small clarity fixes: - 'Flexible capacity' was ambiguous — could be read as 'flexible WRT capacity' rather than 'supports a wide range of capacity-purchase options'. Reword to 'Broad capacity-purchase support', explicit about covering OD / ODCR / Capacity Blocks for ML, and selected per node group. - 'One-click or modular' duplicated the first bullet's one-click claim. Drop the one-click half and lead with the value of the modular path: composing individual stacks for infrastructure reuse / iterating on one piece at a time. * README §7: reuse canonical micro-benchmarks/nccl-tests sbatch instead of inline heredoc The §7 'Running a multi-node GPU job' walk-through inlined a hand-rolled nccl-test.sbatch heredoc, which duplicated the canonical launcher at micro-benchmarks/nccl-tests/slurm/nccl-tests-container.sbatch and could drift from it. Same reuse rule we already applied to tests/ (where 02-nccl-tests.sbatch was deleted and Test 6 redirected to the canonical asset): document only the PCS-specific deltas (login-node import + GPU partition name) and let the canonical sbatch carry the rest. Also pins the NCCL image tag instead of using mutable latest. Pointer to the full Test & Validation Guide added at the bottom. * Default DcgmExporterImage to a DCGM 4.5.2 digest covering all supported GPUs The empty default pushed B300 users into a forced manual override (and a Grafana that quietly stays empty until they figure that out). DCGM 4.5.2 covers Hopper / B200 / B300 with no DCGM_FI_DEV_* field changes vs 4.2.0 per upstream changelog, was validated on real B300 hardware in this round, and the digest pull bypasses the Docker-29.x OCI-index failure on newer NVCR tags. So default to it across the GPU range — every supported family populates Grafana out of the box. The monitoring stack's older 4.2.0 pin remains a one-line override for anyone who needs to match another fleet. Templates (deploy-all + 4 CNGs): Default '' -> the validated digest, parameter description rewritten to lead with the new behavior. README §8 / §4, PARAMETERS, tests/README, OPERATIONS.md §3.1 (renamed: 'B300 needs DcgmExporterImage' -> 'DcgmExporterImage — the default, and when to change it') all updated; cross-references re-pointed at the new anchor. * Strip personal-fork references before merge Two leftovers from the dev / review cycle that shouldn't ship to awslabs/main: - README §8 'Prefer AWS-managed Prometheus/Grafana?' linked to github.com/DaisukeMiyamoto/awsome-distributed-ai/tree/deploy-monitoring/... (a fork URL on the in-flight branch). The path lives on awslabs/main already, so switch to a repo-relative link that resolves correctly under any fork/branch and doesn't bake the dev fork name into mainline docs. - The 4 add-cng*.yaml templates' MonitoringRepo description carried the example 'e.g. DaisukeMiyamoto/aws-parallelcluster-monitoring' as a 'how to test unreleased changes' hint. Defaults are correct (aws-samples/...); drop the personal-fork example from the prose so there's no leftover advertisement of an internal fork. * deploy-all: refresh DcgmExporterImage console label to match new default The Parameters/AWS::CloudFormation::Interface label still said '(empty = default; set for B300)' from the era when the param defaulted to '' and B300 users had to set it explicitly. Now that the default is the DCGM 4.5.2 digest covering Hopper/B200/B300 out of the box, the old label misleads. Update to '(default DCGM 4.5.2 by digest; covers H100/B200/B300)'. * Fix broken ldap_server reference paths The README's Additional Resources list and ROADMAP's User-management item both pointed at `6.ldap_server` / `architectures/6.ldap_server` -- but the actual upstream layout puts it under `1.architectures/6.ldap_server`. From `architectures/aws-pcs/` that resolves through a `../../1.architectures/` relative path. Update both refs so the link doesn't 404 on the rendered PR. * deploy-all: improve 1-click parameter UX -- split AMI build, reorder, slim descriptions Console parameter groupings reorganized for the 1-click case: - New §3 'Container Runtime (Post-install Script)' carries just PostInstallScriptUrl/Args + RootVolumeSize -- the hook every cluster goes through. Previously these were buried alongside Image Builder params, which most users do not touch. - AMI build params (BuildAMI/BaseAmiId/SemanticVersion/BuildSchedule) moved to a new §7 'Custom AMI Build (Optional - skip unless you need a pre-baked DLAMI)', ahead of the Developer group. They are an opt-in path; surfacing 'Optional' in the section title makes that obvious instead of confronting first-time users with SemanticVersion/BuildSchedule with no context. - §2 PCS Cluster: SlurmVersion now precedes LoginNodeInstanceType. The Slurm major is the substantive cluster-shaping choice; the login type is rarely changed from the default. - DcgmExporterImage was orphaned (not in any ParameterGroup, so the console rendered it under a generic 'Other' heading). Now grouped under §8 Developer/ Advanced next to MonitoringVersion / MonitoringRepo where it belongs. Description cleanups -- the long help blocks rendered as a wall of text in the console; trimmed to the parameter-level essentials and pointed at OPERATIONS.md for the rest: - PostInstallScriptUrl: 12 lines -> 5 (kept the BuildAMI=true interaction note, dropped the duplicated 'PCS equivalent of OnNodeConfigured' phrasing). - GrafanaPublicAccessCidr: 14 lines -> 5 (kept the unauthenticated-proxy-paths warning since that is the real footgun, dropped the per-network detail that belongs in OPERATIONS.md §3.2). - BuildAMI: 4-line one-paragraph -> 4 short lines covering the trade-off, the PostInstallScriptUrl pairing, and the single-Slurm-version constraint. - DcgmExporterImage: 6 lines -> 4 -- focused on 'why digest, not tag' since that's the non-obvious part. Migration / override examples stay in OPERATIONS.md §3.1. - PerUnitStorageThroughput / HomeThroughput: tightened to call out that the valid grid is enforced by a Rule, since users see the validation error first. - PrimarySubnetAZ: rewrote from 'AZ id where the subnets will be created' to 'the only required choice' to match the 1-click framing. ParameterLabels updated for the params whose section moved (AMI section now flagged 'optional; auto-resolved if empty' on BaseAmiId so first-time users know to leave it alone). CapacityReservationId label spelled out as 'Capacity Block for ML reservation ID' so it reads correctly in the console even though the parameter name itself is the generic 'CapacityReservationId'. PARAMETERS.md ToC restructured to mirror the new 8 console groups; README §4 text 'all 7 console parameter groups' bumped to 8. No default values changed. Validate-template still reports 42 parameters, no orphans, all relative cross-links resolve. * deploy-all: drop docs/ refs from descriptions; move RootVolumeSize to PCS Cluster Two follow-ups on the 1-click parameter UX: - Parameter Descriptions render as plain text in the …
shimomut
added a commit
that referenced
this pull request
Jul 17, 2026
…et names When NamePrefix was empty the marker bucket name substituted !Ref AWS::StackName. Stack names permit uppercase and up to 128 chars; S3 bucket names permit neither (lowercase only, <=63 chars). A console deploy of the committed template (which the README explicitly invites) with a stack name like 'HyperPodDevOpsAgent' therefore hit InvalidBucketName late in create, triggering a full rollback including AgentSpace/webhook teardown. deploy.sh always passes NamePrefix, so make it required (MinLength: 1, drop the '^$' pattern alternative) — the failure becomes an upfront parameter-validation error rather than a mid-create rollback. Removed the now-dead HasNamePrefix condition and the StackName fallback. Regenerated hyperpod_devops_agent.yaml. Addresses PR #1191 review feedback (KeitaW: #7).
dmvevents
added a commit
to dmvevents/awsome-distributed-ai
that referenced
this pull request
Aug 25, 2026
…rect-SHA fetches, benchmark provenance - kubernetes/vllm-deepep-v2-2node.yaml:145,150 — worker probe bracket form `[v]llm serve` so the exec-shell's own cmdline no longer self-matches pgrep (thread awslabs#1/awslabs#9) - kubernetes/vllm-deepep-v2-2node.yaml:97 — append `|| true` to the LEADER_IP substitution so the FATAL DNS guard can fire under set -e + pipefail (thread awslabs#2) - kubernetes/vllm-deepep-v2-2node.yaml:159 — hugepages-2Mi comment: needs pre-allocated 2Mi hugepages or the pod sits Pending (thread awslabs#11) - kubernetes/vllm-deepep-v2-2node.yaml:115 — DEEPEP_ARCH_LIST commented env now names 10.0 (b200) + 10.3 (b300) so the Blackwell knob is reachable from the manifest (thread awslabs#12) - recipe/verify-image.sh:24,33,34 — convert fi_info efa-direct, ncclGetLsaDevicePointer, ncclGinPlugin checks from `grep -q` to draining `[ "$(... | grep -c X)" -ge 1 ]` (SIGPIPE-141 flake under pipefail), matching Dockerfile Layer 5b (thread awslabs#3) - recipe/run-kernel-test.sh:16 — worker requires an explicit node-rank (`${3:?...}`) instead of defaulting to 0 and colliding with the leader; add the `case "$ROLE"` guard matching serve.sh (thread awslabs#4) - recipe/benchmark_probe.py:18,24 — MODEL reads SERVE_MODEL env with the current value as fallback (+ `import os`) so a non-default model doesn't 100%-fail (thread awslabs#5) - setup_deepep_v2_efa.sh:30,52 — fetch the immutable PR-head SHA directly instead of the moving `refs/pull/N/head` ref for both aws-ofi-nccl #1351 and DeepEP awslabs#612 (thread awslabs#6) - Dockerfile:116 — CMD banner names /opt/serve.sh + /opt/build_deepep.sh (the in-image paths), not recipe/ (thread awslabs#7) - recipe/build_deepep.sh:120 — `tee /tmp/deepep-build.log | tail -30` so a failed build keeps the full nvcc diagnostic (thread awslabs#8) - Dockerfile:60 — gdrcopy pinned by commit SHA (v2.5.2 == c91ad9f) via fetch/checkout instead of the movable `--branch v2.5.2` tag (thread awslabs#13) - recipe/serve.sh:124 — comment on why --trust-remote-code is unconditional (DeepSeek/Kimi models the preflight supports); chmod 755 serve.sh + run-kernel-test.sh in-tree to match the other four scripts (thread awslabs#14) - README.md:114 — in-pod benchmark exec sets OUT_ROOT=/work/benchmarks so results land on the /work volume, not the ephemeral container layer (thread awslabs#16) - setup/env_vars.example:2,12 — header says build-push.sh sources it and recipe/*.sh need `source` first; reword the AWS_OFI_NCCL_PR_SHA note (empty does NOT skip the cherry-pick) (thread awslabs#17) - benchmarks/README.md:88 — caveat: no alternative-backend baseline measured; every table is --all2all-backend deepep_v2 (thread awslabs#18) - recipe/serve.sh:48 — EP_EFA_MAX_QPS comment records the pinned plugin (9c44d34) is 76 commits past the seq-window redesign (6e504db) so the 128-slot cap's precondition no longer holds; default left unchanged (pods down) (thread awslabs#21) Already present at c025946 (prior commits), verified not duplicated: - README shared-experts caveat naming #47785 + DeepSeek (thread awslabs#15) - benchmark_probe.py requests-per-level distribution + unique prompt prefix (threads awslabs#19/awslabs#20) - probes + publishNotReadyAddresses structure (thread awslabs#9) and requests==limits Guaranteed QoS (thread awslabs#10) Signed-off-by: Anton Alexander <dmvevents@gmail.com>
dmvevents
added a commit
to dmvevents/awsome-distributed-ai
that referenced
this pull request
Aug 26, 2026
…2/GIN script Round-3 doc-drift: the VLLM_SHA bump (e2f993dc4 -> 14617c2b, #52632's merge commit) and the appearance of the canonical setup_deepep_gin.sh (2026-08-24) left several docs describing the OLD state. Reconciled WITHOUT fabricating any new measurements — the benchmark tables were measured on the old pin and are now labeled historical, not re-claimed on the new pin. - README setup-script rationale (awslabs#1): reworked to point at the canonical micro-benchmarks/.../deepep-v2-benchmark/setup_deepep_gin.sh and name the three deliberate divergences (unmerged aws-ofi-nccl #1351 param; CPU-proxy vs EFA-GDA; vLLM-wheel torch/NCCL ABI coupling). Kept the correct vendor-sync point (that CI gates only the NVSHMEM setup_deepep_efa.sh). Did NOT adopt the 'delete the DeepEP half + call the canonical' substitution — that needs a docker build to verify and this change set is docs-only. - DeepEP source divergence + Blackwell (awslabs#2): setup_deepep_v2_efa.sh header note + README Known-limitations now state the source is deepseek-ai/DeepEP@b306af06 (not the amazon-contributing fork the canonical pins) and mark the manifest's DEEPEP_ARCH_LIST=10.x knobs documented-but-not-verified — round-1 saw no loadable Blackwell kernel at CUDA 13.0 on this lineage; the fork's st.bulk 64-bit fix (amazon-contributing/DeepEP#3) is the enabling path, to re-verify. - Provenance honesty (awslabs#6/awslabs#9/awslabs#10): benchmarks/README vLLM row + README build note now disclose the shipped image is 0.26.1rc1.dev1000+g14617c2b6 (four minor versions past the measured 0.22.1rc1.dev283+ge2f993dc4); eager table labeled historical to match the non-eager one; 'remain representative' -> 'historical, not what a rebuild produces'; probe-diff note -> 'not directly comparable'. - Saturation framing (awslabs#13): scoped the ~4.75 tok/s per-stream-flat claim to the EAGER sweep; noted the non-eager c=64 wall rise (26.61->34.18s, 3.75 tok/s). - Dockerfile Layer 5 heading (awslabs#7): old-pin identity (PR#41183 first deepep_v2 commit) -> #52632's merge commit, matching the block below it. - env_vars.example tag (awslabs#8): v1-20260818 -> v2-20260825 (postdates the pin bump; IfNotPresent caches by tag). - README serve section (awslabs#4 sub-ask): name SERVE_ENFORCE_EAGER=0 as the knob. Signed-off-by: Anton Alexander <dmvevents@gmail.com>
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Issue #, if available: N/A
Description of changes: Enable text-based Slurm accounting, which stores the data ina text file on head node. No external database required for this simplistic setup.
By submitting this pull request, I confirm that you can use, modify, copy, and redistribute this contribution, under the terms of your choice.