Skip to content

Bump pillow from 9.4.0 to 10.0.1 in /3.test_cases/4.DDP - #8

Merged
perifaws merged 1 commit into
mainfrom
dependabot/pip/3.test_cases/4.DDP/pillow-10.0.1
Oct 4, 2023
Merged

perifaws merged 1 commit into
mainfrom
dependabot/pip/3.test_cases/4.DDP/pillow-10.0.1

Conversation

@dependabot

@dependabot dependabot Bot commented on behalf of github Oct 4, 2023

Copy link
Copy Markdown
Contributor

Bumps pillow from 9.4.0 to 10.0.1.

Release notes

Sourced from pillow's releases.

10.0.1

https://pillow.readthedocs.io/en/stable/releasenotes/10.0.1.html

Changes

10.0.0

https://pillow.readthedocs.io/en/stable/releasenotes/10.0.0.html

Changes

... (truncated)

Changelog

Sourced from pillow's changelog.

10.0.1 (2023-09-15)

  • Updated libwebp to 1.3.2 #7395 [radarhere]

  • Updated zlib to 1.3 #7344 [radarhere]

10.0.0 (2023-07-01)

  • Fixed deallocating mask images #7246 [radarhere]

  • Added ImageFont.MAX_STRING_LENGTH #7244 [radarhere, hugovk]

  • Fix Windows build with pyproject.toml #7230 [hugovk, nulano, radarhere]

  • Do not close provided file handles with libtiff #7199 [radarhere]

  • Convert to HSV if mode is HSV in getcolor() #7226 [radarhere]

  • Added alpha_only argument to getbbox() #7123 [radarhere. hugovk]

  • Prioritise speed in repr_png #7242 [radarhere]

  • Do not use CFFI access by default on PyPy #7236 [radarhere]

  • Limit size even if one dimension is zero in decompression bomb check #7235 [radarhere]

  • Use --config-settings instead of deprecated --global-option #7171 [radarhere]

  • Better C integer definitions #6645 [Yay295, hugovk]

  • Fixed finding dependencies on Cygwin #7175 [radarhere]

  • Changed grabclipboard() to use PNG instead of JPG compression on macOS #7219 [abey79, radarhere]

... (truncated)

Commits

Dependabot compatibility score

Dependabot will resolve any conflicts with this PR as long as you don't alter it yourself. You can also trigger a rebase manually by commenting @dependabot rebase.


Dependabot commands and options

You can trigger Dependabot actions by commenting on this PR:

  • @dependabot rebase will rebase this PR
  • @dependabot recreate will recreate this PR, overwriting any edits that have been made to it
  • @dependabot merge will merge this PR after your CI passes on it
  • @dependabot squash and merge will squash and merge this PR after your CI passes on it
  • @dependabot cancel merge will cancel a previously requested merge and block automerging
  • @dependabot reopen will reopen this PR if it is closed
  • @dependabot close will close this PR and stop Dependabot recreating it. You can achieve the same result by closing it manually
  • @dependabot show <dependency name> ignore conditions will show all of the ignore conditions of the specified dependency
  • @dependabot ignore this major version will close this PR and stop Dependabot creating any more for this major version (unless you reopen the PR or upgrade to it yourself)
  • @dependabot ignore this minor version will close this PR and stop Dependabot creating any more for this minor version (unless you reopen the PR or upgrade to it yourself)
  • @dependabot ignore this dependency will close this PR and stop Dependabot creating any more for this dependency (unless you reopen the PR or upgrade to it yourself)
    You can disable automated security fix PRs for this repo from the Security Alerts page.

Bumps [pillow](https://github.com/python-pillow/Pillow) from 9.4.0 to 10.0.1.
- [Release notes](https://github.com/python-pillow/Pillow/releases)
- [Changelog](https://github.com/python-pillow/Pillow/blob/main/CHANGES.rst)
- [Commits](python-pillow/Pillow@9.4.0...10.0.1)

---
updated-dependencies:
- dependency-name: pillow
  dependency-type: direct:production
...

Signed-off-by: dependabot[bot] <support@github.com>
@dependabot dependabot Bot added the dependencies Pull requests that update a dependency file label Oct 4, 2023
@perifaws
perifaws merged commit 8c4b322 into main Oct 4, 2023
@dependabot
dependabot Bot deleted the dependabot/pip/3.test_cases/4.DDP/pillow-10.0.1 branch October 4, 2023 15:13
KeitaW pushed a commit that referenced this pull request Jun 4, 2024
…DDP/pillow-10.0.1

Bump pillow from 9.4.0 to 10.0.1 in /3.test_cases/4.DDP
awsankur pushed a commit that referenced this pull request Jun 5, 2024
…DDP/pillow-10.0.1

Bump pillow from 9.4.0 to 10.0.1 in /3.test_cases/4.DDP
dongjin-ml pushed a commit to dongjin-ml/awsome-distributed-training that referenced this pull request Feb 20, 2025
…ases/4.DDP/pillow-10.0.1

Bump pillow from 9.4.0 to 10.0.1 in /3.test_cases/4.DDP
KeitaW pushed a commit that referenced this pull request Feb 17, 2026
…DDP/pillow-10.0.1

Bump pillow from 9.4.0 to 10.0.1 in /3.test_cases/4.DDP
KeitaW pushed a commit that referenced this pull request Feb 17, 2026
…DDP/pillow-10.0.1

Bump pillow from 9.4.0 to 10.0.1 in /3.test_cases/4.DDP
DaisukeMiyamoto added a commit to DaisukeMiyamoto/awsome-distributed-ai that referenced this pull request Jun 4, 2026
…gateway too (review awslabs#7/awslabs#8)

Opening 443 to a CIDR exposes more than the password-gated Grafana: the login
node's nginx also reverse-proxies the unauthenticated /prometheus/, /pushgateway/
and /slurmexporter/ paths, so anyone in the CIDR can read all cluster metrics
(and push) without credentials. Document this real exposure in the parameter
description, the README Option B security notes, and docs/PARAMETERS.md.

Keep 0.0.0.0/0 accepted (no hard block) — it's useful for short-lived PoC /
workshop access where per-user local SSM permissions are impractical — but warn
to narrow or clear it when done. (awslabs#7 documented here; the upstream fix to gate
those nginx locations is tracked separately for aws-parallelcluster-monitoring.)
KeitaW added a commit that referenced this pull request Jun 6, 2026
…tainer runtime to the PCS reference cluster (#1120)

* update README for PCS

* add reference cluster with PCS

* Update PCS README: add deployment options, 1-click deploy, and fix S3 bucket name

* Add unified multi-NIC template for P5/P6 instances and update documentation

* Improve PCS deployment: rename to Pseries, update docs and examples

* Add placement group for on-demand P5/P6 instances in add-cng-p5

* Add launch button and refine PCS documentation

* add launch button image

* fix typos

* fix typos

* Upgrade Slurm support to 25.05/25.11, default to 25.11

* Switch default AMI to PCS-specific DLAMI image

* Change mount path of FSxL from /shared to /fsx

* Simplify DLAMI build to use PCS-ready base image with Enroot/Pyxis only

* fix typos

* fix typos

* update readme in pcs

* add BaseAMI parameter for ImageBuilder

* use awsome-distributed-ai bucket

* fix typos

* update architecture

* update architecture image

* Address PR review feedback: fix permissions, add documentation

- Change FSx Lustre mount permissions from 0777 to 1777 (sticky bit)
- Add upstream repository reference to README
- Add Testing and Validation section documenting tested configurations

* Refactor cluster deployment to use modular add-cng templates

- Separate cluster core (Slurm scheduler) from compute nodes
- Use add-cng.yaml for both login and compute node groups
- Make queue creation optional in add-cng templates (empty QueueName for login nodes)
- Remove AmiId parameter from cluster.yaml (only needed for compute nodes)
- Update pcs-ml-cluster-deploy-all.yaml to deploy login and cpu1 as separate nested stacks
- Change default DeployOnDemandCNG to true (cpu1 queue enabled by default)
- Update README with new deployment architecture and examples

* Use variables for AZ ID in deployment examples

- Add AZ_ID variable at the start of each example
- Makes it easier to change AZ consistently across examples
- Aligns with CAPACITY_RESERVATION_ID variable pattern in Example 4

* Add AWS ParallelCluster Monitoring reference

- Add link to aws-parallelcluster-monitoring GitHub repository
- Provides comprehensive monitoring solution for HPC clusters with Prometheus and Grafana

* Add AWS Parallel Computing Service description

- Clarify what AWS PCS is in the introduction
- Mention it's a fully managed service for HPC with Slurm scheduler

* Add AWS Parallel Computing Service to top-level README

- Include AWS PCS in the list of supported compute platforms
- Update 9.aws-pcs description to clarify it uses Slurm scheduler

* Add AWS ParallelCluster Monitoring integration to PCS templates

Integrates aws-parallelcluster-monitoring stack (Prometheus, Grafana, DCGM)
into PCS cluster deployment with optional, backward-compatible configuration.

Changes:
- Add monitoring-iam-policy.yaml: IAM permissions for CloudWatch, SSM, Pricing API
- Update cluster.yaml: Add RoleName output for IAM policy attachment
- Update add-cng.yaml: Add DeployMonitoring/MonitoringVersion parameters and UserData script
- Update add-cng-p5.yaml: Add monitoring parameters and UserData for P5/P6 nodes
- Update pcs-ml-cluster-deploy-all.yaml: Add MonitoringIAMPolicyStack and pass parameters to all CNGs
- Update README.md: Add monitoring access instructions via Session Manager

Monitoring deployment:
- Login node: Prometheus, Grafana, custom metrics exporters
- Compute nodes with GPU: DCGM exporter (auto-detected)
- Compute nodes without GPU: Node exporter

Key features:
- Backward compatible: DeployMonitoring defaults to 'false' in add-cng.yaml
- Enabled by default in pcs-ml-cluster-deploy-all.yaml (DeployMonitoring='true')
- Installs appropriate components based on node type via post-install.sh from upstream
- Access via Session Manager port forwarding (local:8443 -> remote:443)
- Grafana password stored in SSM Parameter Store

* Add Ubuntu 24.04 compatibility workaround for monitoring stack

The upstream aws-parallelcluster-monitoring installer hardcodes
PLATFORM_USER=ec2-user for PCS, but Ubuntu 24.04 uses ubuntu
as the default user.

Changes:
- Add symlink workaround in UserData: /home/ec2-user -> /home/ubuntu
- Apply to both add-cng.yaml and add-cng-p5.yaml
- Document known limitation in README.md

The workaround creates a symlink before running post-install.sh,
ensuring the monitoring stack installs to the correct home directory
on Ubuntu-based AMIs.

* Fix RoleName output in cluster.yaml

The upstream pcs-iip-minimal.yaml does not export RoleName.
Construct RoleName directly using the same pattern as the nested
stack parameter: AWSPCS-pcs-<hash>-<region>-role

* Add monitoring stack integration for AWS PCS

- Add aws-parallelcluster-monitoring integration with IAM policies
- Add Slurm OpenMetrics configuration (MetricsType, CommunicationParameters)
- Add MonitoringRole parameter (login/compute/none) for flexible deployment
- Add monitoring-role and pcs-cluster-id tags for login nodes
- Add InstanceMetadataTags support for IMDS tag access
- Add PCS API permissions (pcs:GetCluster, pcs:ListComputeNodeGroups)
- Fix shell compatibility (== to = for POSIX sh)
- Add test guide for monitoring stack validation
- Add architecture diagram for monitoring components

* Address PR #1109 review feedback (Round 2)

- Fix stray blank line in root README workshop table
- Change ml-cluster-prerequisites.yaml license from Apache-2.0 to MIT-0
- Fix Pyxis Slurm version compatibility: install for both 25.05 and 25.11
- Update PseriesMinCount description to clarify static vs dynamic scaling
- Add AMI build time note (~30 minutes) to README

* Fix monitoring installation bug with robust pattern matching

Improved sed pattern to fix aws-parallelcluster-monitoring upstream bug:
- Old: Line number based (83s/...) - fragile if file changes
- New: Pattern based with whitespace preservation - more robust
- Pattern: s/^\([[:space:]]*\)local login_id$/\1login_id=/g

Also fixed ClusterName parameter in MonitoringIAMPolicyStack to use
ClusterId instead of stack name for correct SSM parameter path.

Tested with successful automated deployment on pcs-monitoring-test-v4.

* Add architectures/aws-pcs directory for new structure

Created new architectures directory structure alongside existing
1.architectures to prepare for repository reorganization.

Contents:
- architectures/aws-pcs/: Complete AWS PCS architecture templates
- All CloudFormation templates with monitoring integration
- README and architecture diagrams

This maintains backward compatibility with 1.architectures/9.aws-pcs
while establishing the new standardized structure.

* Add tests directory to architectures/aws-pcs

Migrated test documentation from 1.architectures/9.aws-pcs/tests/
to new standardized structure.

Contents:
- README.md: Test directory overview
- monitoring-stack-test.md: Monitoring stack validation guide

* Remove old 1.architectures/9.aws-pcs directory

Removed legacy directory structure as content has been migrated
to architectures/aws-pcs/ in the new standardized structure.

All templates, documentation, and tests are now available in:
architectures/aws-pcs/

* Update README.md to reference new architectures/aws-pcs path

Changed link from 1.architectures/9.aws-pcs to architectures/aws-pcs
to reflect the new directory structure.

* Add ClusterName tag to launch templates

Added ClusterName tag to instance TagSpecifications in both
add-cng.yaml and add-cng-p5.yaml for better instance identification
and filtering.

Tag value: ClusterId parameter

* Apply monitoring integration to architectures/aws-pcs

Synced monitoring integration changes to new directory structure:
- Add DeployMonitoring and MonitoringVersion parameters
- Add MonitoringRole parameter (login/compute/none)
- Add ClusterName tag to launch templates
- Add monitoring stack installation with bug fix for upstream issue
- Add Slurm OpenMetrics configuration
- Add monitoring IAM policy and IMDS tag support

* Add monitoring usage documentation to README

* Fix MonitoringIAMPolicy ClusterName parameter and add missing template

* Integrate IAM resources into cluster.yaml to eliminate external dependencies

- Remove nested stack dependency on pcs-iip-minimal.yaml
- Move IAM Role and Instance Profile directly into cluster.yaml
- Integrate MonitoringPolicy as conditional ManagedPolicy in cluster.yaml
- Remove HpcRecipesS3Bucket and HpcRecipesBranch parameters
- Remove MonitoringIAMPolicyStack nested stack from pcs-ml-cluster-deploy-all.yaml
- Simplify deployment by reducing nested stacks from 6 to 4

* Add Prerequisites stack reuse and UserData Enroot/Pyxis installation

This commit implements two major features for faster testing iterations:

1. Prerequisites Stack Reuse
   - Replace 7 Existing* parameters with single PrerequisitesStackName
   - Use CloudFormation Exports/Imports for automatic resource resolution
   - Reduce subsequent cluster deployment time from 20-30 min to 5-10 min
   - CloudFormation enforces deletion protection for shared infrastructure

2. UserData Enroot/Pyxis Installation
   - Add InstallEnrootPyxis parameter for optional container support
   - Install Enroot 3.5.0 and Pyxis v0.20.0 on first boot
   - Support both Slurm 25.05 and 25.11 versions
   - Enable rapid testing without ImageBuilder (8-12 min boot vs 3 min)

Changes:
- pcs-ml-cluster-deploy-all.yaml: Add PrerequisitesStackName parameter and reorganize metadata
- add-cng.yaml, add-cng-p5.yaml: Add UserData installation logic with proper shell escaping
- ml-cluster-prerequisites.yaml: Add Export aliases for compatibility
- architecture-components.md: Document main components for diagram generation

Deployment modes:
- Mode 1 (Production): BuildAMI=true, InstallEnrootPyxis=false
- Mode 2 (Testing): BuildAMI=false, InstallEnrootPyxis=true
- Mode 3 (Minimal): BuildAMI=false, InstallEnrootPyxis=false

* Refactor PCS deployment templates and update documentation

- Refactor pcs-ml-cluster-deploy-all.yaml to call cluster.yaml + add-cng.yaml directly
  - Remove cluster-with-compute-nodes.yaml intermediate layer
  - Add separate stacks: ClusterStack, LoginNodeGroupStack, OnDemandCNGStack, PseriesCNGStack
  - Maintain README compatibility: cluster.yaml can be deployed standalone

- Delete unused monitoring-iam-policy.yaml
  - MonitoringPolicy is now integrated into cluster.yaml

- Update CLAUDE.md with modular deployment guide
  - Add standalone Prerequisites deployment instructions
  - Document environment variable usage for rapid iteration
  - Add Enroot/Pyxis testing procedures
  - Include verification and troubleshooting steps

- Backup cluster-with-compute-nodes.yaml for reference

* Fix shell variable escaping in UserData for Enroot/Pyxis installation

- Change all shell variables from $ to 94274 for proper CloudFormation !Sub escaping
- Fixes 404 error when downloading Enroot packages (variables were not expanded)
- Affects: ENROOT_RELEASE, PYXIS_RELEASE, SLURM_VERSION, arch, distribution, etc.
- Now curl downloads correct URLs like v3.5.0 instead of v$ENROOT_RELEASE

* Externalize Enroot/Pyxis installation to separate script

- Create install-enroot-pyxis.sh script with proper variable handling
- Upload script to S3 (midaisuk-llm-dev/scripts/)
- Simplify UserData to download and execute external script
- Fixes CloudFormation !Sub variable escaping issues
- Benefits: maintainability, testability, no variable escaping problems

* Fix Enroot/Pyxis installation script URL to use GitHub raw URL

- Change S3 URL to GitHub raw URL in add-cng.yaml
- Add security restrictions to CLAUDE.md (no S3 public access)
- Verified installation works on login node i-0738c0e1f93dd452c
- Enroot 3.5.0 and Pyxis v0.20.0 successfully installed

* Generalize Enroot/Pyxis UserData into a post-install script hook

Replace the InstallEnrootPyxis boolean with a generic PostInstallScriptUrl /
PostInstallScriptArgs pair (PCS equivalent of ParallelCluster OnNodeConfigured)
across add-cng.yaml, add-cng-p5.yaml and pcs-ml-cluster-deploy-all.yaml.

- add-cng-p5.yaml: drop the duplicated inline Enroot/Pyxis install block; both
  CNG templates now fetch and run the same external script via the hook.
- deploy-all defaults reflect the most common setup: BuildAMI=false,
  DeployMonitoring=true, PostInstallScriptUrl set to the upstream installer URL.
- DeployMonitoring keeps its platform-specific logic unchanged.
- Document runtime dependencies (Slurm headers, OS family, egress, apt lock,
  GPU toolkit) in the script header and README; post-install log is now
  /var/log/pcs-post-install.log.
- Remove stale cluster-with-compute-nodes.yaml.backup and refresh architecture diagram.

* Add monitoring stack test cluster for Ubuntu 24.04 verification

Minimal single-login-node PCS template (pcluster-monitoring-test-cluster.yaml)
plus a deploy helper and notes, used to verify the aws-parallelcluster-monitoring
stack installs cleanly on the PCS Ubuntu 24.04 DLAMI before the upstream fixes
land. Points at the DaisukeMiyamoto fork branch for testing.

* Pin aws-parallelcluster-monitoring to release tag v2.6.2

Replace the 'latest'/'main' monitoring references with a pinned release tag for
upstream stability. MonitoringVersion now defaults to v2.6.2 and is used both to
fetch post-install.sh from the matching tag and as the version argument passed
to it, across add-cng.yaml, add-cng-p5.yaml and pcs-ml-cluster-deploy-all.yaml.

Note: v2.6.2 still lacks the Ubuntu fixes (upstream PR #44), so the ec2-user and
local-login_id UserData workarounds remain until a release with those fixes can
be pinned.

* Update README to pin MonitoringVersion to v2.6.2

Reflect the version-pinning policy in the monitoring section: default tag
v2.6.2, explain it is used for both the post-install.sh fetch and the version
argument, and note why latest/main are avoided.

* Fix Enroot/Pyxis install on GPU nodes: gpg no-tty and stable nvidia repo path

* Add P5 GPU training validation tests and 300 GiB root volume option

- Add tests/gpu-training-validation/ (nvidia-smi, 2-node NCCL all_reduce both
  AMI-native and via the prebuilt public.ecr.aws/hpc-cloud/nccl-tests image, and
  a Megatron-LM Llama-2 smoke test), plus RESULTS.md documenting the run.
- Add RootVolumeSize parameter (default 300 GiB) + BlockDeviceMappings to
  add-cng.yaml and add-cng-p5.yaml; the ~75 GiB DLAMI default overflows when
  pulling/importing large container images.

* Bump monitoring to v2.6.3 and drop Ubuntu ec2-user workarounds

aws-parallelcluster-monitoring v2.6.3 (PR #44) natively detects the Ubuntu
'ubuntu' user, so the templates no longer need the ec2-user shim or the
'local login_id' sed patch. Pin MonitoringVersion default to v2.6.3 across
add-cng.yaml, add-cng-p5.yaml and pcs-ml-cluster-deploy-all.yaml, simplify the
UserData monitoring block, and update the README and test docs accordingly.

* Tag nodes Name=HeadNode/Compute for monitoring dashboards

The monitoring Grafana dashboards filter on the EC2 Name tag relabeled into the
instance_name label, expecting the literals 'HeadNode' (login/head) and
'Compute' (compute). Derive Name from whether the node group has a Slurm queue
(CreateQueue), and preserve the per-node-group name on a new CngName tag.

* Make RootVolumeSize consistent across CNG templates and deploy-all

- add-cng.yaml: add MinValue: 75 to match add-cng-p5.yaml.
- pcs-ml-cluster-deploy-all.yaml: add a RootVolumeSize parameter (default 300
  GiB) and pass it down to the login, on-demand, and P-series nested stacks so
  the whole-cluster deploy can control root volume size in one place.

* Remove duplicate MinValue in add-cng.yaml RootVolumeSize

* PCS monitoring: free the Name tag, identify node type via monitoring-role, add MonitoringRepo

- Tag login and compute nodes with monitoring-role (login|compute) so the
  monitoring stack identifies node type by that tag instead of Name. Name now
  defaults to PCS-<cngname> and is free for arbitrary use. Replaces the
  IsMonitoringLogin condition with IsMonitoringRoleSet so the tag is emitted for
  both roles; deploy-all passes MonitoringRole=compute to the compute CNGs.
- Add a MonitoringRepo parameter (default aws-samples/aws-parallelcluster-monitoring)
  and pass both the ref and repo to post-install.sh, so a fork + branch can be
  used to test unreleased monitoring changes. MonitoringVersion may now be a branch.
- README: document MonitoringRepo, the monitoring-role node-type convention, and
  fork/branch testing.

* Add P6-B200 and P6-B300 GPU compute node group templates

Add per-family multi-NIC EFA templates for the Blackwell P6 instances, since
their network-card counts differ from P5 (P6-B200 has 8 network cards, P6-B300
has 17, vs 16/32 for P5). Each template hardcodes the right number of EFA
interfaces using the proven P5 pattern (primary NCI0 on DeviceIndex 0, the rest
on DeviceIndex 1, all InterfaceType: efa), and drops the NetworkInterfaceCount
parameter / Use32Interfaces condition since the count is fixed per family. All
other parameters and resources (PCS CNG/Queue, UserData, RootVolumeSize,
monitoring, Capacity Block) are identical to add-cng-p5.yaml.

Not yet validated on real p6 hardware and not yet wired into the deploy-all
orchestration template.

* deploy-all: select P-series NIC template by instance type

PseriesInstanceType now accepts p6-b200.48xlarge and p6-b300.48xlarge in
addition to the p5 family. Conditions (PseriesIsP5/PseriesIsP6B200/
PseriesIsP6B300) derive the family from the instance type and create the
matching nested stack: add-cng-p5.yaml (NetworkInterfaceCount), add-cng-p6-b200
.yaml (8 NICs), or add-cng-p6-b300.yaml (17 NICs). NetworkInterfaceCount is
passed only to the P5 stack; the P6 stacks omit it since their NIC count is
fixed. Note NetworkInterfaceCount as P5-only in its description.

* P6 templates: use EFA 'Use case 2' NIC layout (ENA primary + efa-only/ENA pairs)

Align the P6-B200/P6-B300 network interface config with the EC2 EFA docs
'Use case 2' (maximum EFA + ENA bandwidth), matching the ParallelCluster B300
LaunchTemplateOverrides tutorial:
- Primary NCI 0, DeviceIndex 0: ENA (the primary cannot be EFA-only).
- NCI 1..N-1: an efa-only interface on DeviceIndex 0 and an ENA interface on
  DeviceIndex 1.
This replaces the earlier P5-style layout (all interfaces InterfaceType: efa,
primary on a single card) which is not the AWS-recommended P6 configuration.
P6-B200 = 8 network cards (15 interfaces); P6-B300 = 17 network cards (33).

* P6 templates: use InterfaceType efa (not efa-only) for PCS compatibility

The earlier 'Use case 2' layout (efa-only on dev0 + ENA on dev1) failed to launch
through PCS: PCS does not propagate efa-only network interfaces from the launch
template to the RunInstances call, and instead requested EFA on network card 0
(EFA limit 0 on P6), causing Client.AttachmentLimitExceeded.

Switch to the layout that actually launches via PCS, validated on real
p6-b300.48xlarge hardware (Capacity Block, us-west-2):
- Network card 0, device index 0: plain ENA (P6 NIC 0 does not support EFA).
- Network cards 1..N-1, device index 0: InterfaceType: efa.
P6-B200 = 8 network cards, P6-B300 = 17. The bare-EC2 RunInstances accepts both
efa-only and efa layouts, but PCS only handles 'efa', so we use 'efa'.

* CNG templates: track latest launch template version

CustomLaunchTemplate.Version was hardcoded to 1, so when a template change
produced a new launch template version (e.g. NIC or UserData edits), the compute
node group kept using version 1 and silently ignored the update — observed on
p6-b300 where a fixed NIC layout (v2) never reached the node group. Use
!GetAtt PCSLaunchTemplate.LatestVersionNumber in add-cng.yaml, add-cng-p5.yaml,
add-cng-p6-b200.yaml, and add-cng-p6-b300.yaml so node groups follow template
updates. Verified on real p6-b300: the node group moved to v2 and launched.

* add-cng-p6-b200: use EFA on all 8 network cards

Unlike p6-b300 (network card 0 cannot do EFA; MaxEfaInterfaces 16 of 17),
p6-b200 supports EFA on all 8 network cards (MaxEfaInterfaces=8). Configure every
card with InterfaceType: efa (primary NCI0 on DeviceIndex 0, NCI1-7 on
DeviceIndex 1) to maximize EFA bandwidth, matching the proven p5 all-efa layout.
Not yet validated on real p6-b200 hardware (Capacity Block opens 2026-06-03).

* P6 templates: set family-specific CngName defaults

CngName defaulted to 'gpu-p5' (leftover from copying add-cng-p5.yaml). Set
'gpu-p6-b200' and 'gpu-p6-b300' so standalone deploys get a sensible node group
name. (deploy-all overrides this via PseriesCngName, so it only affects direct
use of these templates.)

* PCS monitoring: retry post-install on transient failure; docs to /opt path

The monitoring tree now installs on node-local /opt (fixed in the
aws-parallelcluster-monitoring fork) instead of the shared NFS /home, which was
racing across nodes at first boot (Stale file handle). As a belt-and-suspenders
safety net, the UserData monitoring step in add-cng.yaml / add-cng-p6-b200.yaml /
add-cng-p6-b300.yaml now retries post-install.sh up to 3x (the installer is
idempotent) so a one-off transient first-boot failure self-heals.

Also update test docs / sample template references from
/home/ubuntu/aws-parallelcluster-monitoring to /opt/aws-parallelcluster-monitoring.

* PCS docs/params: human-friendly 1-click, FSx Region deployment types, README restructure

- deploy-all: add ParameterLabels (friendly console names) + group RootVolumeSize;
  add LustreDeploymentType (PERSISTENT_1/2) and OpenZFSDeploymentType (SINGLE_AZ*)
  so Region-dependent FSx tiers are selectable (defaults unchanged); relax
  PerUnitStorageThroughput to a free-form number.
- ml-cluster-prerequisites: wire the two FSx deployment-type params; emit Lustre
  MetadataConfiguration only for PERSISTENT_2.
- cluster.yaml: remove 5 unused params (VpcId/PublicSubnetId/FSx*); only the
  private subnet + security group are needed by the PCS cluster.
- README: update Key Features/Architecture to the current default (PCS-ready DLAMI,
  no AMI build, built-in monitoring, Enroot/Pyxis first-boot-vs-AMI choice, P5/P6
  multi-NIC auto-select); restructure (consolidate template selection into one
  reference table, trim examples, slim Monitoring); add GPU Node List screenshot.

* PCS GPU CNG: auto-derive EFA NIC count from instance type; clarify CapacityReservationId is Capacity-Block-only

- add-cng-p5: drop the NetworkInterfaceCount parameter; derive the EFA interface
  count from InstanceType (p5en.48xlarge = 16, p5/p5e = 32) via the Use32Interfaces
  condition; constrain InstanceType to the P5 family. deploy-all: remove the
  NetworkInterfaceCount parameter/label/group/passthrough; PseriesInstanceType now
  drives both the template choice and the NIC count.
- CapacityReservationId (all P-series templates + deploy-all): document that SETTING
  it forces MarketType=capacity-block (Capacity Blocks for ML), and that an ODCR ID
  must NOT go here — an open ODCR is consumed automatically by On-Demand launches,
  so leave it empty for On-Demand/ODCR. Relabel and update README capacity guidance.

* deploy-all params: generic PostInstallScriptUrl override note; move MonitoringRepo/Version to developer group

- PostInstallScriptUrl: drop the fork-specific (DaisukeMiyamoto) override example;
  keep the generic 'override with another HTTP(S) URL' note.
- Move MonitoringVersion/MonitoringRepo out of '2. PCS Cluster Configuration' into
  the final developer/advanced group (alongside the nested-template S3 settings),
  since they are developer-facing source overrides.

* README: lead with one-click ML-training-ready; align params to console groups; clarify GPU/EFA, monitoring

- Key Features: lead with 'one click to an ML-training-ready cluster' (the main
  value); demote no-AMI-build/PCS-ready-DLAMI to a supporting note.
- Quick Start + all examples: set AZ_ID in the first line so the one required choice
  (Availability Zone) is explicit.
- Configuration: restructure the parameter tables into the 7 console parameter groups
  (matches deploy-all's AWS::CloudFormation::Interface).
- GPU compute: unify NIC/network-card wording to 'EFA interfaces'; drop the
  misleading p5en-only NVSwitch note (all P5/P6 use NVLink/NVSwitch intra-node).
- Monitoring: reference the GPU dashboard screenshot from the section intro; add a
  pointer to the AWS-managed Prometheus/Grafana option (4.prometheus-grafana).

* README: add a quick NCCL multi-node GPU test (enroot import + Pyxis sbatch)

Adds a 'Running a multi-node GPU job (NCCL test)' section: import nccl-tests to /fsx
via Enroot, submit a 2-node 16-GPU all_reduce_perf via Pyxis, and interpret the result
(EFA provider line, '# Out of bounds values : 0 OK', busbw). Points to the FSDP test
case for full training.

* Bump default MonitoringVersion to v2.6.5 (PRs #48/#49 merged upstream)

aws-parallelcluster-monitoring v2.6.4 (PR #48, /opt install — fixes the shared-/home
Stale-file-handle race) and v2.6.5 (PR #49, dcgm-exporter Docker-29.x tag + throttle
tiles) are now released upstream. Bump the templates' default MonitoringVersion from
v2.6.3 -> v2.6.5 so the default deploy uses the fixed upstream release directly — no
fork/branch override needed. MonitoringRepo already defaults to aws-samples upstream.

* README: reference monitoring v2.6.5 (PCS /opt + DCGM Docker-29.x fixes now released upstream)

* docs: split full parameter reference into PARAMETERS.md; slim README Configuration

Move the 7-group parameter detail tables to a dedicated PARAMETERS.md and leave the
README Configuration section as a short overview + most-used-parameters table + link.
The concept guides (container runtime, FSx Region availability, GPU compute) stay in
the README since they inform choices. README ~460 -> ~412 lines.

* README: run NCCL enroot import on the login node; drop PR references from Monitoring

- NCCL test: import the image directly on the login node (it has Enroot and writes to
  shared /fsx) instead of via srun on a GPU node.
- Monitoring: remove the upstream PR link; describe the v2.6.5 fixes without referencing
  PR numbers.

* README: move Templates after Cleanup; order Configuration as container/GPU/storage; number sections

- Reorder Configuration subsections to Container runtime -> GPU compute -> Storage
  (user-facing choices first).
- Move the Templates reference section to after Cleanup (reference material near the end).
- Number all top-level sections 1-13; fix the affected internal anchor links.

* deploy-all: optional Grafana public access via GrafanaPublicAccessCidr (login-only SG)

Add GrafanaPublicAccessCidr (default empty = SSM-only, unchanged behavior). When set
to a CIDR, deploy-all creates a login-only security group opening HTTPS/443 to that
CIDR and attaches it to the login node via a new add-cng.yaml ExtraSecurityGroupId
parameter — so Grafana is reachable at https://<login-public-ip>/grafana/ without SSM
port forwarding. Crucially the SG is attached ONLY to the login node, not the shared
cluster SG, so compute nodes and FSx are not exposed. Validated; AllowedPattern guards
the CIDR; README + PARAMETERS.md updated.

* README: document Grafana access as two options (SSM port-forward vs GrafanaPublicAccessCidr)

Restructure 'Accessing Grafana' into Option A (SSM port forwarding, default/private) and
Option B (direct public access via GrafanaPublicAccessCidr — login-only SG on 443), with
the shared admin-password step up top and security notes (login-only exposure, tight
CIDR over 0.0.0.0/0, self-signed cert). Update the Key Features monitoring bullet.

* README: link Quick Start to the Grafana public-access option (GrafanaPublicAccessCidr)

* README: fold Cleanup into Quick Start (Console or CLI); merge User Management into Additional Resources; renumber

- Move cleanup to the end of Quick Start; note deletion via CloudFormation Management
  Console (select stack -> Delete) or CLI.
- Merge the User Management LDAP link into Additional Resources.
- Renumber sections (now 1-11).

* docs: move architecture-components.md and PARAMETERS.md into docs/; fix links

Create architectures/aws-pcs/docs/ and relocate the two reference docs there. Update
README links to docs/PARAMETERS.md and PARAMETERS.md's back-links to ../README.md.

* tests: drop obsolete monitoring test assets; refresh tests/README; bump monitoring-stack-test to v2.6.5

- Remove pcluster-monitoring-test-cluster.yaml + deploy-monitoring-test.sh +
  monitoring-ubuntu24-fixes.md (the Ubuntu/PCS fixes are resolved upstream in v2.6.x,
  and deploy-all covers the monitoring test path). Also removed the stale template from
  the test bucket.
- tests/README: drop the Ubuntu-workaround use case, add the GPU Training Validation
  entry, note Grafana public access.
- monitoring-stack-test.md: v2.6.3 -> v2.6.5 references.

* tests: consolidate into a single guide; flatten scripts; cover full matrix

- Merge monitoring-stack-test.md + gpu-training-validation/ into one tests/README.md
  with a coverage matrix (monitoring; Enroot/Pyxis via UserData AND ImageBuilder; CPU;
  G-series single-NIC GPU; P5/P6-b200/p6-b300 multi-NIC; NCCL; FSDP), each test stating
  what to run and the expected result inline.
- Flatten gpu-training-validation/ to tests/ root; replace the Megatron stage with a
  working FSDP smoke test (03-fsdp-llama2.sbatch).
- enroot import now runs directly on the login node (300 GiB root), not a batch job;
  drop the /dev/shm staging and the AMI-native NCCL variant.
- Fix the README link to the new tests guide.

* README: add OpenZFS (/home) throughput row to the FSx storage table

The Storage table listed Lustre PerUnitStorageThroughput but omitted the OpenZFS
HomeThroughput. Add it with the per-deployment-type valid values (SINGLE_AZ_2/HA_2 =
160..10240, SINGLE_AZ_HA_1 = 128..4096, SINGLE_AZ_1 = 64..4096), per the FSx docs.

* tests/README: add a Verified configurations section

Summarize the configs validated on real hardware (2x p6-b300 ~760 GB/s NCCL /
~195 TFLOPS FSDP; 2x p5 ~480 GB/s / ~60 TFLOPS; CPU+monitoring; Grafana public access),
the deploy path used, and the EFA-NIC-count / FSDP-loss notes.

* templates: fix inaccurate top-level Descriptions

- cluster.yaml: was described as a full HPC cluster with login+compute nodes, EFS, and a
  new VPC — none of which it does. Corrected to: PCS scheduler core only (no nodes),
  attaches to an existing subnet/SG, optional monitoring IAM + Slurm accounting.
- add-cng.yaml: expand the terse one-liner to describe single-NIC CNGs (login/CPU/
  single-NIC GPU), FSx mounts, Enroot/Pyxis + monitoring, queue-only-when-QueueName.
- ml-cluster-prerequisites.yaml: mention the FSx for OpenZFS (/home) filesystem it also
  creates (was Lustre-only), and that FSx deployment types/throughput are parameters.

* tests/README: add the validated 2x p6-b200 row (NCCL ~654 GB/s/8 nics, FSDP ~223 TFLOPS/GPU)

* install-enroot-pyxis.sh: wait for apt/dpkg lock + retry apt (fix cold-boot exit 100)

At first boot the script raced Ubuntu's unattended-upgrades for the dpkg lock; the very
first apt-get aborted with 'Could not get lock /var/lib/dpkg/lock-frontend ... held by
(unattended-upgr)' under set -e (observed as post-install exit 100, leaving the node
without Enroot/Pyxis). Add wait_for_apt_lock() (polls the dpkg/apt locks via fuser, up to
300s) and an apt_get() wrapper (DPkg::Lock::Timeout=300 + up to 5 retries); route all apt
calls through it. Verified: a manual rerun after the lock cleared installs cleanly.

* tests/README: flag B300 NCCL bandwidth as needs-larger-message/more-nodes follow-up

Mark the ~760 GB/s as busbw and add a footnote: B300 has 2x the EFA cards of B200 but
the 2-node/16 GiB busbw was only ~1.16x higher, so the 2-node test likely doesn't
saturate all 16 cards — re-test with larger message sizes (~64 GiB) and more nodes
before treating it as B300 peak network bandwidth.

* docs: add ROADMAP.md (future implementation TODO) + link from README

Add architectures/aws-pcs/docs/ROADMAP.md as a markdown checklist of implementation items
under consideration (P6-B300 at-scale NCCL re-test, targeted ODCR, Trn template, Multi-AZ
FSx, P6 in the AMI recipe, DCGM throttle-reason signal, test-matrix automation). GitHub
Issues are disabled on the fork, so an in-repo, version-controlled checklist is the SSOT.
Link it (and PARAMETERS.md) from the README Additional Resources.

* docs/ROADMAP: reset to the agreed open items (Multi-AZ prereq, managed monitoring, Trainium validation, test automation)

* add-cng templates: make AmiId optional (auto-resolve PCS-ready DLAMI from SSM when empty)

Change AmiId in all four add-cng templates from a required AWS::EC2::Image::Id to an
optional String (default ''): when empty, the node group's AmiId resolves
{{resolve:ssm:/aws/service/pcs/ami/dlami-base-ubuntu2404/x86_64/latest/ami-id}} via a
HasAmiId condition; a non-empty ami-xxxx still overrides (e.g. a custom Enroot/Pyxis AMI).
deploy-all is unaffected (it always passes a resolved/built AmiId). Users no longer need
to look up the AMI ID.

* docs/ROADMAP: add user-management backend (LDAP/AD) integration item

* docs/ROADMAP: add P6e-GB200/GB300 (Grace-Blackwell, arm64) support item

* README: warn against combining BuildAMI=true with the default PostInstallScriptUrl (double install)

PostInstallScriptUrl defaults to the Enroot/Pyxis installer and deploy-all passes it to
the CNGs regardless of BuildAMI. So BuildAMI=true + the default URL makes each node both
boot the pre-baked AMI AND re-run the installer at first boot. Document that BuildAMI=true
should be paired with PostInstallScriptUrl=""; reword 'two combinable options' to 'choose one'.

* tests/README: note BuildAMI=true must pair with PostInstallScriptUrl="" (avoid double install)

* tests: record BuildAMI=true deploy-all validation (no double-install)

* aws-pcs: fix PostInstallScriptUrl org (awslabs) and make enroot cache per-user

- PostInstallScriptUrl default + tests README clone URL pointed at aws-samples;
  this repo lives under awslabs, so the raw URLs 404 at first boot. (review #1)
- enroot ENROOT_CACHE_PATH was shared (/tmp/enroot) while data/runtime are
  per-user, so a second user importing an image sharing a cached layer hit an
  unreadable 0640 blob and the job died. Make the cache per-user too. (review #2)

* aws-pcs: validate FSx throughput against deployment type (review #3/#4)

FSx throughput is a discrete grid that depends on the deployment type, but the
two were uncoupled: the default PerUnitStorageThroughput=250 is invalid for the
PERSISTENT_1 fallback the parameter recommends, and HomeThroughput=320 is invalid
for the SINGLE_AZ_1 fallback the description steers users toward -- both surfaced
as an opaque CreateFileSystem failure deep in the nested prerequisites stack.

- Make both throughput params String + AllowedValues (full valid grid -> console
  dropdown, no typos), keeping the existing defaults (250 / 320) unchanged.
- Add a Rules section to both ml-cluster-prerequisites.yaml and
  pcs-ml-cluster-deploy-all.yaml asserting (DeploymentType, throughput) is a valid
  pair, so a mismatch fails at stack-create with a clear message.
- Cross-reference the coupling in both parameter descriptions.

* deploy-all: force PostInstallScriptUrl empty when BuildAMI=true (review #6)

BuildAMI=true bakes Enroot/Pyxis into the custom AMI, but PostInstallScriptUrl
was still passed (defaulted) to every CNG stack, so the installer ran a second
time at first boot. The 'set it empty when BuildAMI=true' rule lived only in the
PR body, not in the console parameter descriptions, and nothing enforced it.

Guard all five CNG passthroughs with !If [BuildAMIEnabled, '', !Ref ...] so the
template enforces it, and cross-reference the interaction in both the BuildAMI
and PostInstallScriptUrl parameter descriptions.

* install-enroot-pyxis.sh: make idempotent + fix slurmd PATH (review #6/#18)

PostInstallScriptUrl is a generic first-boot hook, so the template no longer
forces it empty under BuildAMI=true (reverted). Instead the installer itself is
now safe to re-run on an AMI where Enroot/Pyxis is already baked in:

- Skip the Enroot .deb install when 'enroot version' already equals the target
  (reinstall only on a version mismatch).
- Skip Pyxis per Slurm version when spank_pyxis.so + the plugstack symlink already
  exist, instead of rebuilding from source every boot.
- Derive the slurmd PATH from the slurm-*/bin dirs actually present (a hardcoded
  slurm-25.11 silently mismatched SlurmVersion=25.05 clusters), and guard the
  append with grep -qxF so a re-run does not add a duplicate PATH line.
- Document the execution context (runs before slurmd starts; does not restart it)
  and the idempotency contract in the script header.

* deploy-all: keep PostInstallScriptUrl as a generic hook, warn (not block) on BuildAMI

PostInstallScriptUrl is intended as a generic first-boot hook, not Enroot/Pyxis-
specific, so do NOT force it empty under BuildAMI=true (that would block other
custom first-boot setup). Revert the !If guard to a plain !Ref passthrough and
reword the BuildAMI / PostInstallScriptUrl descriptions to warn about the default
installer reinstalling on a pre-baked AMI. The installer is now idempotent
(separate commit), so a redundant default re-run is a fast no-op anyway.

* docs: GrafanaPublicAccessCidr exposes unauthenticated Prometheus/Pushgateway too (review #7/#8)

Opening 443 to a CIDR exposes more than the password-gated Grafana: the login
node's nginx also reverse-proxies the unauthenticated /prometheus/, /pushgateway/
and /slurmexporter/ paths, so anyone in the CIDR can read all cluster metrics
(and push) without credentials. Document this real exposure in the parameter
description, the README Option B security notes, and docs/PARAMETERS.md.

Keep 0.0.0.0/0 accepted (no hard block) — it's useful for short-lived PoC /
workshop access where per-user local SSM permissions are impractical — but warn
to narrow or clear it when done. (#7 documented here; the upstream fix to gate
those nginx locations is tracked separately for aws-parallelcluster-monitoring.)

* docs: note p6-b300 GPU metrics gap (DCGM 4.2.0 pin); track upstream (review #9)

Monitoring v2.6.5 pins dcgm-exporter to DCGM 4.2.0, which doesn't support B300
(needs >=4.4.0); the pin is forced by the Docker 29.x OCI-index pull issue (#47).
So p6-b300 GPU dashboards stay empty while every other GPU type is fine.

Rather than patch the monitoring internals from the PCS template layer (fragile,
and the fix belongs upstream), filed aws-parallelcluster-monitoring#50 proposing a
configurable dcgm-exporter image (env var, current pin as default) so a B300-capable
build can be supplied by digest without a fork. Document the known limitation in the
README monitoring section and the tests verified-configurations matrix (the B300 row
no longer overclaims '16 GPUs in Grafana') with a link to the upstream issue.

* GPU add-cng: lock InstanceType with AllowedValues + document the per-family split (review #5/#10)

Decided to keep one add-cng template per GPU family rather than consolidate: the
templates are ~85% identical but their NetworkInterfaces EFA layouts genuinely
differ (p5/p6-b200 put EFA on card 0 with DeviceIndex 1; p6-b300 has an ENA-only
card 0 with EFA on cards 1-16 at DeviceIndex 0, 17 cards total). A single template
would need ~32 per-card !If wrappers, or an Fn::ForEach transform whose required
CAPABILITY_AUTO_EXPAND breaks the one-click quick-create links -- and would mean
editing the already-validated p5/b200 blocks.

- #5: add AllowedValues to InstanceType on the p6-b200 and p6-b300 templates (p5
  already had it), so a type whose NIC layout doesn't match the template's block
  fails at stack validation instead of with an opaque launch error.
- #10: document the split rationale in a header comment on all three templates and
  add an Fn::ForEach consolidation item to docs/ROADMAP.md as future work.

* tests: reuse canonical NCCL/FSDP assets instead of shipping copies (review #11-#14)

The tests/ dir duplicated assets the repo already maintains, which would drift and
violated the 'reuse existing assets' contribution rule. Drop the four local scripts
and point the guide at the canonical sources, documenting only the PCS-specific deltas:

- Remove 01-nvidia-smi.sbatch (#13) -> a one-line interactive srun in Test 5.
- Remove 02-nccl-tests.sbatch + 02-import-nccl-image.sh (#11/#14) -> Test 6 uses
  micro-benchmarks/nccl-tests/slurm/nccl-tests-container.sbatch; PCS deltas =
  import on the login node (Lustre can't host the enroot overlay) with a PINNED
  image tag (no more 'latest'), and the GPU queue partition name.
- Remove 03-fsdp-llama2.sbatch (#12) -> Test 7 uses 3.test_cases/pytorch/FSDP
  (create_venv.sh + slurm/llama2_7b-training.sbatch); PCS deltas = venv + HF_HOME
  on /fsx (not NFS /home), and --nodes=2.
- Replace the intro script table with a canonical-asset reference table.

Verified-configurations matrix and footnotes are kept (real-HW results).

* install-enroot-pyxis.sh: document that Slurm 24.11 is out of scope (review #17)

The reviewer noted the script only builds Pyxis for 25.05/25.11 while PCS also
supports 24.11. That scoping is deliberate and already consistent: cluster.yaml's
SlurmVersion AllowedValues are 25.05/25.11 only, so the templates can't create a
24.11 cluster. Rather than add an unvalidated 24.11 build, document in the script
header that 24.11 is intentionally out of scope and to keep SLURM_VERSIONS in sync
with cluster.yaml's AllowedValues.

* docs: AMI-pinning tip for production + align double-install notes with idempotent installer (review #16)

#16: AmiId defaults to the SSM /latest/ PCS-ready DLAMI path (no version-pinned
parameter exists) and CloudFormation re-resolves it on stack updates, so scale-out
nodes can drift to a newer AMI. Add a production tip to README: resolve it once and
pass the result as AmiId so every node is identical.

Also align the BuildAMI + PostInstallScriptUrl notes (README + tests/README) with the
#6 resolution: the installer is now idempotent and PostInstallScriptUrl stays a generic
hook, so leaving the default under BuildAMI=true is a fast no-op (not a conflict);
PostInstallScriptUrl="" is recommended only for the cleanest boot.

* deploy-all + CNG: add DcgmExporterImage param to override dcgm-exporter (review #9, monitoring #50)

Wire an optional DcgmExporterImage through deploy-all -> every CNG stack -> the
monitoring UserData, which exports DCGM_EXPORTER_IMAGE before running post-install.sh.
Empty (default) keeps the monitoring stack's own default dcgm pin (DCGM 4.2.0, up to
B200), so existing behaviour is unchanged. For p6-b300 (needs DCGM >= 4.4.0) set a
newer build, ideally by digest to bypass the Docker 29.x OCI-index pull failure.

Pairs with aws-parallelcluster-monitoring #50 (DCGM_EXPORTER_IMAGE support); lets a
one-click deploy enable B300 GPU dashboards without a monitoring fork.

* tests/README + README: reflect B300 validation findings (review #9 follow-up)

From the live 2x p6-b300 run (Slurm 25.11, Docker 29.5.3, DCGM 4.5.2 by digest):

- B300 GPU metrics DO populate once DcgmExporterImage is set to a B300-capable build
  by digest (bypasses the Docker-29.x OCI-index pull failure). Replaces the earlier
  'GPU metrics empty' note; verified all 16 B300 GPUs in Grafana. Needs monitoring
  DCGM_EXPORTER_IMAGE support (aws-parallelcluster-monitoring#51, for #50).
- NCCL all_reduce scales past the canonical 16 GiB sweep: ~654 GB/s @16 GiB but
  ~751 GB/s @64 GiB on 2x p6-b300 (16 nics) -- so 16 GiB was unsaturated, answering the
  review's 'larger message size' ask. 128/256 GiB OOM (buffer > B300 memory); scale nodes
  not buffer for a higher peak.
- FSDP validated two ways on p6-b300: shared-/fsx venv (~200 TFLOPS/GPU) and Enroot/Pyxis
  container via CONTAINER_IMAGE (~193 TFLOPS/GPU). Test 7 now documents both, plus the venv
  PATH --export needed for torchrun to resolve on every node.
- Updated the verified-configurations matrix and the README NCCL/monitoring notes to match.

* install-enroot-pyxis.sh: pin slurmd PATH to the cluster's Slurm version (fix Pyxis break on multi-version DLAMI)

Regression found on a live p6-b300 cluster: every --container-image job died with
'spank_pyxis.so: task_init() failed'. Root cause: the slurmd PATH line put every
/opt/aws/pcs/scheduler/slurm-*/bin on PATH, and glob order puts the OLDEST (24.11)
first. The Pyxis/Enroot PMI hook runs 'scontrol show config'; the 24.11 scontrol
fatally fails to parse the 25.11 cluster's slurm.conf ('unrecognized key: MetricsType'
-- set by cluster.yaml for the monitoring OpenMetrics endpoint), aborting the hook.

Fix: resolve the cluster's actual Slurm version from the live controller (slurmd has
--conf-server) via a probe 'scontrol show config', then put only that version's bin on
PATH (matched on major.minor). Falls back to the newest installed version if the
controller isn't reachable (e.g. during an AMI build). Idempotent: strips any prior
PATH= line before appending.

Verified on the live p6-b300 cluster (Slurm 25.11): picks slurm-25.11/bin, and NCCL +
FSDP --container-image jobs run cleanly. This supersedes the earlier glob approach that
itself regressed the previously-working hardcoded single-version PATH.

* Default MonitoringVersion to v2.9.1 (DCGM_EXPORTER_IMAGE support for B300)

aws-parallelcluster-monitoring v2.9.1 shipped the DCGM_EXPORTER_IMAGE override
(PR #51, for #50) that lets DcgmExporterImage enable B300 GPU metrics without a
fork, and carries the PCS /opt install + Docker-29.x DCGM fixes from v2.6.4+.
Bump the MonitoringVersion default v2.6.5 -> v2.9.1 across deploy-all + the four
add-cng templates, and update the README / PARAMETERS / tests prose to match.
tests/README also records the v2.9.1 B300 validation (GPU metrics via digest,
NCCL 64 GiB ~751 GB/s, FSDP venv ~205 TFLOPS / container ~193 TFLOPS).

* install-enroot-pyxis.sh: derive slurmd PATH from the slurmd unit, not scontrol (boot-safe)

The previous version detected the cluster Slurm version by running scontrol show
config -- but this script runs at first boot BEFORE slurmd starts, when the config
source isn't established yet, so scontrol prints nothing and (under set -eo pipefail)
the failing command substitution aborted the whole post-install with exit 1. The
PATH line was never written (slurmd still ran via its absolute ExecStart path, but
the post-install reported failure and the Pyxis PMI hook had no usable scontrol on
container nodes).

Derive the version from the slurmd systemd unit's ExecStart instead
(/opt/aws/pcs/scheduler/slurm-<ver>/sbin/slurmd), which PCS sets to the chosen
version and is readable at first boot. Pins PATH to exactly that version's bin;
falls back to the newest installed version if the unit can't be read (AMI build).
Verified on a live p6-b300 (Slurm 25.11): extracts slurm-25.11/bin correctly.

* cluster.yaml: gate MetricsType on Slurm 25.11+ (fix 25.05 cluster create failure)

PCS rejects cluster creation when SlurmCustomSettings includes MetricsType on
Slurm < 25.11 ('MetricsType is only supported for Slurm version 25.11 or later'),
but the template emitted MetricsType=metrics/openmetrics + CommunicationParameters=
enable_http whenever DeployMonitoring=true, regardless of version. So a 25.05 cluster
with monitoring (a valid AllowedValues combo) always failed to create.

Gate those two settings on a new IncludeMetricsConfig condition (monitoring AND
SlurmVersion=25.11). Also redefine IncludeSlurmConfig so monitoring-on-25.05 doesn't
force an otherwise-empty SlurmCustomSettings list (which PCS also rejects): emit
SlurmConfiguration only when accounting, an accounting policy, or the metrics config
actually applies. The Slurm OpenMetrics dashboards simply won't populate on 25.05;
all other monitoring (node/GPU/CloudWatch) is unaffected.

* install-enroot-pyxis.sh: resolve slurmd PATH from PCS profile.d/unit (robust, boot-safe)

slurmd needs the cluster's Slurm bin on its PATH for the Pyxis PMI hook
(50-slurm-pmi.sh runs 'scontrol show config' in the job env; PCS's
/etc/profile.d/slurm.sh only affects login shells, and slurmd's own PATH omits
scontrol -- so without this, --container-image jobs fail with
'Command not found: scontrol').

Resolve the cluster's version from the first source present at post-install time:
(1) the /etc/profile.d/slurm.sh symlink PCS points at the chosen version, then
(2) the slurmd systemd unit ExecStart, then (3) newest installed version. Each step
is guarded with an if + '|| true' so a missing source can't trip set -e (the prior
'[ -n VER ] && ...' form returned non-zero and aborted post-install with exit 1 when
the version source wasn't ready yet at first boot). Pins exactly one version's bin,
so an older scontrol never shadows and breaks on a newer slurm.conf key (MetricsType).
Verified resolution under set -e on a live 25.11 node.

* Pass cluster Slurm version to post-install via PCS_SLURM_VERSION env (fix Pyxis PATH)

The node can't discover the cluster's Slurm version at post-install time (cloud-init,
before slurmd/profile.d/controller exist), and the PCS DLAMI ships multiple versions
where an older scontrol breaks the Pyxis PMI hook on a newer slurm.conf. Pass the
version in explicitly instead of guessing:

- install-enroot-pyxis.sh: read PCS_SLURM_VERSION and pin slurmd's PATH to that
  version's bin; fall back to newest installed if unset (manual run). This is also the
  natural way to get the version once PCS ships a native post-install hook.
- add-cng*.yaml (x4): add a SlurmVersion parameter (Default 25.11, AllowedValues
  25.05/25.11 -- so direct add-cng use or an older deploy-all that omits it still works)
  and export PCS_SLURM_VERSION before running the post-install script.
- deploy-all: pass SlurmVersion through to every CNG stack.
- pcs-ready-dlami (AMI build): the AMI is cluster-agnostic, so list all slurm-*/bin
  newest-first on slurmd's PATH (no older-version shadowing) instead of hardcoding
  25.11; a PostInstallScriptUrl boot then rewrites it to the exact cluster version.

* install-enroot-pyxis.sh: build Pyxis only for the cluster's Slurm version + per-version plugin path

Two linked fixes for the SPANK plugin version mismatch that crashed slurmd on a 25.05
cluster ('Incompatible Slurm plugin version (25.11.2)'):

1. The Pyxis Makefile installs spank_pyxis.so to $libdir/slurm and bakes that absolute
   path into the generated pyxis.conf. The old code used the single shared
   /usr/local/lib for every version, so building 25.05 then 25.11 left only the 25.11
   .so, and the 25.05 plugstack pointed slurmd at it -> slurmd refused to start. Install
   each version to a per-version libdir (/opt/aws/pcs/scheduler/slurm-<ver>/lib) and
   write that exact path into the version's plugstack pyxis.conf.

2. Build ONLY the cluster's version when PCS_SLURM_VERSION is provided (the common
   path), instead of all of them -- no wasted build, no chance of a wrong-version .so.
   Falls back to all supported versions when unset (manual run / cluster-agnostic AMI).

Idempotency check now keys on the per-version .so, not the shared path.

* tests/README: add pre-merge full test matrix + Enroot/Pyxis regression-test rule

Several real bugs this round only appeared in specific combinations (25.05-only
MetricsType rejection, a Pyxis SPANK plugin built for the wrong Slurm version stopping
slurmd, a first-boot post-install failure). Document the bar for merging:

- A 'Pre-merge full test matrix' up top: every Slurm version (25.05 AND 25.11), both
  container-runtime paths (first-boot install AND BuildAMI=true), clean first boot,
  CPU + each GPU family, monitoring, and template lint.
- A regression-test rule in Test 2: ANY change to install-enroot-pyxis.sh must be
  retested on all supported Slurm versions and via a BuildAMI=true run (the AMI path
  carries its own copy of the logic), on a clean first boot. Also corrected the Pyxis
  plugin path to the per-version location.

* AMI build: single-version Pyxis (add SlurmVersion param, plumb from deploy-all)

The pcs-ready-dlami AMI used to bake Pyxis for every supported Slurm version into a
shared /usr/local/lib/slurm path. Each make install clobbered the previous version's
spank_pyxis.so, leaving only the last version (25.11) -- so an AMI used by a 25.05
cluster crashed slurmd at start with 'Incompatible Slurm plugin version'.

Mirror the post-install fix: build Pyxis only for the cluster's Slurm version, passed
in as a new SlurmVersion CFN parameter (Default 25.11, AllowedValues 25.05/25.11).
The slurmd PATH is then a single, unambiguous bin dir for that version -- no
multi-version glob, no version-detection, no .so collision.

- pcs-ready-dlami-with-enroot-pyxis.yaml: add SlurmVersion parameter; switch the
  EnrootPyxisComponent Data to !Sub so ${SlurmVersion} is injected; bash vars in the
  block are escaped to ${!var}; SLURM_VERSIONS pinned to the parameter; slurmd PATH
  set directly to /opt/aws/pcs/scheduler/slurm-${SlurmVersion}/bin.
- pcs-ml-cluster-deploy-all.yaml: pass SlurmVersion through to DLAMIStack so the AMI
  always matches the cluster's Slurm version.

Backward compatible: the AMI build template still has a SlurmVersion default (25.11),
matching the cluster.yaml default.

* docs: add OPERATIONS.md and slim README/tests caveats to one-line pointers

Move the recurring operational caveats (Slurm 25.05 vs 25.11 capability differences,
AMI single-version rule, MonitoringVersion migration notes, B300 DcgmExporterImage
setup, AMI pinning for production, FSx deployment-type/throughput coupling, P6-B300
NIC topology lock-in) into a dedicated docs/OPERATIONS.md so they live in one place
and stop expanding the README. Also document things this round's testing made
explicit:

- AMI build is single-Slurm-version by design (the SPANK plugin ABI is version-locked,
  so SlurmVersion on the DLAMI stack must match the cluster).
- The PostInstall path passes the cluster's Slurm version via the PCS_SLURM_VERSION env
  var (UserData -> install-enroot-pyxis.sh) -- the node can't discover it at first
  boot, before slurmd / profile.d / the controller config exist.
- v2.6.5 -> v2.9.1 brings Grafana 13 and DCGM_EXPORTER_IMAGE support.

README and tests/README keep the structural content; the long caveat blocks become
one-line pointers into OPERATIONS.md sections. tests/README adds measured stack
creation times for the four CPU patterns (PostInstall and BuildAMI x 25.05/25.11) and
the GPU+CB run, and the Pre-merge matrix's BuildAMI row now requires verifying the
SlurmVersion=cluster.SlurmVersion match.

* docs/PARAMETERS.md: add DcgmExporterImage; cross-reference OPERATIONS.md from SlurmVersion + BuildAMI + MonitoringVersion

DcgmExporterImage was missing from the parameter reference (it landed in the
deploy-all template earlier in this branch). Add it under Developer/Advanced with a
pointer to the validated B300 digest in OPERATIONS.md §3.1.

Tighten three existing rows that have non-trivial operational implications now in
OPERATIONS.md:
- SlurmVersion -> §1 (Slurm OpenMetrics is 25.11+ only; the value also drives Pyxis
  build version and the AMI it gets baked into).
- BuildAMI -> §2 (the AMI is single-Slurm-version; SlurmVersion on the AMI build
  must match the cluster's).
- MonitoringVersion -> §3 (v2.6.5 -> v2.9.1 migration notes, Grafana 13).

* docs/OPERATIONS.md: explain why single-version Pyxis is fine across cluster Slurm upgrades

Add §2.3 covering the upgrade-compatibility argument for the single-version pin: per
the Slurm upgrade policy, scontrol/srun from version N interoperate with a slurmctld
of N+1 (or N+2/N+3 on 24.11+), so a cluster upgrade does not break nodes built for the
prior version; advancing means setting the new SlurmVersion and redeploying so Pyxis
is rebuilt for it. Pairs with Reply E to KeitaW's #844 review thread.

* docs/OPERATIONS.md: record the cgroup-v2 prolog race as a known issue

Hit reproducibly during the 4-pattern CPU validation: the very first srun against a
freshly-launched cpu1 node fails with 'Job prolog failed' and drains the node, while
subsequent srun's land on sibling nodes and run cleanly.

Cause is below the templates: PCS bootstrap starts slurmd, then ~9s later systemd
forces a slurmd restart because cgroup-v2 cpuset.cpus/cpuset.mems setup fails ('No
space left on device' is the Linux misleading-error for an empty/invalid cpuset
value). The user srun arrives in that gap, the prolog handshake never reaches the
re-started slurmd, and the controller times out after 420s. Nothing this PR's scripts
modify (cgroup, slurmd unit, post-install timing) is in the path.

Recorded under §7 Known issues with the journal trace and the workaround
('scontrol update nodename=cpu1-N state=resume', or just resubmit and let it land on a
sibling). Recommendations recap renumbered to §8; existing OPERATIONS.md anchors
referenced from README/tests/PARAMETERS are unaffected (#1-#4).

* docs/ROADMAP.md: add Software stack section (Spack, Intel oneAPI, NVIDIA HPC SDK, modules)

The cluster currently ships only the PCS DLAMI's pre-installed CUDA/NCCL/EFA stack
plus Enroot/Pyxis for containers — no native HPC package manager and no first-class
hook for traditional toolchains (Intel oneAPI, NVIDIA HPC SDK). Track these as future
work so users running source-built MPI/BLAS/scientific apps or Fortran workloads have
a documented path; once Spack lands, an Lmod-style module system rounds it out.

* README: surface SlurmVersion in §4 + document the AMI build's pipeline rationale

Two pending edits:

- §4 'Most-used parameters' table didn't list SlurmVersion. It became a structural
  parameter this round (drives Slurm OpenMetrics availability, the Pyxis build
  version, and the AMI's baked Slurm) and deserves a row alongside BuildAMI etc.
  Cross-references OPERATIONS.md §1 for the full version trade-off.
- §4 Container runtime: spell out what the in-stack BuildAMI=true template earns
  beyond a one-shot Image Builder run (managed pipeline: scheduled rebuilds, AMI
  lifecycle deprecation, SSM parameter publishing of the latest AMI ID), so the
  rationale matches Reply F to KeitaW's follow-up review and the next reviewer
  doesn't re-litigate it.

* README: tighten Key Features (capacity scope, drop one-click duplication)

Two small clarity fixes:

- 'Flexible capacity' was ambiguous — could be read as 'flexible WRT capacity'
  rather than 'supports a wide range of capacity-purchase options'. Reword to
  'Broad capacity-purchase support', explicit about covering OD / ODCR / Capacity
  Blocks for ML, and selected per node group.
- 'One-click or modular' duplicated the first bullet's one-click claim. Drop the
  one-click half and lead with the value of the modular path: composing individual
  stacks for infrastructure reuse / iterating on one piece at a time.

* README §7: reuse canonical micro-benchmarks/nccl-tests sbatch instead of inline heredoc

The §7 'Running a multi-node GPU job' walk-through inlined a hand-rolled
nccl-test.sbatch heredoc, which duplicated the canonical launcher at
micro-benchmarks/nccl-tests/slurm/nccl-tests-container.sbatch and could drift from it.
Same reuse rule we already applied to tests/ (where 02-nccl-tests.sbatch was deleted
and Test 6 redirected to the canonical asset): document only the PCS-specific deltas
(login-node import + GPU partition name) and let the canonical sbatch carry the rest.
Also pins the NCCL image tag instead of using mutable latest. Pointer to the full
Test & Validation Guide added at the bottom.

* Default DcgmExporterImage to a DCGM 4.5.2 digest covering all supported GPUs

The empty default pushed B300 users into a forced manual override (and a Grafana that
quietly stays empty until they figure that out). DCGM 4.5.2 covers Hopper / B200 / B300
with no DCGM_FI_DEV_* field changes vs 4.2.0 per upstream changelog, was validated on
real B300 hardware in this round, and the digest pull bypasses the Docker-29.x OCI-index
failure on newer NVCR tags. So default to it across the GPU range — every supported
family populates Grafana out of the box. The monitoring stack's older 4.2.0 pin remains
a one-line override for anyone who needs to match another fleet.

Templates (deploy-all + 4 CNGs): Default '' -> the validated digest, parameter
description rewritten to lead with the new behavior. README §8 / §4, PARAMETERS,
tests/README, OPERATIONS.md §3.1 (renamed: 'B300 needs DcgmExporterImage' ->
'DcgmExporterImage — the default, and when to change it') all updated; cross-references
re-pointed at the new anchor.

* Strip personal-fork references before merge

Two leftovers from the dev / review cycle that shouldn't ship to awslabs/main:

- README §8 'Prefer AWS-managed Prometheus/Grafana?' linked to
  github.com/DaisukeMiyamoto/awsome-distributed-ai/tree/deploy-monitoring/...
  (a fork URL on the in-flight branch). The path lives on awslabs/main already, so
  switch to a repo-relative link that resolves correctly under any fork/branch and
  doesn't bake the dev fork name into mainline docs.
- The 4 add-cng*.yaml templates' MonitoringRepo description carried the example
  'e.g. DaisukeMiyamoto/aws-parallelcluster-monitoring' as a 'how to test unreleased
  changes' hint. Defaults are correct (aws-samples/...); drop the personal-fork
  example from the prose so there's no leftover advertisement of an internal fork.

* deploy-all: refresh DcgmExporterImage console label to match new default

The Parameters/AWS::CloudFormation::Interface label still said
'(empty = default; set for B300)' from the era when the param defaulted to ''
and B300 users had to set it explicitly. Now that the default is the DCGM 4.5.2
digest covering Hopper/B200/B300 out of the box, the old label misleads. Update
to '(default DCGM 4.5.2 by digest; covers H100/B200/B300)'.

* Fix broken ldap_server reference paths

The README's Additional Resources list and ROADMAP's User-management item both
pointed at `6.ldap_server` / `architectures/6.ldap_server` -- but the actual
upstream layout puts it under `1.architectures/6.ldap_server`. From
`architectures/aws-pcs/` that resolves through a `../../1.architectures/`
relative path. Update both refs so the link doesn't 404 on the rendered PR.

* deploy-all: improve 1-click parameter UX -- split AMI build, reorder, slim descriptions

Console parameter groupings reorganized for the 1-click case:

- New §3 'Container Runtime (Post-install Script)' carries just
  PostInstallScriptUrl/Args + RootVolumeSize -- the hook every cluster goes
  through. Previously these were buried alongside Image Builder params, which
  most users do not touch.
- AMI build params (BuildAMI/BaseAmiId/SemanticVersion/BuildSchedule) moved to
  a new §7 'Custom AMI Build (Optional - skip unless you need a pre-baked DLAMI)',
  ahead of the Developer group. They are an opt-in path; surfacing 'Optional'
  in the section title makes that obvious instead of confronting first-time
  users with SemanticVersion/BuildSchedule with no context.
- §2 PCS Cluster: SlurmVersion now precedes LoginNodeInstanceType. The Slurm
  major is the substantive cluster-shaping choice; the login type is rarely
  changed from the default.
- DcgmExporterImage was orphaned (not in any ParameterGroup, so the console
  rendered it under a generic 'Other' heading). Now grouped under §8 Developer/
  Advanced next to MonitoringVersion / MonitoringRepo where it belongs.

Description cleanups -- the long help blocks rendered as a wall of text in the
console; trimmed to the parameter-level essentials and pointed at OPERATIONS.md
for the rest:

- PostInstallScriptUrl: 12 lines -> 5 (kept the BuildAMI=true interaction note,
  dropped the duplicated 'PCS equivalent of OnNodeConfigured' phrasing).
- GrafanaPublicAccessCidr: 14 lines -> 5 (kept the unauthenticated-proxy-paths
  warning since that is the real footgun, dropped the per-network detail that
  belongs in OPERATIONS.md §3.2).
- BuildAMI: 4-line one-paragraph -> 4 short lines covering the trade-off, the
  PostInstallScriptUrl pairing, and the single-Slurm-version constraint.
- DcgmExporterImage: 6 lines -> 4 -- focused on 'why digest, not tag' since
  that's the non-obvious part. Migration / override examples stay in
  OPERATIONS.md §3.1.
- PerUnitStorageThroughput / HomeThroughput: tightened to call out that the
  valid grid is enforced by a Rule, since users see the validation error first.
- PrimarySubnetAZ: rewrote from 'AZ id where the subnets will be created' to
  'the only required choice' to match the 1-click framing.

ParameterLabels updated for the params whose section moved (AMI section now
flagged 'optional; auto-resolved if empty' on BaseAmiId so first-time users
know to leave it alone). CapacityReservationId label spelled out as 'Capacity
Block for ML reservation ID' so it reads correctly in the console even though
the parameter name itself is the generic 'CapacityReservationId'.

PARAMETERS.md ToC restructured to mirror the new 8 console groups; README §4
text 'all 7 console parameter groups' bumped to 8.

No default values changed. Validate-template still reports 42 parameters, no
orphans, all relative cross-links resolve.

* deploy-all: drop docs/ refs from descriptions; move RootVolumeSize to PCS Cluster

Two follow-ups on the 1-click parameter UX:

- Parameter Descriptions render as plain text in the …
DaisukeMiyamoto added a commit to DaisukeMiyamoto/awsome-distributed-ai that referenced this pull request Jul 8, 2026
…pc8a numbers, add C3 Expected

- infra-test.md: repoint the dashboard-access link to README awslabs#82-monitoring
  (the old awslabs#8-monitoring slug no longer exists)
- hpc-efa-test.md: scope the bandwidth-numbers pointer to hpc8a, matching
  what the referenced table actually records
- multi-user-test.md: add the inline Expected for C3 (sreport per-user
  utilization) that previously existed only in the removed verdict checklist
shimomut added a commit that referenced this pull request Jul 17, 2026
With WEBHOOK_LOG_PAYLOAD enabled, the log carried the full HMAC signature, the
signed timestamp, and the exact body — everything needed to replay the POST
verbatim for as long as the receiver accepts the timestamp (CodeQL alert #8
anchors here). The key itself isn't recoverable and the flag is opt-in debug, so
this is a contained exposure, but there's no reason the debug aid needs the full
signature.

Log signature[:12] instead — still enough to correlate a POST with the
receiver's view, but no longer replayable. Clears the last CodeQL clear-text-
logging alert. Regenerated hyperpod_devops_agent.yaml.

Addresses PR #1191 review feedback (KeitaW: #30 + CodeQL).
shimomut added a commit that referenced this pull request Jul 17, 2026
build-skill-uploader always uploaded to a fixed key lambda/skill_uploader.zip,
and the template references it via S3Bucket/S3Key with no S3ObjectVersion. Per
the AWS::Lambda::Function Code docs, CloudFormation does not auto-detect changes
to an S3 deployment package during stack updates — so after the first deploy,
edits to cr_skill_uploader.py (or a newer bundled boto3) were uploaded to S3 but
the Lambda kept running the old package: stale code in the exact custom resource
whose Delete path gates stack teardown.

Emit lambda/skill_uploader-<sha256[:16]>.zip. The zip is already deterministic
(fixed timestamps, sorted members), so the hash moves only when content moves;
a changed package produces a new key -> a template diff -> a function update.
The key flows through deploy.sh's SkillUploaderS3Key param unchanged.

Addresses PR #1191 review feedback (KeitaW: #8).
dmvevents added a commit to dmvevents/awsome-distributed-ai that referenced this pull request Aug 25, 2026
…rect-SHA fetches, benchmark provenance

- kubernetes/vllm-deepep-v2-2node.yaml:145,150 — worker probe bracket form `[v]llm serve` so the exec-shell's own cmdline no longer self-matches pgrep (thread awslabs#1/awslabs#9)
- kubernetes/vllm-deepep-v2-2node.yaml:97 — append `|| true` to the LEADER_IP substitution so the FATAL DNS guard can fire under set -e + pipefail (thread awslabs#2)
- kubernetes/vllm-deepep-v2-2node.yaml:159 — hugepages-2Mi comment: needs pre-allocated 2Mi hugepages or the pod sits Pending (thread awslabs#11)
- kubernetes/vllm-deepep-v2-2node.yaml:115 — DEEPEP_ARCH_LIST commented env now names 10.0 (b200) + 10.3 (b300) so the Blackwell knob is reachable from the manifest (thread awslabs#12)
- recipe/verify-image.sh:24,33,34 — convert fi_info efa-direct, ncclGetLsaDevicePointer, ncclGinPlugin checks from `grep -q` to draining `[ "$(... | grep -c X)" -ge 1 ]` (SIGPIPE-141 flake under pipefail), matching Dockerfile Layer 5b (thread awslabs#3)
- recipe/run-kernel-test.sh:16 — worker requires an explicit node-rank (`${3:?...}`) instead of defaulting to 0 and colliding with the leader; add the `case "$ROLE"` guard matching serve.sh (thread awslabs#4)
- recipe/benchmark_probe.py:18,24 — MODEL reads SERVE_MODEL env with the current value as fallback (+ `import os`) so a non-default model doesn't 100%-fail (thread awslabs#5)
- setup_deepep_v2_efa.sh:30,52 — fetch the immutable PR-head SHA directly instead of the moving `refs/pull/N/head` ref for both aws-ofi-nccl #1351 and DeepEP awslabs#612 (thread awslabs#6)
- Dockerfile:116 — CMD banner names /opt/serve.sh + /opt/build_deepep.sh (the in-image paths), not recipe/ (thread awslabs#7)
- recipe/build_deepep.sh:120 — `tee /tmp/deepep-build.log | tail -30` so a failed build keeps the full nvcc diagnostic (thread awslabs#8)
- Dockerfile:60 — gdrcopy pinned by commit SHA (v2.5.2 == c91ad9f) via fetch/checkout instead of the movable `--branch v2.5.2` tag (thread awslabs#13)
- recipe/serve.sh:124 — comment on why --trust-remote-code is unconditional (DeepSeek/Kimi models the preflight supports); chmod 755 serve.sh + run-kernel-test.sh in-tree to match the other four scripts (thread awslabs#14)
- README.md:114 — in-pod benchmark exec sets OUT_ROOT=/work/benchmarks so results land on the /work volume, not the ephemeral container layer (thread awslabs#16)
- setup/env_vars.example:2,12 — header says build-push.sh sources it and recipe/*.sh need `source` first; reword the AWS_OFI_NCCL_PR_SHA note (empty does NOT skip the cherry-pick) (thread awslabs#17)
- benchmarks/README.md:88 — caveat: no alternative-backend baseline measured; every table is --all2all-backend deepep_v2 (thread awslabs#18)
- recipe/serve.sh:48 — EP_EFA_MAX_QPS comment records the pinned plugin (9c44d34) is 76 commits past the seq-window redesign (6e504db) so the 128-slot cap's precondition no longer holds; default left unchanged (pods down) (thread awslabs#21)

Already present at c025946 (prior commits), verified not duplicated:
- README shared-experts caveat naming #47785 + DeepSeek (thread awslabs#15)
- benchmark_probe.py requests-per-level distribution + unique prompt prefix (threads awslabs#19/awslabs#20)
- probes + publishNotReadyAddresses structure (thread awslabs#9) and requests==limits Guaranteed QoS (thread awslabs#10)

Signed-off-by: Anton Alexander <dmvevents@gmail.com>
dmvevents added a commit to dmvevents/awsome-distributed-ai that referenced this pull request Aug 26, 2026
…2/GIN script

Round-3 doc-drift: the VLLM_SHA bump (e2f993dc4 -> 14617c2b, #52632's merge
commit) and the appearance of the canonical setup_deepep_gin.sh (2026-08-24)
left several docs describing the OLD state. Reconciled WITHOUT fabricating any
new measurements — the benchmark tables were measured on the old pin and are
now labeled historical, not re-claimed on the new pin.

- README setup-script rationale (awslabs#1): reworked to point at the canonical
  micro-benchmarks/.../deepep-v2-benchmark/setup_deepep_gin.sh and name the
  three deliberate divergences (unmerged aws-ofi-nccl #1351 param; CPU-proxy vs
  EFA-GDA; vLLM-wheel torch/NCCL ABI coupling). Kept the correct vendor-sync
  point (that CI gates only the NVSHMEM setup_deepep_efa.sh). Did NOT adopt the
  'delete the DeepEP half + call the canonical' substitution — that needs a
  docker build to verify and this change set is docs-only.
- DeepEP source divergence + Blackwell (awslabs#2): setup_deepep_v2_efa.sh header note
  + README Known-limitations now state the source is deepseek-ai/DeepEP@b306af06
  (not the amazon-contributing fork the canonical pins) and mark the manifest's
  DEEPEP_ARCH_LIST=10.x knobs documented-but-not-verified — round-1 saw no
  loadable Blackwell kernel at CUDA 13.0 on this lineage; the fork's st.bulk
  64-bit fix (amazon-contributing/DeepEP#3) is the enabling path, to re-verify.
- Provenance honesty (awslabs#6/awslabs#9/awslabs#10): benchmarks/README vLLM row + README build note
  now disclose the shipped image is 0.26.1rc1.dev1000+g14617c2b6 (four minor
  versions past the measured 0.22.1rc1.dev283+ge2f993dc4); eager table labeled
  historical to match the non-eager one; 'remain representative' -> 'historical,
  not what a rebuild produces'; probe-diff note -> 'not directly comparable'.
- Saturation framing (awslabs#13): scoped the ~4.75 tok/s per-stream-flat claim to the
  EAGER sweep; noted the non-eager c=64 wall rise (26.61->34.18s, 3.75 tok/s).
- Dockerfile Layer 5 heading (awslabs#7): old-pin identity (PR#41183 first deepep_v2
  commit) -> #52632's merge commit, matching the block below it.
- env_vars.example tag (awslabs#8): v1-20260818 -> v2-20260825 (postdates the pin bump;
  IfNotPresent caches by tag).
- README serve section (awslabs#4 sub-ask): name SERVE_ENFORCE_EAGER=0 as the knob.

Signed-off-by: Anton Alexander <dmvevents@gmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

dependencies Pull requests that update a dependency file

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant