feat(aws-pcs): マルチユーザ(OpenLDAP)・マルチAZ・IAMポリシー・リージョン拡張・SSHアクセス - #1
Closed
DaisukeMiyamoto wants to merge 91 commits into
Closed
DaisukeMiyamoto wants to merge 91 commits into
DaisukeMiyamoto wants to merge 91 commits into
Conversation
Two principals around the PCS reference cluster have very different
responsibilities, so split them into two sample policies:
- cluster-admin-policy.json (16 Sids, ~7.4 KB) — for the deploying
principal. Covers create / update / delete of every resource the
templates provision: CloudFormation parent + nested stacks; EC2
networking lifecycle (VPC, subnet, IGW, NAT, EIP, SG, route table, VPC
endpoint, launch template, placement group); FSx (Lustre + OpenZFS);
pcs:*; IAM role + instance profile lifecycle scoped to *PCS* /
*ImageBuilder* names with iam:PassRole gated by iam:PassedToService;
SSM Parameter Store for Grafana password + DLAMI auto-resolve;
KMS / Secrets Manager / Image Builder / CloudWatch Logs vended
delivery. Derived primarily from the AWS-published "minimum permissions
for an AWS PCS service administrator" policy
(security-min-permissions.html), with VPC + FSx + IAM lifecycle layered
because the all-in-one template provisions those itself.
- cluster-user-policy.json (6 Sids, ~1.5 KB) — for end users (ML
engineers) of an already-deployed cluster. Covers ec2/cfn describe for
login-node discovery; pcs:Get*/List* for cluster status; ssm:GetParameter
scoped to /pcs/*/grafana/*; ssm:StartSession scoped via
ssm:resourceTag/aws:pcs:compute-node-group-name=login* (so users CAN
open shells on login nodes but CANNOT on compute nodes); session
terminate-own only via session/${aws:username}-*. No write access of
any kind to cluster resources.
Both policies validated cleanly (0 findings) by aws accessanalyzer
validate-policy.
iam/README.md walks through what each policy covers, the design
intent (combined CRUD vs phase-split, customer-managed vs inline,
why no AmazonPCSFullAccess), what is intentionally NOT covered (the
compute instance role uses the AWS-managed AWSPCSComputeNodePolicy),
how to attach, AWS-managed policy pairing recipe, and the recommended
refinement loop via aws accessanalyzer start-policy-generation against
CloudTrail history of a sandbox deploy.
Sample-grade — slightly broader than strict least-privilege so they
work end-to-end in typical accounts. Production users should tighten
with resource-level conditions (account ID, region, stack-name prefix,
tag matchers) before adopting.
Two changes that came out of real-world verification:
1) The combined cluster-admin policy was ~7.4 KB, exceeding the IAM
per-policy 6,144-char limit (which applies to BOTH inline and
customer-managed policies — the previous README claim that
customer-managed had a 17,408-char limit was wrong; that 17,408 is
the per-user-aggregate inline-policy ceiling).
Split into:
* cluster-admin-policy.json (5,769 chars, 14 Sids — core CFN/EC2/
FSx/PCS/IAM/SSM/KMS/Secrets/Logs)
* cluster-admin-imagebuilder-policy.json (1,674 chars, 2 Sids —
Image Builder + scoped iam:PassRole to imagebuilder.amazonaws.com)
Both fit inside the 6,144-char limit; most users only need the core.
2) The cluster-user policy used
ssm:resourceTag/aws:pcs:compute-node-group-name to scope login-node
access, but PCS does not actually emit that tag (verified via
describe-instances on a deployed login node). Real tags are:
* aws:pcs:compute-node-group-id (random PCS ID, can't be policy-pinned)
* monitoring-role (semantic for monitoring exporters, not access)
* Name (the all-in-one templates set this to PCS-login on login CNGs
and PCS-<cng-name> on compute CNGs, e.g. PCS-cpu1, PCS-hpc8a)
Switched the SSM-StartSession scope to
ssm:resourceTag/Name=PCS-login*. The Name tag is operator-mutable, so
the README now flags this as a hardening point with a forking option
(add a dedicated IsLoginNode tag to the templates).
Wrap each policy file in a CloudFormation template and put a Deploy
button in iam/README.md:
architectures/aws-pcs/assets/cluster-admin-iam.yaml
Creates 2 customer-managed policies (core + optional Image Builder)
+ an IAM Group with both attached. AttachUsers param adds existing
users to the group; AttachImageBuilderPolicy=true pulls in the IB
statements (default false). Output: GroupName, CorePolicyArn,
ImageBuilderPolicyArn (when applicable).
architectures/aws-pcs/assets/cluster-user-iam.yaml
Creates 1 customer-managed policy + IAM Group. Same AttachUsers
pattern. Output: GroupName, PolicyArn.
YAML templates live in assets/ to match the rest of the reference
architecture's CFN artifacts (same S3 sync target, same Quick-create
URL pattern, same maintenance pipeline). The raw JSON files stay under
iam/ as the source of truth that the YAML embeds.
Verified end-to-end on us-east-2 / 2026-06-10:
- Both IAM CFN stacks reach CREATE_COMPLETE
- Test admin user (only customer-managed policies attached) deploys
pcs-ml-cluster-deploy-all.yaml hpc8a configuration end-to-end with
OnDemandEnableEfa=true to CREATE_COMPLETE in ~28 min, no AccessDenied
in 332+ CloudTrail API calls; UpdateStack on
OnDemandMaxCount→3 reaches UPDATE_COMPLETE; DeleteStack via the
admin user is allowed (simulator: cloudformation:DeleteStack=allowed)
- Test user can pcs:GetCluster, ssm:StartSession on login node (Name=
PCS-login), and is implicitDeny on compute node (Name=PCS-hpc8a)
and on cfn:CreateStack / pcs:CreateCluster / ec2:RunInstances
- aws accessanalyzer validate-policy on all three JSON files: 0 errors
iam/README.md updated to:
- Add Quick-create Deploy buttons pointing at the prod S3 path
awsome-distributed-ai.s3.amazonaws.com/templates/cluster-{admin,user}-iam.yaml
- Document the verification matrix and what the admin user actually
needs (the Name-tag scope rationale + caveat)
- Replace the wrong 17,408-char claim with the correct 6,144-char
explanation and the split rationale
- Keep the manual install path (raw JSON) for CLI-only environments
Templates were also uploaded to the test bucket (midaisuk-llm-dev) for
pre-merge sandbox testing.
…ault VPC name from stack name, label cleanup Non-breaking improvements identified in the parameter audit (43 params, no removals possible without breaking PR awslabs#1124 deployments): 1. **Auto-derive `OnDemandEfaInterfaceCount` from `OnDemandInstanceType`** (mirrors the pattern `add-cng-p5.yaml`/`add-cng-p6-*.yaml` use for GPU families). New default 0 means "auto"; old 1/2 values still work as explicit pins so existing stacks pass through unchanged. Auto map: hpc8a.96xlarge / hpc7a.{96,48,24,12}xlarge / hpc6id.32xlarge → 2 anything else (hpc6a.48xlarge, c7i.metal, etc.) → 1 Removes the 2-NIC HPC tax: hpc8a/hpc7a were the most common EFA targets but had to override the parameter; now the default just works. 2. **`VPCName` default → empty, auto-resolves to `${StackName}-VPC`**. Multiple deployments in one account now get unique VPC names without the user having to pass a different VPCName each time. 3. **ParameterGroup label cleanup**: "6. FSx for Lustre (/fsx) + FSx for OpenZFS (/home) (Advanced)" → "6. FSx for Lustre (/fsx) and FSx for OpenZFS (/home)" (Capacity / LustreDeploymentType / PerUnitStorageThroughput are commonly tuned and Region-dependent; not advanced) "7. Developer / Advanced (Nested Templates & Monitoring Source)" → "7. Developer (do not change for normal use)" 4. **ParameterLabel trims**: "Availability Zone (required)" → "Availability Zone" (CFN console already shows the asterisk for required params) "VPC name" → "VPC name (empty = ${StackName}-VPC)" "AMI ID (empty = latest PCS-Ready Deep Learning AMI from SSM; pin in production)" → "AMI ID (empty = SSM auto-resolve; pin in production)" "dcgm-exporter image (default DCGM 4.5.2 by digest; covers H100/B200/B300)" → "dcgm-exporter image (default covers H100/B200/B300)" "EFA interface count (1 for hpc6a, 2 for hpc8a/hpc7a/hpc6id)" → "EFA interface count (0 = auto from instance type)" PARAMETERS.md updated to match. Backward compatibility: every change preserves prior behavior for any stack that was deployed off the published default — VPCName defaulting from "ML-Cluster-VPC" to ${StackName}-VPC only affects new launches where the user did not pass VPCName explicitly. Old stacks keep their existing VPC name on UPDATE_STACK. (Larger reductions tracked in subagent plan for the next minor: collapse CngName/QueueName pairs, drop *MinCount, drop LustreVersion/Compression, rename Capacity → LustreCapacity for prefix consistency. All of those are renames or removes, so deferred to avoid breaking PR awslabs#1124 users.)
Adds opt-in multi-user support via OpenLDAP running on the login node. Default is off (DeployDirectory=false) — existing single-user (ubuntu) clusters are completely unchanged. When DeployDirectory=true: - Login node: installs slapd, stores DB on shared /home/ldap-db (OpenZFS), so the directory survives login CNG recreation. Admin password is stored in SSM Parameter Store at /pcs/<cluster-id>/ldap/admin-password. - Compute nodes: installs SSSD with ldap provider, discovers login node IP via ec2:DescribeInstances tag query, configures NSS/PAM so LDAP users are visible as POSIX users on all nodes. - Home directories auto-created via pam_mkhomedir on shared /home. - Slurm sees LDAP users transparently via NSS (no Slurm config change). New parameters on add-cng.yaml: - DeployDirectory (true/false, default false) - DirectoryDomainSuffix (e.g. dc=cluster,dc=internal) New scripts: - scripts/setup-openldap-server.sh — slapd install + config on login - scripts/setup-ldap-client.sh — SSSD client config on compute nodes - scripts/ldap-add-user.sh — helper to add POSIX users to the directory Not yet wired into deploy-all.yaml or GPU templates (next commit).
Covers enabling multi-user, managing users (add/remove/list/password/groups), running jobs as specific users, SSH access for LDAP users, UID/GID conventions, SSSD caching behavior, troubleshooting, data persistence, and upgrade path to Simple AD / Managed AD.
- setup-openldap-server.sh: add apt-get update before install (PCS-Ready DLAMI has limited apt sources cached; without update, slapd package is not found) - setup-ldap-client.sh: same fix - tests/multi-user-test.md: comprehensive test suite (5 parts: server health, user lifecycle, Slurm integration, multi-node consistency, resilience). Split from README.md to keep the main test guide concise. - tests/README.md: added Test 11 link to multi-user-test.md
…+ access methods doc - setup-ldap-client.sh: change min_id from 10000 to 1001. GID 3000 (clusterusers) was being filtered as "primary gid out of range" by SSSD, preventing user resolution. min_id=1001 excludes only system users (0-1000) while allowing all LDAP-provisioned GIDs. - setup-openldap-server.sh: persist the admin password to /home/ldap-db/.admin-password (shared OpenZFS). Previously the randomly generated password was lost after the setup script exited, making it impossible to administer the directory after login node replacement. The file is chmod 600 and on shared storage. - USER-MANAGEMENT.md: added "Access methods" section comparing SSM vs SSH-over-SSM vs Direct SSH (pro/con/best-for table), SSH key management approaches, and Slurm accounting with PCS (what's managed by PCS vs what admins still need to do manually with sacctmgr). Verified: getent passwd testuser1 now returns testuser1:*:10001:3000:Test User 1:/home/testuser1:/bin/bash on the login node after SSSD restart.
- pcs-ml-cluster-deploy-all.yaml: add DeployDirectory + DirectoryDomainSuffix params (group 7), forward to both LoginNodeGroupStack and OnDemandCNGStack. Without this wiring, compute nodes never received DeployDirectory=true and SSSD client setup was skipped. - docs/USER-MANAGEMENT.md: add ASCII art template structure diagram showing how DeployDirectory flows from deploy-all through the nested stacks into UserData on login (slapd) and compute (SSSD client) nodes.
…nLDAP-LoginNode)
Rename the boolean parameter to an enum that makes it clear both *what*
directory service is used and *where* it runs:
DirectoryService:
AllowedValues: [none, OpenLDAP-LoginNode]
# Future: SimpleAD, ManagedAD, OpenLDAP-External
This prepares for adding Simple AD or external LDAP options in the future
without a breaking rename. The value "OpenLDAP-LoginNode" communicates
that slapd runs on the login node (not an external server).
Also:
- README.md: added template nesting structure diagram (ASCII art)
showing how DirectoryService flows from deploy-all through login
(slapd server) and compute (SSSD client) CNGs.
- Fixed DeployMonitoring default accidentally changed to 'none' by
bulk sed (reverted to 'false' for add-cng, 'true' for deploy-all).
The bulk sed that renamed DeployDirectory→DirectoryService also changed Default:'false' to Default:'none' on every parameter in both templates. This broke FSxLustreEnableEfa, OnDemandEnableEfa, DeployPseriesCNG (deploy-all) and EnableEfa (add-cng) — their AllowedValues are [true,false] so 'none' is invalid. Reverted to 'false' for all 4. Only DirectoryService (default 'none') and MonitoringRole (default 'none') correctly use 'none'.
…ared /home The admin password was previously written to /home/ldap-db/.admin-password (shared OpenZFS, readable by root on all nodes). This is a security concern in multi-user clusters where compute-node users may have sudo. Now: - setup-openldap-server.sh first tries SSM PutParameter (SecureString) at /pcs/<cluster-id>/ldap/admin-password. SSM is access-controlled via IAM and auditable via CloudTrail. - Falls back to the file only if the instance role lacks ssm:PutParameter (with a WARNING in the install log). - cluster.yaml instance role: broadened the SSM Sid from grafana/* to also cover ldap/* (same pattern, same cluster-scoped resource ARN). - Removed the duplicate SSM put from add-cng.yaml UserData (the script now handles it internally with CLUSTER_ID exported from UserData). - USER-MANAGEMENT.md: added fallback note + clarified SSM is the primary storage mechanism.
…up-directory.sh Single script with role argument (server|client) replaces two separate scripts. Benefits: - One URL to manage, no naming inconsistency - Full picture visible in one file (server + client + shared SSSD config) - Natural extension point for future SimpleAD (add a case branch) - Server role now also configures SSSD on the login node itself (so getent works locally without a separate manual step) UserData in add-cng.yaml simplified to: curl setup-directory.sh if login → bash setup-directory.sh server else → bash setup-directory.sh client Environment variable interface unchanged (LDAP_DOMAIN_SUFFIX, CLUSTER_ID, DIRECTORY_DNS_IPS for future SimpleAD). DIRECTORY_DNS_IPS="" triggers login-IP discovery via ec2:DescribeInstances tag query. Also: README template structure diagram updated with external scripts and SSM parameter references.
…ated DirectoryRole param MonitoringRole is for monitoring exporter placement — reusing it for directory server/client branching was a semantic mismatch. New parameter DirectoryRole (none | server | client): - add-cng.yaml: DirectoryRole param added. UserData passes it directly to setup-directory.sh as the role argument. No more MonitoringRole reference in the directory block. - deploy-all.yaml: forwards DirectoryRole='server' to login CNG and DirectoryRole='client' to compute CNG (both conditional on DirectoryEnabled). New Condition DirectoryEnabled added. - Removed unused Conditions (IsLoginNode, SetupLdapServer, SetupLdapClient) that depended on MonitoringRole for directory logic. MonitoringRole remains unchanged for its original purpose (monitoring exporter placement). The two concerns are now fully independent.
Scripts are now fetched from the same S3 bucket as the nested templates (S3BucketName/S3KeyPrefix + scripts/). This eliminates: - Hardcoded GitHub raw URLs that need branch-switching for testing - Dependency on GitHub availability at instance boot time - Version skew between templates and scripts Changes: - add-cng.yaml: UserData uses `aws s3 cp` from the S3 bucket to fetch setup-directory.sh. New params S3BucketName + S3KeyPrefix forwarded from deploy-all. - deploy-all.yaml: forwards S3BucketName/S3KeyPrefix to both login and compute CNG stacks. - cluster.yaml: instance role gets s3:GetObject on */templates/scripts/* so nodes can fetch scripts at boot. Development workflow: 1. Edit scripts/ locally 2. aws s3 sync scripts/ s3://test-bucket/templates/scripts/ 3. Deploy with S3BucketName=test-bucket 4. Scripts are fetched from test bucket — no branch URL changes needed Production: scripts sync alongside templates to the prod bucket. Future PCS post-install hook: pass the same S3 URL.
…ally
The prod S3 sync pipeline runs:
aws s3 sync assets/ s3://awsome-distributed-ai/templates/
Previously scripts/ was a sibling directory requiring a separate sync
command (which was never added). Moving scripts under assets/ means the
existing one-line sync deploys both templates and scripts with zero
operational change for maintainers.
Result on S3:
s3://bucket/templates/pcs-ml-cluster-deploy-all.yaml
s3://bucket/templates/add-cng.yaml
s3://bucket/templates/scripts/setup-directory.sh ← new
s3://bucket/templates/scripts/install-enroot-pyxis.sh
s3://bucket/templates/scripts/ldap-add-user.sh
UserData fetches via:
aws s3 cp s3://${S3BucketName}/${S3KeyPrefix}scripts/setup-directory.sh
README path references updated.
…README USER-MANAGEMENT.md completely rewritten: - Quick reference table at top (common tasks → one-liner commands) - How-it-works diagram (slapd → SSSD → NSS/PAM flow) - Step-by-step for every operation: add/delete/list users, reset password, create groups, batch add, Slurm accounting registration - Verification steps (how to confirm user is visible on compute nodes) - Troubleshooting section (common errors + fixes) - UID/GID convention table - Data persistence + backup/restore - Template structure (how DirectoryRole flows) - Upgrade path to Simple AD README.md: - Fixed template diagram: old script names → setup-directory.sh server/client - Fixed "fetched from GitHub raw" → "fetched from S3" - Added USER-MANAGEMENT.md + DEPLOY-TESTING.md to Additional Resources
- DeployDirectory → DirectoryService=OpenLDAP-LoginNode in tests/ - scripts/install-enroot-pyxis.sh → assets/scripts/... in template descriptions (add-cng*.yaml), ROADMAP, tests/README.md - PostInstallScriptUrl default URL: .../scripts/... → .../assets/scripts/... (file was git-mv'd, GitHub raw URL must match new repo path) - No remaining references to old paths or old param names
…r-monitoring link
…skips shellscript sections)
Test 12 (accounting-test.md):
- Validates Slurm managed accounting with LDAP multi-user
- Covers: sacctmgr user/account creation, resource limits (GrpTRESRunMins),
job tracking (sacct), reporting (sreport), fairshare, and
AccountingPolicyEnforcement behavior (none vs associations,limits,safe)
- Based on https://aws.amazon.com/blogs/hpc/introducing-managed-accounting-for-aws-parallel-computing-service/
Test 13 (gpu-healthcheck-test.md):
- Integrates 4.validation_and_observability/2.gpu-cluster-healthcheck suite
- Lightweight (checks 0-3: nvidia-smi, DCGM L2, EFA enum, topology) ~15 min
- NCCL multi-node (check 5) with per-instance bandwidth thresholds
- Slurm prolog integration guide for production
- When-to-run decision table (deploy, pre-training, steady-state, quarantine)
README: added GPU Health Check link to Additional Resources.
tests/README.md: added Test 12 + Test 13 entries with links.
README.md was too long — reduced from 842 to 73 lines (index + matrix only). Test procedures moved to per-category files: infra-test.md — Tests 1-3, 8 (monitoring, container runtime, AMI build) compute-test.md — Tests 4-6 (CPU queue, GPU families, NCCL EFA) training-test.md — Test 7 (FSDP Llama-2 7B) hpc-efa-test.md — Test 9 (EFA on CPU HPC, OSU benchmarks) storage-test.md — Test 10 (FSx health + performance regression) multi-user-test.md — Tests 11-12 (OpenLDAP + accounting, merged) gpu-healthcheck-test.md — Test 13 (GPU health check suite) Removed accounting-test.md (merged into multi-user-test.md since they are always tested together).
…reaks cloud-init parsing)
…ion numbering
Before: 13 top-level sections mixing core usage with advanced features and reference.
After: 11 sections with clear separation:
§1-7: Core (Key Features → Running a GPU job)
§8: Advanced Features
8.1 Monitoring
8.2 Pre-baking AMI
8.3 User Management (new — was standalone §11)
8.4 IAM Permissions (new)
§9: Templates (reference)
§10: Testing and Validation (reference)
§11: Additional Resources
This makes it clear that §1-7 is the primary user path (deploy + run jobs),
and everything in §8 is opt-in for production/multi-user/security hardening.
…y from cluster core Before: DeployMonitoring + GrafanaPublicAccessCidr were under "PCS Cluster Configuration" (which is really scheduler/AMI/accounting settings). After: §2. PCS Cluster Configuration — Slurm version, login instance, AMI, accounting §7. Additional Cluster Configuration — Monitoring + Multi-User Directory This separates "what the cluster IS" (§2) from "what extra features are layered on top" (§7). Also renamed §6 from "(Advanced)" to just "FSx Storage" (Capacity/DeploymentType are commonly tuned, not advanced) and §8 to just "Developer / Advanced" (shorter).
…cAccessCidr→GrafanaAccessCidr Three changes in deploy-all parameter interface: 1. **SSHAccessCidr** (new, §2 PCS Cluster Configuration): CIDR to open port 22 on the login node. Empty = SSH over SSM only. For multi-user clusters where users connect via standard SSH. 2. **DeployMonitoring → MonitoringStack** (rename, §7): Boolean true/false → enum: 'Prometheus-LoginNode' (default) | 'none'. Aligns with DirectoryService naming pattern (<what>-<where>). Nested stacks receive DeployMonitoring="true"/"false" via !If [MonitoringEnabled] so add-cng.yaml is unchanged (backward compat within the nest). 3. **GrafanaPublicAccessCidr → GrafanaAccessCidr** (rename, §7): Shorter name, same semantics. Opens HTTPS/443 on login node. Security group redesign: - LoginAccessSecurityGroup now conditionally includes BOTH SSH (from SSHAccessCidr) and HTTPS (from GrafanaAccessCidr) rules via !If. - Created only when at least one CIDR is set (LoginAccessEnabled = OR). - Attached only to login node via ExtraSecurityGroupId (compute unaffected). New Conditions: SSHAccessEnabled, GrafanaAccessEnabled, LoginAccessEnabled (OR of both), MonitoringEnabled.
…esults - training-test.md: Test 7b Megatron-LM GPT-3 TP/PP/DP on p6-b200 x4 (~134 TFLOP/s/GPU, lm loss 10.91->10.46, EFA efa-direct 8 nics) - gpu-healthcheck-test.md: intensive suite EFA loopback PASS on p6-b200; caveat that check 5 (NCCL) needs the ECR '#' URI + in-image binary path, so validate multi-node NCCL via canonical Test 6 - README.md: matrix row 7,7b + Megatron canonical-asset row
Capture the follow-up to make CapacityReservationId work for targeted ODCRs (not just Capacity Blocks): a CapacityReservationType enum (none/capacity-block/targeted-odcr) that branches MarketType + placement group, with none/capacity-block staying backward compatible. Notes how to verify without GPU capacity (MinCount=0 launch-template assertion + cheap-type consumption check).
… slim advanced detail - §4 Configuration: add DirectoryService / SSHAccessCidr / GrafanaAccessCidr / AdditionalSubnetAZ2-3 rows so the new user-facing features are discoverable - §4: fix HomeThroughput AllowedValues (SINGLE_AZ_HA_1 includes 64) - §1: correct capacity wording (open ODCR auto-consumed; targeted ODCR on roadmap) - §9: p6-b300 row shows 17 interfaces (16 EFA + 1 ENA), consistent with §4 - move EFA-on-CPU + FSx-Lustre-EFA into Advanced (§8.5/§8.6); consolidate the DCGM version detail into a note at the end of §8.1 Monitoring - §8.2: note the single-login-node / SPOF constraint for DirectoryService - §7: replace specific NCCL busbw numbers with a pointer to tests/compute-test.md - §4: document OnDemandInstanceType + PseriesInstanceType accepted values
…tale diagram memo - OPERATIONS.md: correct OpenZFS HomeThroughput allowed values (192/384/768 are not valid AWS values); SINGLE_AZ_2 groups with HA_2 (160..10240), SINGLE_AZ_HA_1/SINGLE_AZ_1 = 64..4096 — matches the template Rules - USER-MANAGEMENT.md: note that slurm-25.11 paths must become slurm-25.05 when deployed with SlurmVersion=25.05 - delete docs/architecture-components.md: an orphan diagram-drafting memo (not linked anywhere, the architecture image is already published) with stale content (32x EFA on all GPUs, single private subnet, optional Enroot/Pyxis)
…hitecture, slim test memos README: - §1 Key Features: add multi-user (OpenLDAP), access control (IAM + SSH/Grafana CIDR), and multi-AZ / broad Region coverage - §2 Architecture: subnets across up to 3 AZs, SSH/Grafana CIDR access, and the optional OpenLDAP / IAM add-ons - §4 Configuration: reorder the parameter table to match the deploy-all console parameter-group order (Network → PCS Cluster → On-Demand → GPU → Additional); point OnDemandEfaInterfaceCount at §8.5 tests/README: - drop the one-off 'GPU health check verified on B200' memo and the Mumbai 'exercised most deeply' paragraph (results live in the per-test files) - region-coverage GPU column is now a simple ran-on-reserved-capacity check (us-east-1/us-east-2/us-west-2 verified via Capacity Block for ML); drop the per-instance-type detail from the table
…m earlier rounds - README §1: remove the 'Multi-AZ and broad Region coverage' Key Features bullet (multi-AZ stays documented in §4 / Architecture; not a headline feature) - tests/README region coverage: clarify the GPU column ✅ reflects earlier GPU validation rounds, not necessarily the row's deploy/monitoring verification date
…etail sections Key Features bullets describe capability + link to the relevant section instead of embedding parameter names (MonitoringStack/GrafanaAccessCidr/DirectoryService/ SSHAccessCidr). Detailed config stays in §4 and §8.
…s in the table - replace the two overlapping intro paragraphs with one: only PrimarySubnetAZ is required; the table is the most-used subset grouped by the console's parameter groups (separator rows), with storage in its own table and the full reference in PARAMETERS.md - keep the in-table group separator rows for discoverability
…move FSx after it Container Runtime (PostInstallScriptUrl/Args) is rarely touched directly, so it no longer gets its own console group: folded into '5. Additional Cluster Configuration (Monitoring, Multi-User, Container Runtime)'. FSx Storage moves after Additional (now group 6). New order: Network, PCS Cluster, On-Demand, GPU, Additional, FSx, Developer. README §4 separators and PARAMETERS.md §2b updated to match.
…from §7 - move the PCS-Ready DLAMI 'no AMI build needed' note from Key Features into §2 Architecture, pointing to §4 'AMI and container runtime' for the detail (dedup) - §7: link the GPU Cluster Health Check suite as a pre-long-run check - §8.1 Monitoring: replace the MonitoringRepo/Version/DcgmExporterImage bullets with a one-line pointer to PARAMETERS.md (users rarely change these); move the monitoring-role vs Name-tag note to the end of §8.1
Replace the single long table (with in-table separator rows and verbose Purpose cells) with one small table per console parameter group (Network / PCS cluster / On-Demand / GPU / Additional). Trim each Purpose to one line and move the long PseriesInstanceType accepted-values list to the GPU compute subsection. Easier to scan; group structure is obvious from the subheadings.
…with deploy-small tip - §4 config sub-tables now carry the console group numbers (1 Network … 5 Additional) - move the FSx deployment-types/sizing content out of §4 into a new §8.1 Storage at the top of Advanced Features; renumber the rest (Monitoring 8.2 … FSx-EFA 8.7) and fix all §8.x anchor references in README + PARAMETERS.md - add a 'deploy small, expand after' tip: FSx Lustre/OpenZFS can be grown after create, and a smaller filesystem deploys faster — start near minimum Capacity/HomeCapacity and expand once the stack is CREATE_COMPLETE
…s, refocus DEPLOY-TESTING, move IAM verification to tests - remove --profile claude (project test setting) from DEPLOY-TESTING.md and tests/infra-test.md; generalize fixed bucket/region to placeholders - DEPLOY-TESTING.md rewritten for the real audience: a third party deploying not-yet-published templates by hosting them in their own S3 bucket and pointing S3BucketName/S3KeyPrefix at it - PARAMETERS.md sections + order now match the deploy-all console parameter groups exactly (1 Network, 2 PCS Cluster, 3 On-Demand, 4 GPU, 5 Additional incl. monitoring repo/version/dcgm + container runtime, 6 FSx, 7 Developer); fix stale anchor links - move the IAM verification results out of docs/IAM.md into a reproducible tests/iam-test.md (representative two-role use case); add the row to tests/README
The standalone §8.7 'FSx for Lustre over EFA (GPUDirect Storage)' is storage content, so it's now a #### subsection at the end of §8.1 Storage instead of a separate top-level Advanced feature. Removes the cross-reference note.
…template deploys Adds an Advanced Features subsection that links docs/DEPLOY-TESTING.md (deploying fork/ branch/PR templates from your own S3 bucket via S3BucketName/S3KeyPrefix). This also gives DEPLOY-TESTING.md its README entry point — every docs/ file is now linked from the README.
The '5. Additional' heading said 'container runtime' but the table only listed the common params (MonitoringStack/GrafanaAccessCidr/DirectoryService). Retitle to '(monitoring, multi-user)' and add a one-line note that the rarely-changed monitoring-source + container-runtime params live in the same console group (see PARAMETERS.md).
The 'console's group 5 also holds…' line was redundant with the PARAMETERS.md pointer right below it. Group 5 is just the heading + the 3 common params.
…p names The §4 sub-table headings now use the exact console parameter-group labels (1. Network Configuration … 5. Additional Cluster Configuration (Monitoring, Multi-User, Container Runtime)), and group 5 lists PostInstallScriptUrl so the heading and its rows agree.
The default is no longer the awslabs/main GitHub-raw URL — it's empty, which resolves to s3://<S3BucketName>/<S3KeyPrefix>scripts/install-enroot-pyxis.sh. Rewrite the caveat: when testing unpublished templates, point S3BucketName at your own bucket and sync the scripts there, or first-boot Enroot/Pyxis fetch fails (cluster still CREATE_COMPLETE). Link DEPLOY-TESTING.md.
Generalize the s3:// verification note and Test 8 build/deploy commands to use <bucket>/<prefix> placeholders instead of the project's private midaisuk-llm-dev bucket, matching DEPLOY-TESTING.md.
… docs/CUSTOM-AMI.md - IAM.md: add 1-click Deploy buttons to the two-role summary table (matching the deploy table further down) - move the §8.5 step-by-step pre-bake procedure into docs/CUSTOM-AMI.md; README §8.5 keeps a concise summary + Launch button + link (heading text unchanged so the awslabs#85-... anchor referenced elsewhere stays valid)
…single location Match the custom-AMI style: the cluster-admin / cluster-user Deploy buttons now use the launch-stack.svg image, kept in one place (the 'Deploying the policies' table); drop the duplicate kbd buttons from the role summary table.
…uttons Replace the kbd 🚀 buttons in the Templates table with the launch-stack.svg image, consistent with the Quick Start, §8.5 custom-AMI, and IAM policy Deploy buttons.
Turn the two-role bullet list into a table with per-role Launch-stack Deploy buttons, consistent with §9 Templates, §8.5, and docs/IAM.md.
…fy NIC wording, flag CB billing - tests: number the file headings (Tests 11-12 / Test 13 / Test 14) so they line up with the test matrix in tests/README.md - compute-test: '16 cards' → 'the 16 EFA NICs' to avoid confusion with GPU count - README §4: add a CB-billing⚠️ to the CapacityReservationId row (links to GPU compute)
…ope for this doc)
# Conflicts: # architectures/aws-pcs/tests/README.md
DaisukeMiyamoto
pushed a commit
that referenced
this pull request
Jul 6, 2026
…st case on EKS (awslabs#1146) * feat(dreamzero): add launcher scripts + 14B eval config * feat(dreamzero): add self-contained two-stage Dockerfile (MIT-0) + docker build helpers Two-stage EFA-overlay Dockerfile ported from the validated RLinf-on-eks image (deduplicated venv via upstream embodied-libero, DCP-save hotfix patch applied at build). Build context helpers (install_extras.sh, run_training_eks.sh, the DCP patch) live under docker/ (NOT build/, which the repo .gitignore excludes). * feat(dreamzero): add workflow manifests (RayJob SFT, download, metadata, convert, eval) * fix(dreamzero): update stale parent-repo path refs in script/config comments * feat(dreamzero): add optional FSx storage, HF secret, and env_vars templates * feat(dreamzero): add docker buildx build-push helper * feat(dreamzero): add kaniko in-cluster build alternative (pending live validation) Parameterize the Dockerfile stage-2 base (RLINF_UPSTREAM_IMAGE) so the kaniko two-stage flow can FROM the ECR-pushed stage-1 tag. The docker buildx path (build-push.sh) remains the validated primary. The kaniko path is structurally complete and renders valid, but the in-cluster build has NOT yet been run end-to-end -- README marks it pending live validation. * feat(dreamzero): add diagrams (incl. rendered infra SVG) + assets via Git LFS WAM training/inference + infra-topology draw.io sources and rendered SVGs; rollout mp4 + loss-curve png. Binaries tracked via Git LFS (.gitattributes). Rendered the previously-missing infra-dreamzero-sft SVG with the draw.io CLI. * docs(dreamzero): add kubernetes/libero walkthrough README * docs(dreamzero): add top-level test-case README * chore(dreamzero): add MIT-0 headers to ported scripts/config; de-couple KubeRay prereq comment Final self-contained sweep: add SPDX MIT-0 headers to the 5 ported shell scripts + the eval config (carried over from RLinf-on-eks without headers); change the SFT manifest's KubeRay-prereq comment from the RLinf-on-eks terraform reference to the generic helm install. Repo is now FULLY_SELF_CONTAINED. * fix(dreamzero): correct infra diagram to KubeRay RayJob; drop stale assets The infra-dreamzero-sft topology diagram depicted the abandoned StatefulSet + headless Service design (manual `ray start` head election, CodeBuild build step, leaked namespace), contradicting both the RayJob manifest and the README prose beside it. Rewrite it to the validated KubeRay RayJob topology: operator-managed embedded RayCluster (1 head + 1 worker), no manual head election, tool-neutral "build + push to ECR" step, <NAMESPACE> placeholder, and a bidirectional RayJob<->FSx arrow showing the /fsx mount (reads model/dataset/metadata, writes DCP checkpoints). Re-export SVG (white/dark adaptive background to match the sibling diagrams). Remove the rollout video and loss curve: the 1-step smoke-run rollout (success_once=0, expected) reads as broken in a public PR, and the loss curve is from a pre-refactor DeepSpeed run now disconnected from the FSDP2 config. README text describes the validated scope instead. Drop the now-unused assets/*.mp4 and assets/*.png LFS attributes. * fix(dreamzero): root-cause DCP finalization crash (gloo coordinator PG) + force save_full_model_weights=false Reproduced, root-caused, and fixed the DreamZero 16B SFT checkpoint crash end-to-end on 2x p5en (RayJob SUCCEEDED, 207GB DCP, zero errors). Root cause: on torch 2.6, dcp.save's post-write finalization broadcasts a multi-MB pickled result object over the default (NCCL) process group on CUDA. At the end of a long (~209GB / ~20min) checkpoint write this races with NCCL comm teardown, leaving non-coordinator ranks with an all-zero buffer -> `_pickle.UnpicklingError: invalid load key '\x00'`, AFTER all 16 shards + .metadata are already on disk. Fix (dcp-save-gloo-coordinator.patch): pass a dedicated CPU/gloo process group to dcp.save(..., process_group=gloo_pg) so the finalization object-broadcast runs over gloo (CPU), immune to the CUDA/NCCL teardown race -- the same approach torch 2.7+ takes upstream. Replaces the prior symptom-guard dcp-save-finalize-besteffort.patch (removed). Also force +actor.fsdp_config.save_full_model_weights=false in the launcher: libero_sft_dreamzero_14b.yaml omits the key, so it defaults to True, which on the 16B model hits "Backend nccl does not support allgather_into_tensor_coalesced" during the full-state-dict gather. DCP-only + offline convert is the supported path. Harden the Dockerfile patch-apply loop to be nullglob-safe. Update the kubernetes/libero README + kaniko-build.yaml patch references accordingly. * chore(dreamzero): drop unvalidated in-cluster kaniko build path Remove the experimental kaniko in-cluster build (setup/kaniko-build.yaml) and all references to it. Every other test case in the repo builds images with `docker build`/`docker buildx` and pushes to ECR; the closest analog (openvla-oft LIBERO-on-EKS) uses `docker buildx build --platform linux/amd64`. dreamzero was the only test case introducing kaniko, and that path was unvalidated, carried known build bugs, and pinned a `:latest` executor image (against CONTRIBUTING's "do not use a latest tag" rule). The validated `docker buildx` path (setup/build-push.sh) is now the sole, documented build method. kaniko can return in a follow-up PR once live-validated. - delete kubernetes/libero/setup/kaniko-build.yaml - README.md: build statement now points at build-push.sh - kubernetes/libero/README.md: remove the "Alternative path (kaniko)" block and the kaniko layout entry; reword the primary path as the sole path - Dockerfile: reword RLINF_UPSTREAM_IMAGE comment (kept the ARG; it is a generic stage-1 override the buildx path also uses) * chore(dreamzero): drop stray RLinf-on-EKS references from ported test case Remove source-repo references that leaked into the upstream port: - install_extras.sh / eval config: comment wording - dreamzero-wam{,-inference}.drawio + .svg: drop editor agent metadata * fix(dreamzero): make Dockerfile buildable + simplify build context First successful build of the test-case image (validated L1 CodeBuild + L2 container test, 10/10, on p5en.48xlarge): - Move RLINF_UPSTREAM_IMAGE ARG to global scope (before the first FROM). A per-stage ARG is invisible to a later FROM and resolves blank under BuildKit ("base name should not be blank"), which made the Dockerfile unbuildable by CodeBuild and docker buildx alike. - Drop the EXTRAS framework + install_extras.sh: its only value was an editable RLinf install into venvs the DreamZero workflow never uses; RLinf is imported from cwd (/workspace/RLinf), so the install was dead weight. Removes the misleading single-plugin 'extras' abstraction. - Drop run_training_eks.sh: the generic launcher is unused by this DreamZero-only test case (no manifest references it). - Relocate dcp-save-gloo-coordinator.patch to the test-case root and remove the empty docker/scripts/patches/ tree; COPY *.patch instead. - Correct DCP-fix comments: the sync dcp.save path is byte-identical through >= torch 2.8, so the patch is permanently required (the prior 'torch 2.7+ fixes this' claim was wrong). * chore(dreamzero): move build-push.sh out of one-file setup/ dir Relocate kubernetes/libero/setup/build-push.sh -> kubernetes/libero/build-push.sh and drop the single-file setup/ directory. Matches the sibling openvla-oft/kubernetes/libero/ layout (helper scripts flat in libero/) and colocates the script with the env_vars it sources. - ROOT path ../../.. -> ../.. (one level shallower); verified it still resolves to the test-case root (the buildx build context). - Usage comment: source ../env_vars -> source ./env_vars (now same dir). - Update references in README.md (root + libero walkthrough) and Dockerfile comments. * docs(dreamzero): correct model/transfer claims to match sources Align the READMEs with the DreamZero paper, HF model card, and RLinf docs: - Parameter count in the intros: 16.48B -> 14B (the published headline; the README titles already say 14B). The measured ~16.48B instantiated-model figure is kept where it matters (FSDP sharding / OOM / VRAM sections). - Architecture phrasing: drop the unsourced 'shared causal self-attention space' for 'causal (autoregressive) ... via flow matching', grounded in the CausalWanModel class and the paper. - LIBERO framing: it is a manipulation benchmark on the same Franka arm as DROID, not a new embodiment, so drop 'new embodiment' / 'cross-embodiment transfer' and frame the sample around its real purpose (EKS deployment). - Add the upstream 5B LIBERO-Spatial accuracy (~96.7% success_once at step 18000) as evidence the recipe converges with sufficient steps. * docs(dreamzero): note real->sim domain gap + swap-in-your-own-data caveat Add a closing caveat to the intro: LIBERO is a simulation of the same Franka Panda arm DROID captures in the real world, so warm-starting onto LIBERO bridges a real->sim visual domain gap, and in practice users would substitute their own dataset for task-specific or cross-embodiment fine-tuning. Avoids the term 'negative transfer' (the RLinf docs recommend warm-starting from the released checkpoint, i.e. the prior is beneficial). * docs(dreamzero): link root README to detailed architecture; drop unused inference diagram - Root README Architecture section now links to the full topology + WAM component breakdown in kubernetes/libero/README.md#architecture (previously just a bare image with no pointer to the detail). - Remove dreamzero-wam-inference.drawio + .svg: it depicts the paper's closed-loop real-time inference path, which this SFT+eval test case does not cover. It was referenced by no README in either repo and was already flagged for removal in the RLinf-on-eks rearchitecture plan. * docs(dreamzero): fix broken 1.architectures link + reframe Prerequisites - Fix the relative path to 1.architectures/4.amazon-eks: it needs ../../../ (dreamzero -> pytorch -> 3.test_cases -> root), not ../../ which dead-ends in 3.test_cases/. Both the Prerequisites and References links were wrong; the sibling openvla-oft uses the correct depth. - Reframe Prerequisites around the nodes, not the provisioning mechanism. The RayJob is fixed-size (head 1 + worker replicas/min/max = 1), so GPU autoscaling is a convenience (on-demand p5en provisioning + scale-down), not a requirement -- a static managed node group or Capacity Block works identically. The only Karpenter-ism shipped is the karpenter.sh/do-not-disrupt annotation, which is ignored on non-Karpenter clusters. * docs(dreamzero): fix stale buildspec/CodeBuild references to build-push.sh This test case builds via kubernetes/libero/build-push.sh (docker buildx); it ships no buildspec.yml. Several comments still referenced CodeBuild / a buildspec (carryover from the RLinf-on-eks origin): - Dockerfile: the RLINF_UPSTREAM_IMAGE note and the stage-1 placeholder comment now describe the build-push.sh flow (and fix the local-build example's BUILD_TARGET: embodied-libero, not embodied-maniskill_libero); the 'cloned in pre_build / by the buildspec' notes now say build-push.sh. - dreamzero-eval.yaml: prerequisite note '(examples/buildspec.yml)' -> '(build-push.sh)'. Also improve the 1.architectures link text and target the section README. * docs(dreamzero): clarify eval intro is the pipeline run, not an accuracy result 'evaluate the result in the LIBERO simulator and render in-sim rollout videos' -> 'run the LIBERO simulator eval (which renders in-sim rollout videos)'. The eval + video path is shipped and validated end-to-end, but a 1-step checkpoint yields success_once=0.0 (documented in step 6); this avoids implying the intro promises a meaningful accuracy result. * docs(dreamzero): reword warm-start rationale for clarity Replace 'There is no native LIBERO 14B checkpoint upstream -- warm-starting from DROID is the point' with customer-facing framing: the released DreamZero-DROID checkpoint is the 14B foundation weight you warm-start from, and continue-SFT adapts it to your target data (here LIBERO), the same pattern you'd follow with your own dataset. The old wording used insider framing and was imprecise (a 5B LIBERO checkpoint does exist upstream; only a 14B LIBERO one does not). * fix(dreamzero): run SFT launcher from ConfigMap mount, not deleted scripts dir The RayJob entrypoint copied the launcher to /workspace/eks/scripts/, but that directory no longer exists in the image: an earlier cleanup changed 'COPY docker/scripts/ /workspace/eks/scripts/' to 'COPY *.patch /workspace/eks/patches/', so the image now only creates /workspace/eks/patches. The cp failed with 'No such file or directory' (exit 127) and the RayJob never started. Run the launcher directly from its read-only ConfigMap mount at /tmp/scripts (it is path-independent -- it cd's to /workspace/RLinf itself), with a fallback to /workspace/eks/scripts for images that still bake it in. Caught by a live multi-step SFT run on 2x p5en. * docs(dreamzero): bullet the upstream-projects list in libero README * docs(dreamzero): add real 300-step training-convergence result Both READMEs' validation-scope sections now report the actual multi-step run on 2x p5en (FSDP2 + KubeRay, the shipped stack): train/loss 0.232 -> 0.085 over 300 steps (~6.9 s/step), a 207 GB DCP checkpoint written with zero UnpicklingError (exercising the gloo-coordinator fix at the full 16.48B scale). Framed honestly as 'trains and converges', NOT a task-accuracy claim (300 steps is short; the released 14B trained for 100K). The 1-step success_once=0.0 note and the upstream 5B ~96.7% accuracy reference are retained. * docs(dreamzero): de-jargon the checkpoint-fix mention in validation notes The validation-scope summaries (near the top of both READMEs) were the first place a reader met 'UnpicklingError' / 'gloo-coordinator fix', but the explanation only appears in the troubleshooting table (libero README) and nowhere in the root README. Reword to plain outcome language -- 'no corruption or save-time crashes' -- keeping the patch link (and a 'see Troubleshooting' pointer). UnpicklingError now appears only in the troubleshooting row, where a reader who hits it would look, with full context. Also made the root README's patch reference a clickable link. * docs(dreamzero): link Troubleshooting + explain 14B vs 16.48B - libero README: make 'see Troubleshooting' a real anchor link (#troubleshooting). - Reconcile the 14B/16.48B discrepancy that appeared unexplained: '14B' is the Wan video-diffusion DiT backbone (headline); '16.48B' is the full trainable WAM once the action/state encoders + action head are added (live run reports 16,484,292,448 params). Added a callout box in the libero README and a concise inline gloss in the root README so the two figures are no longer ambiguous. * docs(dreamzero): reconcile 14B / 16.48B / 23B param figures precisely The HF model card publishes the checkpoint as '14B' (23B on-disk safetensors incl. frozen encoders); our live SFT run trains 16,484,292,448 params. Rewrite the callout to cover all three scopes of the SAME checkpoint and stop implying the encoders/head are added on top -- they are part of the published WAM: - 14B = publisher headline (Wan DiT backbone) - 16.48B = trainable params when instantiated for full SFT (backbone + the model's action/state encoders + action head) - 23B = on-disk total (also includes frozen CLIP / UMT5-XXL / Wan VAE) '14B checkpoint' phrasing is kept (it matches the publisher); root README gloss tightened to match and point at the callout. * docs(dreamzero): add non-commercial model-license notice (legal) Per legal review: the GEAR-Dreams/DreamZero-DROID model is released under a non-commercial license (CC-BY-NC-4.0). Add a prominent notice at the top of both READMEs stating that any production use needs additional approvals and that users should review the license terms before using the model to make an informed decision. The notice scopes itself to the MODEL only; the repository code remains MIT-0 (unchanged). * docs(dreamzero): list optional FSX_* vars in env_vars.example The storage/ manifests use FSX_SUBNET_ID, FSX_SECURITY_GROUP_IDS, FSX_FILESYSTEM_ID, FSX_DNS_NAME, and FSX_MOUNT_NAME, but env_vars.example only listed the always-needed vars. Add the FSX_* vars (commented out, marked optional -- only for customers who provision FSx via storage/*.yaml rather than reusing an existing fsx-claim). The README already gives these as inline exports at the storage step; this just makes the env-var reference complete. No manifest or behavior change. * docs(dreamzero): fix comment accuracy + remove personal namespace from examples Comment audit before PR: - Remove personal namespace leak: 4 manifests had 'export NAMESPACE=natharno' and dreamzero-sft.yaml had 'rlinf' in their usage examples. Normalize all to 'dreamzero' (matching env_vars.example). - run_dreamzero_sft_eks.sh: fix the FSx free-space figure (>=200GB -> >=250GB, consistent with the READMEs and manifests); consolidate two overlapping save_full_model_weights/DCP comment blocks into one accurate block (the first vaguely said the gather is 'slow/stalls'; the accurate cause is the NCCL allgather_into_tensor_coalesced error); tighten the metadata comment. Net -8 lines, no behavior change. * docs(dreamzero): note why static PVC uses storageClassName: "" Add a comment explaining the empty-string storageClassName on the static FSx PVC is intentional (disables dynamic provisioning so the PVC binds the named static PV, rather than the cluster default StorageClass provisioning a new volume). Prevents a future reader from 'fixing' it by adding a class, which would break the static bind. * docs(dreamzero): clarify provisioner vs capacity-backing in Prerequisites The previous wording listed 'managed node group, Capacity Block reservation, or Karpenter' as parallel options, conflating two independent axes. Reword to: provisioner (static managed node group OR Karpenter) backed by a capacity reservation (Capacity Block for ML OR ODCR) -- either capacity type works with either provisioner. Drop 'on-demand' as a backing option since 2x p5en on-demand is effectively unobtainable; these instances come from a reservation. * docs(dreamzero): reword loss-curve description 'clean, monotonic-ish curve' -> 'steady downward trend' -- more professional and accurate (the curve declined overall but had step-to-step noise, not strict monotonicity). * docs(dreamzero): deep-link the root README to the 14B-vs-16.48B-vs-23B note The note is a blockquote callout (no auto-generated anchor). Add an explicit <a id="param-counts"></a> anchor before it in the libero README and turn the root README's plain-text reference into a deep link (kubernetes/libero/README.md#param-counts) that jumps straight to it. * fix(dreamzero): drop no-op network-layer topology constraint + false latency claim The topologySpreadConstraints used topologyKey topology.k8s.aws/network-node-layer-2, a label applied only by SageMaker HyperPod -- vanilla EKS nodes (incl. this cluster) do not carry it, so with whenUnsatisfiable: ScheduleAnyway the constraint was a no-op. It was also semantically backwards: a spread constraint spreads pods across domains, while the README claimed it 'prefers co-location ... for lowest NCCL latency' (co-location would need podAffinity, and the inter-node network proximity it targets is a HyperPod-only capability here). Remove the dead constraint from the head + worker (keeping the working podAntiAffinity that pins one pod per node via kubernetes.io/hostname) and drop the inaccurate README sentence. Soft scheduling change only; training correctness unaffected. * docs(dreamzero): spell out acronyms on first use (AWS convention) Expand technical acronyms at first body use (inline, AWS docs/blog style), skipping well-known ones (GPU/CPU/CUDA/NVIDIA) and title occurrences: - Root README: SFT, FSDP, EFA, NCCL, DCP, ECR, DROID. - libero README: FSDP, VRAM, EFA, NCCL, DCP, RDMA, CSI, VPC, IRSA, PVC, VAE, I2V (WAM/DiT/SFT were already spelled out). * docs(dreamzero): fix libero prereq to not require autoscaling/Karpenter Prereq #1 still framed 'GPU autoscaling (e.g. Karpenter)' as a requirement, contradicting the root README fix (commit 1f7ac97). The workload is fixed-size, so autoscaling is a convenience, not a requirement. Reword heading ('provision' -> 'schedule') and body to match: nodes can come from a static managed node group or Karpenter, backed by a Capacity Block for ML or an ODCR. * fix(dreamzero): recommend HF_TOKEN for rate limits + actually wire it in 'Hugging Face auth is optional' was accurate for access but glossed over a real reliability issue: anonymous HF downloads are rate-limited per source IP and the anonymous tier is much stricter than authenticated (HF's #1 cause of 429s). This download is large/multi-file and egresses through a shared EKS NAT gateway, so the cluster's other workloads share the same per-IP anonymous quota. - README: reword the auth note to 'optional, but recommended' and explain the per-IP rate-limit / shared-NAT contention; reword the download step to match. - model-download.yaml: the download script reads os.environ['HF_TOKEN'] but the Job never injected it -- the token could not actually reach the container. Add an env entry sourcing HF_TOKEN from the hf-token Secret with optional: true (Job still runs anonymously if the Secret is absent). Makes the recommendation functional. * docs(dreamzero): fix Step-by-step accuracy (RayJob logs + save_full cause) Final accuracy pass on the Step-by-step section: - Step 4: 'kubectl logs job/dreamzero-sft' was wrong -- dreamzero-sft is a RayJob, not a batch Job. KubeRay creates a submitter Job by that name whose logs only show 'ray job submit' plumbing; the training driver logs are on the Ray head pod. Use 'kubectl logs -l ray.io/node-type=head' instead. - Step 5: align the save_full_model_weights rationale with the accurate cause ('Backend nccl does not support allgather_into_tensor_coalesced') instead of the vaguer 'rank-0 gather stalls'. Verified accurate (no change needed): all job/ConfigMap names, the libero_sft_dreamzero_14b config name, eval LOG_DIR, convert STEP/output path, the video path, total_num_envs=16, and the dataset/metadata notes. * docs(dreamzero): fix stale build-push.sh path in .gitignore comment The script moved out of setup/ to kubernetes/libero/build-push.sh (commit 58477a8); update the comment to match. Ignore rules unchanged. * fix(dreamzero): use immutable IMAGE_TAG instead of mutable :latest A mutable :latest tag makes the pushed image non-reproducible and is the issue CONTRIBUTING calls out ("do not use a latest tag"). build-push.sh now defaults to an immutable IMAGE_TAG (dz-${DREAMZERO_REF}) and pushes ${ECR_URI}:${IMAGE_TAG}. The same IMAGE_TAG is threaded through every manifest's image reference and added to each restricted envsubst allow-list, plus env_vars.example and the README, so the pushed and deployed images are guaranteed to match. * refactor(dreamzero): standardize launcher mount on /opt/scripts SFT mounted its launcher ConfigMap at /tmp/scripts while convert and eval use /opt/scripts. Standardize SFT on /opt/scripts (and the matching entrypoint paths) so the recipe is consistent and easier to copy-adapt between steps.
DaisukeMiyamoto
added a commit
that referenced
this pull request
Jul 8, 2026
…ote #1 The E2E column is the row's primary verification result, so 'Verified' fits better; the old 'Verified' date column becomes 'Date'. Footnotes renumbered in table order (1=Verified/E2E, 2=OpenZFS type, 3=instance types).
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
PR 原案(日本語)— AWS PCS リファレンスアーキテクチャ メジャーアップデート
タイトル
概要
AWS PCS リファレンスアーキテクチャ(
architectures/aws-pcs/)のメジャーアップデート。PR awslabs#1120 の監視スタックの上に、直接ユーザに恩恵のある5つの機能を追加する:
10リージョン・実ハードウェア(hpc6a/hpc8a、p6-b200 ×4)でエンドツーエンド検証済み。
A. 機能追加(直接ユーザに恩恵があるもの)
A-1. マルチユーザ対応(OpenLDAP)
全ノードで同一 UID として解決され、そのまま Slurm ジョブを実行できる。
SSSD で自動参加。ユーザデータベースは OpenZFS 上の
/homeに保存。/home上にあるためログイン障害後の復旧は可能だが、ディレクトリサービスとしてはログインノードが単一障害点(有効時はログインノード1台構成)。
A-2. VPC / サブネットのマルチAZ 対応
特定 AZ のキャパシティ(CB/予約)に合わせたサブネット配置ができる。
primary AZ のみに置く(コスト最小)。
A-3. cluster-admin / cluster-user 用 IAM ポリシー定義
テンプレートとして配布できる。
admin 権限だけで deploy-all が完走する(
iam:CreatePolicy不要 — 付随的変更 B-3 による)。A-4. 対応リージョンの拡張
把握でき、リージョン固有の落とし穴を回避できる。
(ストレージ deployment type・検証日を含む)。
代替値を表に明記。
A-5. ログインノードへの SSH アクセス設定
フォワードに依存しない接続経路)。
該当ポートのみを開放。コンピュートノードは非公開のまま。Grafana 公開(443)も同じ仕組み。
B. 付随的変更(上記の機能追加を成立させるために必要になったもの)
B-1. パラメータ整理(リネーム / 集約 / 削除)
なぜ必要か: 新機能(SSH/Grafana アクセス、監視)に合わせ、deploy-all のパラメータ命名・
粒度を一貫させるため。呼び出し側の修正が必要(詳細は「破壊的変更と移行」表)。
DeployMonitoring(bool) →MonitoringStack(enum) にリネーム(挙動は等価)。GrafanaPublicAccessCidr→GrafanaAccessCidrにリネーム(SSHAccessCidrと対に)。OnDemandEnableEfaを削除しOnDemandEfaInterfaceCount(0/1/2)に集約。VPCNameを削除(スタック名から自動命名)。B-2. VPC プライベートサブネットの CIDR 分割(A-2 マルチAZ に伴う)
なぜ必要か: 追加 AZ 用のサブネットを切り出すため、CIDR の分割数を増やす必要があった。
B-3. PCS インスタンスロール権限の inline 分割(A-3 IAM に伴う)
なぜ必要か: 「admin 権限だけで deploy-all 完走」を成立させるため、デプロイ時に
iam:CreatePolicyを要求するマネージドポリシーを排除する必要があった。B-4. PostInstallScriptUrl の取得方式・配置(A-1/A-4 に伴う)
なぜ必要か: ディレクトリ用スクリプトの配布と、public S3 を持てない環境での取得を
成立させるため、取得元を S3(インスタンスロール経由)に切り替えた。
s3://スキームに対応(private バケットで動作、public S3 不要)。http(s)://も従来通り可。assets/scripts/配下へ移動し、配布バケットの同期対象に含めた。B-5. ドキュメント / テスト整理(全機能に伴う)
なぜ必要か: 追加機能で README が肥大化し、基本フローと上級者向け機能の区別、および
パラメータ齟齬の自動検出が必要になった。
§8 Advanced Featuresに集約。§4 Configuration のパラメータ表は deploy-all のコンソール ParameterGroup と同じ名前・順序)。
docs/に整理:IAM.md/USER-MANAGEMENT.md/CUSTOM-AMI.md(AMI プリベイク)/DEPLOY-TESTING.md(未公開テンプレートを自分の S3 バケットから検証する手順)/PARAMETERS.md/OPERATIONS.md。検証実績は手順としてtests/に集約。gpu-healthcheck/iam)に分割、README は索引 + マトリクス。
tests/lint-docs.sh: リネーム残り・未文書化パラメータ・リンク切れ検出)をpre-merge Test 0 に追加。
C. バグ修正(機能追加中に発見した既存の不具合)
Nameタグがハードコードだったのをスタック名ベースに修正。破壊的変更と移行
DeployMonitoring=true/falseMonitoringStack=Prometheus-LoginNode/noneGrafanaPublicAccessCidrGrafanaAccessCidrOnDemandEnableEfa+OnDemandEfaInterfaceCountOnDemandEfaInterfaceCount=0/1/2VPCName/17/18(+ 追加AZ用に予約)検証
Capacity Block)、ほか計10リージョン。Slurm 25.05 / 25.11 両方。
SSH、監視スタック、OpenLDAP サーバ + コンピュート SSSD 自動参加、LDAP ユーザでの Slurm ジョブ・
複数ノード UID 一致・ユーザ削除伝播、EFA 有効/無効、全テンプレートの CFN lint。
iam:CreatePolicy不要)、cluster-user で SSM 接続のみ・LDAP パスワード不可。
architectures/aws-pcs/tests/配下に記録。メンテナーへの注記
(
S3BucketName/S3KeyPrefix配下、ブートスクリプトはassets/scripts/相当)から取得する。マージ後、
architectures/aws-pcs/assets/配下(テンプレート +scripts/)を配布バケットへコピー(同期)する必要がある。これを行わないと、ネストスタックの取得やノード first-boot の
Enroot/Pyxis・ディレクトリセットアップが失敗する。