Skip to content

feat(aws-pcs): マルチユーザ(OpenLDAP)・マルチAZ・IAMポリシー・リージョン拡張・SSHアクセス - #1

Closed
DaisukeMiyamoto wants to merge 91 commits into
mainfrom
feat/multi-user
Closed

DaisukeMiyamoto wants to merge 91 commits into
mainfrom
feat/multi-user

Conversation

@DaisukeMiyamoto

Copy link
Copy Markdown
Owner

PR 原案(日本語)— AWS PCS リファレンスアーキテクチャ メジャーアップデート

ブランチ feat/multi-user → awslabs/main。PR awslabs#1120(monitoring stack)のマージを前提に
その上へ積み上げる。タイトル/本文はそのまま GitHub PR に転記できる構成。


タイトル

feat(aws-pcs): マルチユーザ(OpenLDAP)・マルチAZ・IAMポリシー・リージョン拡張・SSHアクセス

概要

AWS PCS リファレンスアーキテクチャ(architectures/aws-pcs/)のメジャーアップデート。
PR awslabs#1120 の監視スタックの上に、直接ユーザに恩恵のある5つの機能を追加する:

  1. マルチユーザ対応(OpenLDAP)
  2. VPC / サブネットのマルチAZ 対応
  3. cluster-admin / cluster-user 用 IAM ポリシー定義
  4. 対応リージョンの拡張
  5. ログインノードへの SSH アクセス設定

10リージョン・実ハードウェア(hpc6a/hpc8a、p6-b200 ×4)でエンドツーエンド検証済み。

⚠️ 破壊的変更を含む(パラメータ名・VPC CIDR レイアウト)。後述の「破壊的変更と移行」を参照。


A. 機能追加(直接ユーザに恩恵があるもの)

A-1. マルチユーザ対応(OpenLDAP)

  • 何ができるか: 1つのクラスタを複数ユーザでそれぞれ利用可能にする。LDAP 上のユーザが
    全ノードで同一 UID として解決され、そのまま Slurm ジョブを実行できる。
  • どう構成しているか: OpenLDAP サーバをログインノードで稼働させ、コンピュートノードは
    SSSD で自動参加。ユーザデータベースは OpenZFS 上の /home に保存。
  • 制約: ユーザ DB が /home 上にあるためログイン障害後の復旧は可能だが、ディレクトリ
    サービスとしてはログインノードが単一障害点(有効時はログインノード1台構成)。

A-2. VPC / サブネットのマルチAZ 対応

  • 何ができるか: VPC のプライベートサブネットを最大3AZ に展開でき、可用性の高い構成や、
    特定 AZ のキャパシティ(CB/予約)に合わせたサブネット配置ができる。
  • どう構成しているか: prerequisites に追加 AZ 用サブネットを作成。NAT ゲートウェイは
    primary AZ のみに置く(コスト最小)。
  • 制約: VPC の CIDR レイアウトが変わる(付随的変更 B-2 / 破壊的変更を参照)。

A-3. cluster-admin / cluster-user 用 IAM ポリシー定義

  • 何ができるか: クラスタ管理者・利用者それぞれの最小権限を、既製の CloudFormation
    テンプレートとして配布できる。
  • どう構成しているか: admin / user 用のポリシー+グループを別テンプレートで提供。
    admin 権限だけで deploy-all が完走する(iam:CreatePolicy 不要 — 付随的変更 B-3 による)。
  • 制約: cluster-user は SSM 接続のみで、LDAP 管理パスワードは参照できない。

A-4. 対応リージョンの拡張

  • 何ができるか: どのリージョンで何が動くか(デプロイ/監視/ストレージ/Pyxis/GPU)を
    把握でき、リージョン固有の落とし穴を回避できる。
  • どう構成しているか: 10リージョンで実デプロイ検証し、リージョンカバレッジ表に記録
    (ストレージ deployment type・検証日を含む)。
  • 制約: 一部リージョンは既定のストレージ deployment type が使えない(例: Osaka の OpenZFS)。
    代替値を表に明記。

A-5. ログインノードへの SSH アクセス設定

  • 何ができるか: 指定 CIDR からログインノードへ直接 SSH できる(踏み台や SSM ポート
    フォワードに依存しない接続経路)。
  • どう構成しているか: ログインノード専用のセキュリティグループで、指定 CIDR からの
    該当ポートのみを開放。コンピュートノードは非公開のまま。Grafana 公開(443)も同じ仕組み。
  • 制約: 既定では無効(CIDR 未指定なら追加のセキュリティグループは作られない)。

B. 付随的変更(上記の機能追加を成立させるために必要になったもの)

B-1. パラメータ整理(リネーム / 集約 / 削除)

なぜ必要か: 新機能(SSH/Grafana アクセス、監視)に合わせ、deploy-all のパラメータ命名・
粒度を一貫させるため。呼び出し側の修正が必要(詳細は「破壊的変更と移行」表)。

  • DeployMonitoring(bool) → MonitoringStack(enum) にリネーム(挙動は等価)。
  • GrafanaPublicAccessCidr → GrafanaAccessCidr にリネーム(SSHAccessCidr と対に)。
  • OnDemandEnableEfa を削除し OnDemandEfaInterfaceCount(0/1/2)に集約。
  • VPCName を削除(スタック名から自動命名)。
  • ParameterGroup を再編(監視/ディレクトリ系を1グループに集約)。

B-2. VPC プライベートサブネットの CIDR 分割(A-2 マルチAZ に伴う)

なぜ必要か: 追加 AZ 用のサブネットを切り出すため、CIDR の分割数を増やす必要があった。

  • プライマリサブネットのサイズが変わり、残りを追加 AZ 用に予約する。
  • AZ を追加しない場合でもプライマリサブネットの CIDR が変わる(破壊的)。→ 破壊的変更表参照。

B-3. PCS インスタンスロール権限の inline 分割(A-3 IAM に伴う)

なぜ必要か: 「admin 権限だけで deploy-all 完走」を成立させるため、デプロイ時に
iam:CreatePolicy を要求するマネージドポリシーを排除する必要があった。

  • マネージドポリシーを廃止し、機能ごとの inline ポリシーへ分割(常時 / 監視有効時)。

B-4. PostInstallScriptUrl の取得方式・配置(A-1/A-4 に伴う)

なぜ必要か: ディレクトリ用スクリプトの配布と、public S3 を持てない環境での取得を
成立させるため、取得元を S3(インスタンスロール経由)に切り替えた。

  • s3:// スキームに対応(private バケットで動作、public S3 不要)。http(s):// も従来通り可。
  • ブートスクリプトを assets/scripts/ 配下へ移動し、配布バケットの同期対象に含めた。
  • 既定の取得元が GitHub raw → S3 に変わる(破壊的変更表参照)。

B-5. ドキュメント / テスト整理(全機能に伴う)

なぜ必要か: 追加機能で README が肥大化し、基本フローと上級者向け機能の区別、および
パラメータ齟齬の自動検出が必要になった。

  • README を再構成(基本フローを先頭、上級者向け機能を §8 Advanced Features に集約。
    §4 Configuration のパラメータ表は deploy-all のコンソール ParameterGroup と同じ名前・順序)。
  • 機能別ガイドを docs/ に整理: IAM.md / USER-MANAGEMENT.md / CUSTOM-AMI.md(AMI プリベイク)/
    DEPLOY-TESTING.md(未公開テンプレートを自分の S3 バケットから検証する手順)/ PARAMETERS.md /
    OPERATIONS.md。検証実績は手順として tests/ に集約。
  • tests ガイドをカテゴリ別ファイル(infra/compute/training/hpc-efa/storage/multi-user/
    gpu-healthcheck/iam)に分割、README は索引 + マトリクス。
  • ドキュメント整合 lint(tests/lint-docs.sh: リネーム残り・未文書化パラメータ・リンク切れ検出)を
    pre-merge Test 0 に追加。

C. バグ修正(機能追加中に発見した既存の不具合)

  • prerequisites の VPC Name タグがハードコードだったのをスタック名ベースに修正。
  • first-boot の apt/dpkg ロック競合(unattended-upgrades)への待機・リトライを追加。
  • cloud-init のディレクトリ設定を MIME shellscript → runcmd に移動(PCS は shellscript を実行しない)。

破壊的変更と移行

変更 旧 新 移行
パラメータ名 DeployMonitoring=true/false MonitoringStack=Prometheus-LoginNode/none deploy-all 呼び出しを更新
パラメータ名 GrafanaPublicAccessCidr GrafanaAccessCidr 同上
パラメータ集約 OnDemandEnableEfa + OnDemandEfaInterfaceCount OnDemandEfaInterfaceCount=0/1/2 0=無効、1/2=有効に読み替え
パラメータ削除 VPCName (削除、スタック名から自動) 指定をやめる
VPC CIDR レイアウト primary private subnet /17 /18(+ 追加AZ用に予約) 既存スタックの in-place 更新不可(サブネット再作成)。新規デプロイ推奨
IAM マネージドポリシー inline ポリシー 監視有効スタックの update で IAM リソース差し替え

検証

  • 検証環境: us-east-2(CPU/EFA 一般)、us-west-2(Slurm 版差異)、ap-south-1(GPU: p6-b200 ×4、
    Capacity Block)、ほか計10リージョン。Slurm 25.05 / 25.11 両方。
  • 検証方法:
    • メジャーアップデート e2e(deploy-all 1本): マルチAZ サブネット配置、SSHAccessCidr によるログイン
      SSH、監視スタック、OpenLDAP サーバ + コンピュート SSSD 自動参加、LDAP ユーザでの Slurm ジョブ・
      複数ノード UID 一致・ユーザ削除伝播、EFA 有効/無効、全テンプレートの CFN lint。
    • EFA 実トラフィック(hpc8a 2ノード、OSU ベンチ)。
    • GPU(p6-b200 ×4): NCCL all_reduce、Megatron-LM GPT-3 分散学習、GPU ヘルスチェック。
    • 最小権限 IAM: cluster-admin ロールで deploy-all を作成〜削除(iam:CreatePolicy 不要)、
      cluster-user で SSM 接続のみ・LDAP パスワード不可。
    • リージョンカバレッジ: 10リージョンでデプロイ + 監視 + ストレージ + Pyxis を確認。
  • 詳細な手順・結果は architectures/aws-pcs/tests/ 配下に記録。

メンテナーへの注記

  • 本 PR のテンプレートは、ネストされた CFN テンプレートとブートスクリプトを配布用 S3 バケット
    (S3BucketName/S3KeyPrefix 配下、ブートスクリプトは assets/scripts/ 相当)から取得する。
    マージ後、architectures/aws-pcs/assets/ 配下(テンプレート + scripts/)を配布バケットへ
    コピー(同期)する必要がある
    。これを行わないと、ネストスタックの取得やノード first-boot の
    Enroot/Pyxis・ディレクトリセットアップが失敗する。
  • 前提: PR feat(aws-pcs): add integrated monitoring, P6 GPU node groups, and container runtime to the PCS reference cluster awslabs/awsome-distributed-ai#1120(monitoring stack)。

Two principals around the PCS reference cluster have very different
responsibilities, so split them into two sample policies:

- cluster-admin-policy.json (16 Sids, ~7.4 KB) — for the deploying
  principal. Covers create / update / delete of every resource the
  templates provision: CloudFormation parent + nested stacks; EC2
  networking lifecycle (VPC, subnet, IGW, NAT, EIP, SG, route table, VPC
  endpoint, launch template, placement group); FSx (Lustre + OpenZFS);
  pcs:*; IAM role + instance profile lifecycle scoped to *PCS* /
  *ImageBuilder* names with iam:PassRole gated by iam:PassedToService;
  SSM Parameter Store for Grafana password + DLAMI auto-resolve;
  KMS / Secrets Manager / Image Builder / CloudWatch Logs vended
  delivery. Derived primarily from the AWS-published "minimum permissions
  for an AWS PCS service administrator" policy
  (security-min-permissions.html), with VPC + FSx + IAM lifecycle layered
  because the all-in-one template provisions those itself.

- cluster-user-policy.json (6 Sids, ~1.5 KB) — for end users (ML
  engineers) of an already-deployed cluster. Covers ec2/cfn describe for
  login-node discovery; pcs:Get*/List* for cluster status; ssm:GetParameter
  scoped to /pcs/*/grafana/*; ssm:StartSession scoped via
  ssm:resourceTag/aws:pcs:compute-node-group-name=login* (so users CAN
  open shells on login nodes but CANNOT on compute nodes); session
  terminate-own only via session/${aws:username}-*. No write access of
  any kind to cluster resources.

Both policies validated cleanly (0 findings) by aws accessanalyzer
validate-policy.

iam/README.md walks through what each policy covers, the design
intent (combined CRUD vs phase-split, customer-managed vs inline,
why no AmazonPCSFullAccess), what is intentionally NOT covered (the
compute instance role uses the AWS-managed AWSPCSComputeNodePolicy),
how to attach, AWS-managed policy pairing recipe, and the recommended
refinement loop via aws accessanalyzer start-policy-generation against
CloudTrail history of a sandbox deploy.

Sample-grade — slightly broader than strict least-privilege so they
work end-to-end in typical accounts. Production users should tighten
with resource-level conditions (account ID, region, stack-name prefix,
tag matchers) before adopting.
Two changes that came out of real-world verification:

1) The combined cluster-admin policy was ~7.4 KB, exceeding the IAM
   per-policy 6,144-char limit (which applies to BOTH inline and
   customer-managed policies — the previous README claim that
   customer-managed had a 17,408-char limit was wrong; that 17,408 is
   the per-user-aggregate inline-policy ceiling).

   Split into:
     * cluster-admin-policy.json (5,769 chars, 14 Sids — core CFN/EC2/
       FSx/PCS/IAM/SSM/KMS/Secrets/Logs)
     * cluster-admin-imagebuilder-policy.json (1,674 chars, 2 Sids —
       Image Builder + scoped iam:PassRole to imagebuilder.amazonaws.com)
   Both fit inside the 6,144-char limit; most users only need the core.

2) The cluster-user policy used
   ssm:resourceTag/aws:pcs:compute-node-group-name to scope login-node
   access, but PCS does not actually emit that tag (verified via
   describe-instances on a deployed login node). Real tags are:
     * aws:pcs:compute-node-group-id (random PCS ID, can't be policy-pinned)
     * monitoring-role (semantic for monitoring exporters, not access)
     * Name (the all-in-one templates set this to PCS-login on login CNGs
       and PCS-<cng-name> on compute CNGs, e.g. PCS-cpu1, PCS-hpc8a)

   Switched the SSM-StartSession scope to
   ssm:resourceTag/Name=PCS-login*. The Name tag is operator-mutable, so
   the README now flags this as a hardening point with a forking option
   (add a dedicated IsLoginNode tag to the templates).

Wrap each policy file in a CloudFormation template and put a Deploy
button in iam/README.md:

  architectures/aws-pcs/assets/cluster-admin-iam.yaml
    Creates 2 customer-managed policies (core + optional Image Builder)
    + an IAM Group with both attached. AttachUsers param adds existing
    users to the group; AttachImageBuilderPolicy=true pulls in the IB
    statements (default false). Output: GroupName, CorePolicyArn,
    ImageBuilderPolicyArn (when applicable).

  architectures/aws-pcs/assets/cluster-user-iam.yaml
    Creates 1 customer-managed policy + IAM Group. Same AttachUsers
    pattern. Output: GroupName, PolicyArn.

YAML templates live in assets/ to match the rest of the reference
architecture's CFN artifacts (same S3 sync target, same Quick-create
URL pattern, same maintenance pipeline). The raw JSON files stay under
iam/ as the source of truth that the YAML embeds.

Verified end-to-end on us-east-2 / 2026-06-10:
  - Both IAM CFN stacks reach CREATE_COMPLETE
  - Test admin user (only customer-managed policies attached) deploys
    pcs-ml-cluster-deploy-all.yaml hpc8a configuration end-to-end with
    OnDemandEnableEfa=true to CREATE_COMPLETE in ~28 min, no AccessDenied
    in 332+ CloudTrail API calls; UpdateStack on
    OnDemandMaxCount→3 reaches UPDATE_COMPLETE; DeleteStack via the
    admin user is allowed (simulator: cloudformation:DeleteStack=allowed)
  - Test user can pcs:GetCluster, ssm:StartSession on login node (Name=
    PCS-login), and is implicitDeny on compute node (Name=PCS-hpc8a)
    and on cfn:CreateStack / pcs:CreateCluster / ec2:RunInstances
  - aws accessanalyzer validate-policy on all three JSON files: 0 errors

iam/README.md updated to:
  - Add Quick-create Deploy buttons pointing at the prod S3 path
    awsome-distributed-ai.s3.amazonaws.com/templates/cluster-{admin,user}-iam.yaml
  - Document the verification matrix and what the admin user actually
    needs (the Name-tag scope rationale + caveat)
  - Replace the wrong 17,408-char claim with the correct 6,144-char
    explanation and the split rationale
  - Keep the manual install path (raw JSON) for CLI-only environments

Templates were also uploaded to the test bucket (midaisuk-llm-dev) for
pre-merge sandbox testing.
…ault VPC name from stack name, label cleanup

Non-breaking improvements identified in the parameter audit (43 params,
no removals possible without breaking PR awslabs#1124 deployments):

1. **Auto-derive `OnDemandEfaInterfaceCount` from `OnDemandInstanceType`**
   (mirrors the pattern `add-cng-p5.yaml`/`add-cng-p6-*.yaml` use for GPU
   families). New default 0 means "auto"; old 1/2 values still work as
   explicit pins so existing stacks pass through unchanged. Auto map:

     hpc8a.96xlarge / hpc7a.{96,48,24,12}xlarge / hpc6id.32xlarge → 2
     anything else (hpc6a.48xlarge, c7i.metal, etc.) → 1

   Removes the 2-NIC HPC tax: hpc8a/hpc7a were the most common EFA
   targets but had to override the parameter; now the default just works.

2. **`VPCName` default → empty, auto-resolves to `${StackName}-VPC`**.
   Multiple deployments in one account now get unique VPC names without
   the user having to pass a different VPCName each time.

3. **ParameterGroup label cleanup**:
     "6. FSx for Lustre (/fsx) + FSx for OpenZFS (/home) (Advanced)"
       → "6. FSx for Lustre (/fsx) and FSx for OpenZFS (/home)"
       (Capacity / LustreDeploymentType / PerUnitStorageThroughput are
       commonly tuned and Region-dependent; not advanced)
     "7. Developer / Advanced (Nested Templates & Monitoring Source)"
       → "7. Developer (do not change for normal use)"

4. **ParameterLabel trims**:
     "Availability Zone (required)" → "Availability Zone"
       (CFN console already shows the asterisk for required params)
     "VPC name" → "VPC name (empty = ${StackName}-VPC)"
     "AMI ID (empty = latest PCS-Ready Deep Learning AMI from SSM; pin in production)"
       → "AMI ID (empty = SSM auto-resolve; pin in production)"
     "dcgm-exporter image (default DCGM 4.5.2 by digest; covers H100/B200/B300)"
       → "dcgm-exporter image (default covers H100/B200/B300)"
     "EFA interface count (1 for hpc6a, 2 for hpc8a/hpc7a/hpc6id)"
       → "EFA interface count (0 = auto from instance type)"

PARAMETERS.md updated to match.

Backward compatibility: every change preserves prior behavior for any
stack that was deployed off the published default — VPCName defaulting
from "ML-Cluster-VPC" to ${StackName}-VPC only affects new launches
where the user did not pass VPCName explicitly. Old stacks keep their
existing VPC name on UPDATE_STACK.

(Larger reductions tracked in subagent plan for the next minor: collapse
CngName/QueueName pairs, drop *MinCount, drop LustreVersion/Compression,
rename Capacity → LustreCapacity for prefix consistency. All of those
are renames or removes, so deferred to avoid breaking PR awslabs#1124 users.)
Adds opt-in multi-user support via OpenLDAP running on the login node.
Default is off (DeployDirectory=false) — existing single-user (ubuntu)
clusters are completely unchanged.

When DeployDirectory=true:
- Login node: installs slapd, stores DB on shared /home/ldap-db
  (OpenZFS), so the directory survives login CNG recreation. Admin
  password is stored in SSM Parameter Store at
  /pcs/<cluster-id>/ldap/admin-password.
- Compute nodes: installs SSSD with ldap provider, discovers login
  node IP via ec2:DescribeInstances tag query, configures NSS/PAM
  so LDAP users are visible as POSIX users on all nodes.
- Home directories auto-created via pam_mkhomedir on shared /home.
- Slurm sees LDAP users transparently via NSS (no Slurm config change).

New parameters on add-cng.yaml:
- DeployDirectory (true/false, default false)
- DirectoryDomainSuffix (e.g. dc=cluster,dc=internal)

New scripts:
- scripts/setup-openldap-server.sh — slapd install + config on login
- scripts/setup-ldap-client.sh — SSSD client config on compute nodes
- scripts/ldap-add-user.sh — helper to add POSIX users to the directory

Not yet wired into deploy-all.yaml or GPU templates (next commit).
Covers enabling multi-user, managing users (add/remove/list/password/groups),
running jobs as specific users, SSH access for LDAP users, UID/GID conventions,
SSSD caching behavior, troubleshooting, data persistence, and upgrade path to
Simple AD / Managed AD.
- setup-openldap-server.sh: add apt-get update before install (PCS-Ready
  DLAMI has limited apt sources cached; without update, slapd package is
  not found)
- setup-ldap-client.sh: same fix
- tests/multi-user-test.md: comprehensive test suite (5 parts: server
  health, user lifecycle, Slurm integration, multi-node consistency,
  resilience). Split from README.md to keep the main test guide concise.
- tests/README.md: added Test 11 link to multi-user-test.md
…+ access methods doc

- setup-ldap-client.sh: change min_id from 10000 to 1001. GID 3000
  (clusterusers) was being filtered as "primary gid out of range" by
  SSSD, preventing user resolution. min_id=1001 excludes only system
  users (0-1000) while allowing all LDAP-provisioned GIDs.

- setup-openldap-server.sh: persist the admin password to
  /home/ldap-db/.admin-password (shared OpenZFS). Previously the
  randomly generated password was lost after the setup script exited,
  making it impossible to administer the directory after login node
  replacement. The file is chmod 600 and on shared storage.

- USER-MANAGEMENT.md: added "Access methods" section comparing SSM vs
  SSH-over-SSM vs Direct SSH (pro/con/best-for table), SSH key
  management approaches, and Slurm accounting with PCS (what's managed
  by PCS vs what admins still need to do manually with sacctmgr).

Verified: getent passwd testuser1 now returns
  testuser1:*:10001:3000:Test User 1:/home/testuser1:/bin/bash
on the login node after SSSD restart.
- pcs-ml-cluster-deploy-all.yaml: add DeployDirectory + DirectoryDomainSuffix
  params (group 7), forward to both LoginNodeGroupStack and OnDemandCNGStack.
  Without this wiring, compute nodes never received DeployDirectory=true and
  SSSD client setup was skipped.
- docs/USER-MANAGEMENT.md: add ASCII art template structure diagram showing
  how DeployDirectory flows from deploy-all through the nested stacks into
  UserData on login (slapd) and compute (SSSD client) nodes.
…nLDAP-LoginNode)

Rename the boolean parameter to an enum that makes it clear both *what*
directory service is used and *where* it runs:

  DirectoryService:
    AllowedValues: [none, OpenLDAP-LoginNode]
    # Future: SimpleAD, ManagedAD, OpenLDAP-External

This prepares for adding Simple AD or external LDAP options in the future
without a breaking rename. The value "OpenLDAP-LoginNode" communicates
that slapd runs on the login node (not an external server).

Also:
- README.md: added template nesting structure diagram (ASCII art)
  showing how DirectoryService flows from deploy-all through login
  (slapd server) and compute (SSSD client) CNGs.
- Fixed DeployMonitoring default accidentally changed to 'none' by
  bulk sed (reverted to 'false' for add-cng, 'true' for deploy-all).
The bulk sed that renamed DeployDirectory→DirectoryService also changed
Default:'false' to Default:'none' on every parameter in both templates.
This broke FSxLustreEnableEfa, OnDemandEnableEfa, DeployPseriesCNG
(deploy-all) and EnableEfa (add-cng) — their AllowedValues are
[true,false] so 'none' is invalid.

Reverted to 'false' for all 4. Only DirectoryService (default 'none')
and MonitoringRole (default 'none') correctly use 'none'.
…ared /home

The admin password was previously written to /home/ldap-db/.admin-password
(shared OpenZFS, readable by root on all nodes). This is a security
concern in multi-user clusters where compute-node users may have sudo.

Now:
- setup-openldap-server.sh first tries SSM PutParameter (SecureString)
  at /pcs/<cluster-id>/ldap/admin-password. SSM is access-controlled
  via IAM and auditable via CloudTrail.
- Falls back to the file only if the instance role lacks ssm:PutParameter
  (with a WARNING in the install log).
- cluster.yaml instance role: broadened the SSM Sid from grafana/* to
  also cover ldap/* (same pattern, same cluster-scoped resource ARN).
- Removed the duplicate SSM put from add-cng.yaml UserData (the script
  now handles it internally with CLUSTER_ID exported from UserData).
- USER-MANAGEMENT.md: added fallback note + clarified SSM is the
  primary storage mechanism.
…up-directory.sh

Single script with role argument (server|client) replaces two separate
scripts. Benefits:
- One URL to manage, no naming inconsistency
- Full picture visible in one file (server + client + shared SSSD config)
- Natural extension point for future SimpleAD (add a case branch)
- Server role now also configures SSSD on the login node itself (so
  getent works locally without a separate manual step)

UserData in add-cng.yaml simplified to:
  curl setup-directory.sh
  if login → bash setup-directory.sh server
  else     → bash setup-directory.sh client

Environment variable interface unchanged (LDAP_DOMAIN_SUFFIX, CLUSTER_ID,
DIRECTORY_DNS_IPS for future SimpleAD). DIRECTORY_DNS_IPS="" triggers
login-IP discovery via ec2:DescribeInstances tag query.

Also: README template structure diagram updated with external scripts
and SSM parameter references.
…ated DirectoryRole param

MonitoringRole is for monitoring exporter placement — reusing it for
directory server/client branching was a semantic mismatch.

New parameter DirectoryRole (none | server | client):
- add-cng.yaml: DirectoryRole param added. UserData passes it directly
  to setup-directory.sh as the role argument. No more MonitoringRole
  reference in the directory block.
- deploy-all.yaml: forwards DirectoryRole='server' to login CNG and
  DirectoryRole='client' to compute CNG (both conditional on
  DirectoryEnabled). New Condition DirectoryEnabled added.
- Removed unused Conditions (IsLoginNode, SetupLdapServer, SetupLdapClient)
  that depended on MonitoringRole for directory logic.

MonitoringRole remains unchanged for its original purpose (monitoring
exporter placement). The two concerns are now fully independent.
Scripts are now fetched from the same S3 bucket as the nested templates
(S3BucketName/S3KeyPrefix + scripts/). This eliminates:
- Hardcoded GitHub raw URLs that need branch-switching for testing
- Dependency on GitHub availability at instance boot time
- Version skew between templates and scripts

Changes:
- add-cng.yaml: UserData uses `aws s3 cp` from the S3 bucket to fetch
  setup-directory.sh. New params S3BucketName + S3KeyPrefix forwarded
  from deploy-all.
- deploy-all.yaml: forwards S3BucketName/S3KeyPrefix to both login and
  compute CNG stacks.
- cluster.yaml: instance role gets s3:GetObject on */templates/scripts/*
  so nodes can fetch scripts at boot.

Development workflow:
  1. Edit scripts/ locally
  2. aws s3 sync scripts/ s3://test-bucket/templates/scripts/
  3. Deploy with S3BucketName=test-bucket
  4. Scripts are fetched from test bucket — no branch URL changes needed

Production: scripts sync alongside templates to the prod bucket.
Future PCS post-install hook: pass the same S3 URL.
…ally

The prod S3 sync pipeline runs:
  aws s3 sync assets/ s3://awsome-distributed-ai/templates/

Previously scripts/ was a sibling directory requiring a separate sync
command (which was never added). Moving scripts under assets/ means the
existing one-line sync deploys both templates and scripts with zero
operational change for maintainers.

Result on S3:
  s3://bucket/templates/pcs-ml-cluster-deploy-all.yaml
  s3://bucket/templates/add-cng.yaml
  s3://bucket/templates/scripts/setup-directory.sh    ← new
  s3://bucket/templates/scripts/install-enroot-pyxis.sh
  s3://bucket/templates/scripts/ldap-add-user.sh

UserData fetches via:
  aws s3 cp s3://${S3BucketName}/${S3KeyPrefix}scripts/setup-directory.sh

README path references updated.
…README

USER-MANAGEMENT.md completely rewritten:
- Quick reference table at top (common tasks → one-liner commands)
- How-it-works diagram (slapd → SSSD → NSS/PAM flow)
- Step-by-step for every operation: add/delete/list users, reset password,
  create groups, batch add, Slurm accounting registration
- Verification steps (how to confirm user is visible on compute nodes)
- Troubleshooting section (common errors + fixes)
- UID/GID convention table
- Data persistence + backup/restore
- Template structure (how DirectoryRole flows)
- Upgrade path to Simple AD

README.md:
- Fixed template diagram: old script names → setup-directory.sh server/client
- Fixed "fetched from GitHub raw" → "fetched from S3"
- Added USER-MANAGEMENT.md + DEPLOY-TESTING.md to Additional Resources
- DeployDirectory → DirectoryService=OpenLDAP-LoginNode in tests/
- scripts/install-enroot-pyxis.sh → assets/scripts/... in template
  descriptions (add-cng*.yaml), ROADMAP, tests/README.md
- PostInstallScriptUrl default URL: .../scripts/... → .../assets/scripts/...
  (file was git-mv'd, GitHub raw URL must match new repo path)
- No remaining references to old paths or old param names
Test 12 (accounting-test.md):
  - Validates Slurm managed accounting with LDAP multi-user
  - Covers: sacctmgr user/account creation, resource limits (GrpTRESRunMins),
    job tracking (sacct), reporting (sreport), fairshare, and
    AccountingPolicyEnforcement behavior (none vs associations,limits,safe)
  - Based on https://aws.amazon.com/blogs/hpc/introducing-managed-accounting-for-aws-parallel-computing-service/

Test 13 (gpu-healthcheck-test.md):
  - Integrates 4.validation_and_observability/2.gpu-cluster-healthcheck suite
  - Lightweight (checks 0-3: nvidia-smi, DCGM L2, EFA enum, topology) ~15 min
  - NCCL multi-node (check 5) with per-instance bandwidth thresholds
  - Slurm prolog integration guide for production
  - When-to-run decision table (deploy, pre-training, steady-state, quarantine)

README: added GPU Health Check link to Additional Resources.
tests/README.md: added Test 12 + Test 13 entries with links.
README.md was too long — reduced from 842 to 73 lines (index + matrix only).
Test procedures moved to per-category files:

  infra-test.md      — Tests 1-3, 8 (monitoring, container runtime, AMI build)
  compute-test.md    — Tests 4-6 (CPU queue, GPU families, NCCL EFA)
  training-test.md   — Test 7 (FSDP Llama-2 7B)
  hpc-efa-test.md    — Test 9 (EFA on CPU HPC, OSU benchmarks)
  storage-test.md    — Test 10 (FSx health + performance regression)
  multi-user-test.md — Tests 11-12 (OpenLDAP + accounting, merged)
  gpu-healthcheck-test.md — Test 13 (GPU health check suite)

Removed accounting-test.md (merged into multi-user-test.md since they
are always tested together).
…ion numbering

Before: 13 top-level sections mixing core usage with advanced features and reference.
After: 11 sections with clear separation:

  §1-7: Core (Key Features → Running a GPU job)
  §8:   Advanced Features
        8.1 Monitoring
        8.2 Pre-baking AMI
        8.3 User Management (new — was standalone §11)
        8.4 IAM Permissions (new)
  §9:   Templates (reference)
  §10:  Testing and Validation (reference)
  §11:  Additional Resources

This makes it clear that §1-7 is the primary user path (deploy + run jobs),
and everything in §8 is opt-in for production/multi-user/security hardening.
…y from cluster core

Before: DeployMonitoring + GrafanaPublicAccessCidr were under "PCS Cluster
Configuration" (which is really scheduler/AMI/accounting settings).

After:
  §2. PCS Cluster Configuration — Slurm version, login instance, AMI, accounting
  §7. Additional Cluster Configuration — Monitoring + Multi-User Directory

This separates "what the cluster IS" (§2) from "what extra features are
layered on top" (§7). Also renamed §6 from "(Advanced)" to just "FSx Storage"
(Capacity/DeploymentType are commonly tuned, not advanced) and §8 to just
"Developer / Advanced" (shorter).
…cAccessCidr→GrafanaAccessCidr

Three changes in deploy-all parameter interface:

1. **SSHAccessCidr** (new, §2 PCS Cluster Configuration):
   CIDR to open port 22 on the login node. Empty = SSH over SSM only.
   For multi-user clusters where users connect via standard SSH.

2. **DeployMonitoring → MonitoringStack** (rename, §7):
   Boolean true/false → enum: 'Prometheus-LoginNode' (default) | 'none'.
   Aligns with DirectoryService naming pattern (<what>-<where>).
   Nested stacks receive DeployMonitoring="true"/"false" via !If [MonitoringEnabled]
   so add-cng.yaml is unchanged (backward compat within the nest).

3. **GrafanaPublicAccessCidr → GrafanaAccessCidr** (rename, §7):
   Shorter name, same semantics. Opens HTTPS/443 on login node.

Security group redesign:
- LoginAccessSecurityGroup now conditionally includes BOTH SSH (from
  SSHAccessCidr) and HTTPS (from GrafanaAccessCidr) rules via !If.
- Created only when at least one CIDR is set (LoginAccessEnabled = OR).
- Attached only to login node via ExtraSecurityGroupId (compute unaffected).

New Conditions: SSHAccessEnabled, GrafanaAccessEnabled, LoginAccessEnabled
(OR of both), MonitoringEnabled.
…esults

- training-test.md: Test 7b Megatron-LM GPT-3 TP/PP/DP on p6-b200 x4
  (~134 TFLOP/s/GPU, lm loss 10.91->10.46, EFA efa-direct 8 nics)
- gpu-healthcheck-test.md: intensive suite EFA loopback PASS on p6-b200;
  caveat that check 5 (NCCL) needs the ECR '#' URI + in-image binary path,
  so validate multi-node NCCL via canonical Test 6
- README.md: matrix row 7,7b + Megatron canonical-asset row
Capture the follow-up to make CapacityReservationId work for targeted ODCRs
(not just Capacity Blocks): a CapacityReservationType enum
(none/capacity-block/targeted-odcr) that branches MarketType + placement group,
with none/capacity-block staying backward compatible. Notes how to verify
without GPU capacity (MinCount=0 launch-template assertion + cheap-type
consumption check).
… slim advanced detail

- §4 Configuration: add DirectoryService / SSHAccessCidr / GrafanaAccessCidr /
  AdditionalSubnetAZ2-3 rows so the new user-facing features are discoverable
- §4: fix HomeThroughput AllowedValues (SINGLE_AZ_HA_1 includes 64)
- §1: correct capacity wording (open ODCR auto-consumed; targeted ODCR on roadmap)
- §9: p6-b300 row shows 17 interfaces (16 EFA + 1 ENA), consistent with §4
- move EFA-on-CPU + FSx-Lustre-EFA into Advanced (§8.5/§8.6); consolidate the
  DCGM version detail into a note at the end of §8.1 Monitoring
- §8.2: note the single-login-node / SPOF constraint for DirectoryService
- §7: replace specific NCCL busbw numbers with a pointer to tests/compute-test.md
- §4: document OnDemandInstanceType + PseriesInstanceType accepted values
…tale diagram memo

- OPERATIONS.md: correct OpenZFS HomeThroughput allowed values (192/384/768 are
  not valid AWS values); SINGLE_AZ_2 groups with HA_2 (160..10240),
  SINGLE_AZ_HA_1/SINGLE_AZ_1 = 64..4096 — matches the template Rules
- USER-MANAGEMENT.md: note that slurm-25.11 paths must become slurm-25.05 when
  deployed with SlurmVersion=25.05
- delete docs/architecture-components.md: an orphan diagram-drafting memo (not
  linked anywhere, the architecture image is already published) with stale
  content (32x EFA on all GPUs, single private subnet, optional Enroot/Pyxis)
…hitecture, slim test memos

README:
- §1 Key Features: add multi-user (OpenLDAP), access control (IAM + SSH/Grafana
  CIDR), and multi-AZ / broad Region coverage
- §2 Architecture: subnets across up to 3 AZs, SSH/Grafana CIDR access, and the
  optional OpenLDAP / IAM add-ons
- §4 Configuration: reorder the parameter table to match the deploy-all console
  parameter-group order (Network → PCS Cluster → On-Demand → GPU → Additional);
  point OnDemandEfaInterfaceCount at §8.5

tests/README:
- drop the one-off 'GPU health check verified on B200' memo and the Mumbai
  'exercised most deeply' paragraph (results live in the per-test files)
- region-coverage GPU column is now a simple ran-on-reserved-capacity check
  (us-east-1/us-east-2/us-west-2 verified via Capacity Block for ML); drop the
  per-instance-type detail from the table
…m earlier rounds

- README §1: remove the 'Multi-AZ and broad Region coverage' Key Features bullet
  (multi-AZ stays documented in §4 / Architecture; not a headline feature)
- tests/README region coverage: clarify the GPU column ✅ reflects earlier GPU
  validation rounds, not necessarily the row's deploy/monitoring verification date
…etail sections

Key Features bullets describe capability + link to the relevant section instead
of embedding parameter names (MonitoringStack/GrafanaAccessCidr/DirectoryService/
SSHAccessCidr). Detailed config stays in §4 and §8.
…s in the table

- replace the two overlapping intro paragraphs with one: only PrimarySubnetAZ is
  required; the table is the most-used subset grouped by the console's parameter
  groups (separator rows), with storage in its own table and the full reference
  in PARAMETERS.md
- keep the in-table group separator rows for discoverability
…move FSx after it

Container Runtime (PostInstallScriptUrl/Args) is rarely touched directly, so it no
longer gets its own console group: folded into '5. Additional Cluster Configuration
(Monitoring, Multi-User, Container Runtime)'. FSx Storage moves after Additional
(now group 6). New order: Network, PCS Cluster, On-Demand, GPU, Additional, FSx,
Developer. README §4 separators and PARAMETERS.md §2b updated to match.
…from §7

- move the PCS-Ready DLAMI 'no AMI build needed' note from Key Features into §2
  Architecture, pointing to §4 'AMI and container runtime' for the detail (dedup)
- §7: link the GPU Cluster Health Check suite as a pre-long-run check
- §8.1 Monitoring: replace the MonitoringRepo/Version/DcgmExporterImage bullets
  with a one-line pointer to PARAMETERS.md (users rarely change these); move the
  monitoring-role vs Name-tag note to the end of §8.1
Replace the single long table (with in-table separator rows and verbose Purpose
cells) with one small table per console parameter group (Network / PCS cluster /
On-Demand / GPU / Additional). Trim each Purpose to one line and move the long
PseriesInstanceType accepted-values list to the GPU compute subsection. Easier to
scan; group structure is obvious from the subheadings.
…with deploy-small tip

- §4 config sub-tables now carry the console group numbers (1 Network … 5 Additional)
- move the FSx deployment-types/sizing content out of §4 into a new §8.1 Storage at
  the top of Advanced Features; renumber the rest (Monitoring 8.2 … FSx-EFA 8.7) and
  fix all §8.x anchor references in README + PARAMETERS.md
- add a 'deploy small, expand after' tip: FSx Lustre/OpenZFS can be grown after create,
  and a smaller filesystem deploys faster — start near minimum Capacity/HomeCapacity and
  expand once the stack is CREATE_COMPLETE
…s, refocus DEPLOY-TESTING, move IAM verification to tests

- remove --profile claude (project test setting) from DEPLOY-TESTING.md and
  tests/infra-test.md; generalize fixed bucket/region to placeholders
- DEPLOY-TESTING.md rewritten for the real audience: a third party deploying
  not-yet-published templates by hosting them in their own S3 bucket and pointing
  S3BucketName/S3KeyPrefix at it
- PARAMETERS.md sections + order now match the deploy-all console parameter groups
  exactly (1 Network, 2 PCS Cluster, 3 On-Demand, 4 GPU, 5 Additional incl.
  monitoring repo/version/dcgm + container runtime, 6 FSx, 7 Developer); fix stale
  anchor links
- move the IAM verification results out of docs/IAM.md into a reproducible
  tests/iam-test.md (representative two-role use case); add the row to tests/README
The standalone §8.7 'FSx for Lustre over EFA (GPUDirect Storage)' is storage content,
so it's now a #### subsection at the end of §8.1 Storage instead of a separate top-level
Advanced feature. Removes the cross-reference note.
…template deploys

Adds an Advanced Features subsection that links docs/DEPLOY-TESTING.md (deploying fork/
branch/PR templates from your own S3 bucket via S3BucketName/S3KeyPrefix). This also
gives DEPLOY-TESTING.md its README entry point — every docs/ file is now linked from the
README.
The '5. Additional' heading said 'container runtime' but the table only listed the
common params (MonitoringStack/GrafanaAccessCidr/DirectoryService). Retitle to
'(monitoring, multi-user)' and add a one-line note that the rarely-changed
monitoring-source + container-runtime params live in the same console group (see
PARAMETERS.md).
The 'console's group 5 also holds…' line was redundant with the PARAMETERS.md
pointer right below it. Group 5 is just the heading + the 3 common params.
…p names

The §4 sub-table headings now use the exact console parameter-group labels
(1. Network Configuration … 5. Additional Cluster Configuration (Monitoring,
Multi-User, Container Runtime)), and group 5 lists PostInstallScriptUrl so the
heading and its rows agree.
The default is no longer the awslabs/main GitHub-raw URL — it's empty, which resolves
to s3://<S3BucketName>/<S3KeyPrefix>scripts/install-enroot-pyxis.sh. Rewrite the caveat:
when testing unpublished templates, point S3BucketName at your own bucket and sync the
scripts there, or first-boot Enroot/Pyxis fetch fails (cluster still CREATE_COMPLETE).
Link DEPLOY-TESTING.md.
Generalize the s3:// verification note and Test 8 build/deploy commands to use
<bucket>/<prefix> placeholders instead of the project's private midaisuk-llm-dev
bucket, matching DEPLOY-TESTING.md.
… docs/CUSTOM-AMI.md

- IAM.md: add 1-click Deploy buttons to the two-role summary table (matching the
  deploy table further down)
- move the §8.5 step-by-step pre-bake procedure into docs/CUSTOM-AMI.md; README §8.5
  keeps a concise summary + Launch button + link (heading text unchanged so the
  awslabs#85-... anchor referenced elsewhere stays valid)
…single location

Match the custom-AMI style: the cluster-admin / cluster-user Deploy buttons now use the
launch-stack.svg image, kept in one place (the 'Deploying the policies' table); drop the
duplicate kbd buttons from the role summary table.
…uttons

Replace the kbd 🚀 buttons in the Templates table with the launch-stack.svg image,
consistent with the Quick Start, §8.5 custom-AMI, and IAM policy Deploy buttons.
Turn the two-role bullet list into a table with per-role Launch-stack Deploy buttons,
consistent with §9 Templates, §8.5, and docs/IAM.md.
…fy NIC wording, flag CB billing

- tests: number the file headings (Tests 11-12 / Test 13 / Test 14) so they line up
  with the test matrix in tests/README.md
- compute-test: '16 cards' → 'the 16 EFA NICs' to avoid confusion with GPU count
- README §4: add a CB-billing ⚠️ to the CapacityReservationId row (links to GPU compute)
# Conflicts:
#	architectures/aws-pcs/tests/README.md
DaisukeMiyamoto pushed a commit that referenced this pull request Jul 6, 2026
…st case on EKS (awslabs#1146)

* feat(dreamzero): add launcher scripts + 14B eval config

* feat(dreamzero): add self-contained two-stage Dockerfile (MIT-0) + docker build helpers

Two-stage EFA-overlay Dockerfile ported from the validated RLinf-on-eks image
(deduplicated venv via upstream embodied-libero, DCP-save hotfix patch applied
at build). Build context helpers (install_extras.sh, run_training_eks.sh, the
DCP patch) live under docker/ (NOT build/, which the repo .gitignore excludes).

* feat(dreamzero): add workflow manifests (RayJob SFT, download, metadata, convert, eval)

* fix(dreamzero): update stale parent-repo path refs in script/config comments

* feat(dreamzero): add optional FSx storage, HF secret, and env_vars templates

* feat(dreamzero): add docker buildx build-push helper

* feat(dreamzero): add kaniko in-cluster build alternative (pending live validation)

Parameterize the Dockerfile stage-2 base (RLINF_UPSTREAM_IMAGE) so the kaniko
two-stage flow can FROM the ECR-pushed stage-1 tag. The docker buildx path
(build-push.sh) remains the validated primary. The kaniko path is structurally
complete and renders valid, but the in-cluster build has NOT yet been run
end-to-end -- README marks it pending live validation.

* feat(dreamzero): add diagrams (incl. rendered infra SVG) + assets via Git LFS

WAM training/inference + infra-topology draw.io sources and rendered SVGs;
rollout mp4 + loss-curve png. Binaries tracked via Git LFS (.gitattributes).
Rendered the previously-missing infra-dreamzero-sft SVG with the draw.io CLI.

* docs(dreamzero): add kubernetes/libero walkthrough README

* docs(dreamzero): add top-level test-case README

* chore(dreamzero): add MIT-0 headers to ported scripts/config; de-couple KubeRay prereq comment

Final self-contained sweep: add SPDX MIT-0 headers to the 5 ported shell
scripts + the eval config (carried over from RLinf-on-eks without headers);
change the SFT manifest's KubeRay-prereq comment from the RLinf-on-eks
terraform reference to the generic helm install. Repo is now FULLY_SELF_CONTAINED.

* fix(dreamzero): correct infra diagram to KubeRay RayJob; drop stale assets

The infra-dreamzero-sft topology diagram depicted the abandoned
StatefulSet + headless Service design (manual `ray start` head election,
CodeBuild build step, leaked namespace), contradicting both the RayJob
manifest and the README prose beside it. Rewrite it to the validated
KubeRay RayJob topology: operator-managed embedded RayCluster (1 head +
1 worker), no manual head election, tool-neutral "build + push to ECR"
step, <NAMESPACE> placeholder, and a bidirectional RayJob<->FSx arrow
showing the /fsx mount (reads model/dataset/metadata, writes DCP
checkpoints). Re-export SVG (white/dark adaptive background to match the
sibling diagrams).

Remove the rollout video and loss curve: the 1-step smoke-run rollout
(success_once=0, expected) reads as broken in a public PR, and the loss
curve is from a pre-refactor DeepSpeed run now disconnected from the
FSDP2 config. README text describes the validated scope instead. Drop
the now-unused assets/*.mp4 and assets/*.png LFS attributes.

* fix(dreamzero): root-cause DCP finalization crash (gloo coordinator PG) + force save_full_model_weights=false

Reproduced, root-caused, and fixed the DreamZero 16B SFT checkpoint crash
end-to-end on 2x p5en (RayJob SUCCEEDED, 207GB DCP, zero errors).

Root cause: on torch 2.6, dcp.save's post-write finalization broadcasts a
multi-MB pickled result object over the default (NCCL) process group on CUDA.
At the end of a long (~209GB / ~20min) checkpoint write this races with NCCL
comm teardown, leaving non-coordinator ranks with an all-zero buffer ->
`_pickle.UnpicklingError: invalid load key '\x00'`, AFTER all 16 shards +
.metadata are already on disk.

Fix (dcp-save-gloo-coordinator.patch): pass a dedicated CPU/gloo process group
to dcp.save(..., process_group=gloo_pg) so the finalization object-broadcast
runs over gloo (CPU), immune to the CUDA/NCCL teardown race -- the same
approach torch 2.7+ takes upstream. Replaces the prior symptom-guard
dcp-save-finalize-besteffort.patch (removed).

Also force +actor.fsdp_config.save_full_model_weights=false in the launcher:
libero_sft_dreamzero_14b.yaml omits the key, so it defaults to True, which on
the 16B model hits "Backend nccl does not support allgather_into_tensor_coalesced"
during the full-state-dict gather. DCP-only + offline convert is the supported path.

Harden the Dockerfile patch-apply loop to be nullglob-safe. Update the
kubernetes/libero README + kaniko-build.yaml patch references accordingly.

* chore(dreamzero): drop unvalidated in-cluster kaniko build path

Remove the experimental kaniko in-cluster build (setup/kaniko-build.yaml)
and all references to it. Every other test case in the repo builds images
with `docker build`/`docker buildx` and pushes to ECR; the closest analog
(openvla-oft LIBERO-on-EKS) uses `docker buildx build --platform linux/amd64`.
dreamzero was the only test case introducing kaniko, and that path was
unvalidated, carried known build bugs, and pinned a `:latest` executor image
(against CONTRIBUTING's "do not use a latest tag" rule).

The validated `docker buildx` path (setup/build-push.sh) is now the sole,
documented build method. kaniko can return in a follow-up PR once
live-validated.

- delete kubernetes/libero/setup/kaniko-build.yaml
- README.md: build statement now points at build-push.sh
- kubernetes/libero/README.md: remove the "Alternative path (kaniko)" block
  and the kaniko layout entry; reword the primary path as the sole path
- Dockerfile: reword RLINF_UPSTREAM_IMAGE comment (kept the ARG; it is a
  generic stage-1 override the buildx path also uses)

* chore(dreamzero): drop stray RLinf-on-EKS references from ported test case

Remove source-repo references that leaked into the upstream port:
- install_extras.sh / eval config: comment wording
- dreamzero-wam{,-inference}.drawio + .svg: drop editor agent metadata

* fix(dreamzero): make Dockerfile buildable + simplify build context

First successful build of the test-case image (validated L1 CodeBuild +
L2 container test, 10/10, on p5en.48xlarge):

- Move RLINF_UPSTREAM_IMAGE ARG to global scope (before the first FROM).
  A per-stage ARG is invisible to a later FROM and resolves blank under
  BuildKit ("base name should not be blank"), which made the Dockerfile
  unbuildable by CodeBuild and docker buildx alike.
- Drop the EXTRAS framework + install_extras.sh: its only value was an
  editable RLinf install into venvs the DreamZero workflow never uses;
  RLinf is imported from cwd (/workspace/RLinf), so the install was dead
  weight. Removes the misleading single-plugin 'extras' abstraction.
- Drop run_training_eks.sh: the generic launcher is unused by this
  DreamZero-only test case (no manifest references it).
- Relocate dcp-save-gloo-coordinator.patch to the test-case root and
  remove the empty docker/scripts/patches/ tree; COPY *.patch instead.
- Correct DCP-fix comments: the sync dcp.save path is byte-identical
  through >= torch 2.8, so the patch is permanently required (the prior
  'torch 2.7+ fixes this' claim was wrong).

* chore(dreamzero): move build-push.sh out of one-file setup/ dir

Relocate kubernetes/libero/setup/build-push.sh -> kubernetes/libero/build-push.sh
and drop the single-file setup/ directory. Matches the sibling
openvla-oft/kubernetes/libero/ layout (helper scripts flat in libero/) and
colocates the script with the env_vars it sources.

- ROOT path ../../.. -> ../.. (one level shallower); verified it still resolves
  to the test-case root (the buildx build context).
- Usage comment: source ../env_vars -> source ./env_vars (now same dir).
- Update references in README.md (root + libero walkthrough) and Dockerfile
  comments.

* docs(dreamzero): correct model/transfer claims to match sources

Align the READMEs with the DreamZero paper, HF model card, and RLinf docs:
- Parameter count in the intros: 16.48B -> 14B (the published headline; the
  README titles already say 14B). The measured ~16.48B instantiated-model
  figure is kept where it matters (FSDP sharding / OOM / VRAM sections).
- Architecture phrasing: drop the unsourced 'shared causal self-attention
  space' for 'causal (autoregressive) ... via flow matching', grounded in the
  CausalWanModel class and the paper.
- LIBERO framing: it is a manipulation benchmark on the same Franka arm as
  DROID, not a new embodiment, so drop 'new embodiment' / 'cross-embodiment
  transfer' and frame the sample around its real purpose (EKS deployment).
- Add the upstream 5B LIBERO-Spatial accuracy (~96.7% success_once at step
  18000) as evidence the recipe converges with sufficient steps.

* docs(dreamzero): note real->sim domain gap + swap-in-your-own-data caveat

Add a closing caveat to the intro: LIBERO is a simulation of the same Franka
Panda arm DROID captures in the real world, so warm-starting onto LIBERO bridges
a real->sim visual domain gap, and in practice users would substitute their own
dataset for task-specific or cross-embodiment fine-tuning. Avoids the term
'negative transfer' (the RLinf docs recommend warm-starting from the released
checkpoint, i.e. the prior is beneficial).

* docs(dreamzero): link root README to detailed architecture; drop unused inference diagram

- Root README Architecture section now links to the full topology + WAM
  component breakdown in kubernetes/libero/README.md#architecture (previously
  just a bare image with no pointer to the detail).
- Remove dreamzero-wam-inference.drawio + .svg: it depicts the paper's
  closed-loop real-time inference path, which this SFT+eval test case does not
  cover. It was referenced by no README in either repo and was already flagged
  for removal in the RLinf-on-eks rearchitecture plan.

* docs(dreamzero): fix broken 1.architectures link + reframe Prerequisites

- Fix the relative path to 1.architectures/4.amazon-eks: it needs ../../../
  (dreamzero -> pytorch -> 3.test_cases -> root), not ../../ which dead-ends in
  3.test_cases/. Both the Prerequisites and References links were wrong; the
  sibling openvla-oft uses the correct depth.
- Reframe Prerequisites around the nodes, not the provisioning mechanism. The
  RayJob is fixed-size (head 1 + worker replicas/min/max = 1), so GPU
  autoscaling is a convenience (on-demand p5en provisioning + scale-down), not a
  requirement -- a static managed node group or Capacity Block works identically.
  The only Karpenter-ism shipped is the karpenter.sh/do-not-disrupt annotation,
  which is ignored on non-Karpenter clusters.

* docs(dreamzero): fix stale buildspec/CodeBuild references to build-push.sh

This test case builds via kubernetes/libero/build-push.sh (docker buildx); it
ships no buildspec.yml. Several comments still referenced CodeBuild / a
buildspec (carryover from the RLinf-on-eks origin):
- Dockerfile: the RLINF_UPSTREAM_IMAGE note and the stage-1 placeholder comment
  now describe the build-push.sh flow (and fix the local-build example's
  BUILD_TARGET: embodied-libero, not embodied-maniskill_libero); the
  'cloned in pre_build / by the buildspec' notes now say build-push.sh.
- dreamzero-eval.yaml: prerequisite note '(examples/buildspec.yml)' -> '(build-push.sh)'.
Also improve the 1.architectures link text and target the section README.

* docs(dreamzero): clarify eval intro is the pipeline run, not an accuracy result

'evaluate the result in the LIBERO simulator and render in-sim rollout videos'
-> 'run the LIBERO simulator eval (which renders in-sim rollout videos)'. The
eval + video path is shipped and validated end-to-end, but a 1-step checkpoint
yields success_once=0.0 (documented in step 6); this avoids implying the intro
promises a meaningful accuracy result.

* docs(dreamzero): reword warm-start rationale for clarity

Replace 'There is no native LIBERO 14B checkpoint upstream -- warm-starting from
DROID is the point' with customer-facing framing: the released DreamZero-DROID
checkpoint is the 14B foundation weight you warm-start from, and continue-SFT
adapts it to your target data (here LIBERO), the same pattern you'd follow with
your own dataset. The old wording used insider framing and was imprecise (a 5B
LIBERO checkpoint does exist upstream; only a 14B LIBERO one does not).

* fix(dreamzero): run SFT launcher from ConfigMap mount, not deleted scripts dir

The RayJob entrypoint copied the launcher to /workspace/eks/scripts/, but that
directory no longer exists in the image: an earlier cleanup changed
'COPY docker/scripts/ /workspace/eks/scripts/' to 'COPY *.patch
/workspace/eks/patches/', so the image now only creates /workspace/eks/patches.
The cp failed with 'No such file or directory' (exit 127) and the RayJob never
started. Run the launcher directly from its read-only ConfigMap mount at
/tmp/scripts (it is path-independent -- it cd's to /workspace/RLinf itself),
with a fallback to /workspace/eks/scripts for images that still bake it in.
Caught by a live multi-step SFT run on 2x p5en.

* docs(dreamzero): bullet the upstream-projects list in libero README

* docs(dreamzero): add real 300-step training-convergence result

Both READMEs' validation-scope sections now report the actual multi-step run on
2x p5en (FSDP2 + KubeRay, the shipped stack): train/loss 0.232 -> 0.085 over 300
steps (~6.9 s/step), a 207 GB DCP checkpoint written with zero UnpicklingError
(exercising the gloo-coordinator fix at the full 16.48B scale). Framed honestly
as 'trains and converges', NOT a task-accuracy claim (300 steps is short; the
released 14B trained for 100K). The 1-step success_once=0.0 note and the
upstream 5B ~96.7% accuracy reference are retained.

* docs(dreamzero): de-jargon the checkpoint-fix mention in validation notes

The validation-scope summaries (near the top of both READMEs) were the first
place a reader met 'UnpicklingError' / 'gloo-coordinator fix', but the
explanation only appears in the troubleshooting table (libero README) and
nowhere in the root README. Reword to plain outcome language -- 'no corruption
or save-time crashes' -- keeping the patch link (and a 'see Troubleshooting'
pointer). UnpicklingError now appears only in the troubleshooting row, where a
reader who hits it would look, with full context. Also made the root README's
patch reference a clickable link.

* docs(dreamzero): link Troubleshooting + explain 14B vs 16.48B

- libero README: make 'see Troubleshooting' a real anchor link (#troubleshooting).
- Reconcile the 14B/16.48B discrepancy that appeared unexplained: '14B' is the
  Wan video-diffusion DiT backbone (headline); '16.48B' is the full trainable
  WAM once the action/state encoders + action head are added (live run reports
  16,484,292,448 params). Added a callout box in the libero README and a concise
  inline gloss in the root README so the two figures are no longer ambiguous.

* docs(dreamzero): reconcile 14B / 16.48B / 23B param figures precisely

The HF model card publishes the checkpoint as '14B' (23B on-disk safetensors
incl. frozen encoders); our live SFT run trains 16,484,292,448 params. Rewrite
the callout to cover all three scopes of the SAME checkpoint and stop implying
the encoders/head are added on top -- they are part of the published WAM:
- 14B  = publisher headline (Wan DiT backbone)
- 16.48B = trainable params when instantiated for full SFT (backbone + the
  model's action/state encoders + action head)
- 23B  = on-disk total (also includes frozen CLIP / UMT5-XXL / Wan VAE)
'14B checkpoint' phrasing is kept (it matches the publisher); root README gloss
tightened to match and point at the callout.

* docs(dreamzero): add non-commercial model-license notice (legal)

Per legal review: the GEAR-Dreams/DreamZero-DROID model is released under a
non-commercial license (CC-BY-NC-4.0). Add a prominent notice at the top of both
READMEs stating that any production use needs additional approvals and that users
should review the license terms before using the model to make an informed
decision. The notice scopes itself to the MODEL only; the repository code remains
MIT-0 (unchanged).

* docs(dreamzero): list optional FSX_* vars in env_vars.example

The storage/ manifests use FSX_SUBNET_ID, FSX_SECURITY_GROUP_IDS,
FSX_FILESYSTEM_ID, FSX_DNS_NAME, and FSX_MOUNT_NAME, but env_vars.example only
listed the always-needed vars. Add the FSX_* vars (commented out, marked
optional -- only for customers who provision FSx via storage/*.yaml rather than
reusing an existing fsx-claim). The README already gives these as inline exports
at the storage step; this just makes the env-var reference complete. No manifest
or behavior change.

* docs(dreamzero): fix comment accuracy + remove personal namespace from examples

Comment audit before PR:
- Remove personal namespace leak: 4 manifests had 'export NAMESPACE=natharno'
  and dreamzero-sft.yaml had 'rlinf' in their usage examples. Normalize all to
  'dreamzero' (matching env_vars.example).
- run_dreamzero_sft_eks.sh: fix the FSx free-space figure (>=200GB -> >=250GB,
  consistent with the READMEs and manifests); consolidate two overlapping
  save_full_model_weights/DCP comment blocks into one accurate block (the first
  vaguely said the gather is 'slow/stalls'; the accurate cause is the NCCL
  allgather_into_tensor_coalesced error); tighten the metadata comment. Net -8
  lines, no behavior change.

* docs(dreamzero): note why static PVC uses storageClassName: ""

Add a comment explaining the empty-string storageClassName on the static FSx PVC
is intentional (disables dynamic provisioning so the PVC binds the named static
PV, rather than the cluster default StorageClass provisioning a new volume).
Prevents a future reader from 'fixing' it by adding a class, which would break
the static bind.

* docs(dreamzero): clarify provisioner vs capacity-backing in Prerequisites

The previous wording listed 'managed node group, Capacity Block reservation, or
Karpenter' as parallel options, conflating two independent axes. Reword to:
provisioner (static managed node group OR Karpenter) backed by a capacity
reservation (Capacity Block for ML OR ODCR) -- either capacity type works with
either provisioner. Drop 'on-demand' as a backing option since 2x p5en
on-demand is effectively unobtainable; these instances come from a reservation.

* docs(dreamzero): reword loss-curve description

'clean, monotonic-ish curve' -> 'steady downward trend' -- more professional and
accurate (the curve declined overall but had step-to-step noise, not strict
monotonicity).

* docs(dreamzero): deep-link the root README to the 14B-vs-16.48B-vs-23B note

The note is a blockquote callout (no auto-generated anchor). Add an explicit
<a id="param-counts"></a> anchor before it in the libero README and turn the
root README's plain-text reference into a deep link
(kubernetes/libero/README.md#param-counts) that jumps straight to it.

* fix(dreamzero): drop no-op network-layer topology constraint + false latency claim

The topologySpreadConstraints used topologyKey topology.k8s.aws/network-node-layer-2,
a label applied only by SageMaker HyperPod -- vanilla EKS nodes (incl. this
cluster) do not carry it, so with whenUnsatisfiable: ScheduleAnyway the constraint
was a no-op. It was also semantically backwards: a spread constraint spreads pods
across domains, while the README claimed it 'prefers co-location ... for lowest
NCCL latency' (co-location would need podAffinity, and the inter-node network
proximity it targets is a HyperPod-only capability here). Remove the dead
constraint from the head + worker (keeping the working podAntiAffinity that pins
one pod per node via kubernetes.io/hostname) and drop the inaccurate README
sentence. Soft scheduling change only; training correctness unaffected.

* docs(dreamzero): spell out acronyms on first use (AWS convention)

Expand technical acronyms at first body use (inline, AWS docs/blog style),
skipping well-known ones (GPU/CPU/CUDA/NVIDIA) and title occurrences:
- Root README: SFT, FSDP, EFA, NCCL, DCP, ECR, DROID.
- libero README: FSDP, VRAM, EFA, NCCL, DCP, RDMA, CSI, VPC, IRSA, PVC, VAE,
  I2V (WAM/DiT/SFT were already spelled out).

* docs(dreamzero): fix libero prereq to not require autoscaling/Karpenter

Prereq #1 still framed 'GPU autoscaling (e.g. Karpenter)' as a requirement,
contradicting the root README fix (commit 1f7ac97). The workload is fixed-size,
so autoscaling is a convenience, not a requirement. Reword heading ('provision'
-> 'schedule') and body to match: nodes can come from a static managed node
group or Karpenter, backed by a Capacity Block for ML or an ODCR.

* fix(dreamzero): recommend HF_TOKEN for rate limits + actually wire it in

'Hugging Face auth is optional' was accurate for access but glossed over a real
reliability issue: anonymous HF downloads are rate-limited per source IP and the
anonymous tier is much stricter than authenticated (HF's #1 cause of 429s). This
download is large/multi-file and egresses through a shared EKS NAT gateway, so
the cluster's other workloads share the same per-IP anonymous quota.

- README: reword the auth note to 'optional, but recommended' and explain the
  per-IP rate-limit / shared-NAT contention; reword the download step to match.
- model-download.yaml: the download script reads os.environ['HF_TOKEN'] but the
  Job never injected it -- the token could not actually reach the container. Add
  an env entry sourcing HF_TOKEN from the hf-token Secret with optional: true
  (Job still runs anonymously if the Secret is absent). Makes the recommendation
  functional.

* docs(dreamzero): fix Step-by-step accuracy (RayJob logs + save_full cause)

Final accuracy pass on the Step-by-step section:
- Step 4: 'kubectl logs job/dreamzero-sft' was wrong -- dreamzero-sft is a
  RayJob, not a batch Job. KubeRay creates a submitter Job by that name whose
  logs only show 'ray job submit' plumbing; the training driver logs are on the
  Ray head pod. Use 'kubectl logs -l ray.io/node-type=head' instead.
- Step 5: align the save_full_model_weights rationale with the accurate cause
  ('Backend nccl does not support allgather_into_tensor_coalesced') instead of
  the vaguer 'rank-0 gather stalls'.
Verified accurate (no change needed): all job/ConfigMap names, the
libero_sft_dreamzero_14b config name, eval LOG_DIR, convert STEP/output path,
the video path, total_num_envs=16, and the dataset/metadata notes.

* docs(dreamzero): fix stale build-push.sh path in .gitignore comment

The script moved out of setup/ to kubernetes/libero/build-push.sh
(commit 58477a8); update the comment to match. Ignore rules unchanged.

* fix(dreamzero): use immutable IMAGE_TAG instead of mutable :latest

A mutable :latest tag makes the pushed image non-reproducible and is the
issue CONTRIBUTING calls out ("do not use a latest tag"). build-push.sh
now defaults to an immutable IMAGE_TAG (dz-${DREAMZERO_REF}) and pushes
${ECR_URI}:${IMAGE_TAG}. The same IMAGE_TAG is threaded through every
manifest's image reference and added to each restricted envsubst
allow-list, plus env_vars.example and the README, so the pushed and
deployed images are guaranteed to match.

* refactor(dreamzero): standardize launcher mount on /opt/scripts

SFT mounted its launcher ConfigMap at /tmp/scripts while convert and
eval use /opt/scripts. Standardize SFT on /opt/scripts (and the matching
entrypoint paths) so the recipe is consistent and easier to copy-adapt
between steps.
DaisukeMiyamoto added a commit that referenced this pull request Jul 8, 2026
…ote #1

The E2E column is the row's primary verification result, so 'Verified' fits
better; the old 'Verified' date column becomes 'Date'. Footnotes renumbered
in table order (1=Verified/E2E, 2=OpenZFS type, 3=instance types).
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant