feat(dreamzero): add DreamZero LIBERO 14B SFT (World-Action Model) test case on EKS - #1146
Conversation
…cker build helpers Two-stage EFA-overlay Dockerfile ported from the validated RLinf-on-eks image (deduplicated venv via upstream embodied-libero, DCP-save hotfix patch applied at build). Build context helpers (install_extras.sh, run_training_eks.sh, the DCP patch) live under docker/ (NOT build/, which the repo .gitignore excludes).
…ta, convert, eval)
…e validation) Parameterize the Dockerfile stage-2 base (RLINF_UPSTREAM_IMAGE) so the kaniko two-stage flow can FROM the ECR-pushed stage-1 tag. The docker buildx path (build-push.sh) remains the validated primary. The kaniko path is structurally complete and renders valid, but the in-cluster build has NOT yet been run end-to-end -- README marks it pending live validation.
… Git LFS WAM training/inference + infra-topology draw.io sources and rendered SVGs; rollout mp4 + loss-curve png. Binaries tracked via Git LFS (.gitattributes). Rendered the previously-missing infra-dreamzero-sft SVG with the draw.io CLI.
…le KubeRay prereq comment Final self-contained sweep: add SPDX MIT-0 headers to the 5 ported shell scripts + the eval config (carried over from RLinf-on-eks without headers); change the SFT manifest's KubeRay-prereq comment from the RLinf-on-eks terraform reference to the generic helm install. Repo is now FULLY_SELF_CONTAINED.
…ssets The infra-dreamzero-sft topology diagram depicted the abandoned StatefulSet + headless Service design (manual `ray start` head election, CodeBuild build step, leaked namespace), contradicting both the RayJob manifest and the README prose beside it. Rewrite it to the validated KubeRay RayJob topology: operator-managed embedded RayCluster (1 head + 1 worker), no manual head election, tool-neutral "build + push to ECR" step, <NAMESPACE> placeholder, and a bidirectional RayJob<->FSx arrow showing the /fsx mount (reads model/dataset/metadata, writes DCP checkpoints). Re-export SVG (white/dark adaptive background to match the sibling diagrams). Remove the rollout video and loss curve: the 1-step smoke-run rollout (success_once=0, expected) reads as broken in a public PR, and the loss curve is from a pre-refactor DeepSpeed run now disconnected from the FSDP2 config. README text describes the validated scope instead. Drop the now-unused assets/*.mp4 and assets/*.png LFS attributes.
…G) + force save_full_model_weights=false Reproduced, root-caused, and fixed the DreamZero 16B SFT checkpoint crash end-to-end on 2x p5en (RayJob SUCCEEDED, 207GB DCP, zero errors). Root cause: on torch 2.6, dcp.save's post-write finalization broadcasts a multi-MB pickled result object over the default (NCCL) process group on CUDA. At the end of a long (~209GB / ~20min) checkpoint write this races with NCCL comm teardown, leaving non-coordinator ranks with an all-zero buffer -> `_pickle.UnpicklingError: invalid load key '\x00'`, AFTER all 16 shards + .metadata are already on disk. Fix (dcp-save-gloo-coordinator.patch): pass a dedicated CPU/gloo process group to dcp.save(..., process_group=gloo_pg) so the finalization object-broadcast runs over gloo (CPU), immune to the CUDA/NCCL teardown race -- the same approach torch 2.7+ takes upstream. Replaces the prior symptom-guard dcp-save-finalize-besteffort.patch (removed). Also force +actor.fsdp_config.save_full_model_weights=false in the launcher: libero_sft_dreamzero_14b.yaml omits the key, so it defaults to True, which on the 16B model hits "Backend nccl does not support allgather_into_tensor_coalesced" during the full-state-dict gather. DCP-only + offline convert is the supported path. Harden the Dockerfile patch-apply loop to be nullglob-safe. Update the kubernetes/libero README + kaniko-build.yaml patch references accordingly.
Remove the experimental kaniko in-cluster build (setup/kaniko-build.yaml) and all references to it. Every other test case in the repo builds images with `docker build`/`docker buildx` and pushes to ECR; the closest analog (openvla-oft LIBERO-on-EKS) uses `docker buildx build --platform linux/amd64`. dreamzero was the only test case introducing kaniko, and that path was unvalidated, carried known build bugs, and pinned a `:latest` executor image (against CONTRIBUTING's "do not use a latest tag" rule). The validated `docker buildx` path (setup/build-push.sh) is now the sole, documented build method. kaniko can return in a follow-up PR once live-validated. - delete kubernetes/libero/setup/kaniko-build.yaml - README.md: build statement now points at build-push.sh - kubernetes/libero/README.md: remove the "Alternative path (kaniko)" block and the kaniko layout entry; reword the primary path as the sole path - Dockerfile: reword RLINF_UPSTREAM_IMAGE comment (kept the ARG; it is a generic stage-1 override the buildx path also uses)
… case
Remove source-repo references that leaked into the upstream port:
- install_extras.sh / eval config: comment wording
- dreamzero-wam{,-inference}.drawio + .svg: drop editor agent metadata
First successful build of the test-case image (validated L1 CodeBuild +
L2 container test, 10/10, on p5en.48xlarge):
- Move RLINF_UPSTREAM_IMAGE ARG to global scope (before the first FROM).
A per-stage ARG is invisible to a later FROM and resolves blank under
BuildKit ("base name should not be blank"), which made the Dockerfile
unbuildable by CodeBuild and docker buildx alike.
- Drop the EXTRAS framework + install_extras.sh: its only value was an
editable RLinf install into venvs the DreamZero workflow never uses;
RLinf is imported from cwd (/workspace/RLinf), so the install was dead
weight. Removes the misleading single-plugin 'extras' abstraction.
- Drop run_training_eks.sh: the generic launcher is unused by this
DreamZero-only test case (no manifest references it).
- Relocate dcp-save-gloo-coordinator.patch to the test-case root and
remove the empty docker/scripts/patches/ tree; COPY *.patch instead.
- Correct DCP-fix comments: the sync dcp.save path is byte-identical
through >= torch 2.8, so the patch is permanently required (the prior
'torch 2.7+ fixes this' claim was wrong).
Relocate kubernetes/libero/setup/build-push.sh -> kubernetes/libero/build-push.sh and drop the single-file setup/ directory. Matches the sibling openvla-oft/kubernetes/libero/ layout (helper scripts flat in libero/) and colocates the script with the env_vars it sources. - ROOT path ../../.. -> ../.. (one level shallower); verified it still resolves to the test-case root (the buildx build context). - Usage comment: source ../env_vars -> source ./env_vars (now same dir). - Update references in README.md (root + libero walkthrough) and Dockerfile comments.
Align the READMEs with the DreamZero paper, HF model card, and RLinf docs: - Parameter count in the intros: 16.48B -> 14B (the published headline; the README titles already say 14B). The measured ~16.48B instantiated-model figure is kept where it matters (FSDP sharding / OOM / VRAM sections). - Architecture phrasing: drop the unsourced 'shared causal self-attention space' for 'causal (autoregressive) ... via flow matching', grounded in the CausalWanModel class and the paper. - LIBERO framing: it is a manipulation benchmark on the same Franka arm as DROID, not a new embodiment, so drop 'new embodiment' / 'cross-embodiment transfer' and frame the sample around its real purpose (EKS deployment). - Add the upstream 5B LIBERO-Spatial accuracy (~96.7% success_once at step 18000) as evidence the recipe converges with sufficient steps.
…veat Add a closing caveat to the intro: LIBERO is a simulation of the same Franka Panda arm DROID captures in the real world, so warm-starting onto LIBERO bridges a real->sim visual domain gap, and in practice users would substitute their own dataset for task-specific or cross-embodiment fine-tuning. Avoids the term 'negative transfer' (the RLinf docs recommend warm-starting from the released checkpoint, i.e. the prior is beneficial).
…ed inference diagram - Root README Architecture section now links to the full topology + WAM component breakdown in kubernetes/libero/README.md#architecture (previously just a bare image with no pointer to the detail). - Remove dreamzero-wam-inference.drawio + .svg: it depicts the paper's closed-loop real-time inference path, which this SFT+eval test case does not cover. It was referenced by no README in either repo and was already flagged for removal in the RLinf-on-eks rearchitecture plan.
- Fix the relative path to 1.architectures/4.amazon-eks: it needs ../../../ (dreamzero -> pytorch -> 3.test_cases -> root), not ../../ which dead-ends in 3.test_cases/. Both the Prerequisites and References links were wrong; the sibling openvla-oft uses the correct depth. - Reframe Prerequisites around the nodes, not the provisioning mechanism. The RayJob is fixed-size (head 1 + worker replicas/min/max = 1), so GPU autoscaling is a convenience (on-demand p5en provisioning + scale-down), not a requirement -- a static managed node group or Capacity Block works identically. The only Karpenter-ism shipped is the karpenter.sh/do-not-disrupt annotation, which is ignored on non-Karpenter clusters.
…sh.sh This test case builds via kubernetes/libero/build-push.sh (docker buildx); it ships no buildspec.yml. Several comments still referenced CodeBuild / a buildspec (carryover from the RLinf-on-eks origin): - Dockerfile: the RLINF_UPSTREAM_IMAGE note and the stage-1 placeholder comment now describe the build-push.sh flow (and fix the local-build example's BUILD_TARGET: embodied-libero, not embodied-maniskill_libero); the 'cloned in pre_build / by the buildspec' notes now say build-push.sh. - dreamzero-eval.yaml: prerequisite note '(examples/buildspec.yml)' -> '(build-push.sh)'. Also improve the 1.architectures link text and target the section README.
…acy result 'evaluate the result in the LIBERO simulator and render in-sim rollout videos' -> 'run the LIBERO simulator eval (which renders in-sim rollout videos)'. The eval + video path is shipped and validated end-to-end, but a 1-step checkpoint yields success_once=0.0 (documented in step 6); this avoids implying the intro promises a meaningful accuracy result.
Replace 'There is no native LIBERO 14B checkpoint upstream -- warm-starting from DROID is the point' with customer-facing framing: the released DreamZero-DROID checkpoint is the 14B foundation weight you warm-start from, and continue-SFT adapts it to your target data (here LIBERO), the same pattern you'd follow with your own dataset. The old wording used insider framing and was imprecise (a 5B LIBERO checkpoint does exist upstream; only a 14B LIBERO one does not).
…ripts dir The RayJob entrypoint copied the launcher to /workspace/eks/scripts/, but that directory no longer exists in the image: an earlier cleanup changed 'COPY docker/scripts/ /workspace/eks/scripts/' to 'COPY *.patch /workspace/eks/patches/', so the image now only creates /workspace/eks/patches. The cp failed with 'No such file or directory' (exit 127) and the RayJob never started. Run the launcher directly from its read-only ConfigMap mount at /tmp/scripts (it is path-independent -- it cd's to /workspace/RLinf itself), with a fallback to /workspace/eks/scripts for images that still bake it in. Caught by a live multi-step SFT run on 2x p5en.
Both READMEs' validation-scope sections now report the actual multi-step run on 2x p5en (FSDP2 + KubeRay, the shipped stack): train/loss 0.232 -> 0.085 over 300 steps (~6.9 s/step), a 207 GB DCP checkpoint written with zero UnpicklingError (exercising the gloo-coordinator fix at the full 16.48B scale). Framed honestly as 'trains and converges', NOT a task-accuracy claim (300 steps is short; the released 14B trained for 100K). The 1-step success_once=0.0 note and the upstream 5B ~96.7% accuracy reference are retained.
…otes The validation-scope summaries (near the top of both READMEs) were the first place a reader met 'UnpicklingError' / 'gloo-coordinator fix', but the explanation only appears in the troubleshooting table (libero README) and nowhere in the root README. Reword to plain outcome language -- 'no corruption or save-time crashes' -- keeping the patch link (and a 'see Troubleshooting' pointer). UnpicklingError now appears only in the troubleshooting row, where a reader who hits it would look, with full context. Also made the root README's patch reference a clickable link.
- libero README: make 'see Troubleshooting' a real anchor link (#troubleshooting). - Reconcile the 14B/16.48B discrepancy that appeared unexplained: '14B' is the Wan video-diffusion DiT backbone (headline); '16.48B' is the full trainable WAM once the action/state encoders + action head are added (live run reports 16,484,292,448 params). Added a callout box in the libero README and a concise inline gloss in the root README so the two figures are no longer ambiguous.
The HF model card publishes the checkpoint as '14B' (23B on-disk safetensors incl. frozen encoders); our live SFT run trains 16,484,292,448 params. Rewrite the callout to cover all three scopes of the SAME checkpoint and stop implying the encoders/head are added on top -- they are part of the published WAM: - 14B = publisher headline (Wan DiT backbone) - 16.48B = trainable params when instantiated for full SFT (backbone + the model's action/state encoders + action head) - 23B = on-disk total (also includes frozen CLIP / UMT5-XXL / Wan VAE) '14B checkpoint' phrasing is kept (it matches the publisher); root README gloss tightened to match and point at the callout.
Per legal review: the GEAR-Dreams/DreamZero-DROID model is released under a non-commercial license (CC-BY-NC-4.0). Add a prominent notice at the top of both READMEs stating that any production use needs additional approvals and that users should review the license terms before using the model to make an informed decision. The notice scopes itself to the MODEL only; the repository code remains MIT-0 (unchanged).
The storage/ manifests use FSX_SUBNET_ID, FSX_SECURITY_GROUP_IDS, FSX_FILESYSTEM_ID, FSX_DNS_NAME, and FSX_MOUNT_NAME, but env_vars.example only listed the always-needed vars. Add the FSX_* vars (commented out, marked optional -- only for customers who provision FSx via storage/*.yaml rather than reusing an existing fsx-claim). The README already gives these as inline exports at the storage step; this just makes the env-var reference complete. No manifest or behavior change.
…m examples Comment audit before PR: - Remove personal namespace leak: 4 manifests had 'export NAMESPACE=natharno' and dreamzero-sft.yaml had 'rlinf' in their usage examples. Normalize all to 'dreamzero' (matching env_vars.example). - run_dreamzero_sft_eks.sh: fix the FSx free-space figure (>=200GB -> >=250GB, consistent with the READMEs and manifests); consolidate two overlapping save_full_model_weights/DCP comment blocks into one accurate block (the first vaguely said the gather is 'slow/stalls'; the accurate cause is the NCCL allgather_into_tensor_coalesced error); tighten the metadata comment. Net -8 lines, no behavior change.
Add a comment explaining the empty-string storageClassName on the static FSx PVC is intentional (disables dynamic provisioning so the PVC binds the named static PV, rather than the cluster default StorageClass provisioning a new volume). Prevents a future reader from 'fixing' it by adding a class, which would break the static bind.
…ites The previous wording listed 'managed node group, Capacity Block reservation, or Karpenter' as parallel options, conflating two independent axes. Reword to: provisioner (static managed node group OR Karpenter) backed by a capacity reservation (Capacity Block for ML OR ODCR) -- either capacity type works with either provisioner. Drop 'on-demand' as a backing option since 2x p5en on-demand is effectively unobtainable; these instances come from a reservation.
'clean, monotonic-ish curve' -> 'steady downward trend' -- more professional and accurate (the curve declined overall but had step-to-step noise, not strict monotonicity).
…B note The note is a blockquote callout (no auto-generated anchor). Add an explicit <a id="param-counts"></a> anchor before it in the libero README and turn the root README's plain-text reference into a deep link (kubernetes/libero/README.md#param-counts) that jumps straight to it.
…latency claim The topologySpreadConstraints used topologyKey topology.k8s.aws/network-node-layer-2, a label applied only by SageMaker HyperPod -- vanilla EKS nodes (incl. this cluster) do not carry it, so with whenUnsatisfiable: ScheduleAnyway the constraint was a no-op. It was also semantically backwards: a spread constraint spreads pods across domains, while the README claimed it 'prefers co-location ... for lowest NCCL latency' (co-location would need podAffinity, and the inter-node network proximity it targets is a HyperPod-only capability here). Remove the dead constraint from the head + worker (keeping the working podAntiAffinity that pins one pod per node via kubernetes.io/hostname) and drop the inaccurate README sentence. Soft scheduling change only; training correctness unaffected.
Expand technical acronyms at first body use (inline, AWS docs/blog style), skipping well-known ones (GPU/CPU/CUDA/NVIDIA) and title occurrences: - Root README: SFT, FSDP, EFA, NCCL, DCP, ECR, DROID. - libero README: FSDP, VRAM, EFA, NCCL, DCP, RDMA, CSI, VPC, IRSA, PVC, VAE, I2V (WAM/DiT/SFT were already spelled out).
Prereq #1 still framed 'GPU autoscaling (e.g. Karpenter)' as a requirement, contradicting the root README fix (commit 1f7ac97). The workload is fixed-size, so autoscaling is a convenience, not a requirement. Reword heading ('provision' -> 'schedule') and body to match: nodes can come from a static managed node group or Karpenter, backed by a Capacity Block for ML or an ODCR.
'Hugging Face auth is optional' was accurate for access but glossed over a real reliability issue: anonymous HF downloads are rate-limited per source IP and the anonymous tier is much stricter than authenticated (HF's #1 cause of 429s). This download is large/multi-file and egresses through a shared EKS NAT gateway, so the cluster's other workloads share the same per-IP anonymous quota. - README: reword the auth note to 'optional, but recommended' and explain the per-IP rate-limit / shared-NAT contention; reword the download step to match. - model-download.yaml: the download script reads os.environ['HF_TOKEN'] but the Job never injected it -- the token could not actually reach the container. Add an env entry sourcing HF_TOKEN from the hf-token Secret with optional: true (Job still runs anonymously if the Secret is absent). Makes the recommendation functional.
…ause)
Final accuracy pass on the Step-by-step section:
- Step 4: 'kubectl logs job/dreamzero-sft' was wrong -- dreamzero-sft is a
RayJob, not a batch Job. KubeRay creates a submitter Job by that name whose
logs only show 'ray job submit' plumbing; the training driver logs are on the
Ray head pod. Use 'kubectl logs -l ray.io/node-type=head' instead.
- Step 5: align the save_full_model_weights rationale with the accurate cause
('Backend nccl does not support allgather_into_tensor_coalesced') instead of
the vaguer 'rank-0 gather stalls'.
Verified accurate (no change needed): all job/ConfigMap names, the
libero_sft_dreamzero_14b config name, eval LOG_DIR, convert STEP/output path,
the video path, total_num_envs=16, and the dataset/metadata notes.
mvinci12
left a comment
There was a problem hiding this comment.
Thanks for an exceptionally thorough and well-documented contribution — this is one of the more carefully validated test cases I've reviewed.
Strengths
- Genuinely validated end-to-end. A real 300-step SFT run on 2×
p5en.48xlarge(16× H200, FSDP2full_shardover EFA) reducingtrain/loss0.232 → 0.085, a 207 GB sharded DCP checkpoint written cleanly, converted to.pt, and consumed by the LIBERO sim eval. The KubeRay RayJob reachingSUCCEEDEDis shown, not asserted. - Honest framing. Explicitly not a task-accuracy claim —
eval/success_once = 0.0for a 1-step checkpoint is documented as expected, with the upstream 5B ~96.7%@18k reference provided as the convergence baseline. No fabricated numbers. - Real root-cause engineering. The
dcp-save-gloo-coordinator.patch(gloo PG for thedcp.savefinalization broadcast to dodge the NCCL-teardown race) is correctly diagnosed and the fix is sound; the note that it is not torch-version-specific through ≥2.8 is appreciated. - Clean security posture. No secrets committed (
*.examplefiles,secretKeyRef, gitignoredenv_vars), prominent CC-BY-NC-4.0 model-license notice at the top of both READMEs, code remains MIT-0. - Consistent with repo conventions. Layout mirrors the sibling
openvla-oft/kubernetes/libero/; SPDX headers present; the restricted-envsubstallow-list pattern is correctly documented to protect inline shell vars;set -euo pipefail+PYTHONPATHunbound-guard handled correctly throughout. - Excellent troubleshooting table and configuration deep-dive — the
num_action_per_block, Hydra+prefix, andsave_full_model_weights=falserationales will save the next person hours.
Requesting changes
Three small items before merge (all inline). None are correctness blockers for the validated path, but #2 affects reproducibility and aligns with CONTRIBUTING:
.gitignore— comment references a stalesetup/path.build-push.sh— pushes:latest; recommend an explicit version tag.convert-checkpoint.yaml— launcher mount path inconsistent with the other manifests.
Details inline. Thanks again — happy to re-review promptly.
| @@ -0,0 +1,3 @@ | |||
| # Transient build clones created by kubernetes/libero/setup/build-push.sh | |||
There was a problem hiding this comment.
Stale path in the comment: this references kubernetes/libero/setup/build-push.sh, but the script was moved out of setup/ to kubernetes/libero/build-push.sh (commit 58477a8). The ignore rules below (/RLinf/, /DreamZero/) are correct — just update the comment to drop setup/.
There was a problem hiding this comment.
Fixed in 802bdf5 — dropped setup/ from the comment so it points at kubernetes/libero/build-push.sh. Ignore rules left as-is.
| DREAMZERO_REPO="${DREAMZERO_REPO:-https://github.com/RLinf/dreamzero.git}" | ||
| DREAMZERO_REF="${DREAMZERO_REF:-ab790c198fbc}" | ||
| BUILD_TARGET="${BUILD_TARGET:-embodied-libero}" | ||
| TAG="${TAG:-latest}" |
There was a problem hiding this comment.
TAG defaults to latest, and stage 2 pushes ${ECR_URI}:${TAG} while every manifest pins image: ${ECR_URI}:latest. A mutable :latest tag makes the image non-reproducible and is the same issue CONTRIBUTING calls out ("do not use a latest tag") — it was part of why the kaniko path was dropped (commit dbb6404). Recommend defaulting to an explicit/immutable version tag (e.g. derived from UPSTREAM_REF/DREAMZERO_REF or a date/semver), and updating the ${ECR_URI}:latest references in the manifests (or documenting a single IMAGE_TAG env that flows into both build-push.sh and the rendered YAML) so the pushed image and the deployed image are guaranteed to match.
There was a problem hiding this comment.
Fixed in 3262f60 by going the single IMAGE_TAG route. build-push.sh now defaults IMAGE_TAG to an immutable dz-${DREAMZERO_REF} and pushes ${ECR_URI}:${IMAGE_TAG}. The same IMAGE_TAG is threaded into every manifest's image: reference and added to each restricted envsubst allow-list (so ${IMAGE_TAG} gets substituted alongside ${ECR_URI}/${NAMESPACE}), and it's documented in env_vars.example + the README env-var table. No more :latest anywhere — pushed and deployed images are now guaranteed to match.
| - name: fsx | ||
| mountPath: /fsx | ||
| - name: scripts | ||
| mountPath: /opt/scripts |
There was a problem hiding this comment.
Launcher mount path is inconsistent across the manifests: SFT (dreamzero-sft.yaml) mounts its launcher ConfigMap at /tmp/scripts, while convert and eval use /opt/scripts. It works, but standardizing on one path (and matching command: accordingly) would make the recipe easier to follow and less error-prone when copy-adapting between steps.
There was a problem hiding this comment.
Fixed in b14e0ad — standardized on /opt/scripts (the path convert and eval already use, so SFT was the odd one out). Updated the SFT launcher mountPath, both entrypoint references, and the accompanying comment to match.
The script moved out of setup/ to kubernetes/libero/build-push.sh (commit 58477a8); update the comment to match. Ignore rules unchanged.
A mutable :latest tag makes the pushed image non-reproducible and is the
issue CONTRIBUTING calls out ("do not use a latest tag"). build-push.sh
now defaults to an immutable IMAGE_TAG (dz-${DREAMZERO_REF}) and pushes
${ECR_URI}:${IMAGE_TAG}. The same IMAGE_TAG is threaded through every
manifest's image reference and added to each restricted envsubst
allow-list, plus env_vars.example and the README, so the pushed and
deployed images are guaranteed to match.
SFT mounted its launcher ConfigMap at /tmp/scripts while convert and eval use /opt/scripts. Standardize SFT on /opt/scripts (and the matching entrypoint paths) so the recipe is consistent and easier to copy-adapt between steps.
What
A self-contained test case for continue-SFT (supervised fine-tuning) of the DreamZero 14B World-Action Model (WAM) on the LIBERO manipulation benchmark, on Amazon EKS. It demonstrates multi-node FSDP2 + KubeRay RayJob training over EFA, with sharded DCP (Distributed Checkpoint) saving, an offline DCP→
.ptconversion, and in-simulator LIBERO eval.Everything lives under
3.test_cases/pytorch/dreamzero/(24 files, +2,762 lines, no changes elsewhere). The image is a two-stage build: the upstream RLinfembodied-liberotarget + an EFA/NCCL overlay.Hardware
p5en.48xlarge(8× H200 each = 16 GPUs), 16 EFA NICs/nodep5en.48xlargeValidation status
docker buildx(two-stage) and was validated end-to-end through CodeBuild + a 10/10 container test onp5en.48xlarge(CUDA, EFA device + libs, multi-venv, RLinf import, patch-applied invariants, FSx, Ray, OpenMPI).p5en.48xlargereducedtrain/loss0.232 → 0.085 (~6.9 s/step) and wrote a 207 GB sharded DCP checkpoint with no corruption or save-time crashes. This demonstrates the pipeline trains and converges — it is not a task-accuracy claim (300 steps is short; the released 14B trained for 100K).eval/success_once = 0.0is expected for a short run.Notes for reviewers
GEAR-Dreams/DreamZero-DROIDcheckpoint is released under CC-BY-NC-4.0 (non-commercial). This is called out prominently at the top of both READMEs; production use needs additional approvals. The repository code remains MIT-0..svgdiagrams (diagrams/*.drawio.svgare LFS pointers).envsubst+ env vars; no file editing needed to run the walkthrough. No secrets committed (only*.examplefiles).b3bbabb1f461, dreamzero/grootab790c198fbc, and theGEAR-Dreams/DreamZero-DROIDcheckpoint.dcp-save-gloo-coordinator.patch— a DCP-finalization fix (passes a gloo process group todcp.saveso the post-write object-broadcast avoids the NCCL teardown race). Verified necessary regardless of torch version (the synchronousdcp.savepath is unchanged through ≥ torch 2.8).Layout
Dockerfile+dcp-save-gloo-coordinator.patch— two-stage RLinf + EFA overlay imagekubernetes/libero/— the EKS recipe:build-push.sh,model-download.yaml,generate-metadata.yaml,dreamzero-sft.yaml(RayJob),convert-checkpoint.yaml,dreamzero-eval.yaml, launcher scripts, optional FSx storage + HF-token templates, and a full step-by-stepREADME.mddiagrams/— WAM architecture + multi-node infra topology