Skip to content

feat(dreamzero): add DreamZero LIBERO 14B SFT (World-Action Model) test case on EKS - #1146

Merged
bluecrayon52 merged 45 commits into
mainfrom
feat/dreamzero-test-case
Jun 23, 2026
Merged

bluecrayon52 merged 45 commits into
mainfrom
feat/dreamzero-test-case

Conversation

@bluecrayon52

Copy link
Copy Markdown
Contributor

What

A self-contained test case for continue-SFT (supervised fine-tuning) of the DreamZero 14B World-Action Model (WAM) on the LIBERO manipulation benchmark, on Amazon EKS. It demonstrates multi-node FSDP2 + KubeRay RayJob training over EFA, with sharded DCP (Distributed Checkpoint) saving, an offline DCP→.pt conversion, and in-simulator LIBERO eval.

Everything lives under 3.test_cases/pytorch/dreamzero/ (24 files, +2,762 lines, no changes elsewhere). The image is a two-stage build: the upstream RLinf embodied-libero target + an EFA/NCCL overlay.

Hardware

  • Training: 2× p5en.48xlarge (8× H200 each = 16 GPUs), 16 EFA NICs/node
  • Eval: 1× p5en.48xlarge
  • Shared storage: FSx for Lustre (≥250 GB free; a 14B DCP checkpoint is ~140–206 GB)

Validation status

  • Build (L1/L2): builds via docker buildx (two-stage) and was validated end-to-end through CodeBuild + a 10/10 container test on p5en.48xlarge (CUDA, EFA device + libs, multi-venv, RLinf import, patch-applied invariants, FSx, Ray, OpenMPI).
  • Training: a real 300-step SFT run on 2× p5en.48xlarge reduced train/loss 0.232 → 0.085 (~6.9 s/step) and wrote a 207 GB sharded DCP checkpoint with no corruption or save-time crashes. This demonstrates the pipeline trains and converges — it is not a task-accuracy claim (300 steps is short; the released 14B trained for 100K).
  • Eval: the LIBERO simulator eval runs and renders an in-sim rollout video; eval/success_once = 0.0 is expected for a short run.

Notes for reviewers

  • Non-commercial model license: the GEAR-Dreams/DreamZero-DROID checkpoint is released under CC-BY-NC-4.0 (non-commercial). This is called out prominently at the top of both READMEs; production use needs additional approvals. The repository code remains MIT-0.
  • Git LFS required to view the .svg diagrams (diagrams/*.drawio.svg are LFS pointers).
  • Fully parameterized — envsubst + env vars; no file editing needed to run the walkthrough. No secrets committed (only *.example files).
  • Upstream pins: RLinf b3bbabb1f461, dreamzero/groot ab790c198fbc, and the GEAR-Dreams/DreamZero-DROID checkpoint.
  • Ships dcp-save-gloo-coordinator.patch — a DCP-finalization fix (passes a gloo process group to dcp.save so the post-write object-broadcast avoids the NCCL teardown race). Verified necessary regardless of torch version (the synchronous dcp.save path is unchanged through ≥ torch 2.8).

Layout

  • Dockerfile + dcp-save-gloo-coordinator.patch — two-stage RLinf + EFA overlay image
  • kubernetes/libero/ — the EKS recipe: build-push.sh, model-download.yaml, generate-metadata.yaml, dreamzero-sft.yaml (RayJob), convert-checkpoint.yaml, dreamzero-eval.yaml, launcher scripts, optional FSx storage + HF-token templates, and a full step-by-step README.md
  • diagrams/ — WAM architecture + multi-node infra topology

…cker build helpers

Two-stage EFA-overlay Dockerfile ported from the validated RLinf-on-eks image
(deduplicated venv via upstream embodied-libero, DCP-save hotfix patch applied
at build). Build context helpers (install_extras.sh, run_training_eks.sh, the
DCP patch) live under docker/ (NOT build/, which the repo .gitignore excludes).
…e validation)

Parameterize the Dockerfile stage-2 base (RLINF_UPSTREAM_IMAGE) so the kaniko
two-stage flow can FROM the ECR-pushed stage-1 tag. The docker buildx path
(build-push.sh) remains the validated primary. The kaniko path is structurally
complete and renders valid, but the in-cluster build has NOT yet been run
end-to-end -- README marks it pending live validation.
… Git LFS

WAM training/inference + infra-topology draw.io sources and rendered SVGs;
rollout mp4 + loss-curve png. Binaries tracked via Git LFS (.gitattributes).
Rendered the previously-missing infra-dreamzero-sft SVG with the draw.io CLI.
…le KubeRay prereq comment

Final self-contained sweep: add SPDX MIT-0 headers to the 5 ported shell
scripts + the eval config (carried over from RLinf-on-eks without headers);
change the SFT manifest's KubeRay-prereq comment from the RLinf-on-eks
terraform reference to the generic helm install. Repo is now FULLY_SELF_CONTAINED.
…ssets

The infra-dreamzero-sft topology diagram depicted the abandoned
StatefulSet + headless Service design (manual `ray start` head election,
CodeBuild build step, leaked namespace), contradicting both the RayJob
manifest and the README prose beside it. Rewrite it to the validated
KubeRay RayJob topology: operator-managed embedded RayCluster (1 head +
1 worker), no manual head election, tool-neutral "build + push to ECR"
step, <NAMESPACE> placeholder, and a bidirectional RayJob<->FSx arrow
showing the /fsx mount (reads model/dataset/metadata, writes DCP
checkpoints). Re-export SVG (white/dark adaptive background to match the
sibling diagrams).

Remove the rollout video and loss curve: the 1-step smoke-run rollout
(success_once=0, expected) reads as broken in a public PR, and the loss
curve is from a pre-refactor DeepSpeed run now disconnected from the
FSDP2 config. README text describes the validated scope instead. Drop
the now-unused assets/*.mp4 and assets/*.png LFS attributes.
…G) + force save_full_model_weights=false

Reproduced, root-caused, and fixed the DreamZero 16B SFT checkpoint crash
end-to-end on 2x p5en (RayJob SUCCEEDED, 207GB DCP, zero errors).

Root cause: on torch 2.6, dcp.save's post-write finalization broadcasts a
multi-MB pickled result object over the default (NCCL) process group on CUDA.
At the end of a long (~209GB / ~20min) checkpoint write this races with NCCL
comm teardown, leaving non-coordinator ranks with an all-zero buffer ->
`_pickle.UnpicklingError: invalid load key '\x00'`, AFTER all 16 shards +
.metadata are already on disk.

Fix (dcp-save-gloo-coordinator.patch): pass a dedicated CPU/gloo process group
to dcp.save(..., process_group=gloo_pg) so the finalization object-broadcast
runs over gloo (CPU), immune to the CUDA/NCCL teardown race -- the same
approach torch 2.7+ takes upstream. Replaces the prior symptom-guard
dcp-save-finalize-besteffort.patch (removed).

Also force +actor.fsdp_config.save_full_model_weights=false in the launcher:
libero_sft_dreamzero_14b.yaml omits the key, so it defaults to True, which on
the 16B model hits "Backend nccl does not support allgather_into_tensor_coalesced"
during the full-state-dict gather. DCP-only + offline convert is the supported path.

Harden the Dockerfile patch-apply loop to be nullglob-safe. Update the
kubernetes/libero README + kaniko-build.yaml patch references accordingly.
Remove the experimental kaniko in-cluster build (setup/kaniko-build.yaml)
and all references to it. Every other test case in the repo builds images
with `docker build`/`docker buildx` and pushes to ECR; the closest analog
(openvla-oft LIBERO-on-EKS) uses `docker buildx build --platform linux/amd64`.
dreamzero was the only test case introducing kaniko, and that path was
unvalidated, carried known build bugs, and pinned a `:latest` executor image
(against CONTRIBUTING's "do not use a latest tag" rule).

The validated `docker buildx` path (setup/build-push.sh) is now the sole,
documented build method. kaniko can return in a follow-up PR once
live-validated.

- delete kubernetes/libero/setup/kaniko-build.yaml
- README.md: build statement now points at build-push.sh
- kubernetes/libero/README.md: remove the "Alternative path (kaniko)" block
  and the kaniko layout entry; reword the primary path as the sole path
- Dockerfile: reword RLINF_UPSTREAM_IMAGE comment (kept the ARG; it is a
  generic stage-1 override the buildx path also uses)
… case

Remove source-repo references that leaked into the upstream port:
- install_extras.sh / eval config: comment wording
- dreamzero-wam{,-inference}.drawio + .svg: drop editor agent metadata
First successful build of the test-case image (validated L1 CodeBuild +
L2 container test, 10/10, on p5en.48xlarge):

- Move RLINF_UPSTREAM_IMAGE ARG to global scope (before the first FROM).
  A per-stage ARG is invisible to a later FROM and resolves blank under
  BuildKit ("base name should not be blank"), which made the Dockerfile
  unbuildable by CodeBuild and docker buildx alike.
- Drop the EXTRAS framework + install_extras.sh: its only value was an
  editable RLinf install into venvs the DreamZero workflow never uses;
  RLinf is imported from cwd (/workspace/RLinf), so the install was dead
  weight. Removes the misleading single-plugin 'extras' abstraction.
- Drop run_training_eks.sh: the generic launcher is unused by this
  DreamZero-only test case (no manifest references it).
- Relocate dcp-save-gloo-coordinator.patch to the test-case root and
  remove the empty docker/scripts/patches/ tree; COPY *.patch instead.
- Correct DCP-fix comments: the sync dcp.save path is byte-identical
  through >= torch 2.8, so the patch is permanently required (the prior
  'torch 2.7+ fixes this' claim was wrong).
Relocate kubernetes/libero/setup/build-push.sh -> kubernetes/libero/build-push.sh
and drop the single-file setup/ directory. Matches the sibling
openvla-oft/kubernetes/libero/ layout (helper scripts flat in libero/) and
colocates the script with the env_vars it sources.

- ROOT path ../../.. -> ../.. (one level shallower); verified it still resolves
  to the test-case root (the buildx build context).
- Usage comment: source ../env_vars -> source ./env_vars (now same dir).
- Update references in README.md (root + libero walkthrough) and Dockerfile
  comments.
Align the READMEs with the DreamZero paper, HF model card, and RLinf docs:
- Parameter count in the intros: 16.48B -> 14B (the published headline; the
  README titles already say 14B). The measured ~16.48B instantiated-model
  figure is kept where it matters (FSDP sharding / OOM / VRAM sections).
- Architecture phrasing: drop the unsourced 'shared causal self-attention
  space' for 'causal (autoregressive) ... via flow matching', grounded in the
  CausalWanModel class and the paper.
- LIBERO framing: it is a manipulation benchmark on the same Franka arm as
  DROID, not a new embodiment, so drop 'new embodiment' / 'cross-embodiment
  transfer' and frame the sample around its real purpose (EKS deployment).
- Add the upstream 5B LIBERO-Spatial accuracy (~96.7% success_once at step
  18000) as evidence the recipe converges with sufficient steps.
…veat

Add a closing caveat to the intro: LIBERO is a simulation of the same Franka
Panda arm DROID captures in the real world, so warm-starting onto LIBERO bridges
a real->sim visual domain gap, and in practice users would substitute their own
dataset for task-specific or cross-embodiment fine-tuning. Avoids the term
'negative transfer' (the RLinf docs recommend warm-starting from the released
checkpoint, i.e. the prior is beneficial).
…ed inference diagram

- Root README Architecture section now links to the full topology + WAM
  component breakdown in kubernetes/libero/README.md#architecture (previously
  just a bare image with no pointer to the detail).
- Remove dreamzero-wam-inference.drawio + .svg: it depicts the paper's
  closed-loop real-time inference path, which this SFT+eval test case does not
  cover. It was referenced by no README in either repo and was already flagged
  for removal in the RLinf-on-eks rearchitecture plan.
- Fix the relative path to 1.architectures/4.amazon-eks: it needs ../../../
  (dreamzero -> pytorch -> 3.test_cases -> root), not ../../ which dead-ends in
  3.test_cases/. Both the Prerequisites and References links were wrong; the
  sibling openvla-oft uses the correct depth.
- Reframe Prerequisites around the nodes, not the provisioning mechanism. The
  RayJob is fixed-size (head 1 + worker replicas/min/max = 1), so GPU
  autoscaling is a convenience (on-demand p5en provisioning + scale-down), not a
  requirement -- a static managed node group or Capacity Block works identically.
  The only Karpenter-ism shipped is the karpenter.sh/do-not-disrupt annotation,
  which is ignored on non-Karpenter clusters.
…sh.sh

This test case builds via kubernetes/libero/build-push.sh (docker buildx); it
ships no buildspec.yml. Several comments still referenced CodeBuild / a
buildspec (carryover from the RLinf-on-eks origin):
- Dockerfile: the RLINF_UPSTREAM_IMAGE note and the stage-1 placeholder comment
  now describe the build-push.sh flow (and fix the local-build example's
  BUILD_TARGET: embodied-libero, not embodied-maniskill_libero); the
  'cloned in pre_build / by the buildspec' notes now say build-push.sh.
- dreamzero-eval.yaml: prerequisite note '(examples/buildspec.yml)' -> '(build-push.sh)'.
Also improve the 1.architectures link text and target the section README.
…acy result

'evaluate the result in the LIBERO simulator and render in-sim rollout videos'
-> 'run the LIBERO simulator eval (which renders in-sim rollout videos)'. The
eval + video path is shipped and validated end-to-end, but a 1-step checkpoint
yields success_once=0.0 (documented in step 6); this avoids implying the intro
promises a meaningful accuracy result.
Replace 'There is no native LIBERO 14B checkpoint upstream -- warm-starting from
DROID is the point' with customer-facing framing: the released DreamZero-DROID
checkpoint is the 14B foundation weight you warm-start from, and continue-SFT
adapts it to your target data (here LIBERO), the same pattern you'd follow with
your own dataset. The old wording used insider framing and was imprecise (a 5B
LIBERO checkpoint does exist upstream; only a 14B LIBERO one does not).
…ripts dir

The RayJob entrypoint copied the launcher to /workspace/eks/scripts/, but that
directory no longer exists in the image: an earlier cleanup changed
'COPY docker/scripts/ /workspace/eks/scripts/' to 'COPY *.patch
/workspace/eks/patches/', so the image now only creates /workspace/eks/patches.
The cp failed with 'No such file or directory' (exit 127) and the RayJob never
started. Run the launcher directly from its read-only ConfigMap mount at
/tmp/scripts (it is path-independent -- it cd's to /workspace/RLinf itself),
with a fallback to /workspace/eks/scripts for images that still bake it in.
Caught by a live multi-step SFT run on 2x p5en.
Both READMEs' validation-scope sections now report the actual multi-step run on
2x p5en (FSDP2 + KubeRay, the shipped stack): train/loss 0.232 -> 0.085 over 300
steps (~6.9 s/step), a 207 GB DCP checkpoint written with zero UnpicklingError
(exercising the gloo-coordinator fix at the full 16.48B scale). Framed honestly
as 'trains and converges', NOT a task-accuracy claim (300 steps is short; the
released 14B trained for 100K). The 1-step success_once=0.0 note and the
upstream 5B ~96.7% accuracy reference are retained.
…otes

The validation-scope summaries (near the top of both READMEs) were the first
place a reader met 'UnpicklingError' / 'gloo-coordinator fix', but the
explanation only appears in the troubleshooting table (libero README) and
nowhere in the root README. Reword to plain outcome language -- 'no corruption
or save-time crashes' -- keeping the patch link (and a 'see Troubleshooting'
pointer). UnpicklingError now appears only in the troubleshooting row, where a
reader who hits it would look, with full context. Also made the root README's
patch reference a clickable link.
- libero README: make 'see Troubleshooting' a real anchor link (#troubleshooting).
- Reconcile the 14B/16.48B discrepancy that appeared unexplained: '14B' is the
  Wan video-diffusion DiT backbone (headline); '16.48B' is the full trainable
  WAM once the action/state encoders + action head are added (live run reports
  16,484,292,448 params). Added a callout box in the libero README and a concise
  inline gloss in the root README so the two figures are no longer ambiguous.
The HF model card publishes the checkpoint as '14B' (23B on-disk safetensors
incl. frozen encoders); our live SFT run trains 16,484,292,448 params. Rewrite
the callout to cover all three scopes of the SAME checkpoint and stop implying
the encoders/head are added on top -- they are part of the published WAM:
- 14B  = publisher headline (Wan DiT backbone)
- 16.48B = trainable params when instantiated for full SFT (backbone + the
  model's action/state encoders + action head)
- 23B  = on-disk total (also includes frozen CLIP / UMT5-XXL / Wan VAE)
'14B checkpoint' phrasing is kept (it matches the publisher); root README gloss
tightened to match and point at the callout.
Per legal review: the GEAR-Dreams/DreamZero-DROID model is released under a
non-commercial license (CC-BY-NC-4.0). Add a prominent notice at the top of both
READMEs stating that any production use needs additional approvals and that users
should review the license terms before using the model to make an informed
decision. The notice scopes itself to the MODEL only; the repository code remains
MIT-0 (unchanged).
The storage/ manifests use FSX_SUBNET_ID, FSX_SECURITY_GROUP_IDS,
FSX_FILESYSTEM_ID, FSX_DNS_NAME, and FSX_MOUNT_NAME, but env_vars.example only
listed the always-needed vars. Add the FSX_* vars (commented out, marked
optional -- only for customers who provision FSx via storage/*.yaml rather than
reusing an existing fsx-claim). The README already gives these as inline exports
at the storage step; this just makes the env-var reference complete. No manifest
or behavior change.
…m examples

Comment audit before PR:
- Remove personal namespace leak: 4 manifests had 'export NAMESPACE=natharno'
  and dreamzero-sft.yaml had 'rlinf' in their usage examples. Normalize all to
  'dreamzero' (matching env_vars.example).
- run_dreamzero_sft_eks.sh: fix the FSx free-space figure (>=200GB -> >=250GB,
  consistent with the READMEs and manifests); consolidate two overlapping
  save_full_model_weights/DCP comment blocks into one accurate block (the first
  vaguely said the gather is 'slow/stalls'; the accurate cause is the NCCL
  allgather_into_tensor_coalesced error); tighten the metadata comment. Net -8
  lines, no behavior change.
Add a comment explaining the empty-string storageClassName on the static FSx PVC
is intentional (disables dynamic provisioning so the PVC binds the named static
PV, rather than the cluster default StorageClass provisioning a new volume).
Prevents a future reader from 'fixing' it by adding a class, which would break
the static bind.
…ites

The previous wording listed 'managed node group, Capacity Block reservation, or
Karpenter' as parallel options, conflating two independent axes. Reword to:
provisioner (static managed node group OR Karpenter) backed by a capacity
reservation (Capacity Block for ML OR ODCR) -- either capacity type works with
either provisioner. Drop 'on-demand' as a backing option since 2x p5en
on-demand is effectively unobtainable; these instances come from a reservation.
'clean, monotonic-ish curve' -> 'steady downward trend' -- more professional and
accurate (the curve declined overall but had step-to-step noise, not strict
monotonicity).
…B note

The note is a blockquote callout (no auto-generated anchor). Add an explicit
<a id="param-counts"></a> anchor before it in the libero README and turn the
root README's plain-text reference into a deep link
(kubernetes/libero/README.md#param-counts) that jumps straight to it.
…latency claim

The topologySpreadConstraints used topologyKey topology.k8s.aws/network-node-layer-2,
a label applied only by SageMaker HyperPod -- vanilla EKS nodes (incl. this
cluster) do not carry it, so with whenUnsatisfiable: ScheduleAnyway the constraint
was a no-op. It was also semantically backwards: a spread constraint spreads pods
across domains, while the README claimed it 'prefers co-location ... for lowest
NCCL latency' (co-location would need podAffinity, and the inter-node network
proximity it targets is a HyperPod-only capability here). Remove the dead
constraint from the head + worker (keeping the working podAntiAffinity that pins
one pod per node via kubernetes.io/hostname) and drop the inaccurate README
sentence. Soft scheduling change only; training correctness unaffected.
Expand technical acronyms at first body use (inline, AWS docs/blog style),
skipping well-known ones (GPU/CPU/CUDA/NVIDIA) and title occurrences:
- Root README: SFT, FSDP, EFA, NCCL, DCP, ECR, DROID.
- libero README: FSDP, VRAM, EFA, NCCL, DCP, RDMA, CSI, VPC, IRSA, PVC, VAE,
  I2V (WAM/DiT/SFT were already spelled out).
Prereq #1 still framed 'GPU autoscaling (e.g. Karpenter)' as a requirement,
contradicting the root README fix (commit 1f7ac97). The workload is fixed-size,
so autoscaling is a convenience, not a requirement. Reword heading ('provision'
-> 'schedule') and body to match: nodes can come from a static managed node
group or Karpenter, backed by a Capacity Block for ML or an ODCR.
'Hugging Face auth is optional' was accurate for access but glossed over a real
reliability issue: anonymous HF downloads are rate-limited per source IP and the
anonymous tier is much stricter than authenticated (HF's #1 cause of 429s). This
download is large/multi-file and egresses through a shared EKS NAT gateway, so
the cluster's other workloads share the same per-IP anonymous quota.

- README: reword the auth note to 'optional, but recommended' and explain the
  per-IP rate-limit / shared-NAT contention; reword the download step to match.
- model-download.yaml: the download script reads os.environ['HF_TOKEN'] but the
  Job never injected it -- the token could not actually reach the container. Add
  an env entry sourcing HF_TOKEN from the hf-token Secret with optional: true
  (Job still runs anonymously if the Secret is absent). Makes the recommendation
  functional.
…ause)

Final accuracy pass on the Step-by-step section:
- Step 4: 'kubectl logs job/dreamzero-sft' was wrong -- dreamzero-sft is a
  RayJob, not a batch Job. KubeRay creates a submitter Job by that name whose
  logs only show 'ray job submit' plumbing; the training driver logs are on the
  Ray head pod. Use 'kubectl logs -l ray.io/node-type=head' instead.
- Step 5: align the save_full_model_weights rationale with the accurate cause
  ('Backend nccl does not support allgather_into_tensor_coalesced') instead of
  the vaguer 'rank-0 gather stalls'.
Verified accurate (no change needed): all job/ConfigMap names, the
libero_sft_dreamzero_14b config name, eval LOG_DIR, convert STEP/output path,
the video path, total_num_envs=16, and the dataset/metadata notes.
@bluecrayon52
bluecrayon52 requested a review from mvinci12 June 21, 2026 19:58

@mvinci12 mvinci12 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Thanks for an exceptionally thorough and well-documented contribution — this is one of the more carefully validated test cases I've reviewed.

Strengths

  • Genuinely validated end-to-end. A real 300-step SFT run on 2× p5en.48xlarge (16× H200, FSDP2 full_shard over EFA) reducing train/loss 0.232 → 0.085, a 207 GB sharded DCP checkpoint written cleanly, converted to .pt, and consumed by the LIBERO sim eval. The KubeRay RayJob reaching SUCCEEDED is shown, not asserted.
  • Honest framing. Explicitly not a task-accuracy claim — eval/success_once = 0.0 for a 1-step checkpoint is documented as expected, with the upstream 5B ~96.7%@18k reference provided as the convergence baseline. No fabricated numbers.
  • Real root-cause engineering. The dcp-save-gloo-coordinator.patch (gloo PG for the dcp.save finalization broadcast to dodge the NCCL-teardown race) is correctly diagnosed and the fix is sound; the note that it is not torch-version-specific through ≥2.8 is appreciated.
  • Clean security posture. No secrets committed (*.example files, secretKeyRef, gitignored env_vars), prominent CC-BY-NC-4.0 model-license notice at the top of both READMEs, code remains MIT-0.
  • Consistent with repo conventions. Layout mirrors the sibling openvla-oft/kubernetes/libero/; SPDX headers present; the restricted-envsubst allow-list pattern is correctly documented to protect inline shell vars; set -euo pipefail + PYTHONPATH unbound-guard handled correctly throughout.
  • Excellent troubleshooting table and configuration deep-dive — the num_action_per_block, Hydra + prefix, and save_full_model_weights=false rationales will save the next person hours.

Requesting changes

Three small items before merge (all inline). None are correctness blockers for the validated path, but #2 affects reproducibility and aligns with CONTRIBUTING:

  1. .gitignore — comment references a stale setup/ path.
  2. build-push.sh — pushes :latest; recommend an explicit version tag.
  3. convert-checkpoint.yaml — launcher mount path inconsistent with the other manifests.

Details inline. Thanks again — happy to re-review promptly.

@@ -0,0 +1,3 @@
# Transient build clones created by kubernetes/libero/setup/build-push.sh

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Stale path in the comment: this references kubernetes/libero/setup/build-push.sh, but the script was moved out of setup/ to kubernetes/libero/build-push.sh (commit 58477a8). The ignore rules below (/RLinf/, /DreamZero/) are correct — just update the comment to drop setup/.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 802bdf5 — dropped setup/ from the comment so it points at kubernetes/libero/build-push.sh. Ignore rules left as-is.

DREAMZERO_REPO="${DREAMZERO_REPO:-https://github.com/RLinf/dreamzero.git}"
DREAMZERO_REF="${DREAMZERO_REF:-ab790c198fbc}"
BUILD_TARGET="${BUILD_TARGET:-embodied-libero}"
TAG="${TAG:-latest}"

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

TAG defaults to latest, and stage 2 pushes ${ECR_URI}:${TAG} while every manifest pins image: ${ECR_URI}:latest. A mutable :latest tag makes the image non-reproducible and is the same issue CONTRIBUTING calls out ("do not use a latest tag") — it was part of why the kaniko path was dropped (commit dbb6404). Recommend defaulting to an explicit/immutable version tag (e.g. derived from UPSTREAM_REF/DREAMZERO_REF or a date/semver), and updating the ${ECR_URI}:latest references in the manifests (or documenting a single IMAGE_TAG env that flows into both build-push.sh and the rendered YAML) so the pushed image and the deployed image are guaranteed to match.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in 3262f60 by going the single IMAGE_TAG route. build-push.sh now defaults IMAGE_TAG to an immutable dz-${DREAMZERO_REF} and pushes ${ECR_URI}:${IMAGE_TAG}. The same IMAGE_TAG is threaded into every manifest's image: reference and added to each restricted envsubst allow-list (so ${IMAGE_TAG} gets substituted alongside ${ECR_URI}/${NAMESPACE}), and it's documented in env_vars.example + the README env-var table. No more :latest anywhere — pushed and deployed images are now guaranteed to match.

- name: fsx
mountPath: /fsx
- name: scripts
mountPath: /opt/scripts

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Launcher mount path is inconsistent across the manifests: SFT (dreamzero-sft.yaml) mounts its launcher ConfigMap at /tmp/scripts, while convert and eval use /opt/scripts. It works, but standardizing on one path (and matching command: accordingly) would make the recipe easier to follow and less error-prone when copy-adapting between steps.

Copy link
Copy Markdown
Contributor Author

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Fixed in b14e0ad — standardized on /opt/scripts (the path convert and eval already use, so SFT was the odd one out). Updated the SFT launcher mountPath, both entrypoint references, and the accompanying comment to match.

The script moved out of setup/ to kubernetes/libero/build-push.sh
(commit 58477a8); update the comment to match. Ignore rules unchanged.
A mutable :latest tag makes the pushed image non-reproducible and is the
issue CONTRIBUTING calls out ("do not use a latest tag"). build-push.sh
now defaults to an immutable IMAGE_TAG (dz-${DREAMZERO_REF}) and pushes
${ECR_URI}:${IMAGE_TAG}. The same IMAGE_TAG is threaded through every
manifest's image reference and added to each restricted envsubst
allow-list, plus env_vars.example and the README, so the pushed and
deployed images are guaranteed to match.
SFT mounted its launcher ConfigMap at /tmp/scripts while convert and
eval use /opt/scripts. Standardize SFT on /opt/scripts (and the matching
entrypoint paths) so the recipe is consistent and easier to copy-adapt
between steps.
@bluecrayon52
bluecrayon52 requested a review from mvinci12 June 23, 2026 14:46

@mvinci12 mvinci12 left a comment

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LGTM!

@bluecrayon52
bluecrayon52 merged commit 96acdde into main Jun 23, 2026
5 checks passed
@bluecrayon52
bluecrayon52 deleted the feat/dreamzero-test-case branch June 23, 2026 16:08
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants