fix(scripts): nuke-and-rebuild self-bootstraps templates; add E2E test - #2122
Merged
hongmingwang-moleculeai merged 1 commit intoApr 26, 2026
Merged
Conversation
HongmingWang-Rabbit
requested a review
from hongmingwang-moleculeai
as a code owner
April 26, 2026 21:22
5 tasks
Two paper cuts the fix addresses:
1. nuke-and-rebuild.sh wipes the compose stack but never re-populates
workspace-configs-templates/, org-templates/, or plugins/. Those dirs
are .gitignored — the curated set lives in manifest.json as external
repos cloned via clone-manifest.sh (idempotent). Without that step,
a fresh checkout or a post-deletion run leaves the dirs empty, which
silently hides the entire template palette in Canvas + falls back to
bare default workspace provisioning. Symptom: "Deploy your first
agent" shows zero templates.
2. The existing ws-* container reap was already in the script (good),
but it only fires when this script runs. Folks running `docker compose
down -v` directly leave orphan ws-* containers behind. Documented
that explicitly in the script comment so future readers understand
why those lines are critical.
The fix is just `bash clone-manifest.sh` added to the script. clone-
manifest.sh is idempotent — populated dirs short-circuit, so a re-nuke
on a healthy machine pays only a few stat calls.
scripts/test-nuke-and-rebuild.sh exercises the canonical workflow end-
to-end:
- plants a fake orphan ws-* container, then asserts it gets reaped
- renames the manifest dirs to simulate a fresh checkout, then
asserts they get repopulated
- waits for /health and asserts the platform sees the same template
count on disk as via /configs in the container (catches bind-mount
drift)
- asserts the image-auto-refresh watcher (PR #2114) starts, since
that's load-bearing for the CD chain users now rely on
The test pre-flights port 5432/6379/8080 and exits 0 with a SKIP
message if a non-target compose project is holding them — common when
parallel monorepo checkouts coexist on one Docker daemon.
scripts/ is intentionally outside CI shellcheck per ci.yml comment, but
both files pass `shellcheck --severity=warning` anyway.
Defers but does not solve the runtime root-cause for orphan ws-* after
plain `docker compose down -v`: the orphan-sweeper in the platform only
reaps containers whose workspace row says status='removed', so a wiped
DB → no row → sweeper ignores them. Proper fix needs container labels
keyed to a per-platform-instance UUID so the sweeper can confidently
reap "containers I provisioned that aren't in my DB anymore" without
nuking a sibling platform's containers on a shared daemon. Tracked as
task #109's follow-up; out of scope for this PR.
Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
HongmingWang-Rabbit
force-pushed
the
fix/nuke-and-rebuild-self-bootstraps
branch
from
April 26, 2026 21:37
d118d54 to
44d0444
Compare
HongmingWang-Rabbit
pushed a commit
that referenced
this pull request
Jun 12, 2026
PR #2122 makes workspace Pause/Resume cascade opt-in via ?cascade=true. Without this parameter, pausing/resuming a workspace with descendants returns 409 Conflict, breaking the Canvas ContextMenu and batch-pause flows that currently depend on implicit cascade behavior. Add ?cascade=true to all Canvas callers so the existing UX behavior is preserved when #2122 lands. Refs: #2122 /sop-ack
HongmingWang-Rabbit
pushed a commit
that referenced
this pull request
Jun 12, 2026
…via ?cascade=true' (#2122) from fix/pause-resume-cascade-opt-in-1991 into main
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
What this fixes
Two paper cuts surfaced today during a real `bash scripts/nuke-and-rebuild.sh` flow:
What's in the PR
Pre-flight check
The test exits `0` with a clear SKIP message when ports 5432/6379/8080 are held by a non-target compose project (common when parallel monorepo checkouts coexist on one Docker daemon).
Test plan
What this PR does NOT solve (deliberately)
Plain `docker compose down -v` (without invoking the canonical script) still leaves `ws-*` orphans behind. The runtime root cause is in `internal/registry/orphan_sweeper.go` — it only reaps containers whose workspace row has `status='removed'`. A wiped DB has no row at all, so the sweeper ignores them.
Proper fix needs container labels keyed to a per-platform-instance UUID so the sweeper can reap "containers I provisioned that aren't in my DB anymore" without nuking a sibling platform's containers on a shared daemon. Tracked as a follow-up under task #109; deferred from this PR to keep scope tight.
🤖 Generated with Claude Code