Skip to content

fix(docker): revert frontend image to node:22-alpine (unblock docker-publish) - #846

Merged
LucasSantana-Dev merged 1 commit into
mainfrom
fix/dockerfile-frontend-node22
May 13, 2026
Merged

LucasSantana-Dev merged 1 commit into
mainfrom
fix/dockerfile-frontend-node22

Conversation

@LucasSantana-Dev

@LucasSantana-Dev LucasSantana-Dev commented May 13, 2026 •

Copy link
Copy Markdown
Owner

Production-blocker fix

docker-publish.yml has failed every run since 2026-05-10 (when Dependabot #831 bumped Dockerfile.frontend from node:22-alpine → node:26-alpine). v2.10.0's container images never built — the GitHub release shipped a tag but not a deployed image.

Root cause

Dockerfile.frontend's builder stage runs npm ci --legacy-peer-deps in the workspace root, which installs all workspace deps — including @discordjs/opus from the bot workspace.

@discordjs/opus@0.10.0 uses @discordjs/node-pre-gyp@0.4.5 for prebuilt binaries. That version doesn't ship a Node 26 ABI prebuilt. The fallback path would be to compile from source, but node:26-alpine doesn't include the C toolchain, so:

npm error node-pre-gyp ERR! cwd /app/node_modules/@discordjs/opus
npm error node-pre-gyp ERR! node -v v26.1.0
npm error node-pre-gyp ERR! not ok
ERROR: failed to build: failed to solve: process "/bin/sh -c YOUTUBE_DL_SKIP_DOWNLOAD=1 ... npm ci ..."
       did not complete successfully: exit code: 1

Why it slipped past PR #831's CI

PR #831's CI ran unit tests on GitHub-hosted Ubuntu runners (toolchain present). docker-publish runs full image builds post-merge via workflow_run trigger — it never gated #831's merge.

Why this is a 1-line revert (not a Node 26 fix)

Three viable fixes, in increasing scope:

Option Scope Risk
Revert to node:22-alpine (this PR) 1 line None — root Dockerfile already uses 22-alpine via ARG
Add apk add --no-cache make g++ python3 to the frontend builder 4 lines Slower builds (~30s/build), but works on any Node version
Bump @discordjs/opus to a version with Node 26 prebuilts Multiple files No such version exists yet on npm (latest is 0.10.0, published a year ago)

This PR takes Option 1 because it matches the rest of Lucky's images and unblocks shipping immediately.

Follow-up issue worth filing

When @discordjs/opus ships Node 24/26 ABI prebuilts (or migrates to Node-API), revisit Option 3. Until then, Lucky stays on Node 22 LTS.

Related work in this session

Dockerfile.frontend ran npm ci in the workspace root, which pulls in
@discordjs/opus from the bot workspace. On node:26-alpine the prebuilt
opus binaries are missing (node-pre-gyp v0.4.5 doesnt ship Node 26 ABI)
and the alpine image lacks the C toolchain to compile from source, so
docker-publish has failed every run since #831 merged on 2026-05-10.

Reverting just Dockerfile.frontend to node:22-alpine matches the root
Dockerfile (ARG NODE_VERSION=22-alpine). The package.json engines field
still allows >=22 <27, so this is just rolling the container base back
to the version with prebuilt native binaries.

Follow-up issue should track moving to Node 24 alpine once @discordjs/opus
ships matching prebuilts, OR adding the build-base apk package to the
frontend builder stage so native compilation works without prebuilts.

Closes the actual audit RED finding (docker-publish, not deploy.yml).

Refs:
- PR #831 (introduced the bump)
- PR #845 (fixed deploy.yml typo; was a real bug but not THE blocker)
- audit_deep_lucky_2026-05-13.md (RED finding)

Co-Authored-By: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
@vercel

vercel Bot commented May 13, 2026 •

Copy link
Copy Markdown
Contributor

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated (UTC)
lucky Ready Ready Preview, Comment May 13, 2026 9:45pm

Request Review

@coderabbitai

coderabbitai Bot commented May 13, 2026

Copy link
Copy Markdown

Warning

Rate limit exceeded

@LucasSantana-Dev has exceeded the limit for the number of commits that can be reviewed per hour. Please wait 11 minutes and 10 seconds before requesting another review.

You’ve run out of usage credits. Purchase more in the billing tab.

⌛ How to resolve this issue?

After the wait time has elapsed, a review can be triggered using the @coderabbitai review command as a PR comment. Alternatively, push new commits to this PR.

We recommend that you space out your commits to avoid hitting the rate limit.

🚦 How do rate limits work?

CodeRabbit enforces hourly rate limits for each developer per organization.

Our paid plans have higher rate limits than the trial, open-source and free plans. In all cases, we re-allow further reviews after a brief timeout.

Please see our FAQ for further information.

ℹ️ Review info
⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: CHILL

Plan: Pro

Run ID: f890a8ff-eacb-4374-bdf4-1df755c354aa

📥 Commits

Reviewing files that changed from the base of the PR and between 4c14ff2 and 7e0f827.

📒 Files selected for processing (1)
  • Dockerfile.frontend
✨ Finishing Touches
🧪 Generate unit tests (beta)
  • Create PR with unit tests
  • Commit unit tests in branch fix/dockerfile-frontend-node22

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands and usage tips.

@greptile-apps greptile-apps Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

LucasSantana-Dev has reached the 50-review limit for trial accounts. To continue receiving code reviews, upgrade your plan.

@github-actions

Copy link
Copy Markdown

Failed to generate code suggestions for PR

@sonarqubecloud

Copy link
Copy Markdown

@LucasSantana-Dev
LucasSantana-Dev merged commit 03cac3e into main May 13, 2026
20 checks passed
LucasSantana-Dev added a commit that referenced this pull request May 13, 2026
Dockerfile.frontend ran npm ci in the workspace root, which pulls in
@discordjs/opus from the bot workspace. On node:26-alpine the prebuilt
opus binaries are missing (node-pre-gyp v0.4.5 doesnt ship Node 26 ABI)
and the alpine image lacks the C toolchain to compile from source, so
docker-publish has failed every run since #831 merged on 2026-05-10.

Reverting just Dockerfile.frontend to node:22-alpine matches the root
Dockerfile (ARG NODE_VERSION=22-alpine). The package.json engines field
still allows >=22 <27, so this is just rolling the container base back
to the version with prebuilt native binaries.

Follow-up issue should track moving to Node 24 alpine once @discordjs/opus
ships matching prebuilts, OR adding the build-base apk package to the
frontend builder stage so native compilation works without prebuilts.

Closes the actual audit RED finding (docker-publish, not deploy.yml).

Refs:
- PR #831 (introduced the bump)
- PR #845 (fixed deploy.yml typo; was a real bug but not THE blocker)
- audit_deep_lucky_2026-05-13.md (RED finding)

Co-authored-by: Claude Opus 4.7 (1M context) <noreply@anthropic.com>
LucasSantana-Dev added a commit that referenced this pull request May 14, 2026
…s, hardening) (#848)

* fix(docker): exclude worktrees + agent dirs from build context

Build context was including `.worktrees/` (3.4GB), `worktrees/` (707MB),
`.wt-specs/` (41MB), `.claude/` (115MB), and `.agents/` (183MB) — totalling
~4.5GB of duplicated trees and AI agent state shipped to the Docker daemon
on every build. None of this is needed inside any image.

Also exclude `archive/` and `downloads/` (host-only state).

* fix(docker): real bot healthcheck via Redis TCP PING

The previous HEALTHCHECK was `node -e "console.log('Service is running')"`,
which always exits 0 regardless of bot state — orchestrators could never
detect a wedged or disconnected bot.

New check opens a raw TCP socket to Redis ($REDIS_HOST / $REDIS_PORT) and
sends the RESP PING command. Pass = +PONG within 3s; anything else fails.
This confirms (1) node can execute inside the container and (2) the bot's
critical Redis dependency is reachable from this container. No new deps.

Start period bumped 5s → 30s to account for `prisma migrate deploy` running
before the bot process starts.

* fix(docker): add development stage for dev compose

`docker-compose.dev.yml` referenced `target: development` but no such
stage existed in `Dockerfile` — `docker compose -f docker-compose.dev.yml
up --build` would fail with 'failed to find target development'.

New `development` stage derives from `base-runtime` (already has ffmpeg /
opus / yt-dlp), adds native build tools, and runs `tsx watch` via
`npm run dev --workspace=packages/bot`. Compose still bind-mounts host
source over `/app`; node_modules installed on first run to populate the
anonymous volume.

* chore(docker): non-root nginx (UID 101) + healthchecks

Switch `Dockerfile.nginx` and `Dockerfile.frontend` from `nginx:alpine`
(runs as root to bind port 80) to `nginxinc/nginx-unprivileged:1.27-alpine`
(UID 101, listens on 8080 by default — no NET_BIND_SERVICE capability
needed).

Changes:
- nginx confs (frontend + reverse proxy): `listen 80` → `listen 8080`
- nginx reverse-proxy upstream: `http://frontend:80` → `http://frontend:8080`
- Dockerfile.frontend + Dockerfile.nginx: pinned image, EXPOSE 8080, added
  HEALTHCHECK via `wget --spider` (busybox wget ships in the base image).
- docker-compose.yml: port mapping `${NGINX_PORT:-8080}:80` → `:8080`.

DEPLOY ACTION REQUIRED: update `cloudflared/config-lucky.yml` on the
homelab so the tunnel ingress points at `http://nginx:8080` instead of
`http://nginx:80` before merging this PR. Host-side `NGINX_PORT`
default is unchanged (8080).

* chore(docker): add resource limits + env_file fallback

Compose previously had no memory or CPU limits on any service — a runaway
bot or backend process could OOM-kill the homelab host. Adds tiered
defaults via YAML anchors:

  small-svc  (frontend, nginx, webhook, cloudflared)   128m  / 0.25 cpu
  medium-svc (redis, backend)                          512m  / 0.5  cpu
  large-svc  (postgres, bot)                           1g    / 1.0  cpu

Bot + backend also get `env_file: .env` so any var missing from the
explicit `environment:` block falls back to .env at startup. Explicit
entries still take precedence, so behavior is unchanged for vars already
listed.

Cloudflared `user: root` removed — the official image's nonroot default
is sufficient for `tunnel run` with a mounted config dir.

`docker-compose config -q` passes.

* chore(docker): pin webhook base to almir/webhook:2.8.3

The deploy webhook container has `/var/run/docker.sock` bind-mounted in
`docker-compose.yml`, which gives anything inside it effective root on the
host. Pulling `almir/webhook:latest` (last published 2026-02-12) every
rebuild made that surface vulnerable to silent upstream changes.

Pin to `2.8.3` (current latest, identical digest as `latest` at time of
this change). Also fix the inline-comment placement on `USER root` — it
was on the same line which is parsed differently across Docker versions.

* chore(docker): consolidate stages, venv yt-dlp, align frontend dev to node 22

Three small but durable cleanups:

1. Drop the no-op `base-runtime-backend` stage. `production-backend` now
   derives directly from `node:${NODE_VERSION}` — the intermediate stage
   only set WORKDIR, which the production stage already does.

2. Replace `pip3 install --break-system-packages` for yt-dlp with a
   proper PEP-668-compliant venv at `/opt/ytdlp`, symlinked into
   `/usr/local/bin`. Same behavior, no warning suppression, easier to
   audit and upgrade. /root/.cache is cleaned in the same layer.

3. `packages/frontend/Dockerfile.dev` was using `node:24-alpine` while
   production frontend is pinned to `node:22-alpine` (PR #846). Aligned
   to 22 to avoid silent native-module drift. Also removed
   `npm cache clean --force` which defeated the BuildKit cache mount.

* docs(docker): ADR for chore/docker-overhaul

Captures the 8-commit Docker surface overhaul: motivation, decisions per
commit, consequences (including the cloudflared config-lucky.yml port
edit required on the homelab before merge), out-of-scope items, and
revisit triggers.

* fix(docker): address CodeRabbit findings on #848

- deploy/Dockerfile: pin almir/webhook:2.8.3 by manifest digest
  sha256:f77cc91c91d1527b48052280af38b50791e842ad8dd291e9b360a0c13c9ca991.
  Docker Hub tags are mutable by default; the digest makes the supply-chain
  story (this container holds docker.sock) actually defensible.
- docs/decisions/2026-05-13-docker-overhaul.md: remove dangling reference
  to a separate audit doc that was never committed; ADR contains audit
  findings inline.
LucasSantana-Dev added a commit that referenced this pull request May 14, 2026
…imeout (#850)

* fix(docker): exclude worktrees + agent dirs from build context

Build context was including `.worktrees/` (3.4GB), `worktrees/` (707MB),
`.wt-specs/` (41MB), `.claude/` (115MB), and `.agents/` (183MB) — totalling
~4.5GB of duplicated trees and AI agent state shipped to the Docker daemon
on every build. None of this is needed inside any image.

Also exclude `archive/` and `downloads/` (host-only state).

* fix(docker): real bot healthcheck via Redis TCP PING

The previous HEALTHCHECK was `node -e "console.log('Service is running')"`,
which always exits 0 regardless of bot state — orchestrators could never
detect a wedged or disconnected bot.

New check opens a raw TCP socket to Redis ($REDIS_HOST / $REDIS_PORT) and
sends the RESP PING command. Pass = +PONG within 3s; anything else fails.
This confirms (1) node can execute inside the container and (2) the bot's
critical Redis dependency is reachable from this container. No new deps.

Start period bumped 5s → 30s to account for `prisma migrate deploy` running
before the bot process starts.

* fix(docker): add development stage for dev compose

`docker-compose.dev.yml` referenced `target: development` but no such
stage existed in `Dockerfile` — `docker compose -f docker-compose.dev.yml
up --build` would fail with 'failed to find target development'.

New `development` stage derives from `base-runtime` (already has ffmpeg /
opus / yt-dlp), adds native build tools, and runs `tsx watch` via
`npm run dev --workspace=packages/bot`. Compose still bind-mounts host
source over `/app`; node_modules installed on first run to populate the
anonymous volume.

* chore(docker): non-root nginx (UID 101) + healthchecks

Switch `Dockerfile.nginx` and `Dockerfile.frontend` from `nginx:alpine`
(runs as root to bind port 80) to `nginxinc/nginx-unprivileged:1.27-alpine`
(UID 101, listens on 8080 by default — no NET_BIND_SERVICE capability
needed).

Changes:
- nginx confs (frontend + reverse proxy): `listen 80` → `listen 8080`
- nginx reverse-proxy upstream: `http://frontend:80` → `http://frontend:8080`
- Dockerfile.frontend + Dockerfile.nginx: pinned image, EXPOSE 8080, added
  HEALTHCHECK via `wget --spider` (busybox wget ships in the base image).
- docker-compose.yml: port mapping `${NGINX_PORT:-8080}:80` → `:8080`.

DEPLOY ACTION REQUIRED: update `cloudflared/config-lucky.yml` on the
homelab so the tunnel ingress points at `http://nginx:8080` instead of
`http://nginx:80` before merging this PR. Host-side `NGINX_PORT`
default is unchanged (8080).

* chore(docker): add resource limits + env_file fallback

Compose previously had no memory or CPU limits on any service — a runaway
bot or backend process could OOM-kill the homelab host. Adds tiered
defaults via YAML anchors:

  small-svc  (frontend, nginx, webhook, cloudflared)   128m  / 0.25 cpu
  medium-svc (redis, backend)                          512m  / 0.5  cpu
  large-svc  (postgres, bot)                           1g    / 1.0  cpu

Bot + backend also get `env_file: .env` so any var missing from the
explicit `environment:` block falls back to .env at startup. Explicit
entries still take precedence, so behavior is unchanged for vars already
listed.

Cloudflared `user: root` removed — the official image's nonroot default
is sufficient for `tunnel run` with a mounted config dir.

`docker-compose config -q` passes.

* chore(docker): pin webhook base to almir/webhook:2.8.3

The deploy webhook container has `/var/run/docker.sock` bind-mounted in
`docker-compose.yml`, which gives anything inside it effective root on the
host. Pulling `almir/webhook:latest` (last published 2026-02-12) every
rebuild made that surface vulnerable to silent upstream changes.

Pin to `2.8.3` (current latest, identical digest as `latest` at time of
this change). Also fix the inline-comment placement on `USER root` — it
was on the same line which is parsed differently across Docker versions.

* chore(docker): consolidate stages, venv yt-dlp, align frontend dev to node 22

Three small but durable cleanups:

1. Drop the no-op `base-runtime-backend` stage. `production-backend` now
   derives directly from `node:${NODE_VERSION}` — the intermediate stage
   only set WORKDIR, which the production stage already does.

2. Replace `pip3 install --break-system-packages` for yt-dlp with a
   proper PEP-668-compliant venv at `/opt/ytdlp`, symlinked into
   `/usr/local/bin`. Same behavior, no warning suppression, easier to
   audit and upgrade. /root/.cache is cleaned in the same layer.

3. `packages/frontend/Dockerfile.dev` was using `node:24-alpine` while
   production frontend is pinned to `node:22-alpine` (PR #846). Aligned
   to 22 to avoid silent native-module drift. Also removed
   `npm cache clean --force` which defeated the BuildKit cache mount.

* docs(docker): ADR for chore/docker-overhaul

Captures the 8-commit Docker surface overhaul: motivation, decisions per
commit, consequences (including the cloudflared config-lucky.yml port
edit required on the homelab before merge), out-of-scope items, and
revisit triggers.

* fix(docker): address CodeRabbit findings on #848

- deploy/Dockerfile: pin almir/webhook:2.8.3 by manifest digest
  sha256:f77cc91c91d1527b48052280af38b50791e842ad8dd291e9b360a0c13c9ca991.
  Docker Hub tags are mutable by default; the digest makes the supply-chain
  story (this container holds docker.sock) actually defensible.
- docs/decisions/2026-05-13-docker-overhaul.md: remove dangling reference
  to a separate audit doc that was never committed; ADR contains audit
  findings inline.

* feat(reliability): yt-dlp 3× exponential backoff + snapshot restore timeout

- Replace single-shot yt-dlp execution with 3-attempt retry loop using
  exponential backoff (30s → 45s → 60s). Each attempt has its own
  independent timeout; a settled flag prevents double-resolve when the
  process closes after a timeout kill.
- Wrap restoreSnapshot() in Promise.race() with a 2-second deadline so a
  slow or hung DB call cannot block the entire queue restore path. Logs
  a warning and falls through to an empty queue on timeout.

* test(ytdlp): retry behavior coverage + ADR for integration testing strategy

- Add 4 retry-behavior tests: first-attempt success, retry-on-failure,
  all-attempts-exhausted (expects throw), timeout-kills-and-retries.
  Moved discord-player mock to file scope to eliminate jest.resetModules()
  worker crashes.
- Add ADR 2026-05-14: defer in-CI voice integration testing; yt-dlp retry
  + lifecycle snapshot timeout + backend scaffolding are higher ROI.
LucasSantana-Dev added a commit that referenced this pull request May 14, 2026
The standalone `Dockerfile.frontend` re-ran `npm ci` for the entire
workspace just to build the frontend SPA, which:

1. Wasted ~10-12s per build duplicating the dep install.
2. Exposed the frontend image build to the bot's native deps
   (@discordjs/opus). PR #846 hit this when Dependabot bumped to
   `node:26-alpine` — Alpine without a C toolchain couldn't build
   opus from source, breaking `docker-publish` for 4 days.

This refactor merges frontend building into the main `Dockerfile` as
two new stages:

- `build-frontend` — inherits from existing `build` stage (which
  already has `build-base + python3-dev + opus-dev`), copies frontend
  source, runs `npm run build --workspace=packages/frontend`. Native
  compilation works here even if prebuilts are missing for the chosen
  Node major.
- `production-frontend` — `nginxinc/nginx-unprivileged:1.27-alpine`,
  copies dist from `build-frontend`, serves on :8080, includes
  HEALTHCHECK.

Wired through:
- `docker-compose.yml` frontend service → `target: production-frontend`
- `.github/workflows/docker-publish.yml` matrix entry updated
- `docs/DOCKER.md` updated
- `Dockerfile.frontend` deleted

Overrides the 12-month defer on
`docs/decisions/2026-05-13-frontend-dockerfile-keep-separate.md`. ADR
update follows in a separate commit.
LucasSantana-Dev added a commit that referenced this pull request May 14, 2026
The 12-month-defer ADR was hedging against a hypothetical break. PR #846
already fired the hedge's trigger condition. Consolidating now: 5 files,
~30 LOC delta, no behavior change, eliminates the structural cause.
LucasSantana-Dev added a commit that referenced this pull request May 14, 2026
Three linked decisions from `/research-and-decide` composite on chore/docker-overhaul:

1. `2026-05-13-frontend-dockerfile-keep-separate.md` — Keep separate (deferred
   consolidation). 12-month review trigger + 3 hard escalation triggers.

2. `2026-05-13-orchestration-stay-on-compose.md` — Stay on docker-compose v2.
   Explicit acknowledgement of docker.sock blast radius (per critic). 5 hard
   revisit triggers including 90-day observability-data checkpoint.

3. `2026-05-13-base-image-stay-on-alpine.md` — Flipped from Phase-1's
   recommendation (bookworm-slim). Critic surfaced that PR #846's root cause
   was prebuilt-binary availability for @discordjs/opus, not musl vs glibc.
   Migration would not have prevented the break.

Phase-1 research dispatched 3 parallel agents (general-purpose). Phase-2 critic
(Opus) stress-tested all three leaders + identified cross-decision interactions.
All revisit triggers are concrete + measurable.

Branch parks on docs/docker-decision-adrs because PR #848 is open against
release/v2.11.0 and feedback_no_pr_stacking_2026-05-09 limits to 1 PR per base.
LucasSantana-Dev added a commit that referenced this pull request May 14, 2026
* refactor(docker): consolidate Dockerfile.frontend into main Dockerfile

The standalone `Dockerfile.frontend` re-ran `npm ci` for the entire
workspace just to build the frontend SPA, which:

1. Wasted ~10-12s per build duplicating the dep install.
2. Exposed the frontend image build to the bot's native deps
   (@discordjs/opus). PR #846 hit this when Dependabot bumped to
   `node:26-alpine` — Alpine without a C toolchain couldn't build
   opus from source, breaking `docker-publish` for 4 days.

This refactor merges frontend building into the main `Dockerfile` as
two new stages:

- `build-frontend` — inherits from existing `build` stage (which
  already has `build-base + python3-dev + opus-dev`), copies frontend
  source, runs `npm run build --workspace=packages/frontend`. Native
  compilation works here even if prebuilts are missing for the chosen
  Node major.
- `production-frontend` — `nginxinc/nginx-unprivileged:1.27-alpine`,
  copies dist from `build-frontend`, serves on :8080, includes
  HEALTHCHECK.

Wired through:
- `docker-compose.yml` frontend service → `target: production-frontend`
- `.github/workflows/docker-publish.yml` matrix entry updated
- `docs/DOCKER.md` updated
- `Dockerfile.frontend` deleted

Overrides the 12-month defer on
`docs/decisions/2026-05-13-frontend-dockerfile-keep-separate.md`. ADR
update follows in a separate commit.

* docs(adr): supersede frontend-dockerfile-keep-separate

The 12-month-defer ADR was hedging against a hypothetical break. PR #846
already fired the hedge's trigger condition. Consolidating now: 5 files,
~30 LOC delta, no behavior change, eliminates the structural cause.

* refactor(compose): dedupe bot+backend env vars via YAML anchor

Bot and backend share ~25 env vars (NODE_ENV, DISCORD auth, Sentry stack,
DATABASE_URL, Redis config, last.fm, Spotify, webapp secrets). Previously
each service repeated the full list with identical `${VAR:-default}`
patterns.

Extracts `x-common-app-env: &common-app-env` as a top-level mapping;
both services consume it via `environment:\n  <<: *common-app-env` +
service-specific overrides inline.

Trade: converts env from list-of-strings (`- KEY=value`) to mapping
(`KEY: value`). YAML merge keys (`<<:`) require mapping style. Compose
accepts both, no behavior change verified via `docker-compose config`.

LOC delta: docker-compose.yml shrinks ~7% (282→263). The real win is
that adding a new shared var now touches one place instead of two.
LucasSantana-Dev added a commit that referenced this pull request May 14, 2026
…split) (#853)

* docs(adr): research-and-decide outputs for Docker surface (3 ADRs)

Three linked decisions from `/research-and-decide` composite on chore/docker-overhaul:

1. `2026-05-13-frontend-dockerfile-keep-separate.md` — Keep separate (deferred
   consolidation). 12-month review trigger + 3 hard escalation triggers.

2. `2026-05-13-orchestration-stay-on-compose.md` — Stay on docker-compose v2.
   Explicit acknowledgement of docker.sock blast radius (per critic). 5 hard
   revisit triggers including 90-day observability-data checkpoint.

3. `2026-05-13-base-image-stay-on-alpine.md` — Flipped from Phase-1's
   recommendation (bookworm-slim). Critic surfaced that PR #846's root cause
   was prebuilt-binary availability for @discordjs/opus, not musl vs glibc.
   Migration would not have prevented the break.

Phase-1 research dispatched 3 parallel agents (general-purpose). Phase-2 critic
(Opus) stress-tested all three leaders + identified cross-decision interactions.
All revisit triggers are concrete + measurable.

Branch parks on docs/docker-decision-adrs because PR #848 is open against
release/v2.11.0 and feedback_no_pr_stacking_2026-05-09 limits to 1 PR per base.

* docs(adr): mark frontend-dockerfile-keep-separate as superseded by PR #851
@LucasSantana-Dev
LucasSantana-Dev deleted the fix/dockerfile-frontend-node22 branch May 23, 2026 02:21
LucasSantana-Dev added a commit that referenced this pull request Jun 12, 2026
… fallback (#1310)

## What

Closes #1309 — **pipeline blocker**: every backend/bot image build fails
since ~00:28Z today.

## Root cause

The floating `node:22-alpine` base bumped Alpine/musl (musl 1.2.6, node
22.22.3) → `@discordjs/opus`'s expected musl prebuilt name changed to
one that **doesn't exist** (verified 404) → `node-pre-gyp
--fallback-to-build` runs in the `deps-production` stage → fails,
because that stage only had `python3` while the `build` stage carries
the full toolchain. Exact failure class documented inline at Dockerfile
L77-81 from PR #846.

## Fix

Mirror the build stage's toolchain in deps-production: `build-base
python3 python3-dev opus-dev`. Fallback now compiles regardless of
prebuilt availability — robust against future base-image bumps (relates
#847).

## Verification

This PR's own docker-build matrix is the proof — it exercises the
failing stage directly. (A rerun of the failed jobs on #1308 reproduced
the failure, confirming determinism.)

<!-- This is an auto-generated description by cubic. -->
---
## Summary by cubic
Add the C toolchain to the Docker deps-production stage so
`@discordjs/opus` can source-build when its `musl` prebuilt is missing
after the `node:22-alpine` bump. Closes #1309 and unblocks all
backend/bot image builds.

- **Bug Fixes**
- Mirror the build stage toolchain in deps-production: `build-base`,
`python3`, `python3-dev`, `opus-dev`.
- `node-pre-gyp --fallback-to-build` now succeeds even if the prebuilt
404s.

<sup>Written for commit 975c0f5.
Summary will update on new commits.</sup>

<a
href="https://cubic.dev/pr/LucasSantana-Dev/Lucky/pull/1310?utm_source=github"
target="_blank" rel="noopener noreferrer"
data-no-image-dialog="true"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://cubic.dev/buttons/review-in-cubic-dark.svg"><source
media="(prefers-color-scheme: light)"
srcset="https://cubic.dev/buttons/review-in-cubic-light.svg"><img
alt="Review in cubic"
src="https://cubic.dev/buttons/review-in-cubic-dark.svg"></picture></a>

<!-- End of auto-generated description by cubic. -->
LucasSantana-Dev added a commit that referenced this pull request Jul 12, 2026
…#1809)

## Root cause (corrected — it was NOT the BuildKit cache)
Every Docker `Build` job failed deterministically:
```
npm error code EUSAGE
npm error Invalid: lock file's eslint@10.6.0 does not satisfy eslint@10.7.0
```
**Reproduced locally with `npm@12.0.1`** (CI installs `npm@12.0.0`). npm
12's `npm ci` rejects a lockfile whose resolved version isn't the **max
satisfying the range** once a newer in-range version publishes.
`eslint@10.7.0` published; range is `^10.4.0`, lock pinned `10.6.0` →
rejected. Local npm 10.9.4 is lenient (why `npm ci` passed locally); a
fresh BuildKit cache (`v5-`) still failed, disproving the cache theory.

## Fix
`npm update eslint --package-lock-only` → lock's eslint `10.6.0 →
10.7.0`. **package-lock.json only**, no `package.json` range change,
41-line diff. Verified: `npm@12.0.1 ci --dry-run` passes.

## Scope note
Minimal by design — unblocks release #1775 + the whole queue now.
Deliberately does NOT bundle the toolchain-modernization asks
(npm→latest, node→latest, TS7): a node base-image bump is high-risk here
(@discordjs/opus source-compiles; node:26-alpine broke #846) and belongs
in its own validated PR. Recurrence (next in-range publish re-breaks npm
12 ci) is a process fix — Renovate lockfile-maintenance — tracked
separately.
LucasSantana-Dev added a commit that referenced this pull request Jul 12, 2026
## Why
Systemic follow-up to the eslint lock break (#1809). npm 12's strict
`npm ci` rejects a lockfile whose resolved version isn't the max
satisfying the range once a newer in-range version publishes. Root
prevention = keep the lock fresh.

## Changes
1. **Renovate `lockFileMaintenance`** (weekly, auto-merge) — regenerates
`package-lock.json` to max-in-range so strict `npm ci` never sees a
stale lock. This is the recurrence fix (`config:recommended` does not
enable it).
2. **Dockerfile: `npm@12.0.0` → `npm@latest`** (3 stages) per your 'npm
always latest stable'.

## ⚠ Tradeoff for your call
`npm@latest` floats — builds are **non-reproducible** (a future npm
major could change behavior mid-builds) and it keeps the strict-ci
behavior. If you'd rather have reproducibility, say so and I'll pin
exact (`npm@12.0.1`) and let Renovate keep it current — the repo's
existing pin-and-Renovate pattern. Lockfile-maintenance still has a
residual <1-week window between an in-range publish and the Monday run.

## Deferred to their own validated PRs (not bundled — too risky for a
hotfix-adjacent change)
- **node 24 → latest**: `@discordjs/opus` source-compiles per ABI; the
Dockerfile documents `node:26-alpine` **broke #846**. Also **CI jobs run
node 22 while Docker builds node 24** — reconcile in that PR.
- **TypeScript v7**: folds into the existing in-flight TS7 migration.

## Dependency
Stacks on **#1809** (eslint lock fix). This branch's Docker Build
inherits main's stale lock until #1809 merges; rebase after and it goes
green.

<!-- This is an auto-generated description by cubic. -->
---
## Summary by cubic
Keeps the lockfile fresh and switches all Docker stages to `npm@latest`
to prevent strict `npm ci` failures when newer in-range versions
publish.

- **Dependencies**
- Enable Renovate `lockFileMaintenance` with a Monday-before-6am
schedule and auto-merge (incl. platform) to keep `package-lock.json` at
max-in-range.
- Install `npm@latest` in all Docker stages. Note: floating `npm`
reduces build reproducibility.

<sup>Written for commit 0b9f0a9.
Summary will update on new commits.</sup>

<a
href="https://cubic.dev/pr/LucasSantana-Dev/Lucky/pull/1810?utm_source=github"
target="_blank" rel="noopener noreferrer"
data-no-image-dialog="true"><picture><source
media="(prefers-color-scheme: dark)"
srcset="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"><source
media="(prefers-color-scheme: light)"
srcset="https://www.cubic.dev/buttons/review-in-cubic-light.svg"><img
alt="Review in cubic"
src="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"></picture></a>

<!-- End of auto-generated description by cubic. -->

This branch was successfully deployed

1 active deployment
Preview — 7e0f8274 Deployed May 13, 2026 by vercel[bot]
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant