Repository navigation
fix(docker): chown @prisma so bot can write migrate engine (#1734) - #1735
Conversation
Fixes #1734 (P1: bot crash-loop, 262 restarts, prod down 2026-07-09). production-bot copies node_modules from deps-production (no prisma generate) and the COPY leaves /app/node_modules/@prisma root-owned. Boot CMD 'npx prisma migrate deploy' downloads the schema-engine into node_modules/@prisma/engines at runtime, but the container runs as uid 1001 (bot) -> write denied: Error: Can't write to /app/node_modules/@prisma/engines ... right permissions. Client query-engine is fine (baked at packages/shared/src/generated/prisma via custom output). Only the migrate schema-engine download was blocked. Adding @prisma to the chown lets it write as the bot user. Remove the host docker-compose.override.yml (bot user:0 emergency mitigation) after deploy.
|
Warning Review limit reachedYou’ve reached a temporary PR review limit under our Fair Usage Limits Policy. Next review available in: 36 minutes Your organization has reached its usage spending cap. Adjust your spending cap in the billing tab. How can I continue?After more reviews become available, a review can be triggered using the To avoid repeated limits, reduce automatic review volume by pausing incremental auto-reviews earlier, using label-based review opt-in, excluding WIP or generated PR titles, or requesting reviews manually when the PR is ready. If your team needs uninterrupted high-volume reviews, an organization admin can enable usage-based reviews. How do review limits work?CodeRabbit enforces per-developer PR review limits for each organization. Most developers receive the normal plan review availability. For paid Pro and Pro+ PR reviews, CodeRabbit uses adaptive limits for sustained high-volume activity. When a developer's recent PR review activity reaches the 95th percentile or higher among CodeRabbit users, additional reviews become available more gradually as earlier reviews age out of the rolling window. Please refer docs for additional details. ✨ Finishing Touches🧪 Generate unit tests (beta)
Comment |
There was a problem hiding this comment.
No issues found across 1 file
Auto-approved: One-line chown addition in Dockerfile so bot user owns @prisma directory, fixing runtime schema-engine download and production crash-loop.
Re-trigger cubic
Bakes engines so migrate deploy skips the boot download (#1734).
|
Failed to generate code suggestions for PR |
There was a problem hiding this comment.
0 issues found across 1 file (changes from recent commits).
Auto-approved: Dockerfile fix to bake Prisma engines and chown directory. Low-risk, isolated build change.
Re-trigger cubic
|
🤖 I have created a release *beep* *boop* --- <details><summary>2.33.1</summary> ## [2.33.1](v2.33.0...v2.33.1) (2026-07-09) ### Bug Fixes * **backend:** await session destroy before logout response ([#1721](#1721)) ([b1f0c19](b1f0c19)) * **docker:** chown [@prisma](https://github.com/prisma) so bot can write migrate engine ([#1734](#1734)) ([#1735](#1735)) ([901e0fd](901e0fd)) * update CONTEXT.md reference in domain.md ([#1744](#1744)) ([c92aa26](c92aa26)) * **webhook:** add curl to webhook container ([#1742](#1742)) ([50d9500](50d9500)) </details> --- This PR was generated with [Release Please](https://github.com/googleapis/release-please). See [documentation](https://github.com/googleapis/release-please#release-please).
…ploy pipeline blocker (#1758) ## Problem Discovered while trying to deploy #1757 (the opus playback fix): every deploy currently fails at the migration step. \`scripts/deploy.sh\` runs \`prisma migrate deploy\` via the **backend** image (\`docker_compose run --rm --no-deps backend ...\`), which fails: \`\`\` Error: Can't write to /app/node_modules/@prisma/engines please make sure you install "prisma" with the right permissions. [deploy] ERROR: MIGRATION_FAILED (prisma migrate deploy) \`\`\` ## Root cause \`deps-production\`'s \`npm ci --omit=dev\` ships \`node_modules\` **without** the Prisma migrate schema-engine, since \`prisma\`/\`@prisma/engines\` are devDependencies. The **bot** Dockerfile stage already has a fix for exactly this (#1734/#1735): it copies the full \`@prisma\` from the \`build\` stage and chowns it to the non-root runtime user. The **backend** stage — which is what actually runs migrations — never got the same fix. It only chowns \`/app/packages/backend/dist /app/prisma\`, not \`/app/node_modules/@prisma\`, and never copies the working engines in from \`build\` at all. ## Fix Mirrors the bot stage's existing pattern onto backend: \`\`\`dockerfile COPY --from=build /app/node_modules/@prisma ./node_modules/@prisma ... chown -R backend:nodejs /app/packages/backend/dist /app/prisma /app/node_modules/@prisma \`\`\` ## Verification Built \`production-backend\` in isolation: \`\`\` $ docker run --rm backend-fix-test sh -c 'ls -ld /app/node_modules/@prisma' drwxr-xr-x 1 backend nodejs ... /app/node_modules/@prisma $ docker run --rm backend-fix-test sh -c 'find /app/node_modules/@prisma -iname "*schema-engine*"' -rwxr-xr-x 1 backend nodejs 22280136 ... /app/node_modules/@prisma/engines/schema-engine-linux-musl-openssl-3.0.x $ docker run --rm backend-fix-test sh -lc 'npx prisma migrate status ...' Error: Connection url is empty. # <- only fails on missing DATABASE_URL (expected, no real env in this isolated test) — permission error is gone \`\`\` ## Impact This is currently blocking **all** deploys to the homelab, including the P0 opus-playback fix (#1757, already merged). Needs to merge and deploy ASAP. <!-- This is an auto-generated description by cubic. --> --- ## Summary by cubic Fixes the production backend Docker image to include and chown Prisma engines so `prisma migrate deploy` runs without permission errors. Unblocks the deploy pipeline. - **Bug Fixes** - Copy `@prisma` from the build stage into `/app/node_modules/@prisma` in the backend image and `chown -R backend:nodejs`. - Prevents the write error caused by `npm ci --omit=dev` omitting `@prisma/engines` and running migrations as the non-root `backend` user. <sup>Written for commit ec5a733. Summary will update on new commits.</sup> <a href="https://cubic.dev/pr/LucasSantana-Dev/Lucky/pull/1758?utm_source=github" target="_blank" rel="noopener noreferrer" data-no-image-dialog="true"><picture><source media="(prefers-color-scheme: dark)" srcset="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"><source media="(prefers-color-scheme: light)" srcset="https://www.cubic.dev/buttons/review-in-cubic-light.svg"><img alt="Review in cubic" src="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"></picture></a> <!-- End of auto-generated description by cubic. --> <!-- This is an auto-generated comment: release notes by coderabbit.ai --> ## Summary by CodeRabbit - **Bug Fixes** - Improved production container setup to ensure Prisma database operations run correctly at runtime. - Updated file permissions so backend processes can access required database engine and schema files without elevated privileges. <!-- end of auto-generated comment: release notes by coderabbit.ai -->
A green docker build proves layers stack, not that @discordjs/opus loads. #1735 shipped a Prisma engine that crash-looped 262x behind a passing build. Load the production-bot image and require the native addon for real.
#1838) ## Problem The v2.35.1 deploy **failed in production** (2026-07-16 02:29 UTC) at: ``` Applying migration `20260713000000_support_session_status_enum` Error: P3018 / code 42883 ERROR: operator does not exist: "SupportSessionStatus" = text ``` Production was never impacted (containers never swapped — still healthy on v2.35.0), but the failed migration **wedged all future prod migrations** until manually resolved (done: `migrate resolve --rolled-back` + dropped the orphan enum type; prod `migrate status` = "up to date"). ## Root cause `20260712000000_support_sessions/migration.sql` creates a **partial unique index** whose predicate compares `status` to a text literal: ```sql CREATE UNIQUE INDEX "support_sessions_one_open_per_user" ON "support_sessions"("guildId", "requestorId") WHERE "status" = 'open'; ``` The enum migration drops the CHECK constraint but not this index. `ALTER COLUMN "status" TYPE "SupportSessionStatus"` forces Postgres to re-validate the index predicate against the new type → `SupportSessionStatus = text` → no such operator. **Why nothing caught it:** partial indexes aren't expressible in `schema.prisma`, so Prisma's generator never emitted a drop/recreate and there is no CI drift check on migration-only objects. The original migration even documents this ("migration-only, no CI drift check runs"). Same class as #1735/#1734 — a defect only reachable when the SQL actually runs against a real DB. ## Fix Drop the partial index before the type change, recreate it after (predicate is type-consistent against the enum): ```sql DROP INDEX "support_sessions_one_open_per_user"; CREATE TYPE ... ; ALTER TABLE ... DROP CONSTRAINT ...; ALTER COLUMN ... TYPE ...; CREATE UNIQUE INDEX "support_sessions_one_open_per_user" ... WHERE "status" = 'open'; ``` Editing the migration in place is correct: it never applied successfully anywhere (rolled back in prod, never ran in staging — both still `text`), so there is no checksum drift. ## Verification Applied the **full 45-migration chain** to a scratch `postgres:18-alpine` (matching prod's PG 18.3), with an injection control: | chain | result | |---|---| | **broken** migration | reproduces prod's exact `operator does not exist: "SupportSessionStatus" = text` | | **fixed** migration | 45/45 apply clean; `status` udt = `SupportSessionStatus`; partial index recreated | The harness fails on the bug and passes on the fix — same harness, only the migration content differs. ## Follow-ups (tracked in #1837, not in this PR) - `scripts/deploy.sh:358` — `curl: command not found` in the failure-notification path, so homelab deploy failures present as SHA-mismatch timeouts instead of reporting the real error. This is why the failure cause had to be dug out of the container log. - **Prevention:** wire the scratch-DB full-chain apply above into CI so migration-only SQL defects can't merge again. Closes #1837 <!-- This is an auto-generated description by cubic. --> --- ## Summary by cubic Dropped and recreated the partial unique index `support_sessions_one_open_per_user` around the `support_sessions.status` enum migration to stop the “operator does not exist: SupportSessionStatus = text” error and unblock production migrations. Closes #1837. - **Bug Fixes** - Drop the index before the type change; recreate it after with the same predicate (`WHERE "status" = 'open'`). - Avoids predicate revalidation mismatch when `status` changes from `text` to the enum. - Verified by running the full migration chain on Postgres 18: broken chain fails; fixed chain succeeds. <sup>Written for commit 933d0e4. Summary will update on new commits.</sup> <a href="https://cubic.dev/pr/LucasSantana-Dev/Lucky/pull/1838?utm_source=github" target="_blank" rel="noopener noreferrer" data-no-image-dialog="true"><picture><source media="(prefers-color-scheme: dark)" srcset="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"><source media="(prefers-color-scheme: light)" srcset="https://www.cubic.dev/buttons/review-in-cubic-light.svg"><img alt="Review in cubic" src="https://www.cubic.dev/buttons/review-in-cubic-dark.svg"></picture></a> <!-- End of auto-generated description by cubic. -->



Fixes #1734 — P1 bot crash-loop (262 restarts, prod down 2026-07-09)
Root cause
production-botcopiesnode_modulesfromdeps-production(runsnpm ci, notprisma generate); the COPY leaves/app/node_modules/@prismaroot-owned. Boot CMDnpx prisma migrate deployneeds the schema-engine, downloaded intonode_modules/@prisma/enginesat runtime — container runs as uid 1001 (bot) → write denied → crash-loop.The client query-engine is fine (baked via custom
output = packages/shared/src/generated/prisma). Only the migrate schema-engine download was blocked.Fix
Add
/app/node_modules/@prismato theproduction-botchown → runtime engine download succeeds as the bot user. One line.Cleanup after deploy
Host has an emergency
docker-compose.override.ymlrunning the bot asuser: "0:0"(restored prod during the outage). Remove it once this image deploys so the bot runs as uid 1001.Follow-up (not here)
Bake the schema-engine at build to drop the boot-time Prisma-CDN dependency; add a bot-container liveness dead-man (ran 262 restarts before a human noticed).
Summary by cubic
Bakes
@prismaengines from the build stage into the image and chowns/app/node_modules/@prismato thebotuser sonpx prisma migrate deployruns without a runtime download, permission errors, or a Prisma CDN dependency. Fixes #1734.docker-compose.override.ymlthat runs the bot asuser: "0:0".Written for commit ef8862f. Summary will update on new commits.