From 954c11a2fdfdec8f8e65772cec20c3cf2ce30a2a Mon Sep 17 00:00:00 2001 From: NubsCarson Date: Thu, 2 Jul 2026 12:12:42 +0000 Subject: [PATCH] ci(cloud-cf-deploy): make migrate-db runner configurable to unblock prod deploys when hosted queue is jammed MIME-Version: 1.0 Content-Type: text/plain; charset=UTF-8 Content-Transfer-Encoding: 8bit The migrate-db gate (every deploy job `needs` it) was hardcoded runs-on: ubuntu-latest. When the org-wide GitHub-hosted queue backs up for hours (the recurring zombie-freeze), migrate-db never gets a runner and ALL prod deploys stall — even though the self-hosted hetzner-robot pool is idle (it runs the deploy jobs fine). Today the develop→main promote (#11434, money-gate restore) sat queued 1.5h+ on this exact gate while prod stayed stale. Make it configurable via CLOUD_CF_MIGRATE_RUNNER_JSON (defaults to ["ubuntu-latest"], preserving the deliberate hosted default). Ops can fail migrations over to the robot pool when hosted is jammed. Migrate is a light `bun install` + `db:cloud:migrate`, not the heavy working-tree materialization that motivated keeping deploys off the self-hosted fleet, so the robots handle it safely. actionlint clean. --- .github/workflows/cloud-cf-deploy.yml | 14 +++++++++++--- 1 file changed, 11 insertions(+), 3 deletions(-) diff --git a/.github/workflows/cloud-cf-deploy.yml b/.github/workflows/cloud-cf-deploy.yml index da61e3127e1f2..a0d740aa35f70 100644 --- a/.github/workflows/cloud-cf-deploy.yml +++ b/.github/workflows/cloud-cf-deploy.yml @@ -105,9 +105,17 @@ jobs: migrate-db: name: Run Database Migrations if: github.event_name != 'pull_request' - # GitHub-hosted on purpose: migrations must not inherit the shared - # self-hosted deploy fleet's failure modes (mirrors cloud-deploy-backend). - runs-on: ubuntu-latest + # GitHub-hosted by DEFAULT so migrations don't inherit the shared + # self-hosted deploy fleet's heavy-checkout failure modes (mirrors + # cloud-deploy-backend). But the hosted ubuntu-latest pool periodically + # backs up for hours (org-wide zombie-freeze), and since every deploy job + # `needs: migrate-db`, a starved hosted queue blocks ALL prod deploys even + # while the self-hosted robot pool sits idle. CLOUD_CF_MIGRATE_RUNNER_JSON + # lets ops fail migrations over to the robot pool (e.g. + # ["self-hosted","Linux","X64","hetzner-robot"]) when hosted is jammed; + # migrate is a light `bun install` + `db:cloud:migrate`, not the heavy + # working-tree materialization that motivated keeping deploys off-hosted. + runs-on: ${{ fromJSON(vars.CLOUD_CF_MIGRATE_RUNNER_JSON || '["ubuntu-latest"]') }} # Job-level concurrency groups are repo-wide, so this serializes against # cloud-deploy-backend's migrate-db too: overlapping triggers (a develop # push fires both workflows) can never run the migrator concurrently