Skip to content

ci(docker): fail fast when the omni-build runner is offline - #14701

Merged
diegosouzapw merged 3 commits into
diegosouzapw:release/v3.8.52from
gonisulaimann:fix/docker-publish-runner-preflight
Oct 7, 2026
Merged

diegosouzapw merged 3 commits into
diegosouzapw:release/v3.8.52from
gonisulaimann:fix/docker-publish-runner-preflight

Conversation

@gonisulaimann

@gonisulaimann gonisulaimann commented Sep 23, 2026 •

Copy link
Copy Markdown
Contributor

Closes #14560.

timeout-minutes cannot fix the silent publish stall because the clock only starts once a runner picks the job up — queue time is unbounded. So the fix is a gate before the build instead of a timeout around it:

  • runner-preflight job (needs prepare, 5 min, hosted): when USE_VPS_RUNNER=true, lists repo runners via the Actions API and fails the run if no online runner carries the omni-build label. The error message names the problem and the fallbacks (restart the self-hosted listener, or USE_VPS_RUNNER=false — noting hosted dies ResourceExhausted per ci: docker-publish OOM (ResourceExhausted) on every release/v3.8.51 push — :next stale since 08-22 and the v3.8.51 tag will not build on the hosted runner #11976). A publish that used to queue forever now goes red in minutes with the verdict in the job name.
  • timeout-minutes: 350 on the build job as a backstop for a wedged build step (Docker/BuildKit hang) — ~2h50m observed build + headroom. Deliberately not sized for queue time since it can't bound that.
  • Nightly runner-alert (schedule + workflow_dispatch): mirrors the nightly-release-green issue loop — opens a deduplicated tracking issue when the omni-build pool has zero online runners, comments with fresh evidence while it persists, auto-closes on recovery. Publish jobs skip themselves on the schedule event (prepare/preflight/build/merge all gate on github.event_name != 'schedule'), so the cron exists only to drive the alert.

Verified: workflow YAML parses, all five jobs' if/needs edges re-checked after the schedule addition (a skipped prepare yields empty outputs.skip, so the build/merge guards needed the explicit event gate — added), preflight jq filter tested locally against the API shape.


⚠️ base-red inherited: #14547 — the failing checks (API Route Typecheck, Docs Gates, Unit fast-path shards) also fail on release/v3.8.51 tip; none touch this PR's scope. (PR #14693 fixes two of the stale-test failures; #14683 owns the cliproxy typecheck + env-doc pair.)

Maintainer rework (merge-batch 2026-09-24)

Merged the current release/v3.8.51 tip (real merge, your commit untouched) and added one commit on top:

  • Token / fail-open. GET /repos/{owner}/{repo}/actions/runners needs the repository Administration: read permission, and a job's permissions: block can't give that to GITHUB_TOKEN. As written, both new jobs would get a 403, so every publish would go red at the preflight. The runner listing now uses an optional RUNNER_STATUS_TOKEN secret, falling back to github.token. If the listing fails, both jobs log a ::warning:: and exit 0. A missing or under-scoped token therefore can't block a publish, and the check turns on once the secret is provisioned.
  • Skipped-preflight skipped the build. build now needs: runner-preflight, and that job is skipped when USE_VPS_RUNNER is off. A plain if: adds an implicit success() check, so the whole build would have been skipped too. The fix: !cancelled() plus an explicit needs.runner-preflight.result == 'success' || 'skipped' check.
  • Nightly run cancelled publishes. The schedule run executes on the default-branch ref and shared docker-publish-${{ github.ref }} with cancel-in-progress: true, so every night it would cancel any release-branch publish still running. It now gets its own concurrency group.
  • Added tests/unit/docker-publish-runner-preflight-14560.test.ts, which covers all three cases. It fails 3/3 on your original head and passes 3/3 here. The existing docker-publish suites still pass (6/6 total). actionlint shows nothing new (the only finding is a pre-existing SC2034 in the merge job).

Still open, needs an operator decision: provision RUNNER_STATUS_TOKEN (a fine-grained PAT with Administration: read on this repo, or a classic PAT with repo). Until that secret exists, the preflight and the alert only log a warning. The end-to-end behavior (preflight turns red when ONLINE=0) can only be confirmed by a real workflow_dispatch after the secret is added.

@diegosouzapw

Copy link
Copy Markdown
Owner

This is exactly the right shape of fix — a preflight gate instead of a timeout around an unbounded queue. One thing worth double-checking before merge: runner-preflight and runner-alert both call GET /repos/{owner}/{repo}/actions/runners with ${{ github.token }} under only contents: read (issues: write for the alert job). Listing self-hosted runners typically needs an org-level "Administration"/"Self-hosted runners" read permission (or admin:org for a classic PAT) — neither is assignable to GITHUB_TOKEN via a job's permissions: block. Could you do a quick workflow_dispatch dry run (or point me at one) to confirm the API call actually succeeds with the default token here, rather than 403ing? If it does 403, swapping to an admin-scoped PAT secret for just that step should fix it.

…lishes

- Listing self-hosted runners needs the repository "Administration: read"
  permission, which GITHUB_TOKEN can never be granted via `permissions:`.
  Use an optional RUNNER_STATUS_TOKEN secret and fail OPEN with a warning
  when the listing is not accessible, in both runner-preflight and
  runner-alert, so a missing token can never turn the publish red.
- build needs runner-preflight, which is skipped when USE_VPS_RUNNER is off;
  guard with !cancelled() + an explicit result check so the build is not
  skipped along with it.
- Give the nightly schedule run its own concurrency group: it runs on the
  default-branch ref and cancel-in-progress would otherwise kill an
  in-flight release-branch publish every night.
- Add structural regression tests for the three cases.
@gonisulaimann

Copy link
Copy Markdown
Contributor Author

Confirmed — your rework (286d6a3) resolves the permission question the right way: RUNNER_STATUS_TOKEN fail-open with a warning beats burning a PAT slot on an advisory check, and the skip-condition fix on the build dependency closes the real footgun. Ready for merge.

@diegosouzapw diegosouzapw changed the title ci(docker): fail fast when the omni-build runner is offline [defer] ci(docker): fail fast when the omni-build runner is offline Sep 29, 2026
@diegosouzapw diegosouzapw added the deferred-v3.8.52 Grande demais / suspeito para o lote atual; precisa de sessão dedicada no ciclo v3.8.52 label Sep 29, 2026
@diegosouzapw
diegosouzapw changed the base branch from release/v3.8.51 to release/v3.8.52 September 29, 2026 11:20
@diegosouzapw

Copy link
Copy Markdown
Owner

Re-homed to release/v3.8.52: v3.8.51 entered its release freeze, so the branch now belongs to the release captain and development continues on the next cycle. Nothing is wrong with this PR — it just needed a live base. No action needed from you; CI will re-run against the new base.

@diegosouzapw diegosouzapw removed the deferred-v3.8.52 Grande demais / suspeito para o lote atual; precisa de sessão dedicada no ciclo v3.8.52 label Oct 1, 2026
@diegosouzapw diegosouzapw changed the title [defer] ci(docker): fail fast when the omni-build runner is offline ci(docker): fail fast when the omni-build runner is offline Oct 1, 2026
@diegosouzapw
diegosouzapw merged commit 04e6f67 into diegosouzapw:release/v3.8.52 Oct 7, 2026
9 of 16 checks passed
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

fix(docker): a single offline omni-build runner stalls the publish silently — add a job timeout and a runner-availability alert

2 participants