Repository navigation
ci(docker): fail fast when the omni-build runner is offline - #14701
diegosouzapw merged 3 commits into
Conversation
|
This is exactly the right shape of fix — a preflight gate instead of a timeout around an unbounded queue. One thing worth double-checking before merge: |
…lishes - Listing self-hosted runners needs the repository "Administration: read" permission, which GITHUB_TOKEN can never be granted via `permissions:`. Use an optional RUNNER_STATUS_TOKEN secret and fail OPEN with a warning when the listing is not accessible, in both runner-preflight and runner-alert, so a missing token can never turn the publish red. - build needs runner-preflight, which is skipped when USE_VPS_RUNNER is off; guard with !cancelled() + an explicit result check so the build is not skipped along with it. - Give the nightly schedule run its own concurrency group: it runs on the default-branch ref and cancel-in-progress would otherwise kill an in-flight release-branch publish every night. - Add structural regression tests for the three cases.
|
Confirmed — your rework (286d6a3) resolves the permission question the right way: RUNNER_STATUS_TOKEN fail-open with a warning beats burning a PAT slot on an advisory check, and the skip-condition fix on the build dependency closes the real footgun. Ready for merge. |
|
Re-homed to |
04e6f67
into
diegosouzapw:release/v3.8.52
Closes #14560.
timeout-minutescannot fix the silent publish stall because the clock only starts once a runner picks the job up — queue time is unbounded. So the fix is a gate before the build instead of a timeout around it:runner-preflightjob (needsprepare, 5 min, hosted): whenUSE_VPS_RUNNER=true, lists repo runners via the Actions API and fails the run if no online runner carries theomni-buildlabel. The error message names the problem and the fallbacks (restart the self-hosted listener, orUSE_VPS_RUNNER=false— noting hosted dies ResourceExhausted per ci: docker-publish OOM (ResourceExhausted) on every release/v3.8.51 push — :next stale since 08-22 and the v3.8.51 tag will not build on the hosted runner #11976). A publish that used to queue forever now goes red in minutes with the verdict in the job name.timeout-minutes: 350on the build job as a backstop for a wedged build step (Docker/BuildKit hang) — ~2h50m observed build + headroom. Deliberately not sized for queue time since it can't bound that.runner-alert(schedule+workflow_dispatch): mirrors the nightly-release-green issue loop — opens a deduplicated tracking issue when theomni-buildpool has zero online runners, comments with fresh evidence while it persists, auto-closes on recovery. Publish jobs skip themselves on the schedule event (prepare/preflight/build/mergeall gate ongithub.event_name != 'schedule'), so the cron exists only to drive the alert.Verified: workflow YAML parses, all five jobs'
if/needsedges re-checked after the schedule addition (a skippedprepareyields emptyoutputs.skip, so the build/merge guards needed the explicit event gate — added), preflight jq filter tested locally against the API shape.release/v3.8.51tip; none touch this PR's scope. (PR #14693 fixes two of the stale-test failures; #14683 owns the cliproxy typecheck + env-doc pair.)Maintainer rework (merge-batch 2026-09-24)
Merged the current
release/v3.8.51tip (real merge, your commit untouched) and added one commit on top:GET /repos/{owner}/{repo}/actions/runnersneeds the repository Administration: read permission, and a job'spermissions:block can't give that toGITHUB_TOKEN. As written, both new jobs would get a 403, so every publish would go red at the preflight. The runner listing now uses an optionalRUNNER_STATUS_TOKENsecret, falling back togithub.token. If the listing fails, both jobs log a::warning::and exit 0. A missing or under-scoped token therefore can't block a publish, and the check turns on once the secret is provisioned.buildnowneeds: runner-preflight, and that job is skipped whenUSE_VPS_RUNNERis off. A plainif:adds an implicitsuccess()check, so the whole build would have been skipped too. The fix:!cancelled()plus an explicitneeds.runner-preflight.result == 'success' || 'skipped'check.schedulerun executes on the default-branch ref and shareddocker-publish-${{ github.ref }}withcancel-in-progress: true, so every night it would cancel any release-branch publish still running. It now gets its own concurrency group.tests/unit/docker-publish-runner-preflight-14560.test.ts, which covers all three cases. It fails 3/3 on your original head and passes 3/3 here. The existing docker-publish suites still pass (6/6 total). actionlint shows nothing new (the only finding is a pre-existing SC2034 in the merge job).Still open, needs an operator decision: provision
RUNNER_STATUS_TOKEN(a fine-grained PAT with Administration: read on this repo, or a classic PAT withrepo). Until that secret exists, the preflight and the alert only log a warning. The end-to-end behavior (preflight turns red whenONLINE=0) can only be confirmed by a realworkflow_dispatchafter the secret is added.