Skip to content
Merged
Show file tree
Hide file tree
Changes from all commits
Commits
File filter

Filter by extension

Filter by extension


Conversations
Failed to load comments.
Loading
Jump to
Jump to file
Failed to load files.
Loading
Diff view
Diff view
7 changes: 5 additions & 2 deletions .github/workflows/ci-macos.yml
Original file line number Diff line number Diff line change
Expand Up @@ -486,9 +486,11 @@ jobs:
# `warm`: the SwiftPM cache restore and the seed are skipped, and the
# compile rebuilds only what changed since the last job on this Mac. Any
# mismatch, miss or failure falls back to the ephemeral flow below.
# Main's full-suite dispatch may run here too (pr_runner_pool.py), and
# the main commit it keeps is what the next pull requests merge onto.
# scripts/ci/owned_build_state.py has the rules; nothing is uploaded.
id: owned-state
if: steps.reuse-products.outputs.hit != 'true' && github.event_name == 'pull_request' && startsWith(env.CMUX_PRODUCT_RUNNER, 'glaeda-')
if: steps.reuse-products.outputs.hit != 'true' && (github.event_name == 'pull_request' || github.event_name == 'workflow_dispatch' && github.ref == 'refs/heads/main') && startsWith(env.CMUX_PRODUCT_RUNNER, 'glaeda-')
continue-on-error: true
run: |
set -euo pipefail
Expand All @@ -511,7 +513,8 @@ jobs:
timeout-minutes: 3
env:
SEED_PREFIX: admission-derived-data-v1-${{ runner.os }}-${{ runner.arch }}-${{ steps.owned-state.outputs.fingerprint }}-
MERGED_ONTO: ${{ inputs.source_parent1 || github.event.pull_request.base.sha }}
# Like the seed adoption below: a main dispatch starts at its own commit.
MERGED_ONTO: ${{ github.event_name == 'pull_request' && (inputs.source_parent1 || github.event.pull_request.base.sha) || github.sha }}
MAX_DISTANCE: ${{ vars.CI_OWNED_PREFER_SEED }}
CMUX_SEED_LOCAL_CACHE: ${{ env.CMUX_OWNED_STATE_ROOT }}/seeds
GH_TOKEN: ${{ github.token }}
Expand Down
5 changes: 3 additions & 2 deletions .github/workflows/ci-owned-pool-rescue.yml
Original file line number Diff line number Diff line change
Expand Up @@ -20,8 +20,9 @@ run-name: owned-pool-rescue-${{ inputs.run_id || github.event.workflow_run.id }}
# from 2026-09-24 to 09-25 05:00 UTC, 361 of which were skipped (fork pull
# requests) or found no owned pool. The script reads the run from the API and
# applies the checks the event filter did (ci.yml, a same-repository pull
# request run, or a dispatch of a DISPATCH_WORKFLOW_PATHS workflow; attempt
# 1), so a dispatch naming any other run does nothing. Only attempt 1 is
# request run or main's full-suite dispatch, which pr_runner_pool.py places
# like one, or a dispatch of a DISPATCH_WORKFLOW_PATHS workflow; attempt 1),
# so a dispatch naming any other run does nothing. Only attempt 1 is
# watched, so the re-run this workflow starts is never watched twice.
#
# Two kinds of source keep the workflow_run trigger:
Expand Down
26 changes: 18 additions & 8 deletions .github/workflows/ci.yml
Original file line number Diff line number Diff line change
Expand Up @@ -603,11 +603,12 @@ jobs:
id: route-token
# The org's manaflow-glaeda-route App, read-only on runners, so the
# pool choice below sees which owned Macs are idle now. Same-repository
# pull requests only, on this ephemeral Linux runner: a fork run never
# sees the key, and no Mac ever holds it. Any failure (no key yet, the
# App not installed) leaves the token empty, and the choice falls back
# to the janitor snapshot and CI_OWNED_POOL_SLOTS.
if: github.event_name == 'pull_request' && github.event.pull_request.head.repo.full_name == github.repository && vars.CI_PR_POOL_OWNED == '1' && vars.GLAEDA_ROUTE_APP_ID != ''
# pull requests and main's full-suite dispatch only, on this ephemeral
# Linux runner: a fork run never sees the key, and no Mac ever holds
# it. Any failure (no key yet, the App not installed) leaves the token
# empty, and the choice falls back to the janitor snapshot and
# CI_OWNED_POOL_SLOTS.
if: (github.event_name == 'pull_request' && github.event.pull_request.head.repo.full_name == github.repository || github.event_name == 'workflow_dispatch' && github.ref == 'refs/heads/main') && vars.CI_PR_POOL_OWNED == '1' && vars.GLAEDA_ROUTE_APP_ID != ''
continue-on-error: true
timeout-minutes: 1
uses: actions/create-github-app-token@bcd2ba49218906704ab6c1aa796996da409d3eb1 # v3.2.0
Expand All @@ -619,6 +620,11 @@ jobs:
- name: Choose the pull request macOS pool
id: macos-pool
# Fail-safe: any error leaves the outputs empty, which is today's route.
# Main's full-suite dispatch (ci-main-full-suite.yml) is routed too,
# onto the owned pools only, placed like a pull request unless
# CI_OWNED_MAIN_RESERVE holds machines back for pull requests
# (pr_runner_pool.py, "Main's full suite"); ci-macos.yml already reads
# these outputs for a workflow_dispatch on main.
continue-on-error: true
timeout-minutes: 1
env:
Expand All @@ -643,7 +649,9 @@ jobs:
OWNED_LIGHT_RETRY: ${{ vars.CI_OWNED_LIGHT_RETRY }}
# 0 keeps GUI jobs (app-host shards, tests-build-and-lag) off them.
POOL_OWNED_GUI: ${{ vars.CI_PR_POOL_OWNED_GUI }}
CMUX_CI_XCODE_APP_PR: ${{ github.event.pull_request.head.repo.full_name == github.repository && vars.CMUX_CI_XCODE_APP_PR || '' }}
# Machines and root runners main's dispatch leaves free (default 0).
OWNED_MAIN_RESERVE: ${{ vars.CI_OWNED_MAIN_RESERVE }}
CMUX_CI_XCODE_APP_PR: ${{ (github.event.pull_request.head.repo.full_name == github.repository || github.event_name == 'workflow_dispatch' && github.ref == 'refs/heads/main') && vars.CMUX_CI_XCODE_APP_PR || '' }}
# Handed to the macOS 15 pool's jobs only when the run lands there.
CMUX_CI_XCODE_APP_MACOS_15: ${{ vars.CMUX_CI_XCODE_APP_MACOS_15 }}
# This run's macOS jobs, from the routing above, so an owned pool is
Expand Down Expand Up @@ -993,10 +1001,12 @@ jobs:
# this workflow with Actions write, and it runs no repository code:
# no checkout, one API call. The watcher re-checks the run it is given.
# Fail-safe: if the dispatch fails the run is only not watched, and
# ci-status does not wait for this job.
# ci-status does not wait for this job. Main's full-suite dispatch
# (ci-main-full-suite.yml) is placed on the owned pools like a
# same-repository pull request, so it is watched too.
name: Watch owned pool jobs
needs: changes
if: ${{ needs.changes.outputs.macos_pr_owned_jobs != '' && github.event_name == 'pull_request' && github.event.pull_request.head.repo.full_name == github.repository && github.run_attempt == 1 && vars.CI_PR_POOL_OWNED == '1' && (vars.CI_OWNED_POOL_RESCUE || '1') != '0' }}
if: ${{ needs.changes.outputs.macos_pr_owned_jobs != '' && (github.event_name == 'pull_request' && github.event.pull_request.head.repo.full_name == github.repository || github.event_name == 'workflow_dispatch' && github.ref == 'refs/heads/main') && github.run_attempt == 1 && vars.CI_PR_POOL_OWNED == '1' && (vars.CI_OWNED_POOL_RESCUE || '1') != '0' }}
runs-on: ubuntu-24.04 # github-hosted-required: one API call; keeps CI's Linux pool free
timeout-minutes: 5
# Job level, like reverse-test-impact, so macos-admission-gate never
Expand Down
38 changes: 32 additions & 6 deletions docs/ci-runners.md
Original file line number Diff line number Diff line change
Expand Up @@ -210,9 +210,29 @@ names no owned pool.
| --- | --- | --- |
| `CI_PR_POOL_OWNED` | unset (off) | `1` puts owned pools first and turns on the rescue below |
| `CI_OWNED_POOL_SLOTS` | unset (no slots) | JSON, owned pool label to machine count, the `conforming_count` from `glaeda-mini-fleet pools --json`: `{"glaeda-std-xcode-26.6": 12, "glaeda-light-xcode-26.6": 2}`. A class (`{"std": 12, "light": 2}`) or a bare count (`12`, the std class) means that class at the lane's Xcode pin |
| `GLAEDA_ROUTE_APP_ID` + secret `GLAEDA_ROUTE_APP_KEY` | unset (snapshot only) | the org's `manaflow-glaeda-route` App. `ci.yml`'s `changes` job mints a token with `administration: read` for same-repository pull requests only, on its ephemeral Linux runner, and the picker lists the repository's runners: the online, idle runners carrying an owned label are that pool's free machines, less what runs of the last `LIVE_WINDOW_MINUTES` took. That replaces `CI_OWNED_POOL_SLOTS` and the snapshot's owned counts and age. Any failure falls back to them |
| `CI_OWNED_MAIN_RESERVE` | `0` | machines, and root runners, main's full-suite dispatch leaves free for pull requests; above 0 it takes an owned pool only whole (below) |
| `GLAEDA_ROUTE_APP_ID` + secret `GLAEDA_ROUTE_APP_KEY` | unset (snapshot only) | the org's `manaflow-glaeda-route` App. `ci.yml`'s `changes` job mints a token with `administration: read` for same-repository pull requests and main's full-suite dispatch only, on its ephemeral Linux runner, and the picker lists the repository's runners: the online, idle runners carrying an owned label are that pool's free machines, less what runs of the last `LIVE_WINDOW_MINUTES` took. That replaces `CI_OWNED_POOL_SLOTS` and the snapshot's owned counts and age. Any failure falls back to them |
| `CI_OWNED_LIGHT_RETRY` | unset (off) | `1` lets attempt 2, the full re-run the rescue starts for a job stuck on a full `std` pool, take the `light` pool when the run's whole owned peak is free there and `github-actions[bot]` started the re-run (a person's re-run of attempt 2 stays on Blacksmith). The rescue watches that attempt like attempt 1, and a job stuck or refused there goes to Blacksmith on attempt 3. Only while it is on do the janitor and the picker look up attempt 2's marker. Order: std, light, Blacksmith |

Main's full suite: `ci-main-full-suite.yml` dispatches `ci.yml` on main about
32 times a day, each a full suite. Until this change every one ran compile
admission and the seven app-host shards on Blacksmith (p50 wall about 37
minutes; admission 887 s on 6vcpu against 663 s on a mini with a 2 s queue).
The dispatch runs main's own code, so the picker places it like a
same-repository pull request, split and the `CI_PR_POOL_QUEUE_ROUNDS` queue
allowance included, on the owned pools only: jobs that do not fit keep
`MACOS_RUNNER_PR` as before. Its run holds 9 root runners
at peak (admission's, then 7 shards, tests-build-and-lag and cli-product-tests);
the Claude wrapper and remote daemon lanes route only for pull requests, so they
are not counted. `CI_OWNED_MAIN_RESERVE` above 0 holds that many machines and
root runners back for pull requests, and then main takes the pool only whole
and only while its peak is free now, with no queue allowance. Main's CI concurrency group
runs one dispatch at a time, so main never holds more than one run's machines.
Its marker, the janitor's `committed` count, the route replay in newer picks,
the owned Mac's kept build state and the rescue all treat it like a
same-repository pull request run. A retry attempt, a dispatch on another
branch, or `CI_PR_POOL_OWNED` off keeps its route.

Each entry of `CI_OWNED_POOL_SLOTS` that is not an owned label or class with a
positive whole number of machines counts as none. A full label wins over its
class. While owned pools are on, the `changes` job raises a workflow error
Expand Down Expand Up @@ -261,7 +281,11 @@ If one of its jobs waits for a persistent runner longer than
`CI_PR_POOL_QUEUE_ROUNDS` round for a CI run whose picker placed a job in the
queue allowance (it uploads a `macos-pool-queued-<run>-<attempt>-owned`
marker), the watcher confirms the
pull request head has not moved, cancels the run, and re-runs it. A retry
pull request head has not moved, cancels the run, and re-runs it. For main's
full-suite dispatch it checks main's HEAD instead: once main has moved past
the run's commit, the run is cancelled but not re-run, because its completion
makes `ci-main-full-suite.yml` dispatch the newer HEAD; a refused job on
main's run gets its failed jobs re-run whether or not main moved. A retry
attempt never takes a persistent pool, so the re-run lands on Blacksmith as a
whole, and so does a manual "Re-run all jobs".

Expand Down Expand Up @@ -307,8 +331,8 @@ minutes on the ephemeral flow instead (run 36048804178). Any mismatch or miss
falls back to that flow, and a DerivedData over 40 GB is dropped. Only a
successful compile's DerivedData is kept, cloned right after the compile,
before the staging and packaging steps rewrite Build/Products. Only pull
request runs keep or read this state, so main's full-suite dispatch never
builds on it. Moves are renames on one volume, glaeda's host lock keeps one job
request runs and main's full-suite dispatch keep or read this state; a main
run leaves the Mac warm for the main commit the next pull requests merge onto. Moves are renames on one volume, glaeda's host lock keeps one job
per Mac, and nothing is uploaded: an owned run writes only its own Mac's state.
The four steps are non-product recipe steps (`product_input_identity.py`), so
no pool's product key changes. Blacksmith and fork runs never take them.
Expand All @@ -319,7 +343,7 @@ every sweep whatever the queue, which frees minis for current work. The other
categories cancel only while more than `CI_JANITOR_QUEUE_THRESHOLD` jobs queue
on a pool the run holds, owned pools included. With `CI_PR_POOL_OWNED=1` the
janitor also lists the artifacts of each in-flight attempt-1, same-repository
pull request CI run (one request per run, more only past 100 artifacts) to
pull request CI run or main full-suite dispatch (one request per run, more only past 100 artifacts) to
read its marker's peak into `committed`. A run's other macOS jobs do not rule
it out: `swift-package-tests` always runs on Blacksmith beside a full suite on
an owned pool.
Expand All @@ -338,7 +362,9 @@ The guard keeps the picker the only way onto an owned pool.
text, and `runner_label_policy.py` refuses one in any `*RUNNER*` variable, so
neither a workflow edit nor `MACOS_RUNNER_PR` can send a job there.
`check_owned_pools_route_through_picker` requires the picked label to reach
jobs only as `pr_runner` or on a `pull_request` `runs-on` branch.
jobs only as `pr_runner` or on a `pull_request` `runs-on` branch. `ci-macos.yml`
reads those inputs for a `workflow_dispatch` on `refs/heads/main` too, which
the picker fills only for main's full-suite dispatch.
`CI_PR_POOL_ORDER` is the one variable that may name owned labels (the guard's
`owned` pattern, which must match `pr_runner_pool.OWNED_LABEL`), and the CI
health report checks every other entry in it against the workflow policy.
Expand Down
Loading
Loading