Skip to content

docs: propose routing CI by capability instead of by vendor - #14010

Merged
teamleaderleo merged 2 commits into
mainfrom
docs/ci-runner-capability-labels
Sep 24, 2026
Merged

teamleaderleo merged 2 commits into
mainfrom
docs/ci-runner-capability-labels

Conversation

@teamleaderleo

@teamleaderleo teamleaderleo commented Sep 23, 2026 •

Copy link
Copy Markdown
Collaborator

Every macOS job in cmux names a vendor's machine label instead of what the job needs. tests-build-and-lag needs a macOS machine with a foreground Aqua login session; its runs-on says this instead:

runs-on: ${{ github.event_name == 'pull_request' && (vars.MACOS_RUNNER_PR || 'blacksmith-6vcpu-macos-15') || vars.CI_PAID_MACOS_OVERFLOW == '1' && vars.MACOS_RUNNER_DISPLAY || 'blacksmith-6vcpu-macos-15' }}

This PR adds docs/ci-runner-capability-labels.md, a proposal to have jobs declare capabilities (runs-on: [self-hosted, macos-26, gui]) and have runners advertise them. Adding a machine, such as a Mac Ultra, then means registering it with the right labels. It needs no workflow edit, repository variable, or guard change. Documentation only: no workflow, test, script, or variable changes, and nothing in the document is implemented.

Why

On main at 02972b7a73, workflows read 9 vars.MACOS_RUNNER_* variables, plus CI_PAID_MACOS_OVERFLOW and CMUX_CI_XCODE_APP_PR. Their fallbacks use three macOS labels: blacksmith-6vcpu-macos-15, blacksmith-6vcpu-macos-26, and blacksmith-12vcpu-macos-26. tests/test_ci_self_hosted_guard.sh defines 35 check_* functions and calls them 40 times. Ten of them only police which vendor string may appear where. #14002 found a case of this drift: one variable served two jobs with incompatible macOS requirements, and the fallbacks disagreed.

What the document proposes

  • The rule. macOS version, SDKs, GUI session, simulator, and disk are job properties that change and get reviewed in a PR. Vendor and cost are operator choices.
  • A vocabulary built from jobs on main: macos-14/15/26, x86_64, sdk-15, gui, ios-simulator, and a large-machine label.
  • Deletions. The runner variables collapse to one translation table plus one cost control. Ten guard functions and two helpers become unrepresentable. On current main they cover 15 of 40 calls, and four more checks shrink.
  • A six-step migration with no flag day. It starts with one non-required job (build-ghosttykit.yml) and ends with tests-build-and-lag.

Settled since the draft

The PR comments settled two points the first draft listed as open, and 3b8514a writes the answers into the document:

  • runs-on accepts a computed label array (${{ fromJSON(...) }}), and a runner must carry all listed labels. This was step 0(a). GitHub's syntax reference does not show the array form (github/docs#20495). Community reports (#78674, #50172) describe dynamic multi-label jobs that queue forever. Step 0 therefore needs a deliberate no-match run, not only a syntax check.
  • Blacksmith offers only its fixed tags (blacksmith-{6,12}vcpu-macos-{15,26,27,latest}) and no customer-defined labels (instance types). The vendor translation table stays for as long as Blacksmith capacity is used. Owned machines still pass their labels through unchanged. The proposal still stands.
  • cpu-12 and disk-large are one Blacksmith tier (12 vCPU, 48 GB, 250 GB), so they become one macos-large label. release-build now runs on the 6-vCPU pool, so only the universal Nightly build uses that tier. Its threshold is still Blacksmith's tier, not a measured need.

Still open

  • Nothing detects a capability set that no runner satisfies. Such a job queues with no error, and a runtime preflight cannot catch it.
  • Runners with different speeds behind one label break timing jobs such as tests-build-and-lag. That lane moves last.
  • CMUX_PRODUCT_RUNNER would need to be replaced by machine-reported identity. What owned machines set ImageOS to has not been established.
  • Label matching has no preference order, so "prefer free, fall back to paid" stays one explicit operator control, as it is today.

Validation

At 3b8514a: python3 scripts/ci/run_python_test_lane.py --lane linux-guard exited 0 with no FAIL: lines, and bash tests/test_ci_self_hosted_guard.sh passed. These show that the new doc breaks no guard. They do not test the proposal. The counts above come from git grep over .github/workflows and the guard file at origin/main 02972b7a73.

🤖 Generated with Claude Code

Jobs name the vendor that owns the machine (`blacksmith-6vcpu-macos-26`),
not what they require. Eleven `MACOS_RUNNER_*` variables carry three
distinct label values, and ten of the guard's thirty-four check functions
exist only to police which vendor string appears where.

This proposes jobs declaring capabilities (`[self-hosted, macos-26, gui]`)
and capacity advertising them, with a single vendor translation layer for
providers that own their label names. Proposal only; nothing implemented.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>
@coderabbitai

coderabbitai Bot commented Sep 23, 2026 •

Copy link
Copy Markdown

Warning

Review limit reached

Next included review available in 3 minutes.

Check out review usage here.

View limit details

Limit details: You’ve used all 10 included reviews currently available.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Repository: manaflow-ai/cmux/.coderabbit.yaml

Review profile: ASSERTIVE

Plan: Advanced

Run ID: 83f1c029-08c7-40d3-bc21-ecc3d6f6335b

📥 Commits

Reviewing files that changed from the base of the PR and between 3c58b03 and 3b8514a.

📒 Files selected for processing (1)
  • docs/ci-runner-capability-labels.md

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

Copy link
Copy Markdown
Contributor

All contributors have signed the CLA ✍️ ✅
Posted by the CLA Assistant Lite bot.

@teamleaderleo

Copy link
Copy Markdown
Collaborator Author

Step 0 resolved: runs-on does accept a label array from an expression

The document makes this its step 0, on the grounds that no cmux workflow does it today and the design fails without it. It works, and it is a common pattern:

runs-on: ${{ fromJSON('["self-hosted", "macos-26", "gui"]') }}
runs-on: ${{ inputs.runner == 'self-hosted' && fromJSON('["self-hosted","linux","docker"]') || inputs.runner }}

A job output parsed with fromJSON works the same way, which is the shape a translation layer would use. This also confirms the cumulative-AND reading: a self-hosted runner must carry all listed labels to be eligible.

Note GitHub's own docs do not show an array example for runs-on — github/docs#20495 is an open request for exactly that. The behaviour is real but underdocumented, which is worth stating in the doc rather than citing the syntax reference, since the reference does not cover it.

The more important finding: open question 3 is not hypothetical

The document asks what detects "no runner satisfies this capability set", and notes a missing label yields an indefinitely queued job with no error. That failure is specifically reported against this pattern:

  • community#78674 — "dynamic runs-on label definition causes job being stuck in queue for ever"
  • community#49302 — "Dynamic runs-on labeling does not react to multi labels"
  • community#50172 — "Self-Hosted runners: runs-on tag with multiple labels does not work as documented"

So the risk is not merely "a typo in a label queues forever" — it is that dynamic multi-label selection has reported matching failures of its own. That raises the cost of step 0 from "confirm the syntax parses" to "confirm matching behaves under our label sets, with a deliberate no-match case", and it strengthens the argument for a scheduling-side detector rather than a runtime preflight: a preflight runs on a runner, and this failure means never reaching one.

Worth connecting to docs/ci-runners.md's own cost guidance: a permanently queued job reports its queue time as duration, so this failure reads as a long, expensive job that actually burned nothing. Whatever detects it should key on runner_name and steps being empty, the same discriminator that section already establishes.

Sources: GitHub Docs — using self-hosted runners in a workflow, GitHub Docs — self-hosted runners reference.

@teamleaderleo

teamleaderleo commented Sep 23, 2026 •

Copy link
Copy Markdown
Collaborator Author

Open question 1 answered: Blacksmith does not support customer-defined labels

From Blacksmith's instance-type reference: only fixed tag names, blacksmith-{6,12}vcpu-macos-{15,26,27,latest}. Custom labels are not offered.

So the vendor translation layer is permanent, not transitional — for as long as any sponsored Blacksmith capacity is in use. The document should say that outright rather than listing it as unresolved. It does not invalidate the design: owned hardware still gets capability labels, and the translation table stays small because Blacksmith exposes exactly three macOS distinctions (version, 6 vs 12 vCPU). But "eventually this goes away" is not true of it.

cpu-12 and disk-large are the same SKU

The document flags both as its weakest entries. They are not two capabilities — on Blacksmith they are one tier:

Tag vCPU RAM Storage
blacksmith-6vcpu-macos-* 6 24 GB 150 GB
blacksmith-12vcpu-macos-* 12 48 GB 250 GB

cpu-12 (nightly universal build) and disk-large (release-build) both resolve to the 12-vCPU tier, so the vocabulary loses a label: one macos-large covers both, and neither smuggles a specific vCPU count into the job description.

It also supplies the figure the document says is missing for "disk-heavy". CLAUDE.md puts the CMUX workload default at 120 GiB, against 150 GB on the 6-vCPU tier — which is why release-build reclaims disk before large cache restores, and why the 250 GB tier is the real requirement rather than a preference.

Two further facts worth folding in:

  • blacksmith-12vcpu-macos-15 exists and nothing in the repo uses it. If a macOS 15 lane is ever disk- or CPU-bound, that tier is available today.
  • macOS 26 images ship both Xcode 26 (default) and Xcode 27, switchable via DEVELOPER_DIR. Relevant to how sdk-15 is framed: the dual-Xcode requirement here is SDK 15 + SDK 26, which this does not satisfy, but it does mean "which Xcode" is already an image property the pool varies independently of the macOS version.

Sources: Blacksmith instance types, Blacksmith quickstart, Blacksmith pricing.

Counts now match main at 02972b7: 9 MACOS_RUNNER_* variables after
#14002 and the MACOS_RUNNER_26_LARGE merge, and 35 guard checks called
40 times, 15 of them deletable. Step 0 and the Blacksmith label question
are settled, so the document states the answers. cpu-12 and disk-large
become one macos-large label because Blacksmith sells them as one tier.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>
@teamleaderleo
teamleaderleo merged commit adddb59 into main Sep 24, 2026
32 checks passed
rustybret pushed a commit to rustybret/bmux that referenced this pull request Sep 24, 2026
9567d6e refactor: give About and Licenses windows explicit ownership (manaflow-ai#13148)
fae46b6 ci: stop retrying a missing cmux-tui manifest (manaflow-ai#14168)
8421357 Point PR checklist and welcome note at the hidden Review Trigger block (manaflow-ai#14167)
b9415db ci: run app-host product consumers on compile admission's pool and Xcode (manaflow-ai#14163)
ac0ceae fix(sidebar): order panels without reading split-container geometry (manaflow-ai#13931)
37edc16 ci: judge Web complexity's trusted files in the pull request's merge (manaflow-ai#14018)
679f4e2 ci: leave three-day-old queued ghosts to GitHub instead of retrying them (manaflow-ai#14166)
07a2e22 fix(web): enumerate complexity-gate sources with git ls-files -z (manaflow-ai#13682)
c72f659 cloud: share concurrent VM stats reads (manaflow-ai#13327)
aa51f16 ci: trim package setup before the macOS compile admission build (manaflow-ai#14160)
82ea1ed ci: land the fleet review fixes manaflow-ai#14159 merged without (manaflow-ai#14165)
adddb59 docs: propose routing CI by capability instead of by vendor (manaflow-ai#14010)
f862390 ci: fix three fleet command gaps from the manaflow-ai#14159 review (manaflow-ai#14164)
77d56b3 agent-chat: make installed harnesses first-class (manaflow-ai#13347)
7dc57f6 Clarify writing guidance for issue and PR descriptions (manaflow-ai#13275)
ccf4963 ci: name the hung test when a Swift package test step stalls (manaflow-ai#14055)
9fca985 ci: guard the fleet routing switch, Xcode pin and quarantine (manaflow-ai#14159)
02972b7 fix: thin around and Developer ID sign the bundled cmux-tui SSH payloads (manaflow-ai#14154)

# Conflicts:
#	.github/workflows/app-host-test-rerun.yml
#	.github/workflows/ci-guards.yml
#	.github/workflows/ci-macos.yml
#	.github/workflows/test-ios.yml
#	.github/workflows/web-complexity-trusted.yml
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant