Skip to content

ci: smoke-test zero-config fork runners - #96

Closed
teamleaderleo wants to merge 106 commits into
mainfrom
ci/fork-runner-smoke-14023
Closed

teamleaderleo wants to merge 106 commits into
mainfrom
ci/fork-runner-smoke-14023

Conversation

@teamleaderleo

Copy link
Copy Markdown
Owner

Temporary validation PR for manaflow-ai#14023. The acceptance condition is that the normal CI workflow starts on GitHub-hosted runners in this fork with no runner variables or Blacksmith installation. Close after the runner-routing evidence is captured.

teamleaderleo and others added 30 commits September 23, 2026 07:05
…nt (manaflow-ai#13907)

* test(simulator): bound the panel waits by a deadline, not a yield count

`"A surviving Simulator host keeps framebuffer publication active"` failed
on manaflow-ai#13414's app-host shard 1 with `frameTransport → nil` at
SimulatorPanelThemeTests.swift:95 and :101, while passing on `main` and on
manaflow-ai#13752's shards. Nothing in that PR touches this subsystem.

The waits here spun a fixed `for _ in 0..<100 { ... await Task.yield() }`
before asserting. A yield count is not a deadline: `Task.yield()` gives the
scheduler a chance to run something else, it does not wait for anything, so
100 yields is however long 100 reschedules happen to take. The bound
tightens exactly when the runner is busy, which is the "fails on correct
code under load" shape `.github/review-bot-rules/test-determinism.md` bans:

  the deadline bounds the FAILURE path only, so load can make a pass slower
  but never turn a pass into a fail

These 10 sites now poll the same predicate against a `ContinuousClock`
deadline. They return the instant the condition holds, so a passing run is
no slower, and only a genuinely broken one waits out the 10 s. No assertion
changed.

Two sites are deliberately left alone: the bare
`for _ in 0..<100 { await Task.yield() }` at
SimulatorPanelIntegrationTests.swift:187 and :252 assert an *absence*
(`discoveryCount == 0`, `!isCompleted`). A deadline cannot bound a wait for
a non-event; those need a positive completion signal to assert after, which
is a change to what the test observes rather than how long it waits. Their
failure direction is also the opposite one — a short spin makes a false
green, not a false red.

Scope note: 17 more yield-count polls remain in 12 other `cmuxTests` files,
and `scripts/check-test-determinism.py` reports 0 findings across the tree
because it only matches sleep call sites. Tracked in manaflow-ai#13903; this commit is
the cluster with the observed failure.

Verified: `swiftc -parse` on both files; `check-test-determinism.py
--roots cmuxTests`, `test_ci_pbxproj_test_wiring.sh` and
`validate_test_execution_registry.py` pass; 127 of the 128 `ci-guards.yml`
commands pass, the exception being `test_ghostty_zig_version_sync.sh`,
which needs the ghostty submodule this worktree does not check out. Not
run on a macOS runner — these tests need one, so CI is their first
execution.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(simulator): wait on a real signal, and share one deadline helper

Review of the first revision found two things the yield-count -> deadline
conversion did not fix.

`SimulatorPaneCoordinator.receive(.frameTransport:)` stores the descriptor
only `if frameIsVisible`, and nothing re-sends it. Polling for
`coordinator.frameTransport != nil` therefore cannot recover a descriptor
that arrived early: it turns a permanent drop into a ten-second wait for a
latch that will never flip. The surviving-host test now waits for
`.setFramebufferPublishing(true)` — the message `reconcileFramePublication`
enqueues when it sets `frameIsVisible` — before emitting the descriptor, so
the drop window is closed rather than polled across.

The ten converted waits also each ended in a bare `break`, so a wait that ran
out fell through into whatever assertion came next; one of them
(`cancelledApplicationTerminationRestoresPanel`) had no assertion on its own
predicate at all and reported a confusing downstream failure instead. They
now share a `waitUntil` helper that requires its predicate at the deadline,
which makes that shape impossible to write and matches the spelling already
used in CloudTerminalCardFlashRegressionTests and CloudTunnelLaunchGateTests.
The helper sleeps 5 ms between probes rather than spinning `Task.yield()`, so
a failing wait no longer pegs the main actor for its full budget.

The two bare settle spins are left alone: both assert an absence, so a
deadline would only make them slower without adding signal.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* Revert the publication gate; CI disproved it

Run 35831156475 on macOS 26 failed at the gate itself: the wait for
`.setFramebufferPublishing(true)` ran its full 10 s and never saw the
message, while every other wait in both suites passed.

`reconcileFramePublication` returns early unless `status == .streaming`
(SimulatorPaneCoordinator+FrameVisibility.swift:72), so the publication
message is not the reliable precursor to the descriptor I took it for, and
gating the emit on it only deadlocks the test against a message that this
path does not produce here. The premise was inferred, not measured; the
measurement says it is wrong.

The census also does not support calling this a real defect. "A surviving
Simulator host keeps framebuffer publication active" passes on main at
f3d204a, a9b0329 and ce1c55c, and passed both earlier dispatches of
this branch. Its one observed failure is on manaflow-ai#13414's run, which had 148
failing tests and corresponding load — which is exactly the shape a fixed
yield count fails under, and exactly what the deadline conversion fixes.

So the original premise stands on its own and the escalation was
unnecessary. The `waitUntil` helper stays: the full "Simulator panel
integration" suite passed in the same run, covering all seven converted
sites there.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…anaflow-ai#13951)

* coderouter: accept chatmux per-VM tokens with team-shared account access

chatmux (chatmux.dev) signs a one-hour ES256 token for each of its Freestyle
VMs with an HSM key; the Freestyle edge injects it as
x-chatmux-vm-authorization, so the guest never holds it. coderouter verifies
it against chatmux's JWKS (issuer allowlist, audience "coderouter", one-hour
maximum lifetime, required claims) with no database lookup. When the header
is present it is the only credential considered and a bad token fails closed.

A chatmux machine gets a new access kind, team-machine: only accounts its
Hexclave team shares (visibility "team"), never anyone's private account.
It has no pool and no cloud_vms row.

Off until CODEROUTER_CHATMUX_JWKS_URL (https) and CODEROUTER_CHATMUX_ISSUERS
are set.

* coderouter: keep chatmux machines on the data plane

A chatmux VM token now fails in the account control plane (it could add
or remove team accounts through a route-token header) and in every Cloud
VM-only route (VM principal, /api/vm/self, vm-usage/self, subrouter teams).
Also: no control-character regex, and a pinned clock in the verifier tests.
…low-ai#13949)

The workspace-rooted path this PR moves was invisible to every check.
`check_every_app_host_home_is_identified_and_cleaned` already finds each
job that prepares an app-host home and holds it to that contract, but it
asserted only the two preconditions `prepare` needs, not the one
`cleanup` enforces at runtime: DerivedData under RUNNER_TEMP.

Extending that loop catches the shipped bug at its pattern rather than
its instance. Reverting the workflow to the pre-fix path fails it, as
does moving ci-macos.yml's shard DerivedData out of RUNNER_TEMP or
dropping the publishing step.

The test job's own `Clean owned DerivedData` arm also had no coverage:
`step()` resolves an ambiguous name to the build job, so a typo in the
test job's ownership pattern shipped green and failed only on a runner,
under `if: always()`, after the tests had passed.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…incomplete (manaflow-ai#13962)

* test: pin that an incomplete app-host run still names its failures

On main's full suite at f3d204a (run 35788803553) all six app-host
shards returned at the incompleteness gate in check_run(), so not one
RATCHET_NEW_FAILURE line was printed across the entire run -- even though
the logs carried real assertion failures. A red suite that names no
regression cannot tell anyone whether a fix landed.

This commit adds the failing test only. It asserts that a run with one
missing terminal result still reports the new and known failures it did
record, and a companion asserting no RATCHET_ noise when nothing failed.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(ci): report recorded failures even when the xcresult is incomplete

check_run() returned at the missing-execution gate, discarding the ratchet
comparison for the whole shard. One selected test with no terminal result
was enough to stop app-host-known-failures.json being consulted at all, so
a shard whose remaining tests regressed and one whose remaining tests went
green printed the same "incomplete" line.

The verdict stays fail-closed -- an incomplete run is still a failed run.
Only the diagnostics change: the failures that were recorded are now named
alongside the incompleteness report.

recorded_failure_diagnostics() is deliberately not the ratchet's own logic.
The ratchet fails fast, reporting new failures and returning without
mentioning known ones because the verdict is already settled; a run being
reported for some other reason wants the complete picture instead.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(ci): summarize the incomplete path's verdicts like the complete one

Review point from the consuming session: the complete path ends with
"known-main failures tolerated: N; typed test cases: M", and without an
equivalent here a reader scanning shard output for evidence that the
accounting ran still sees silence on the incomplete path.

Emitted only when something was recorded, so the exact-match assertion in
test_missing_selected_test_result_never_passes keeps holding.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* test: cover optimistic dashboard team switching

* fix: make dashboard team switching optimistic

* fix: keep a catalog snapshot for rollback

* fix(web): keep optimistic team switches ordered

* test(web): cover overlapping team switches

* fix(web): snapshot legacy team scope through parser

* fix: serialize overlapping team persistence

* fix: preserve rapid team-switch ordering

* test: cover overlapping team-switch failures

* fix(web): keep the team-scope probe's declared type through the reset

`renderReadyScope` assigns `probedScope = undefined` and then renders the
Probe, which reassigns it from inside a closure. Control flow analysis cannot
see that write, so it held the variable at `undefined` for the rest of the
function and `!scope` narrowed the union away entirely — every `scope.switchTeam`
in the suite failed `web-typecheck` with "does not exist on type 'never'".

Read the probe through a function so the declared type survives, and give
`renderReadyScope` an explicit ready-variant return type so the call sites do
not depend on inference through the same reset.

Types only; the 12 tests in this file passed before and after.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Leo <cheerleaderleo@outlook.com>
Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…ai#13966)

* ci: report how many runner minutes came from paid capacity

The scheduled CI health report totals runner minutes by label four times a
day, but every label reads the same: Blacksmith, GitHub-hosted and WarpBuild
sit in one table with nothing marking which of them bill. Blacksmith and
GitHub-hosted are free to this repository; Warp is metered per minute, at
roughly double the rate on its 12-vCPU labels. A lane that drifts onto paid
capacity therefore renders as an ordinary row and nobody notices until an
invoice arrives.

`MACOS_RUNNER_15`, `MACOS_RUNNER_DISPLAY`, `MACOS_RUNNER_DUAL_XCODE`,
`MACOS_RUNNER_26_RELEASE` and `MACOS_RUNNER_26_NIGHTLY_BUILD` have all been
pointing at Warp since 2026-09-19/20, so main and the merge queue run on
metered capacity while pull requests run free. That is a legitimate overflow
response to a saturated pool, but it is invisible, and `docs/ci-runners.md`
records a different intended steady state for each of those variables.

Split the paid labels out into their own line with a per-label breakdown, so
the report says how many metered minutes the window actually contained and
points at the doc that says what the steady state should be.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* ci: count every metered provider, and stop contradicting the runner docs

An independent review found three problems with the first cut.

**Depot was uncounted.** `PAID_RUNNER_PREFIX` was the single string `"warp-"`,
but `depot-*` is permitted by `tests/test_ci_self_hosted_guard.sh` and
documented alongside Warp. Pin a variable to Depot and the report totals zero
and prints "none in the window" -- exactly the silent drift this line exists to
catch. Now a tuple of prefixes, with a test that fails on the warp-only form.

**The render path was untested.** Deleting the entire `if paid_jobs:` block
left all 76 tests green: the accumulator was covered, the rendering was not,
and `test_the_waste_patterns_reach_the_output` asserts exactly this for every
other pattern. Added there; verified the deletion now fails.

**The line contradicted the doc it told you to read.** It claimed every other
label "is free to this repository", while `docs/ci-runners.md` calls
`blacksmith-*` a paid-provider label and `docs/ci/workflow-inventory.md` says
Blacksmith and GitHub-hosted bill at different rates. Both are defensible
readings of "paid" and that is the problem: a report that ends with "check this
against docs/ci-runners.md" cannot disagree with it.

Resolved in favour of precision on both sides. The report now says what is
actually true and checkable -- these are the metered third-party labels;
Blacksmith is sponsored for this organization and GitHub-hosted is free on a
public repo -- and `docs/ci-runners.md` now distinguishes "not a GitHub-hosted
free runner" from "bills this repository per minute". The header no longer says
"(WarpBuild)" now that it covers more than Warp.

Also documented the pattern in `docs/ci/health-report.md`, which enumerates
every other one.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* test: reproduce OMP auto-naming flag rejection

* fix: use OMP-supported auto-naming isolation flags
…anaflow-ai#13952)

The lane guards added alongside the routing fix name the three macOS jobs
they cover. A macOS job added next month inherits none of them: it can pin
CMUX_CI_XCODE_APP_MACOS_15 while its pool follows MACOS_RUNNER_PR, and
scripts/select-ci-xcode.sh hard-exits on a pinned path the image lacks, so
the job fails at Xcode selection the first time the lane moves.

Invert it. Every env value under .github/workflows that names the macos-15
Xcode must also read the pull-request variant, unless its exact
(file, job, key) is listed in EXEMPT with a written reason. Three sites are
exempt today: the streamed-validation lane and the two SDK 15 pins in
swift-package-tests, which is deliberately carved out onto the dual-Xcode
pool.

Keyed per (file, job, key) rather than per file so exempting
swift-package-tests does not silently exempt every other job in
ci-macos.yml.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…ow-ai#13972)

Answering "what would this shard have reported, and did this test even
run?" has required a Mac or a CI round trip. Both are scarce: app-host
shards only run under full-ci, and a shard's own output can omit the
tests that never produced a terminal result.

Everything needed is already uploaded. cmux-app-host-diagnostics-shard-N
carries the typed xcresult JSON and the captured log, and each batch's
.meta records the argv it ran, so the exact -only-testing selectors are
recoverable. cmux-app-host-test-inventory is about 1 MB. This runs the
real check_run() over them on any machine.

--accounting points the replay at a different copy of the module, which
is how manaflow-ai#13962 was measured: origin/main emitted 0 RATCHET lines on main's
run 35788803553 shard 6 where the fixed module emitted 9, on byte-identical
inputs.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
)

* Welcome first-time contributors and credit their work

Of 824 open PRs from outside the team, 779 had only bot comments.
First-time contributors now get one short human-written note on their first
PR, and CLAUDE.md tells agents to land or credit an outside PR before
writing their own fix for the same problem.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Exclude team and PR-farm accounts from the outside-PR count

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* Greet only actual first PRs, and stop contradicting the PR template

An independent review found the gate does not do what its name says, and
measured the consequence rather than theorising it.

`author_association` answers "has no merged commit in this repo", not "this is
their first pull request". cmux merges almost no outside PRs -- that is this
change's own premise -- so a persistent contributor stays
FIRST_TIME_CONTRIBUTOR indefinitely. Of 205 open PRs carrying that association,
only 145 are distinct authors: 29% would have been greeted at least twice.
`brodynies` has 16 open PRs and would have received 16 copies of "thanks for
opening your first cmux pull request". `danielraffel` would have got a sixth
greeting three and a half months after the first. That is precisely the
"you are only talking to a bot" experience this workflow exists to fix, aimed
at the outside contributors who kept showing up anyway.

Count the author's pull requests instead; this one is always included, so more
than one means it is not their first. An unusable answer is treated as "skip"
rather than as an error: the search API can rate-limit or 422, and a
non-numeric reply would otherwise abort the step under `set -e` and leave a red
X on a newcomer's first PR. A missed greeting is recoverable; a wrong greeting
or a red check is not. A marker comment makes a duplicate impossible even when
the count is stale, which search results can be for PRs opened in a burst.

The note also contradicted the template the contributor had just filled in. It
said not to @mention review bots and named two of them, while
`pull_request_template.md` ships a "Review Trigger" block of four mentions to
paste as a comment, and the checklist asks the contributor to confirm they did.
Point at the template instead of against it.

Finally, the note promised "a person on the team reads every outside PR" while
765 of 810 open outside PRs have only bot comments. Posting that to every new
contributor is a promise the repo does not keep, and it invited a bump comment
on nearly all of them. Say what is true: the queue is long, a reply can take a
while, and a nudge is welcome.

Also declare the `contents: read` the API read needs, pin `branches: [main]` to
match every other `pull_request_target` workflow here, and read the note from
`github.workflow_sha` rather than `base.sha`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…anaflow-ai#13959)

* fix: honor CMUX_SSH_RECONNECT_LIMIT above 20 instead of discarding it

The SSH PTY attach supervisor accepted only 1-20 attempts and rewrote
everything else to 20 without a word, so `CMUX_SSH_RECONNECT_LIMIT=50`
silently became 20. The same name defaulted to 86400 in the freestyle
supervisors, so one variable carried two answers 4300x apart.

SSHReconnectBudget now owns both numbers: 20 when the variable is unset
or unusable, 86400 as the ceiling and as the freestyle default. The
attach supervisor and the CLI startup script generate their clamp from
it, and a rejected value prints one line to stderr naming the value it
refused and the value it used.

Failing closed is preserved. Non-digits, empty text, and 0 still fall
back to 20; an oversized count is rejected by digit length before any
`[ ... -gt ... ]` can error out on it instead of answering.

The reconnect delay knobs keep their own hand-written clamps: they need
a ceiling too, but that changes backoff behavior and belongs with manaflow-ai#10419
and manaflow-ai#10422.

Fixes manaflow-ai#13922

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* Make the reconnect budget an instantiated value

`lint-ios-conventions-diff.sh` rejects a new all-static public enum as a
namespace type, asking for "an extension on the receiver type or an
instantiated value with injected dependencies":

    NEW  namespace-type  SSHReconnectBudget.swift
         enum SSHReconnectBudget (all-static public surface, not instantiable)

`SSHReconnectBudget` is now a struct with `limitEnvironmentName`,
`fallbackLimit` and `maximumLimit` as stored properties defaulted in the
initializer, and `limitNormalizationShellLines` as an instance method. Call
sites construct it. The `fallback` parameter became `Int?` because a default
argument can no longer name another stored property of the same value.

Behaviour is unchanged. Compiling the type on Linux, emitting the shell and
running it over the same table gives: unset and `20` resolve to 20 silently,
`21`/`50`/`86400` resolve to themselves, `007` resolves to 7, `86401` and a
20-digit value resolve to 86400 with the notice, and `abc`/`-5`/`0` resolve to
20 with the notice.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* api: report literal text delivery outcome

* test: cover paste deferral and rejected delivery

* test: verify paste acknowledgement against raw PTY bytes
…anaflow-ai#13299)

* feat(tui): adopt Herdr-style agent plugin architecture

* perf(cloud): move VM guest setup into snapshot contract

* docs: preserve Herdr handoff for startup work

* test(cloud): assert snapshot guest contract

* fix(cloud): race first terminal seed with prompt identity

* chore(cloud): promote snapshot-v2 desktop ladder

* perf(cloud): warm snapshot terminal and preserve session journal

* test(cloud): verify warm prompt contract

* fix(cloud): derive seeded terminal from session snapshot

* fix(cloud): tolerate empty warm session before seed

* fix(cloud): verify warm terminal through terminal list

* chore(cloud): promote exact warm Herdr snapshot

* fix(cloud): reuse terminal-list ids during prompt seed

* fix(cloud): inspect terminal list for existing seed

* chore(cloud): promote prompt-seed-safe Herdr snapshot

* fix(cloud): persist terminal id after workspace creation

* chore(cloud): promote final warm Herdr snapshot

* Restore exact warm snapshot manifest after main merge

* chore(cloud): keep snapshot source provenance exact

* chore(cloud): record exact cmux-tui pin in manifest

* test(cloud): cover warm snapshot prompt contract

* chore(cloud): promote merged-source warm snapshot

* test(cloud): align contracts with snapshot-v2 attach

* fix(merge): retain current Workspace cloud mirror helpers

* fix(merge): preserve cloud terminal icon helpers

* fix(merge): restore device manual mirror helper

* fix(cloud): use current Stack SDK in startup benchmark

* feat(mac): add tag-scoped runtime config profiles

* test(cloud): New Machine open must link a just-created VM

Covers the regression where cmux vm open read the cached catalog for a
VM the app had not discovered yet and failed with "The machine's
sessions are unavailable".

* fix(cloud): link a new machine before resolving its first terminal

New Machine open dropped refresh from surface.catalog to save work, but
a VM created a moment ago has no provider or link in the app, so the
cached read had no graph and open failed with "sessions are
unavailable". Add an ensure_linked read mode: discover the machine if
missing and join the provider's current refresh pass only when the
machine is not already linked. Reopen of a linked machine costs no
network; New Machine pays exactly one connect plus one graph read.

* chore(cloud): delete dead Freestyle heal helpers and guard the no-work paths

Create, restore, resume, and attach no longer install, heal, or resize,
so the daemon health/start/pin-check commands, the grow-only resize
helpers, and the driver's daemon manifest dependency had no production
caller. Remove them and their tests, fix comments that still described
attach-time healing, and document the NO-WORK INVARIANT at the top of
the driver so new guest work goes into the snapshot.

* docs(cloud): record the ensure_linked open requirement

* perf(cloud): open a new machine from its create receipt

Measured New Machine (5.9 s) spent 2.2 s on POST /attach-endpoint and
~0.3 s re-reading the fleet list, both for data the create already had,
and ~0.7 s holding the first terminal behind a cosmetic stats read.

- POST /api/vm returns the private address and cmuxTuiContract.
- The app registers the provider straight from that receipt and answers
  the first vm.cmux_remote_info locally for a snapshot-v2 machine.
- Provider refresh publishes the graph before stats; gauges fill in
  when the stats read lands.

* perf(cloud): redial a fresh machine and link it from the create receipt

- CloudHubConnector redials each address every 200 ms (capped) instead of
  one attempt riding TCP backoff; a new VM is reachable ~0.4 s after the
  create response but single attempts waited ~3.7 s or failed at 15 s.
- The registry starts the new machine's first link and graph read as soon
  as the create receipt registers it; the open joins that pass.
- vm.cmux_remote_info keeps a created receipt's declared route instead of
  a second connect race.
- Provider refresh publishes the graph before the guest port scan.
- DEBUG log of the create response's Server-Timing.

* perf(cloud): take create usage events off the critical path; recover tunnel after 5xx

- vm.create.requested runs beside model-plane provisioning and the
  provider call, joined before any failure event or return.
- vm.created is written after the response via after() on the create
  route; restore, fork, and base keep it inline.
- createTunnel bounds tunnels.create at 10 s and recovers the same-key
  tunnel after a 409, any 5xx, or no response. First enrollment hit a
  Freestyle 503 after ~25 s although the tunnel existed (502 in 30.9 s).

* perf(cloud): detect clones within 50 ms, announce first, and quiet resume

- cmux-devbox-boot ticks every 50 ms while parked (every snapshot is taken
  parked), sends the VPC announce before anything else on a clone, and
  parks housekeeping timers and service watchdogs until 10 min after bind.
- The bake zeroes the workqueue watchdog; verify-devbox-image fails on any
  resume-time watchdog kill, lockup, or catch-up timer.
- Rebaked every size (cmux-devbox-pr13299-fastboot3, daemon 01dc721);
  guest resume-to-listening 1.10 s -> 0.62-0.70 s.

* docs(cloud): record New Machine latency floor and open decisions

* chore(cloud): raw Freestyle create-to-reachable benchmark

Creates VMs straight from a devbox snapshot on an existing private network
and times until the daemon accepts through the Mac's WireGuard hub, with no
cmux backend in the path. Measured 616-739 ms total (create 293-335 ms,
reachable +305-405 ms, n=5).

* perf(cloud): never hold a provider refresh on stats or the port scan

The New Machine open joins the provider's first refresh pass, and that
pass still awaited the stats HTTP read (~0.8 s) and the guest port scan
after publishing, so the terminal pane appeared 70-560 ms after the link
was up. Both now run as fenced follow-up tasks. Redial every 50 ms (cap
3 s) so reachability is detected within 50 ms instead of 200 ms.

* perf(cloud): start the clone's daemon before re-keying and prompt sync

Measured on a raw clone: bind at +0.13 s after resume, daemon listening
0.27 s later while ssh-keygen -A (RSA) and the prompt-sync Python start
competed for the clone's 2 vCPUs. The daemon now starts right after the
bind; prompt sync, host re-keying (nice 19, idle I/O), and the timer
re-arm follow it.

* chore(cloud): rebake every size with daemon-first clone boot (fastboot5)

* perf(cloud): first link to a new trusted machine skips the attach request

The app's link manager asked the control plane (POST /attach-endpoint,
Mac-to-backend round trip plus a provider status read) before its first
dial to any machine it had not linked before, so New Machine saw the
daemon ~0.3 s later than a raw probe. A snapshot-v2 create receipt now
records the carrier marker, so the first link dials --carrier directly.

* chore(cloud): probe raw-create reachability every 50 ms

* docs(cloud): record v9 New Machine measurements and remaining blocks

* fix(cloud): new machines show their name in the prompt in every environment

Removing the create-time guest exec left the prompt name to the guest's
reflection fetch through the Freestyle edge, which can only reach a public
origin; every tailnet dev backend returned 401 and VMs showed cmux@cmux.
The control plane now pushes the name once with the rename writer after
the create response (deferred, never awaited), and prompt sync treats a
name appearing in vm-name as published so it refreshes the first prompt.

* chore(cloud): rebake every size with the prompt-name refresh (fastboot6)

* fix(cloud): keep attaching machines created before the snapshot-v2 marker

openCmuxRemote refused every row without cmuxTuiContract, which is every
production machine created before this change: a Mac, CLI, or iOS client
without a saved device could no longer open them. Rows without the
marker now get the same private route (images since 2026-09-06, manaflow-ai#12042,
serve the trusted listener; older ones fail at connect rather than being
healed). A row that never recorded its addresses pays one provider read,
which the workflow persists. No guest exec either way.

* chore: drop root working notes from the branch

Their durable content lives in code comments (NO-WORK INVARIANT, boot
supervisor, catalog read) and the PR description.

* test: drop the source-shape plugin artifact test

It only asserted that workflow files contain certain strings (repository
policy: no source-shape tests), and it had no execution registry lane.
The hosted cmux-tui verification builds and uploads the detector artifact.

* fix(cloud): make the published port list a constant (warning budget)
…manaflow-ai#13958)

Bind workflow-dispatch producers to the actual E2E build recipe GitHub ran, keep dispatch-to-PR trust one-way under realistic PR metadata, and reuse compatible compiled app-host products instead of rebuilding.
…13968)

Record the desired terminal focus mirror when AppKit grants first responder even if the Ghostty runtime is still starting, so creation-time reconciliation can converge correctly.
Make the local-Undo routing regression independent of headless NSApp.keyWindow target resolution by targeting the editable responder explicitly.
…solve (manaflow-ai#13974)

* fix(ci): unquote replayed argv so post-manaflow-ai#13831 selectors resolve

The replay tool reads each batch's -only-testing selectors from the
argv that run-app-host-xcodebuild.sh records with printf '%q', then
unquoted them with .strip("'"). %q escapes shell metacharacters, and
since manaflow-ai#13831 gave Swift Testing selectors their trailing parens there
is now something to escape: the recorded line reads

  arg=-only-testing:cmuxTests/Suite/testFoo\(\)

145 of 235 selector lines in one real batch carry those escapes. Read
literally they match nothing, so check_run returns at its first gate
with "selector matched zero built tests" and never reaches the
analysis the tool exists to produce. Replaying run 35828371879 shard 2
against main emits 278 such lines and no verdict.

This was invisible when the tool was written because it was validated
against a pre-manaflow-ai#13831 run, where selectors had no parens and %q had
nothing to escape. It is broken on every run from here on.

Unquote with shlex, and fail loudly on an unbalanced quote rather than
skipping the line, since a silently shorter selector set is the exact
failure mode this tool exists to expose.

Also derive each batch's xcode status from its own log. The .meta does
not record it and the default was 65 for every batch, but check_run
treats the status as evidence: a batch that really exited 0 replayed
as 65 reports "xcodebuild exited 65 without a typed failed Test Case",
a contradiction that never occurred. xcodebuild's closing banner
carries it. --xcode-status stays as an override.

Before: 278 "selector matched zero built tests", no verdict.
After:  RATCHET_NEW_FAILURE SidebarHiddenPresentationTests/
          visibilityToggleKeepsAppKitTableContainerMounted()
        missing typed test result: TabManagerSessionSnapshotTests/
          testGhosttyFocusSurfaceIdRecordsMappedPanelInFocusHistory()
        RATCHET_NEW_FAILURE WorkspaceContentViewVisibilityTests/
          testMinimalModeToggleDoesNotReevaluateChromeHeavyBodies()

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix(ci): refuse argv forms the replay parser cannot decode

Two hardening findings from review, both verified rather than reasoned.

printf %q falls back to ANSI-C $'...' when a value holds a
non-printable character, and shlex neither decodes that form nor
raises on it. Checked against a real bash subprocess rather than an
assumed encoding:

  printf %q -> "$'-only-testing:cmuxTests/S/testA\tB()'"
  shlex     -> ["$-only-testing:cmuxTests/S/testA\\tB()"]

The leading $ and the literal backslash make it fail the
-only-testing prefix test, so the selector is dropped with no error.
That is the silent selector-set shrink the unbalanced-quote guard was
added to prevent, arriving through a path that guard does not cover.
Unreachable for today's identifiers, which is why it is refused rather
than decoded.

%q output is also one token by construction, so more than one token
means the recorded argv is not what this parser assumes. Refuse that
too instead of silently taking the first.

Re-verified end to end against run 35828371879 shard 2: same two
batches, same recovered verdicts as before the hardening.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test: pin the replay parser against real printf %q output

The decoding here is easy to get wrong by reasoning about it -- both of us
did, in opposite directions, before anyone ran bash. So two of these build
their fixture by shelling out to `printf %q` rather than asserting an
assumed encoding, which is what would have caught the escaped parens at
review time instead of after the tool shipped non-functional.

The other four pin the refusals: ANSI-C `$'...'`, an unbalanced quote, and
multi-token argv. None is reachable for today's identifiers, but each would
otherwise drop a selector without a word, and a selector that silently
vanishes reads exactly like a test the batch never ran -- the finding this
tool exists to report. Five of the six fail against the parser as merged in
manaflow-ai#13972; the bare-identifier case passes both and is the control.

Registered in tests/test-execution.toml and run from the app-host-process
guard group, so the execution registry stays complete.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…3990)

* ci: mirror app-host batch output without racing a tail

The app-host unit-test step writes each batch to a regular file and mirrored
it into the Actions log with a background `tail -f`. When the batch exited,
the step slept 0.2s and killed the tail. That margin is what the batch's last
write had to beat, so a loaded runner could drop the final lines -- the ones
that say why a batch died.

`tests/test_ci_change_areas.py` drives this step with a fake `sleep`, which
makes the margin zero and the race plainly visible. Under 16 busy cores the
guard file failed 4 of 4 runs before this change and passes 4 of 4 after it,
with no change idle. Measured on the single test underneath, it passed 6 of
30 before and 30 of 30 after: the batch reached its crash branch every time,
but the message announcing it did not reach the log.

The step now tracks how many bytes of the batch file it has emitted and
flushes the remainder itself, in the polling loop and once after the batch
exits. There is no margin left to miss and one less background process. The
property the tail existed for is kept: the batch still writes to a regular
file, so detached test descendants cannot retain the CI capture pipe after
xcodebuild exits.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* fix: bound each app-host log flush to its size snapshot

* test: require app-host crash output exactly once

* fix: keep the exact-range reader inside the workflow block

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…w-ai#13973)

Five runner variables can name metered WarpBuild capacity, and between
2026-09-19 and 2026-09-23 all five did: `MACOS_RUNNER_15`,
`MACOS_RUNNER_DISPLAY` and `MACOS_RUNNER_DUAL_XCODE` pointed at
`warp-macos-15-arm64-6x`, `MACOS_RUNNER_26_RELEASE` and
`MACOS_RUNNER_26_NIGHTLY_BUILD` at `warp-macos-26-arm64-12x`. Pull requests
route through `MACOS_RUNNER_PR` and everything else through those, so main and
the merge queue ran on metered capacity while pull requests ran free, and
neither gate validated the other's pool.

Nothing in the repository could see it. A variable's value is not reviewable,
and every guard here reads workflow text -- which is also why
`warp-macos-26-arm64-12x` has been live for three days even though
`check_no_self_hosted_fleet_runners` rejects that exact label on sight.

Read the five through `CI_PAID_MACOS_OVERFLOW` instead. Unset, they are not
read at all and every lane takes its free Blacksmith fallback. Set to `1`, they
select the pool exactly as before. That restores the polarity
`tests/test_ci_repo_variable_defaults.py` already assumes everywhere else:
unset is the cheap reading. Turning paid capacity on now takes two admin
actions; turning it off takes either one, including a pull request anyone with
push access can merge.

`MACOS_RUNNER_26_NIGHTLY_BUILD` also had a fallback weaker than its documented
steady state -- `blacksmith-6vcpu-macos-26` against an intended
`blacksmith-12vcpu-macos-26` -- so gating it would have quietly halved the
nightly universal builder. The fallback now matches the intent.

The three guards that pinned an exact `runs-on:` expression are updated to the
gated form; two of them already matched loosely and needed no change.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…13969)

A pull request that trips a fast Linux guard is red about a minute in, but
its macOS compile admission is queued anyway and starts on a Mac half an
hour later. Over 40 failing CI runs in one 2.3-hour window, 25 of the 27
that reached macOS admission were already red when the debounce checkpoint
finished, and 416 of their 656 macOS job-minutes went to shards that had
not started yet.

The debounce job already waits on a cheap Linux runner, re-reads the pull
request head, and fails the run rather than admitting macOS when the head
moved. It now fails the same way when a job in this run has concluded
`failure`: the verdict is settled on this head, and the fix push opens a
new run.

Declining by failing is load-bearing rather than incidental. A job output
would read better -- the debounce job would stay green and the run would
carry one red job instead of two -- but `macos` is skipped either way, and
only a *failed* dependency makes "Re-run failed jobs" re-run it. An output
would strand that run red until someone re-ran everything.

Tradeoffs: a red run no longer reports whether macOS would also have
failed, so a pull request broken on both platforms needs a second round
trip. Re-run attempts and CI_MACOS_ADMISSION_DEBOUNCE_SECONDS=0 skip the
wait and with it this check, which is how you ask for those results anyway.
Every unreadable case admits -- a failed call, an empty reply, an error
body, or a first page that misses a later job.

macOS admission gains no dependency edge on the Linux suites, which
`test_ci_change_areas.py` forbids, but it does now depend on their results
in substance, and racily: whether a run collects macOS results depends on
whether a Linux guard fails before or after the debounce window closes.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…anaflow-ai#13994)

Follow-ups from the review of manaflow-ai#13973.

The gate guard matched the text before each read with a suffix check, so
`inputs.CI_PAID_MACOS_OVERFLOW == '1' && vars.MACOS_RUNNER_15` passed. It
now requires `vars.` itself. It also pins each gated variable's fallback
literal to its free steady state, so the nightly builder cannot drift
back to 6vcpu without failing CI.

docs/ci-runners.md's Tart restore recipe sets MACOS_RUNNER_15 and
MACOS_RUNNER_DISPLAY, which the gate now ignores unless
CI_PAID_MACOS_OVERFLOW=1. Say so where an admin will read it.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* docs: record how to read CI cost measurements

Three numbers were misread during E2E cost work today, each in a way that
would have sent the next session down a wrong path: a cancelled job's
duration taken as spend when the job never got a runner, scheme selection
treated as a compile lever when the app scheme is 94% of the build, and a
cache hit treated as a warm build when one drifted source file costs 457 s.

Each claim here is a measurement over a 98-run window, with the number
that produced it, so the next person can re-measure rather than trust it.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* docs: do not name an unmeasured cause for the one-file rebuild

Debug builds are not whole-module (SWIFT_COMPILATION_MODE is set only in
Release), so the timing difference stands but the mechanism does not.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…anaflow-ai#14000)

manaflow-ai#13962 made an incomplete xcresult still report the failures it recorded,
but an interrupted run (app-host restart, outer or idle timeout) returned
one gate earlier and stayed silent. On main's run 35862070143, shard 4
recorded two failures and named neither.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…low-ai#13996)

Compile admission builds the test bundle and stops, so a pull request that
edits cmuxTests/ runs none of the code it changed. `suite-coverage` refuses
that run and offers one escape: `full-ci`. That label is all-or-nothing. It
unlocks app-host unit tests, the package lane, the lag lane, release
admission and a Release build on a paid runner.

The routing already knows why the suite was wanted. UNJUDGED_BY_COMPILE_PREFIXES
names cmuxTests/ and cmuxUITests/ exactly, and app-host unit tests reads the
product admission already built, so the one job that judges a cmuxTests/ diff
is also the cheapest thing `full-ci` turns on. Everything else is a rider.

Add `unit-ci`, which asks for compile admission plus that job. The full suite
still implies it, so nothing about `full-ci` changes. `swift-package-tests`
already carried its own path-derived route, so this follows a shape the
workflow had.

Two couplings this has to respect:

- `tests` restates the macOS `if:` as a result contract, and says to keep
  them in sync. A routed unit-ci run must not read as legitimately unrouted,
  so macos_work_required accounts for it.
- The compile-only fork product is skipped as `unused_fork_product`, which
  was true only while no consumer ran under the policy. unit-ci makes app-host
  unit tests a consumer, so the product must still be published for it.
  test_ci_product_publication caught this; its model now covers the tier.

Defaults are unchanged: unit_suite is optional and empty for any caller that
predates it, and coverage_gap still refuses a cmuxUITests/ diff however the
run is routed, because no pull request job executes those.

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…-ai#14008)

unit-ci clears suite-coverage on the strength of `app-host unit tests`,
but it left the build-input reuse steps live. The usual way to use the
label is to add it after a first push, and that rerun finds an earlier
run that compiled the same inputs, sets compile_admitted, and skips
compile admission. App-host runs only behind an admission that
succeeded, so it skipped too, and macOS status did not require it
outside the full suite. The PR came out green having run no tests.

Skip reuse under unit-ci the way the full suite already does, and
make macOS status require app-host under unit-ci, so any other route
to a skip fails instead of passing.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…supported (manaflow-ai#13963)

* test(settings): enforce that advertised cmux.json paths are actually supported

`CmuxSettingsFileStore+SupportedPaths.swift` says of its set: "Settings UI rows
validate against this set so new persisted settings need an explicit cmux.json
review." Nothing enforced it. A row could declare
`configurationReview: .json("some.path")` while the store rejected that path, so
the row displayed a cmux.json key that silently did nothing when a user wrote
it, and the toggle never round-tripped into cmux.json.

I hit this writing a new terminal toggle and only caught it in review, which is
what prompted looking for the general case. The scan found seven more rows
already in this state, across five sections -- so it is a recurring failure mode,
not a one-off slip.

The guard is a source scan: every string literal passed to
`configurationReview: .json(...)` under `CmuxSettingsUI` must appear in the
supported set, or descend from an entry there, since object-valued settings are
listed at their root. Non-literal forms such as `.json(catalog.app.foo.id)` are
skipped because they name a catalog id that cannot be read without type
information. Symbolic entries in the supported set are resolved back to their
`static let` string so they count.

The seven existing offenders are recorded in `KNOWN_UNSUPPORTED` rather than
fixed here. Each needs its own mapping entry and belongs with whoever owns that
section; bundling them would make this an unreviewable change. A second test
asserts the list has no stale entries, so fixing a row requires deleting its
line and the list can only shrink.

Verified the guard is not vacuous: removing `app.minimalMode` from the supported
set makes it fail and name `AppSection.swift:283`.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* ci: run the settings configuration-review guard in the preflight lane

Registering a test in `tests/test-execution.toml` under `linux-guard` is not
enough: `validate_test_execution_registry.py` requires that a `linux-guard`
entry's path appear literally in a workflow, since that lane means "some
workflow runs this directly". Without the invocation the registry validation
fails with "linux-guard lane is not run by any workflow", which is what CI
reported.

Add the invocation to `ci-guards.yml` beside the other structural guards.

Verified with actionlint and by running the registry validator from this
worktree. My earlier local check was invalid: I ran the validator via an
absolute path into the main checkout, so it resolved its repo root there and
inspected a tree without this test at all, and reported success.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

* test(settings): say what this guard actually proves, and stop two false positives

An independent review found the guard's premise overstated and two of its seven
allowlist entries wrong. Both confirmed directly.

`supportedSettingsJSONPaths` has **no production consumer**. Repo-wide it is
read by this guard and one assertion in `FocusHistoryScopeTests`, and by nothing
that parses cmux.json; what actually accepts a key is the hand-written section
parsers in `KeyboardShortcutSettingsFileStore.swift` and
`CmuxSettingsFileStore+AppSection.swift`. So this guard compares two
declarations -- the UI row and the documented set -- and cannot prove that
writing a key does anything. The docstring now says so.

That gap is live in both directions. `app.globalFontMagnification` and
`shortcuts.showModifierHoldHints` were listed as "real defects" but are fully
parsed and applied (`CmuxSettingsFileStore+AppSection.swift:48` and
`KeyboardShortcutSettingsFileStore.swift:929`) -- they were simply missing from
the documented set, so they are added to it and dropped from the allowlist. The
remaining five were re-checked against the parsers by hand and are genuine:
`cloud`, `computerUse` and `customSidebars` have no top-level case in the
section dispatch, and the parsed `automation` section has no `codexIntegration`
key.

In the other direction, `canvas.paneGap` and `canvas.snappingEnabled` are
advertised, are in the supported set, and so pass this guard -- while
`root["canvas"]` is read by no parser at all. Filed separately; the docstring
names it as the class this oracle cannot catch.

Two mechanical fixes behind those:

`_resolve_symbol` matched the first file that merely *mentioned* the type name,
in `rglob` order, so it was both wrong under collision and machine-dependent.
`settingsPath` is already declared by two types. It now requires the file to
declare the type and reports ambiguity instead of guessing -- verified by
appending a decoy `static let settingsPath` to a file that only references
`SessionContentWidthSettings`, which previously hijacked resolution and now
does not.

The "can only shrink" claim was false: nothing compares the list to a baseline,
so a new failure could be parked in the same change that introduced it. The
comment now says that plainly rather than implying a ratchet that does not
exist. The staleness test also used exact membership while the main test used
ancestor matching, which stranded any entry that became supported via an
ancestor as permanently un-reapable; it now mirrors the ancestor match.

Mutation-verified after the change: the decoy no longer hijacks resolution;
making `computerUse` supported now reaps both allowlist entries; and deleting
`app.confirmQuit` from the set still fails with the exact advertising file:line.

Co-Authored-By: Claude Opus 5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…flow-ai#13998)

A `.tab(workspaceID:paneID:index:)` destination again places the new
mirror tab at the requested index. The manaflow-ai#13299 merge replaced
insertCloudManualMirrorTab with a copy that added iconAssetName but
dropped the index parameter and its reorder, so every mirror tab
landed after the selected tab. nativeMirrorTabInsertionHonorsTheSourceOrder
has failed on main since that merge (run 35862070143).

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…ader test (manaflow-ai#13999)

`asyncReaderSurvivesManyShortChunksAheadOfTheConsumer` polled for 200
serial drains against one 5 s deadline. Each poll sleeps 1 ms, and on a
loaded host each sleep wakes much later, so the shared budget ran out
partway through and the test failed at `drainedEveryWrite` with nothing
wrong in the reader. Each write now gets its own 5 s, so a reader that
stops draining still fails on the write it stalls on.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
`unit_test_suites` accepted only a bare suite name, so the cheapest way to
run one app-host case on a quiet runner was the whole suite, with every
earlier case in the same process. That is the wrong tool for a question
like "does this test fail from a fresh process", and the only alternative
was a shared Mac where a 5 s wait can fail from load alone.

It now also accepts `Suite/testName` (XCTest) and `Suite/testName()`
(Swift Testing). The selector stays restricted to identifier characters
around one slash and optional trailing parens; the log file is named with
the slash turned into a dot. A paren-less Swift Testing selector still
fails the run through the existing positive-summary check.

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
@teamleaderleo
teamleaderleo deployed to cloud-vm-image-checks September 23, 2026 17:39 — with GitHub Actions Active
@teamleaderleo
teamleaderleo deployed to cloud-vm-image-checks September 23, 2026 17:41 — with GitHub Actions Active
@teamleaderleo
teamleaderleo deployed to cloud-vm-image-checks September 23, 2026 17:46 — with GitHub Actions Active

This branch was successfully deployed

1 active deployment
cloud-vm-image-checks — cbd4032b Deployed Sep 23, 2026 by teamleaderleo via reachable #8
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

4 participants