Repository navigation
ci: smoke-test zero-config fork runners - #96
Closed
teamleaderleo wants to merge 106 commits into
Closed
teamleaderleo wants to merge 106 commits into
teamleaderleo wants to merge 106 commits into
Conversation
…nt (manaflow-ai#13907) * test(simulator): bound the panel waits by a deadline, not a yield count `"A surviving Simulator host keeps framebuffer publication active"` failed on manaflow-ai#13414's app-host shard 1 with `frameTransport → nil` at SimulatorPanelThemeTests.swift:95 and :101, while passing on `main` and on manaflow-ai#13752's shards. Nothing in that PR touches this subsystem. The waits here spun a fixed `for _ in 0..<100 { ... await Task.yield() }` before asserting. A yield count is not a deadline: `Task.yield()` gives the scheduler a chance to run something else, it does not wait for anything, so 100 yields is however long 100 reschedules happen to take. The bound tightens exactly when the runner is busy, which is the "fails on correct code under load" shape `.github/review-bot-rules/test-determinism.md` bans: the deadline bounds the FAILURE path only, so load can make a pass slower but never turn a pass into a fail These 10 sites now poll the same predicate against a `ContinuousClock` deadline. They return the instant the condition holds, so a passing run is no slower, and only a genuinely broken one waits out the 10 s. No assertion changed. Two sites are deliberately left alone: the bare `for _ in 0..<100 { await Task.yield() }` at SimulatorPanelIntegrationTests.swift:187 and :252 assert an *absence* (`discoveryCount == 0`, `!isCompleted`). A deadline cannot bound a wait for a non-event; those need a positive completion signal to assert after, which is a change to what the test observes rather than how long it waits. Their failure direction is also the opposite one — a short spin makes a false green, not a false red. Scope note: 17 more yield-count polls remain in 12 other `cmuxTests` files, and `scripts/check-test-determinism.py` reports 0 findings across the tree because it only matches sleep call sites. Tracked in manaflow-ai#13903; this commit is the cluster with the observed failure. Verified: `swiftc -parse` on both files; `check-test-determinism.py --roots cmuxTests`, `test_ci_pbxproj_test_wiring.sh` and `validate_test_execution_registry.py` pass; 127 of the 128 `ci-guards.yml` commands pass, the exception being `test_ghostty_zig_version_sync.sh`, which needs the ghostty submodule this worktree does not check out. Not run on a macOS runner — these tests need one, so CI is their first execution. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(simulator): wait on a real signal, and share one deadline helper Review of the first revision found two things the yield-count -> deadline conversion did not fix. `SimulatorPaneCoordinator.receive(.frameTransport:)` stores the descriptor only `if frameIsVisible`, and nothing re-sends it. Polling for `coordinator.frameTransport != nil` therefore cannot recover a descriptor that arrived early: it turns a permanent drop into a ten-second wait for a latch that will never flip. The surviving-host test now waits for `.setFramebufferPublishing(true)` — the message `reconcileFramePublication` enqueues when it sets `frameIsVisible` — before emitting the descriptor, so the drop window is closed rather than polled across. The ten converted waits also each ended in a bare `break`, so a wait that ran out fell through into whatever assertion came next; one of them (`cancelledApplicationTerminationRestoresPanel`) had no assertion on its own predicate at all and reported a confusing downstream failure instead. They now share a `waitUntil` helper that requires its predicate at the deadline, which makes that shape impossible to write and matches the spelling already used in CloudTerminalCardFlashRegressionTests and CloudTunnelLaunchGateTests. The helper sleeps 5 ms between probes rather than spinning `Task.yield()`, so a failing wait no longer pegs the main actor for its full budget. The two bare settle spins are left alone: both assert an absence, so a deadline would only make them slower without adding signal. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Revert the publication gate; CI disproved it Run 35831156475 on macOS 26 failed at the gate itself: the wait for `.setFramebufferPublishing(true)` ran its full 10 s and never saw the message, while every other wait in both suites passed. `reconcileFramePublication` returns early unless `status == .streaming` (SimulatorPaneCoordinator+FrameVisibility.swift:72), so the publication message is not the reliable precursor to the descriptor I took it for, and gating the emit on it only deadlocks the test against a message that this path does not produce here. The premise was inferred, not measured; the measurement says it is wrong. The census also does not support calling this a real defect. "A surviving Simulator host keeps framebuffer publication active" passes on main at f3d204a, a9b0329 and ce1c55c, and passed both earlier dispatches of this branch. Its one observed failure is on manaflow-ai#13414's run, which had 148 failing tests and corresponding load — which is exactly the shape a fixed yield count fails under, and exactly what the deadline conversion fixes. So the original premise stands on its own and the escalation was unnecessary. The `waitUntil` helper stays: the full "Simulator panel integration" suite passed in the same run, covering all seven converted sites there. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…anaflow-ai#13951) * coderouter: accept chatmux per-VM tokens with team-shared account access chatmux (chatmux.dev) signs a one-hour ES256 token for each of its Freestyle VMs with an HSM key; the Freestyle edge injects it as x-chatmux-vm-authorization, so the guest never holds it. coderouter verifies it against chatmux's JWKS (issuer allowlist, audience "coderouter", one-hour maximum lifetime, required claims) with no database lookup. When the header is present it is the only credential considered and a bad token fails closed. A chatmux machine gets a new access kind, team-machine: only accounts its Hexclave team shares (visibility "team"), never anyone's private account. It has no pool and no cloud_vms row. Off until CODEROUTER_CHATMUX_JWKS_URL (https) and CODEROUTER_CHATMUX_ISSUERS are set. * coderouter: keep chatmux machines on the data plane A chatmux VM token now fails in the account control plane (it could add or remove team accounts through a route-token header) and in every Cloud VM-only route (VM principal, /api/vm/self, vm-usage/self, subrouter teams). Also: no control-character regex, and a pinned clock in the verifier tests.
…low-ai#13949) The workspace-rooted path this PR moves was invisible to every check. `check_every_app_host_home_is_identified_and_cleaned` already finds each job that prepares an app-host home and holds it to that contract, but it asserted only the two preconditions `prepare` needs, not the one `cleanup` enforces at runtime: DerivedData under RUNNER_TEMP. Extending that loop catches the shipped bug at its pattern rather than its instance. Reverting the workflow to the pre-fix path fails it, as does moving ci-macos.yml's shard DerivedData out of RUNNER_TEMP or dropping the publishing step. The test job's own `Clean owned DerivedData` arm also had no coverage: `step()` resolves an ambiguous name to the build job, so a typo in the test job's ownership pattern shipped green and failed only on a runner, under `if: always()`, after the tests had passed. Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…incomplete (manaflow-ai#13962) * test: pin that an incomplete app-host run still names its failures On main's full suite at f3d204a (run 35788803553) all six app-host shards returned at the incompleteness gate in check_run(), so not one RATCHET_NEW_FAILURE line was printed across the entire run -- even though the logs carried real assertion failures. A red suite that names no regression cannot tell anyone whether a fix landed. This commit adds the failing test only. It asserts that a run with one missing terminal result still reports the new and known failures it did record, and a companion asserting no RATCHET_ noise when nothing failed. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(ci): report recorded failures even when the xcresult is incomplete check_run() returned at the missing-execution gate, discarding the ratchet comparison for the whole shard. One selected test with no terminal result was enough to stop app-host-known-failures.json being consulted at all, so a shard whose remaining tests regressed and one whose remaining tests went green printed the same "incomplete" line. The verdict stays fail-closed -- an incomplete run is still a failed run. Only the diagnostics change: the failures that were recorded are now named alongside the incompleteness report. recorded_failure_diagnostics() is deliberately not the ratchet's own logic. The ratchet fails fast, reporting new failures and returning without mentioning known ones because the verdict is already settled; a run being reported for some other reason wants the complete picture instead. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(ci): summarize the incomplete path's verdicts like the complete one Review point from the consuming session: the complete path ends with "known-main failures tolerated: N; typed test cases: M", and without an equivalent here a reader scanning shard output for evidence that the accounting ran still sees silence on the incomplete path. Emitted only when something was recorded, so the exact-match assertion in test_missing_selected_test_result_never_passes keeps holding. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* test: cover optimistic dashboard team switching * fix: make dashboard team switching optimistic * fix: keep a catalog snapshot for rollback * fix(web): keep optimistic team switches ordered * test(web): cover overlapping team switches * fix(web): snapshot legacy team scope through parser * fix: serialize overlapping team persistence * fix: preserve rapid team-switch ordering * test: cover overlapping team-switch failures * fix(web): keep the team-scope probe's declared type through the reset `renderReadyScope` assigns `probedScope = undefined` and then renders the Probe, which reassigns it from inside a closure. Control flow analysis cannot see that write, so it held the variable at `undefined` for the rest of the function and `!scope` narrowed the union away entirely — every `scope.switchTeam` in the suite failed `web-typecheck` with "does not exist on type 'never'". Read the probe through a function so the declared type survives, and give `renderReadyScope` an explicit ready-variant return type so the call sites do not depend on inference through the same reset. Types only; the 12 tests in this file passed before and after. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Leo <cheerleaderleo@outlook.com> Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…ai#13966) * ci: report how many runner minutes came from paid capacity The scheduled CI health report totals runner minutes by label four times a day, but every label reads the same: Blacksmith, GitHub-hosted and WarpBuild sit in one table with nothing marking which of them bill. Blacksmith and GitHub-hosted are free to this repository; Warp is metered per minute, at roughly double the rate on its 12-vCPU labels. A lane that drifts onto paid capacity therefore renders as an ordinary row and nobody notices until an invoice arrives. `MACOS_RUNNER_15`, `MACOS_RUNNER_DISPLAY`, `MACOS_RUNNER_DUAL_XCODE`, `MACOS_RUNNER_26_RELEASE` and `MACOS_RUNNER_26_NIGHTLY_BUILD` have all been pointing at Warp since 2026-09-19/20, so main and the merge queue run on metered capacity while pull requests run free. That is a legitimate overflow response to a saturated pool, but it is invisible, and `docs/ci-runners.md` records a different intended steady state for each of those variables. Split the paid labels out into their own line with a per-label breakdown, so the report says how many metered minutes the window actually contained and points at the doc that says what the steady state should be. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * ci: count every metered provider, and stop contradicting the runner docs An independent review found three problems with the first cut. **Depot was uncounted.** `PAID_RUNNER_PREFIX` was the single string `"warp-"`, but `depot-*` is permitted by `tests/test_ci_self_hosted_guard.sh` and documented alongside Warp. Pin a variable to Depot and the report totals zero and prints "none in the window" -- exactly the silent drift this line exists to catch. Now a tuple of prefixes, with a test that fails on the warp-only form. **The render path was untested.** Deleting the entire `if paid_jobs:` block left all 76 tests green: the accumulator was covered, the rendering was not, and `test_the_waste_patterns_reach_the_output` asserts exactly this for every other pattern. Added there; verified the deletion now fails. **The line contradicted the doc it told you to read.** It claimed every other label "is free to this repository", while `docs/ci-runners.md` calls `blacksmith-*` a paid-provider label and `docs/ci/workflow-inventory.md` says Blacksmith and GitHub-hosted bill at different rates. Both are defensible readings of "paid" and that is the problem: a report that ends with "check this against docs/ci-runners.md" cannot disagree with it. Resolved in favour of precision on both sides. The report now says what is actually true and checkable -- these are the metered third-party labels; Blacksmith is sponsored for this organization and GitHub-hosted is free on a public repo -- and `docs/ci-runners.md` now distinguishes "not a GitHub-hosted free runner" from "bills this repository per minute". The header no longer says "(WarpBuild)" now that it covers more than Warp. Also documented the pattern in `docs/ci/health-report.md`, which enumerates every other one. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* test: reproduce OMP auto-naming flag rejection * fix: use OMP-supported auto-naming isolation flags
…anaflow-ai#13952) The lane guards added alongside the routing fix name the three macOS jobs they cover. A macOS job added next month inherits none of them: it can pin CMUX_CI_XCODE_APP_MACOS_15 while its pool follows MACOS_RUNNER_PR, and scripts/select-ci-xcode.sh hard-exits on a pinned path the image lacks, so the job fails at Xcode selection the first time the lane moves. Invert it. Every env value under .github/workflows that names the macos-15 Xcode must also read the pull-request variant, unless its exact (file, job, key) is listed in EXEMPT with a written reason. Three sites are exempt today: the streamed-validation lane and the two SDK 15 pins in swift-package-tests, which is deliberately carved out onto the dual-Xcode pool. Keyed per (file, job, key) rather than per file so exempting swift-package-tests does not silently exempt every other job in ci-macos.yml. Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…ow-ai#13972) Answering "what would this shard have reported, and did this test even run?" has required a Mac or a CI round trip. Both are scarce: app-host shards only run under full-ci, and a shard's own output can omit the tests that never produced a terminal result. Everything needed is already uploaded. cmux-app-host-diagnostics-shard-N carries the typed xcresult JSON and the captured log, and each batch's .meta records the argv it ran, so the exact -only-testing selectors are recoverable. cmux-app-host-test-inventory is about 1 MB. This runs the real check_run() over them on any machine. --accounting points the replay at a different copy of the module, which is how manaflow-ai#13962 was measured: origin/main emitted 0 RATCHET lines on main's run 35788803553 shard 6 where the fixed module emitted 9, on byte-identical inputs. Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
) * Welcome first-time contributors and credit their work Of 824 open PRs from outside the team, 779 had only bot comments. First-time contributors now get one short human-written note on their first PR, and CLAUDE.md tells agents to land or credit an outside PR before writing their own fix for the same problem. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Exclude team and PR-farm accounts from the outside-PR count Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> * Greet only actual first PRs, and stop contradicting the PR template An independent review found the gate does not do what its name says, and measured the consequence rather than theorising it. `author_association` answers "has no merged commit in this repo", not "this is their first pull request". cmux merges almost no outside PRs -- that is this change's own premise -- so a persistent contributor stays FIRST_TIME_CONTRIBUTOR indefinitely. Of 205 open PRs carrying that association, only 145 are distinct authors: 29% would have been greeted at least twice. `brodynies` has 16 open PRs and would have received 16 copies of "thanks for opening your first cmux pull request". `danielraffel` would have got a sixth greeting three and a half months after the first. That is precisely the "you are only talking to a bot" experience this workflow exists to fix, aimed at the outside contributors who kept showing up anyway. Count the author's pull requests instead; this one is always included, so more than one means it is not their first. An unusable answer is treated as "skip" rather than as an error: the search API can rate-limit or 422, and a non-numeric reply would otherwise abort the step under `set -e` and leave a red X on a newcomer's first PR. A missed greeting is recoverable; a wrong greeting or a red check is not. A marker comment makes a duplicate impossible even when the count is stale, which search results can be for PRs opened in a burst. The note also contradicted the template the contributor had just filled in. It said not to @mention review bots and named two of them, while `pull_request_template.md` ships a "Review Trigger" block of four mentions to paste as a comment, and the checklist asks the contributor to confirm they did. Point at the template instead of against it. Finally, the note promised "a person on the team reads every outside PR" while 765 of 810 open outside PRs have only bot comments. Posting that to every new contributor is a promise the repo does not keep, and it invited a bump comment on nearly all of them. Say what is true: the queue is long, a reply can take a while, and a nudge is welcome. Also declare the `contents: read` the API read needs, pin `branches: [main]` to match every other `pull_request_target` workflow here, and read the note from `github.workflow_sha` rather than `base.sha`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…anaflow-ai#13959) * fix: honor CMUX_SSH_RECONNECT_LIMIT above 20 instead of discarding it The SSH PTY attach supervisor accepted only 1-20 attempts and rewrote everything else to 20 without a word, so `CMUX_SSH_RECONNECT_LIMIT=50` silently became 20. The same name defaulted to 86400 in the freestyle supervisors, so one variable carried two answers 4300x apart. SSHReconnectBudget now owns both numbers: 20 when the variable is unset or unusable, 86400 as the ceiling and as the freestyle default. The attach supervisor and the CLI startup script generate their clamp from it, and a rejected value prints one line to stderr naming the value it refused and the value it used. Failing closed is preserved. Non-digits, empty text, and 0 still fall back to 20; an oversized count is rejected by digit length before any `[ ... -gt ... ]` can error out on it instead of answering. The reconnect delay knobs keep their own hand-written clamps: they need a ceiling too, but that changes backoff behavior and belongs with manaflow-ai#10419 and manaflow-ai#10422. Fixes manaflow-ai#13922 Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * Make the reconnect budget an instantiated value `lint-ios-conventions-diff.sh` rejects a new all-static public enum as a namespace type, asking for "an extension on the receiver type or an instantiated value with injected dependencies": NEW namespace-type SSHReconnectBudget.swift enum SSHReconnectBudget (all-static public surface, not instantiable) `SSHReconnectBudget` is now a struct with `limitEnvironmentName`, `fallbackLimit` and `maximumLimit` as stored properties defaulted in the initializer, and `limitNormalizationShellLines` as an instance method. Call sites construct it. The `fallback` parameter became `Int?` because a default argument can no longer name another stored property of the same value. Behaviour is unchanged. Compiling the type on Linux, emitting the shell and running it over the same table gives: unset and `20` resolve to 20 silently, `21`/`50`/`86400` resolve to themselves, `007` resolves to 7, `86401` and a 20-digit value resolve to 86400 with the notice, and `abc`/`-5`/`0` resolve to 20 with the notice. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
* api: report literal text delivery outcome * test: cover paste deferral and rejected delivery * test: verify paste acknowledgement against raw PTY bytes
…anaflow-ai#13299) * feat(tui): adopt Herdr-style agent plugin architecture * perf(cloud): move VM guest setup into snapshot contract * docs: preserve Herdr handoff for startup work * test(cloud): assert snapshot guest contract * fix(cloud): race first terminal seed with prompt identity * chore(cloud): promote snapshot-v2 desktop ladder * perf(cloud): warm snapshot terminal and preserve session journal * test(cloud): verify warm prompt contract * fix(cloud): derive seeded terminal from session snapshot * fix(cloud): tolerate empty warm session before seed * fix(cloud): verify warm terminal through terminal list * chore(cloud): promote exact warm Herdr snapshot * fix(cloud): reuse terminal-list ids during prompt seed * fix(cloud): inspect terminal list for existing seed * chore(cloud): promote prompt-seed-safe Herdr snapshot * fix(cloud): persist terminal id after workspace creation * chore(cloud): promote final warm Herdr snapshot * Restore exact warm snapshot manifest after main merge * chore(cloud): keep snapshot source provenance exact * chore(cloud): record exact cmux-tui pin in manifest * test(cloud): cover warm snapshot prompt contract * chore(cloud): promote merged-source warm snapshot * test(cloud): align contracts with snapshot-v2 attach * fix(merge): retain current Workspace cloud mirror helpers * fix(merge): preserve cloud terminal icon helpers * fix(merge): restore device manual mirror helper * fix(cloud): use current Stack SDK in startup benchmark * feat(mac): add tag-scoped runtime config profiles * test(cloud): New Machine open must link a just-created VM Covers the regression where cmux vm open read the cached catalog for a VM the app had not discovered yet and failed with "The machine's sessions are unavailable". * fix(cloud): link a new machine before resolving its first terminal New Machine open dropped refresh from surface.catalog to save work, but a VM created a moment ago has no provider or link in the app, so the cached read had no graph and open failed with "sessions are unavailable". Add an ensure_linked read mode: discover the machine if missing and join the provider's current refresh pass only when the machine is not already linked. Reopen of a linked machine costs no network; New Machine pays exactly one connect plus one graph read. * chore(cloud): delete dead Freestyle heal helpers and guard the no-work paths Create, restore, resume, and attach no longer install, heal, or resize, so the daemon health/start/pin-check commands, the grow-only resize helpers, and the driver's daemon manifest dependency had no production caller. Remove them and their tests, fix comments that still described attach-time healing, and document the NO-WORK INVARIANT at the top of the driver so new guest work goes into the snapshot. * docs(cloud): record the ensure_linked open requirement * perf(cloud): open a new machine from its create receipt Measured New Machine (5.9 s) spent 2.2 s on POST /attach-endpoint and ~0.3 s re-reading the fleet list, both for data the create already had, and ~0.7 s holding the first terminal behind a cosmetic stats read. - POST /api/vm returns the private address and cmuxTuiContract. - The app registers the provider straight from that receipt and answers the first vm.cmux_remote_info locally for a snapshot-v2 machine. - Provider refresh publishes the graph before stats; gauges fill in when the stats read lands. * perf(cloud): redial a fresh machine and link it from the create receipt - CloudHubConnector redials each address every 200 ms (capped) instead of one attempt riding TCP backoff; a new VM is reachable ~0.4 s after the create response but single attempts waited ~3.7 s or failed at 15 s. - The registry starts the new machine's first link and graph read as soon as the create receipt registers it; the open joins that pass. - vm.cmux_remote_info keeps a created receipt's declared route instead of a second connect race. - Provider refresh publishes the graph before the guest port scan. - DEBUG log of the create response's Server-Timing. * perf(cloud): take create usage events off the critical path; recover tunnel after 5xx - vm.create.requested runs beside model-plane provisioning and the provider call, joined before any failure event or return. - vm.created is written after the response via after() on the create route; restore, fork, and base keep it inline. - createTunnel bounds tunnels.create at 10 s and recovers the same-key tunnel after a 409, any 5xx, or no response. First enrollment hit a Freestyle 503 after ~25 s although the tunnel existed (502 in 30.9 s). * perf(cloud): detect clones within 50 ms, announce first, and quiet resume - cmux-devbox-boot ticks every 50 ms while parked (every snapshot is taken parked), sends the VPC announce before anything else on a clone, and parks housekeeping timers and service watchdogs until 10 min after bind. - The bake zeroes the workqueue watchdog; verify-devbox-image fails on any resume-time watchdog kill, lockup, or catch-up timer. - Rebaked every size (cmux-devbox-pr13299-fastboot3, daemon 01dc721); guest resume-to-listening 1.10 s -> 0.62-0.70 s. * docs(cloud): record New Machine latency floor and open decisions * chore(cloud): raw Freestyle create-to-reachable benchmark Creates VMs straight from a devbox snapshot on an existing private network and times until the daemon accepts through the Mac's WireGuard hub, with no cmux backend in the path. Measured 616-739 ms total (create 293-335 ms, reachable +305-405 ms, n=5). * perf(cloud): never hold a provider refresh on stats or the port scan The New Machine open joins the provider's first refresh pass, and that pass still awaited the stats HTTP read (~0.8 s) and the guest port scan after publishing, so the terminal pane appeared 70-560 ms after the link was up. Both now run as fenced follow-up tasks. Redial every 50 ms (cap 3 s) so reachability is detected within 50 ms instead of 200 ms. * perf(cloud): start the clone's daemon before re-keying and prompt sync Measured on a raw clone: bind at +0.13 s after resume, daemon listening 0.27 s later while ssh-keygen -A (RSA) and the prompt-sync Python start competed for the clone's 2 vCPUs. The daemon now starts right after the bind; prompt sync, host re-keying (nice 19, idle I/O), and the timer re-arm follow it. * chore(cloud): rebake every size with daemon-first clone boot (fastboot5) * perf(cloud): first link to a new trusted machine skips the attach request The app's link manager asked the control plane (POST /attach-endpoint, Mac-to-backend round trip plus a provider status read) before its first dial to any machine it had not linked before, so New Machine saw the daemon ~0.3 s later than a raw probe. A snapshot-v2 create receipt now records the carrier marker, so the first link dials --carrier directly. * chore(cloud): probe raw-create reachability every 50 ms * docs(cloud): record v9 New Machine measurements and remaining blocks * fix(cloud): new machines show their name in the prompt in every environment Removing the create-time guest exec left the prompt name to the guest's reflection fetch through the Freestyle edge, which can only reach a public origin; every tailnet dev backend returned 401 and VMs showed cmux@cmux. The control plane now pushes the name once with the rename writer after the create response (deferred, never awaited), and prompt sync treats a name appearing in vm-name as published so it refreshes the first prompt. * chore(cloud): rebake every size with the prompt-name refresh (fastboot6) * fix(cloud): keep attaching machines created before the snapshot-v2 marker openCmuxRemote refused every row without cmuxTuiContract, which is every production machine created before this change: a Mac, CLI, or iOS client without a saved device could no longer open them. Rows without the marker now get the same private route (images since 2026-09-06, manaflow-ai#12042, serve the trusted listener; older ones fail at connect rather than being healed). A row that never recorded its addresses pays one provider read, which the workflow persists. No guest exec either way. * chore: drop root working notes from the branch Their durable content lives in code comments (NO-WORK INVARIANT, boot supervisor, catalog read) and the PR description. * test: drop the source-shape plugin artifact test It only asserted that workflow files contain certain strings (repository policy: no source-shape tests), and it had no execution registry lane. The hosted cmux-tui verification builds and uploads the detector artifact. * fix(cloud): make the published port list a constant (warning budget)
…manaflow-ai#13958) Bind workflow-dispatch producers to the actual E2E build recipe GitHub ran, keep dispatch-to-PR trust one-way under realistic PR metadata, and reuse compatible compiled app-host products instead of rebuilding.
…13968) Record the desired terminal focus mirror when AppKit grants first responder even if the Ghostty runtime is still starting, so creation-time reconciliation can converge correctly.
Make the local-Undo routing regression independent of headless NSApp.keyWindow target resolution by targeting the editable responder explicitly.
…solve (manaflow-ai#13974) * fix(ci): unquote replayed argv so post-manaflow-ai#13831 selectors resolve The replay tool reads each batch's -only-testing selectors from the argv that run-app-host-xcodebuild.sh records with printf '%q', then unquoted them with .strip("'"). %q escapes shell metacharacters, and since manaflow-ai#13831 gave Swift Testing selectors their trailing parens there is now something to escape: the recorded line reads arg=-only-testing:cmuxTests/Suite/testFoo\(\) 145 of 235 selector lines in one real batch carry those escapes. Read literally they match nothing, so check_run returns at its first gate with "selector matched zero built tests" and never reaches the analysis the tool exists to produce. Replaying run 35828371879 shard 2 against main emits 278 such lines and no verdict. This was invisible when the tool was written because it was validated against a pre-manaflow-ai#13831 run, where selectors had no parens and %q had nothing to escape. It is broken on every run from here on. Unquote with shlex, and fail loudly on an unbalanced quote rather than skipping the line, since a silently shorter selector set is the exact failure mode this tool exists to expose. Also derive each batch's xcode status from its own log. The .meta does not record it and the default was 65 for every batch, but check_run treats the status as evidence: a batch that really exited 0 replayed as 65 reports "xcodebuild exited 65 without a typed failed Test Case", a contradiction that never occurred. xcodebuild's closing banner carries it. --xcode-status stays as an override. Before: 278 "selector matched zero built tests", no verdict. After: RATCHET_NEW_FAILURE SidebarHiddenPresentationTests/ visibilityToggleKeepsAppKitTableContainerMounted() missing typed test result: TabManagerSessionSnapshotTests/ testGhosttyFocusSurfaceIdRecordsMappedPanelInFocusHistory() RATCHET_NEW_FAILURE WorkspaceContentViewVisibilityTests/ testMinimalModeToggleDoesNotReevaluateChromeHeavyBodies() Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix(ci): refuse argv forms the replay parser cannot decode Two hardening findings from review, both verified rather than reasoned. printf %q falls back to ANSI-C $'...' when a value holds a non-printable character, and shlex neither decodes that form nor raises on it. Checked against a real bash subprocess rather than an assumed encoding: printf %q -> "$'-only-testing:cmuxTests/S/testA\tB()'" shlex -> ["$-only-testing:cmuxTests/S/testA\\tB()"] The leading $ and the literal backslash make it fail the -only-testing prefix test, so the selector is dropped with no error. That is the silent selector-set shrink the unbalanced-quote guard was added to prevent, arriving through a path that guard does not cover. Unreachable for today's identifiers, which is why it is refused rather than decoded. %q output is also one token by construction, so more than one token means the recorded argv is not what this parser assumes. Refuse that too instead of silently taking the first. Re-verified end to end against run 35828371879 shard 2: same two batches, same recovered verdicts as before the hardening. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test: pin the replay parser against real printf %q output The decoding here is easy to get wrong by reasoning about it -- both of us did, in opposite directions, before anyone ran bash. So two of these build their fixture by shelling out to `printf %q` rather than asserting an assumed encoding, which is what would have caught the escaped parens at review time instead of after the tool shipped non-functional. The other four pin the refusals: ANSI-C `$'...'`, an unbalanced quote, and multi-token argv. None is reachable for today's identifiers, but each would otherwise drop a selector without a word, and a selector that silently vanishes reads exactly like a test the batch never ran -- the finding this tool exists to report. Five of the six fail against the parser as merged in manaflow-ai#13972; the bare-identifier case passes both and is the control. Registered in tests/test-execution.toml and run from the app-host-process guard group, so the execution registry stays complete. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…3990) * ci: mirror app-host batch output without racing a tail The app-host unit-test step writes each batch to a regular file and mirrored it into the Actions log with a background `tail -f`. When the batch exited, the step slept 0.2s and killed the tail. That margin is what the batch's last write had to beat, so a loaded runner could drop the final lines -- the ones that say why a batch died. `tests/test_ci_change_areas.py` drives this step with a fake `sleep`, which makes the margin zero and the race plainly visible. Under 16 busy cores the guard file failed 4 of 4 runs before this change and passes 4 of 4 after it, with no change idle. Measured on the single test underneath, it passed 6 of 30 before and 30 of 30 after: the batch reached its crash branch every time, but the message announcing it did not reach the log. The step now tracks how many bytes of the batch file it has emitted and flushes the remainder itself, in the polling loop and once after the batch exits. There is no margin left to miss and one less background process. The property the tail existed for is kept: the batch still writes to a regular file, so detached test descendants cannot retain the CI capture pipe after xcodebuild exits. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * fix: bound each app-host log flush to its size snapshot * test: require app-host crash output exactly once * fix: keep the exact-range reader inside the workflow block --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…w-ai#13973) Five runner variables can name metered WarpBuild capacity, and between 2026-09-19 and 2026-09-23 all five did: `MACOS_RUNNER_15`, `MACOS_RUNNER_DISPLAY` and `MACOS_RUNNER_DUAL_XCODE` pointed at `warp-macos-15-arm64-6x`, `MACOS_RUNNER_26_RELEASE` and `MACOS_RUNNER_26_NIGHTLY_BUILD` at `warp-macos-26-arm64-12x`. Pull requests route through `MACOS_RUNNER_PR` and everything else through those, so main and the merge queue ran on metered capacity while pull requests ran free, and neither gate validated the other's pool. Nothing in the repository could see it. A variable's value is not reviewable, and every guard here reads workflow text -- which is also why `warp-macos-26-arm64-12x` has been live for three days even though `check_no_self_hosted_fleet_runners` rejects that exact label on sight. Read the five through `CI_PAID_MACOS_OVERFLOW` instead. Unset, they are not read at all and every lane takes its free Blacksmith fallback. Set to `1`, they select the pool exactly as before. That restores the polarity `tests/test_ci_repo_variable_defaults.py` already assumes everywhere else: unset is the cheap reading. Turning paid capacity on now takes two admin actions; turning it off takes either one, including a pull request anyone with push access can merge. `MACOS_RUNNER_26_NIGHTLY_BUILD` also had a fallback weaker than its documented steady state -- `blacksmith-6vcpu-macos-26` against an intended `blacksmith-12vcpu-macos-26` -- so gating it would have quietly halved the nightly universal builder. The fallback now matches the intent. The three guards that pinned an exact `runs-on:` expression are updated to the gated form; two of them already matched loosely and needed no change. Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…13969) A pull request that trips a fast Linux guard is red about a minute in, but its macOS compile admission is queued anyway and starts on a Mac half an hour later. Over 40 failing CI runs in one 2.3-hour window, 25 of the 27 that reached macOS admission were already red when the debounce checkpoint finished, and 416 of their 656 macOS job-minutes went to shards that had not started yet. The debounce job already waits on a cheap Linux runner, re-reads the pull request head, and fails the run rather than admitting macOS when the head moved. It now fails the same way when a job in this run has concluded `failure`: the verdict is settled on this head, and the fix push opens a new run. Declining by failing is load-bearing rather than incidental. A job output would read better -- the debounce job would stay green and the run would carry one red job instead of two -- but `macos` is skipped either way, and only a *failed* dependency makes "Re-run failed jobs" re-run it. An output would strand that run red until someone re-ran everything. Tradeoffs: a red run no longer reports whether macOS would also have failed, so a pull request broken on both platforms needs a second round trip. Re-run attempts and CI_MACOS_ADMISSION_DEBOUNCE_SECONDS=0 skip the wait and with it this check, which is how you ask for those results anyway. Every unreadable case admits -- a failed call, an empty reply, an error body, or a first page that misses a later job. macOS admission gains no dependency edge on the Linux suites, which `test_ci_change_areas.py` forbids, but it does now depend on their results in substance, and racily: whether a run collects macOS results depends on whether a Linux guard fails before or after the debounce window closes. Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…anaflow-ai#13994) Follow-ups from the review of manaflow-ai#13973. The gate guard matched the text before each read with a suffix check, so `inputs.CI_PAID_MACOS_OVERFLOW == '1' && vars.MACOS_RUNNER_15` passed. It now requires `vars.` itself. It also pins each gated variable's fallback literal to its free steady state, so the nightly builder cannot drift back to 6vcpu without failing CI. docs/ci-runners.md's Tart restore recipe sets MACOS_RUNNER_15 and MACOS_RUNNER_DISPLAY, which the gate now ignores unless CI_PAID_MACOS_OVERFLOW=1. Say so where an admin will read it. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* docs: record how to read CI cost measurements Three numbers were misread during E2E cost work today, each in a way that would have sent the next session down a wrong path: a cancelled job's duration taken as spend when the job never got a runner, scheme selection treated as a compile lever when the app scheme is 94% of the build, and a cache hit treated as a warm build when one drifted source file costs 457 s. Each claim here is a measurement over a 98-run window, with the number that produced it, so the next person can re-measure rather than trust it. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * docs: do not name an unmeasured cause for the one-file rebuild Debug builds are not whole-module (SWIFT_COMPILATION_MODE is set only in Release), so the timing difference stands but the mechanism does not. Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…anaflow-ai#14000) manaflow-ai#13962 made an incomplete xcresult still report the failures it recorded, but an interrupted run (app-host restart, outer or idle timeout) returned one gate earlier and stayed silent. On main's run 35862070143, shard 4 recorded two failures and named neither. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…low-ai#13996) Compile admission builds the test bundle and stops, so a pull request that edits cmuxTests/ runs none of the code it changed. `suite-coverage` refuses that run and offers one escape: `full-ci`. That label is all-or-nothing. It unlocks app-host unit tests, the package lane, the lag lane, release admission and a Release build on a paid runner. The routing already knows why the suite was wanted. UNJUDGED_BY_COMPILE_PREFIXES names cmuxTests/ and cmuxUITests/ exactly, and app-host unit tests reads the product admission already built, so the one job that judges a cmuxTests/ diff is also the cheapest thing `full-ci` turns on. Everything else is a rider. Add `unit-ci`, which asks for compile admission plus that job. The full suite still implies it, so nothing about `full-ci` changes. `swift-package-tests` already carried its own path-derived route, so this follows a shape the workflow had. Two couplings this has to respect: - `tests` restates the macOS `if:` as a result contract, and says to keep them in sync. A routed unit-ci run must not read as legitimately unrouted, so macos_work_required accounts for it. - The compile-only fork product is skipped as `unused_fork_product`, which was true only while no consumer ran under the policy. unit-ci makes app-host unit tests a consumer, so the product must still be published for it. test_ci_product_publication caught this; its model now covers the tier. Defaults are unchanged: unit_suite is optional and empty for any caller that predates it, and coverage_gap still refuses a cmuxUITests/ diff however the run is routed, because no pull request job executes those. Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…-ai#14008) unit-ci clears suite-coverage on the strength of `app-host unit tests`, but it left the build-input reuse steps live. The usual way to use the label is to add it after a first push, and that rerun finds an earlier run that compiled the same inputs, sets compile_admitted, and skips compile admission. App-host runs only behind an admission that succeeded, so it skipped too, and macOS status did not require it outside the full suite. The PR came out green having run no tests. Skip reuse under unit-ci the way the full suite already does, and make macOS status require app-host under unit-ci, so any other route to a skip fails instead of passing. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…supported (manaflow-ai#13963) * test(settings): enforce that advertised cmux.json paths are actually supported `CmuxSettingsFileStore+SupportedPaths.swift` says of its set: "Settings UI rows validate against this set so new persisted settings need an explicit cmux.json review." Nothing enforced it. A row could declare `configurationReview: .json("some.path")` while the store rejected that path, so the row displayed a cmux.json key that silently did nothing when a user wrote it, and the toggle never round-tripped into cmux.json. I hit this writing a new terminal toggle and only caught it in review, which is what prompted looking for the general case. The scan found seven more rows already in this state, across five sections -- so it is a recurring failure mode, not a one-off slip. The guard is a source scan: every string literal passed to `configurationReview: .json(...)` under `CmuxSettingsUI` must appear in the supported set, or descend from an entry there, since object-valued settings are listed at their root. Non-literal forms such as `.json(catalog.app.foo.id)` are skipped because they name a catalog id that cannot be read without type information. Symbolic entries in the supported set are resolved back to their `static let` string so they count. The seven existing offenders are recorded in `KNOWN_UNSUPPORTED` rather than fixed here. Each needs its own mapping entry and belongs with whoever owns that section; bundling them would make this an unreviewable change. A second test asserts the list has no stale entries, so fixing a row requires deleting its line and the list can only shrink. Verified the guard is not vacuous: removing `app.minimalMode` from the supported set makes it fail and name `AppSection.swift:283`. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * ci: run the settings configuration-review guard in the preflight lane Registering a test in `tests/test-execution.toml` under `linux-guard` is not enough: `validate_test_execution_registry.py` requires that a `linux-guard` entry's path appear literally in a workflow, since that lane means "some workflow runs this directly". Without the invocation the registry validation fails with "linux-guard lane is not run by any workflow", which is what CI reported. Add the invocation to `ci-guards.yml` beside the other structural guards. Verified with actionlint and by running the registry validator from this worktree. My earlier local check was invalid: I ran the validator via an absolute path into the main checkout, so it resolved its repo root there and inspected a tree without this test at all, and reported success. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> * test(settings): say what this guard actually proves, and stop two false positives An independent review found the guard's premise overstated and two of its seven allowlist entries wrong. Both confirmed directly. `supportedSettingsJSONPaths` has **no production consumer**. Repo-wide it is read by this guard and one assertion in `FocusHistoryScopeTests`, and by nothing that parses cmux.json; what actually accepts a key is the hand-written section parsers in `KeyboardShortcutSettingsFileStore.swift` and `CmuxSettingsFileStore+AppSection.swift`. So this guard compares two declarations -- the UI row and the documented set -- and cannot prove that writing a key does anything. The docstring now says so. That gap is live in both directions. `app.globalFontMagnification` and `shortcuts.showModifierHoldHints` were listed as "real defects" but are fully parsed and applied (`CmuxSettingsFileStore+AppSection.swift:48` and `KeyboardShortcutSettingsFileStore.swift:929`) -- they were simply missing from the documented set, so they are added to it and dropped from the allowlist. The remaining five were re-checked against the parsers by hand and are genuine: `cloud`, `computerUse` and `customSidebars` have no top-level case in the section dispatch, and the parsed `automation` section has no `codexIntegration` key. In the other direction, `canvas.paneGap` and `canvas.snappingEnabled` are advertised, are in the supported set, and so pass this guard -- while `root["canvas"]` is read by no parser at all. Filed separately; the docstring names it as the class this oracle cannot catch. Two mechanical fixes behind those: `_resolve_symbol` matched the first file that merely *mentioned* the type name, in `rglob` order, so it was both wrong under collision and machine-dependent. `settingsPath` is already declared by two types. It now requires the file to declare the type and reports ambiguity instead of guessing -- verified by appending a decoy `static let settingsPath` to a file that only references `SessionContentWidthSettings`, which previously hijacked resolution and now does not. The "can only shrink" claim was false: nothing compares the list to a baseline, so a new failure could be parked in the same change that introduced it. The comment now says that plainly rather than implying a ratchet that does not exist. The staleness test also used exact membership while the main test used ancestor matching, which stranded any entry that became supported via an ancestor as permanently un-reapable; it now mirrors the ancestor match. Mutation-verified after the change: the decoy no longer hijacks resolution; making `computerUse` supported now reaps both allowlist entries; and deleting `app.confirmQuit` from the set still fails with the exact advertising file:line. Co-Authored-By: Claude Opus 5 <noreply@anthropic.com> --------- Co-authored-by: Claude Opus 5 <noreply@anthropic.com>
…flow-ai#13998) A `.tab(workspaceID:paneID:index:)` destination again places the new mirror tab at the requested index. The manaflow-ai#13299 merge replaced insertCloudManualMirrorTab with a copy that added iconAssetName but dropped the index parameter and its reorder, so every mirror tab landed after the selected tab. nativeMirrorTabInsertionHonorsTheSourceOrder has failed on main since that merge (run 35862070143). Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…ader test (manaflow-ai#13999) `asyncReaderSurvivesManyShortChunksAheadOfTheConsumer` polled for 200 serial drains against one 5 s deadline. Each poll sleeps 1 ms, and on a loaded host each sleep wakes much later, so the shared budget ran out partway through and the test failed at `drainedEveryWrite` with nothing wrong in the reader. Each write now gets its own 5 s, so a reader that stops draining still fails on the write it stalls on. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
`unit_test_suites` accepted only a bare suite name, so the cheapest way to run one app-host case on a quiet runner was the whole suite, with every earlier case in the same process. That is the wrong tool for a question like "does this test fail from a fresh process", and the only alternative was a shared Mac where a 5 s wait can fail from load alone. It now also accepts `Suite/testName` (XCTest) and `Suite/testName()` (Swift Testing). The selector stays restricted to identifier characters around one slash and optional trailing parens; the log file is named with the slash turned into a dot. A paren-less Swift Testing selector still fails the run through the existing positive-summary check. Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
teamleaderleo
deployed
to
cloud-vm-image-checks
September 23, 2026 17:39 — with
GitHub Actions
Active
teamleaderleo
had a problem deploying
to
cloud-vm-image-checks
September 23, 2026 17:41 — with
GitHub Actions
Error
teamleaderleo
deployed
to
cloud-vm-image-checks
September 23, 2026 17:41 — with
GitHub Actions
Active
teamleaderleo
deployed
to
cloud-vm-image-checks
September 23, 2026 17:46 — with
GitHub Actions
Active
This branch was successfully deployed
This file contains hidden or bidirectional Unicode text that may be interpreted or compiled differently than what appears below. To review, open the file in an editor that reveals hidden Unicode characters.
Learn more about bidirectional Unicode characters
Sign up for free
to join this conversation on GitHub.
Already have an account?
Sign in to comment
Add this suggestion to a batch that can be applied as a single commit.This suggestion is invalid because no changes were made to the code.Suggestions cannot be applied while the pull request is closed.Suggestions cannot be applied while viewing a subset of changes.Only one suggestion per line can be applied in a batch.Add this suggestion to a batch that can be applied as a single commit.Applying suggestions on deleted lines is not supported.You must change the existing code in this line in order to create a valid suggestion.Outdated suggestions cannot be applied.This suggestion has been applied or marked resolved.Suggestions cannot be applied from pending reviews.Suggestions cannot be applied on multi-line comments.Suggestions cannot be applied while the pull request is queued to merge.Suggestion cannot be applied right now. Please check back later.
Temporary validation PR for manaflow-ai#14023. The acceptance condition is that the normal CI workflow starts on GitHub-hosted runners in this fork with no runner variables or Blacksmith installation. Close after the runner-routing evidence is captured.