Check that the shipped devbox image is reachable, only when the pair changed - #12132
Conversation
…changed The bake smokes the daemon once, when an image is made. Nothing asks whether a client can reach the daemon in the image we are shipping now, using the cmux-tui build we are shipping now. Those drift apart with no commit at all: files.cmux.com moves to a new client on every cmux-tui release while the manifest keeps pinning yesterday's snapshot. The check boots the manifest's current default and runs a real round trip (enroll a device, connect over ws://[::1]:1337/v1/link, open a workspace, run a process, read its output). When the live client differs from the baked one it repeats the whole trip with the live binary, so a protocol change that would break a freshly built app fails here instead of in production. It costs 17s with both legs, 9s when the client is the baked one, and one small VM that is deleted on every path. Running it when nothing changed is waste, so the job keys a cache on the default image ids plus the client sha256 and skips a pair that has already passed. That makes the daily schedule, which exists to catch the client moving without a commit, free on a day when it did not.
|
Warning Review limit reachedNext included review available in 3 minutes. View limit detailsLimit details: You’ve used all 10 included reviews currently available. You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository. Review configuration: ⚙️ Run configurationConfiguration used: Path: .coderabbit.yaml Review profile: ASSERTIVE Plan: Advanced Run ID: 📒 Files selected for processing (3)
Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out. Comment |
|
All contributors have signed the CLA ✍️ ✅ |
Run 34210445406 booted freestyle/ubuntu-sm, waited out the 120s daemon budget, reported "UNREACHABLE: the baked daemon never came up", deleted the machine anyway, and exited 1, so the job is red when the image is broken and leaves nothing running behind it.
There was a problem hiding this comment.
Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.
❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.
Reviewed by Cursor Bugbot for commit 95a83a5. Configure here.
|
|
||
| concurrency: | ||
| group: cloud-vm-image-reachability-${{ github.ref }} | ||
| cancel-in-progress: true |
There was a problem hiding this comment.
Cancelled runs can leak VMs
Medium Severity
Concurrency cancels an in-progress reachability job when a new run starts on the same ref. The guest is only deleted in the script finally block, which does not run on SIGTERM, so the sm VM can stay billed after the job is cancelled.
Additional Locations (1)
Reviewed by Cursor Bugbot for commit 95a83a5. Configure here.
31cdb0a Check that the shipped devbox image is reachable, only when the pair changed (manaflow-ai#12132) 93cb5be Polish Find and Vault search spacing, corners, and hover (manaflow-ai#12109) bf8ae03 cloud: give devbox panes the Ghostty terminal identity (manaflow-ai#12125) f65fe42 web: run Cloud VM routes on one Effect runtime with a tag-keyed error table (manaflow-ai#12122)
`import.meta.main` (added on main in #12132) fails `tsgo --noEmit` because the web tsconfig does not include bun-types, which turns the CI cheap layer red and skips every macOS job. Cast the meta object so Bun's runtime flag still gates main() while the typecheck passes. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa
`import.meta.main` is Bun-only and the web typecheck program has no Bun typings, so `bun run typecheck` fails on main since #12132: scripts/check-devbox-image-reachable.ts(148,17): error TS2339: Property 'main' does not exist on type 'ImportMeta'. That PR never ran the web typecheck because ci.yml only fires for a short path list. Guard the entrypoint by comparing import.meta.url with argv[1], the pattern the other scripts under web/scripts already use. Direct runs still execute main (`--print-key` prints the key) and importing the module from tests stays side-effect free. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
) `import.meta.main` is Bun-only and is not part of TypeScript's ImportMeta type, so `bun run typecheck` has failed on main since #12132: scripts/check-devbox-image-reachable.ts(148,17): error TS2339: Property 'main' does not exist on type 'ImportMeta'. Use the portable `fileURLToPath(import.meta.url)` vs `process.argv[1]` check that the sibling devbox scripts already rely on. Behavior is unchanged: `bun scripts/check-devbox-image-reachable.ts --print-key` still runs main, and importing `reachabilityKey` from the test does not. Closes #12147 Claude-Session: https://claude.ai/code/session_01Qiy38q5XFYQU3CeccDENty Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* test: cover creation initial command plumbing * test: isolate creation command regression * feat: add initial commands to terminal creation * test: require creation request dispatch * test(web): preserve hosted config overrides * test(cli): expect OMP restore path * test(session): expect codex wrapper shim * docs(cli): sync restore help contract * test(cli): preserve interactive shell for creation command * fix(cli): inject creation commands into interactive shells * fix(cli): preserve terminal creation command text * test(cli): reject invalid creation command input * fix(cli): validate terminal creation input boundaries * fix(cli): resolve creation validation in app context * test(cli): cover review-found creation boundaries * test(cli): import welcome setting contract * test(cli): wait for terminal readiness before welcome assertion * test(cli): exercise focused workspace welcome path * fix(cli): close terminal creation review gaps * test: cover app-host temp path aliases * ci: accept validated macOS temp aliases * Fix Xcode 26 warning regressions * test: keep install command fixture inert * test: repair app-host CLI fixtures * fix: remove duplicate dock initial input plumbing * test: bound workspace readiness regression wait * fix: order dock creation arguments * fix: order dock input after startup options * fix: preserve null terminal creation types * fix: keep readiness wait actor-safe * test: isolate creation command socket environment * test: keep cloud splits local for initial input * fix: keep initial input out of cloud split routing * ci: route browser skill contract through linux runner * test: isolate cloud routing fixture from local surfaces * refactor: keep terminal creation checks within file budgets * fix: wire terminal creation helpers into app target * fix: repair current main compile errors * docs: document --command on terminal creation commands Cover the new --command flag on new-split, new-pane, and new-surface (and the existing one on new-workspace) in the cmux and cmux-workspace skills, the CLI contract (table rows, an Initial terminal command section, and --help probes), and the web API docs in English and Japanese. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa * fix: silence unmutated payload warning in browser key replay CI's Swift warning budget fails on main because the delivered-key payload in TerminalController.swift is declared var but never mutated. Use let so the tests-build-and-lag job can pass on this branch. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa * fix(web): keep devbox reachability script typechecking without bun-types `import.meta.main` (added on main in #12132) fails `tsgo --noEmit` because the web tsconfig does not include bun-types, which turns the CI cheap layer red and skips every macOS job. Cast the meta object so Bun's runtime flag still gates main() while the typecheck passes. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa * fix: repair package test import and Swift warnings inherited from main - FakeTerminalEngine.swift (from #10564) uses UUID without importing Foundation, which breaks `swift test` for CmuxTerminal in the swift-package-tests job. - Parenthesize the two compactMap trailing closures in the pane memory guardrail guard and hop onto the main actor before reconcilePresentation in the session index table observer; both were new warnings over the cmux-owned Swift warning budget in tests-build-and-lag. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa * fix: clear two more Swift warnings inherited from main Drop the unreachable default branch in the cmux-tui snapshot parser's exhaustive resource-kind switch and stop binding the unused rowID shadow in SurfaceCatalogModel, so the cmux-owned warning budget passes. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa * test: bring MachinesPanelModelTests in line with the current cloud tree and catalog Four tests in the app-host shard were stale against main: - a sleeping or broken machine now keeps a Ports group whose status row explains how to discover ports (#12051), and an unregistered machine shows a Connecting placeholder instead of no children; - display rows placed inside a remote workspace carry their tab id like every other placement; - the catalog drops writes for a cloud machine with no registered provider, so the two workspace-group tests register the file's GroupFakeProvider before replacing resources. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa * chore: normalize project.pbxproj after main merges The auto-merged project file drifted from scripts/normalize-pbxproj.py output, which fails the workflow-guard pbxproj check and gates every macOS CI job behind it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa * test: assert pane re-homing through the journal event for started turns #11976 moved the pane-scoped notification clear for UserPromptSubmit and PreToolUse out of the Claude hook and into the app's journal reconciler (clearInvalidatedNotifications keys off the event's workspace and surface). The two moved-pane tests still looked for the hook's old clear_notifications send and turned the focused notification-routing step red on main. They now assert the agent_journal_append event carries the re-homed workspace and surface (and, for PreToolUse, the running phase), and keep the guards against whole-workspace or stale-workspace clears. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa * test: bound WebKit page-load waits in the design-mode screenshot suite On CI app hosts the WebContent process sometimes disappears mid-batch ("Could not signal service com.apple.WebKit.WebContent"), after which a loadHTMLString navigation never reports didFinish. The screenshot evaluator tests awaited that signal without a bound, so one lost process wedged the whole app-host batch until the 30-minute timeout (shard 4 on runs 34232451577 and 34235694241 both stalled in smoothScrollingPageCapturesRequestedRegionAndRestoresOffset). Race every real page-load wait against a 30-second budget and fail the test fast instead, so the rest of the batch still runs and reports. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa * test: tolerate an already-closed socket when bounding PTY bridge reads testPTYBridgeDefersHalfCloseUntilAttachCompletes half-closes its bridge socket, and the bridge answers and closes so quickly that the follow-up SO_RCVTIMEO setsockopt sees a torn-down connection and fails with EINVAL. That thrown error is counted as an unexpected failure and fails the whole app-host batch (it reproduces on main's own shard 6 run 34221588627). The timeout only bounds the read that follows, which returns EOF at once on such a socket, so skip the option instead of throwing. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa * ci: sample the wedged app host before killing a timed-out unit-test batch App-host unit-test batches on main and on PRs intermittently sit at the 30-minute batch timeout with only "started" lines for whichever tests were in flight, which says nothing about what blocked the main actor. Sample the app host process before terminating xcodebuild and print the leading call graph in a log group, so the next hang names its stuck frames. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa * test: queue the first Codex prompt before asserting the second is rejected testCodexInputQueueBeforeThreadIsBounded spawned the first submit in a Task and immediately submitted a second prompt on the same main actor. The second call ran first, took the single pre-thread queue slot, and then awaited a thread the test only starts afterwards, so the test hung until the 30-minute app-host batch timeout (shard 2 on runs 34235694241 and 34245949340, and main's own shard 2). Yield to the spawned task before the second submit so the first prompt holds the slot and the second is rejected as intended. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa * test: gate the concurrent file-preview save on isSaving instead of a FIFO testSaveTextContentIgnoresConcurrentSaveRequest replaced the previewed file with a FIFO so the first write would stay in flight. Since the preview panel re-opens its watched path for change monitoring, a FIFO with no writer can block the app host's main thread, and this test was the unfinished XCTest in two hung app-host shards (shard 1 on run 34232451577, shard 5 on run 34245949340). saveTextContent() sets isSaving synchronously and completes on a later main-actor hop, so the second request is always observed while the first is still saving without any FIFO. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa * ci: match the app host for hang sampling even when launched with arguments Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa * test: wait for the Codex thread/start request before answering it; per-test CI allowance Every CodexAppServerSessionTests case fed the thread/start response after a single Task.yield() following the initialize response. The session sends `initialized` and then `thread/start` from a spawned main-actor task, so when that task had not reached the request yet the response was dropped as unknown, every later submit waited for a thread forever, and the reentrant turn test spun on Task.yield() until the batch timeout (shard 2 on runs 34245949340 and 34256455336; main shows the same). - CodexAppServerSession gains two read-only test seams (isAwaitingThreadStart, hasThread). - The tests wait for the request before feeding its response, and the pending-write spin loop is bounded. - CI passes XCTest's per-test execution allowance (300s by default) to the app-host batches so a single parked test fails with a spindump instead of taking the batch to its 30-minute timeout, and the hang sample prints more threads. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa * test: replace mobile lane timing wait with completion signal --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* ci: persist nightly Xcode compilation caches * test(ci): require nightly caches to prune dead CAS generations before saving The nightly workflow test now demands that both nightly cache jobs prune dead Xcode CAS generations before measuring the 5 GiB bound, save on a miss from the bound step's verdict instead of rescanning with hashFiles, and keep the cache key prefix identical to the PR release build that restores it. It also adds a behavioural test for the pruning helper and wires it into the workflow guard job. Both fail until the next commit adds the helper and the workflow changes. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01R7fdh4BL4KeoUxNukaSdTs * ci: prune dead Xcode CAS generations so warm nightly caches persist Nightly never published a warm compilation cache. Xcode's CAS chains v1.N generations and only deletes the dead one at its next open, so a warm full build leaves two full generations behind: 6.3 GiB on main against the 5 GiB bound, which then removed the directory and the save step found nothing to upload ("Path Validation Error" in every warm run). Only cold builds, at about 3.1 GiB, ever saved, so main kept restoring a days-old entry. Prune the dead generations before measuring, the same way the CAS's own garbage collection would, in both nightly cache jobs and in the PR release build that restores the same cache. Save explicitly on a miss from the bound step's verdict instead of rescanning the directory with hashFiles, which the runner aborts after 120 seconds. Keep the cache key prefix as it was: PR release builds restore nightly's cache by that prefix, so the v2 key bump would have cut them off from the cache warmed from main. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01R7fdh4BL4KeoUxNukaSdTs * fix(web): guard the devbox reachability entrypoint without Bun typings `import.meta.main` is Bun-only and the web typecheck program has no Bun typings, so `bun run typecheck` fails on main since #12132: scripts/check-devbox-image-reachable.ts(148,17): error TS2339: Property 'main' does not exist on type 'ImportMeta'. That PR never ran the web typecheck because ci.yml only fires for a short path list. Guard the entrypoint by comparing import.meta.url with argv[1], the pattern the other scripts under web/scripts already use. Direct runs still execute main (`--print-key` prints the key) and importing the module from tests stays side-effect free. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(ci): require non-fatal cache pruning and pin the shared cache key prefix The nightly workflow test now demands that the pruner call in both nightly cache jobs and the PR release build cannot fail the build, and that every Xcode compilation cache key and restore-key in nightly.yml and ci.yml carries the one shared prefix, so renaming a key on either side (or a key but not its restore-keys) is caught. Mutation-tested: dropping the prune, moving it after the measurement, restoring the hashFiles gate, removing the miss gate, the bound step id, or the save= output, and every prefix rename now fail. Fails until the next commit makes the prune call non-fatal. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * ci: never fail a build because cache pruning failed Pruning is an optimisation. With `set -euo pipefail`, a pruner that cannot start (missing interpreter, syntax error) aborted the bound step and the whole nightly. Log a workflow warning and measure the directory unpruned instead, which at worst reproduces the old behaviour of skipping the save. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(ci): pruner must survive root-level generations and unlistable CAS dirs Pruning a generation that sits directly under the cache root removed a directory the candidate list still named, so the next iteration raised FileNotFoundError and the remaining CAS directories went unpruned for that run. A CAS directory the runner cannot list raised as well. Both now have to be skipped with a message while the other directories are still pruned. Fails until the next commit hardens the pruner. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * ci: keep pruning the remaining CAS dirs when one vanishes or cannot be listed Skip a candidate that an earlier root-level prune already removed, and treat an OSError while inspecting a CAS directory as "leave this one alone" instead of aborting the walk, so one odd directory cannot leave the others unpruned. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(ci): nightly guard requires a build-only measurement lane and an oversize-cache warning A full branch dispatch of nightly.yml signs and notarizes under the release identity, so the only safe way to time the nightly build job from a branch is a run that stops at the unsigned universal build. The guard now requires `build_only` and `cold_cache` workflow_dispatch inputs that skip the helper, signing, notarization, dSYM upload and publication jobs, keep should_publish false, and use their own concurrency group so a measurement run can never cancel a pending publish. It also requires the bound step to report an oversize cache as a workflow warning, since that path silently freezes the cache at the last saved entry. This commit adds the test only; it fails until the workflow changes land. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018wsDzi4Q5uSdU9AYq75tmE * ci: add a build-only nightly measurement lane and warn when the cache bound skips a save nightly.yml gains two workflow_dispatch inputs. `build_only` runs decide and the unsigned universal app build and stops: the helper, signing, notarization, dSYM upload and publication jobs are gated off, should_publish is forced false, and the run gets its own concurrency group so it can never cancel a pending publish. `cold_cache` (only honoured with build_only) skips the compilation cache restore so the same commit can be measured as a cache miss; the missing restore output reads as a miss, so the cold build still saves. This is the only way to time the nightly build job from a branch without notarizing under the release identity. The bound step now reports an oversize cache as a workflow warning in nightly.yml and ci.yml: nothing is saved on that path, so every later build restores the same older entry and the cache silently freezes. The comments also state the real rotation rule from LLVM's UnifiedOnDiskCache: a new primary generation is started when the current one ends a build above half of COMPILATION_CACHE_LIMIT_SIZE, which is 1.5 GiB against a 3.3 GB working set, so every warm build rotates and leaves a dead generation behind. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018wsDzi4Q5uSdU9AYq75tmE * ci: prune-xcode-compilation-cache measures only the generations it removes The pruner walked every generation, including the two live ones, to print per-generation sizes, and the workflow then ran `du -sk` over the retained cache: two full traversals of up to 5 GiB for a log line (CodeRabbit on PR #12039). Stale generations are now selected by name first and only those are measured before removal; the retained size comes from the workflow's single `du`. The speculative handling of generations placed directly under the cache root is gone: Xcode never writes that layout, and a layout the pruner does not recognise is left alone and measured unpruned rather than guessed at. An unlistable CAS directory is still skipped with a message while the remaining directories are pruned. Both behaviours keep their tests. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018wsDzi4Q5uSdU9AYq75tmE * test(ci): build_only must always build the universal app; pruner must leave unlocked directories alone Review findings on PR #12039. The nightly guard now matches each job's complete job-level `if:` verbatim, so the build_only exclusion can only be a conjunctive clause, and requires build_only to imply should_build (a measurement dispatch on main must not depend on the nightly tag) and to disable the fast arm64 path (a measurement always builds the production universal workload). The pruner test requires a CAS directory without a `lock` file to be left alone, since nothing can be locked there, and only exercises the unlistable-directory path when permission bits actually took hold. Test-only commit; it fails until the fixes land. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * ci: build_only always builds the universal app; pruner leaves unlocked CAS directories alone A build_only dispatch is a measurement of the production nightly build, so it now implies should_build (on main it no longer depends on whether the nightly tag already matches HEAD) and ignores the fast arm64 input instead of quietly measuring a one-architecture build. The pruner only ever deletes generations while holding the CAS directory's `lock` exclusively. A directory without a `lock` file has never been opened by the toolchain, so there is nothing to lock; it is now reported and left alone rather than pruned unlocked. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * web(tests): carry the bun:test type shim fix so web-typecheck passes until main has it main at 8ba29ea fails `bun run typecheck`: the tests added by #12257 (billing-alerts, cron-alerts) and #12250 (billing-purchase `.mock.calls`, vm-devbox-image `import.meta.dir`) do not type against the repository's own `web/tests/bun-test.d.ts` shim, so `ci-status` is red for every PR that does not carry a fix. This takes the shim and the two one-line test edits verbatim from PR #12105's branch (7df81ef), whose CI passes on the same base, so a later merge of that fix is conflict-free. No behaviour change: the shim is a `.d.ts` and the test edits keep their assertions. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test: reject non-keyboard events in file explorer shortcuts * fix: unblock app-host tests at event and async wait boundaries * ci: keep focused macOS tests on the required SDK 26 toolchain * test(ci): cap selected SDKs while preserving legacy runner fallback * ci: cap focused-test SDK selection without dropping older runners * test: install a valid shortcut before checking file explorer routing * test: fence detach command capture without blocking the main actor --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
…changed (manaflow-ai#12132) * cloud: prove the shipped devbox image is reachable, and only when it changed The bake smokes the daemon once, when an image is made. Nothing asks whether a client can reach the daemon in the image we are shipping now, using the cmux-tui build we are shipping now. Those drift apart with no commit at all: files.cmux.com moves to a new client on every cmux-tui release while the manifest keeps pinning yesterday's snapshot. The check boots the manifest's current default and runs a real round trip (enroll a device, connect over ws://[::1]:1337/v1/link, open a workspace, run a process, read its output). When the live client differs from the baked one it repeats the whole trip with the live binary, so a protocol change that would break a freshly built app fails here instead of in production. It costs 17s with both legs, 9s when the client is the baked one, and one small VM that is deleted on every path. Running it when nothing changed is waste, so the job keys a cache on the default image ids plus the client sha256 and skips a pair that has already passed. That makes the daily schedule, which exists to catch the client moving without a commit, free on a day when it did not. * TEMPORARY: point the reachability check at a stock image to prove it fails * Restore the manifest default after proving the check fails Run 34210445406 booted freestyle/ubuntu-sm, waited out the 120s daemon budget, reported "UNREACHABLE: the baked daemon never came up", deleted the machine anyway, and exited 1, so the job is red when the image is broken and leaves nothing running behind it.
…aflow-ai#12154) `import.meta.main` is Bun-only and is not part of TypeScript's ImportMeta type, so `bun run typecheck` has failed on main since manaflow-ai#12132: scripts/check-devbox-image-reachable.ts(148,17): error TS2339: Property 'main' does not exist on type 'ImportMeta'. Use the portable `fileURLToPath(import.meta.url)` vs `process.argv[1]` check that the sibling devbox scripts already rely on. Behavior is unchanged: `bun scripts/check-devbox-image-reachable.ts --print-key` still runs main, and importing `reachabilityKey` from the test does not. Closes manaflow-ai#12147 Claude-Session: https://claude.ai/code/session_01Qiy38q5XFYQU3CeccDENty Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* test: cover creation initial command plumbing * test: isolate creation command regression * feat: add initial commands to terminal creation * test: require creation request dispatch * test(web): preserve hosted config overrides * test(cli): expect OMP restore path * test(session): expect codex wrapper shim * docs(cli): sync restore help contract * test(cli): preserve interactive shell for creation command * fix(cli): inject creation commands into interactive shells * fix(cli): preserve terminal creation command text * test(cli): reject invalid creation command input * fix(cli): validate terminal creation input boundaries * fix(cli): resolve creation validation in app context * test(cli): cover review-found creation boundaries * test(cli): import welcome setting contract * test(cli): wait for terminal readiness before welcome assertion * test(cli): exercise focused workspace welcome path * fix(cli): close terminal creation review gaps * test: cover app-host temp path aliases * ci: accept validated macOS temp aliases * Fix Xcode 26 warning regressions * test: keep install command fixture inert * test: repair app-host CLI fixtures * fix: remove duplicate dock initial input plumbing * test: bound workspace readiness regression wait * fix: order dock creation arguments * fix: order dock input after startup options * fix: preserve null terminal creation types * fix: keep readiness wait actor-safe * test: isolate creation command socket environment * test: keep cloud splits local for initial input * fix: keep initial input out of cloud split routing * ci: route browser skill contract through linux runner * test: isolate cloud routing fixture from local surfaces * refactor: keep terminal creation checks within file budgets * fix: wire terminal creation helpers into app target * fix: repair current main compile errors * docs: document --command on terminal creation commands Cover the new --command flag on new-split, new-pane, and new-surface (and the existing one on new-workspace) in the cmux and cmux-workspace skills, the CLI contract (table rows, an Initial terminal command section, and --help probes), and the web API docs in English and Japanese. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa * fix: silence unmutated payload warning in browser key replay CI's Swift warning budget fails on main because the delivered-key payload in TerminalController.swift is declared var but never mutated. Use let so the tests-build-and-lag job can pass on this branch. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa * fix(web): keep devbox reachability script typechecking without bun-types `import.meta.main` (added on main in manaflow-ai#12132) fails `tsgo --noEmit` because the web tsconfig does not include bun-types, which turns the CI cheap layer red and skips every macOS job. Cast the meta object so Bun's runtime flag still gates main() while the typecheck passes. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa * fix: repair package test import and Swift warnings inherited from main - FakeTerminalEngine.swift (from manaflow-ai#10564) uses UUID without importing Foundation, which breaks `swift test` for CmuxTerminal in the swift-package-tests job. - Parenthesize the two compactMap trailing closures in the pane memory guardrail guard and hop onto the main actor before reconcilePresentation in the session index table observer; both were new warnings over the cmux-owned Swift warning budget in tests-build-and-lag. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa * fix: clear two more Swift warnings inherited from main Drop the unreachable default branch in the cmux-tui snapshot parser's exhaustive resource-kind switch and stop binding the unused rowID shadow in SurfaceCatalogModel, so the cmux-owned warning budget passes. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa * test: bring MachinesPanelModelTests in line with the current cloud tree and catalog Four tests in the app-host shard were stale against main: - a sleeping or broken machine now keeps a Ports group whose status row explains how to discover ports (manaflow-ai#12051), and an unregistered machine shows a Connecting placeholder instead of no children; - display rows placed inside a remote workspace carry their tab id like every other placement; - the catalog drops writes for a cloud machine with no registered provider, so the two workspace-group tests register the file's GroupFakeProvider before replacing resources. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa * chore: normalize project.pbxproj after main merges The auto-merged project file drifted from scripts/normalize-pbxproj.py output, which fails the workflow-guard pbxproj check and gates every macOS CI job behind it. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa * test: assert pane re-homing through the journal event for started turns manaflow-ai#11976 moved the pane-scoped notification clear for UserPromptSubmit and PreToolUse out of the Claude hook and into the app's journal reconciler (clearInvalidatedNotifications keys off the event's workspace and surface). The two moved-pane tests still looked for the hook's old clear_notifications send and turned the focused notification-routing step red on main. They now assert the agent_journal_append event carries the re-homed workspace and surface (and, for PreToolUse, the running phase), and keep the guards against whole-workspace or stale-workspace clears. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa * test: bound WebKit page-load waits in the design-mode screenshot suite On CI app hosts the WebContent process sometimes disappears mid-batch ("Could not signal service com.apple.WebKit.WebContent"), after which a loadHTMLString navigation never reports didFinish. The screenshot evaluator tests awaited that signal without a bound, so one lost process wedged the whole app-host batch until the 30-minute timeout (shard 4 on runs 34232451577 and 34235694241 both stalled in smoothScrollingPageCapturesRequestedRegionAndRestoresOffset). Race every real page-load wait against a 30-second budget and fail the test fast instead, so the rest of the batch still runs and reports. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa * test: tolerate an already-closed socket when bounding PTY bridge reads testPTYBridgeDefersHalfCloseUntilAttachCompletes half-closes its bridge socket, and the bridge answers and closes so quickly that the follow-up SO_RCVTIMEO setsockopt sees a torn-down connection and fails with EINVAL. That thrown error is counted as an unexpected failure and fails the whole app-host batch (it reproduces on main's own shard 6 run 34221588627). The timeout only bounds the read that follows, which returns EOF at once on such a socket, so skip the option instead of throwing. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa * ci: sample the wedged app host before killing a timed-out unit-test batch App-host unit-test batches on main and on PRs intermittently sit at the 30-minute batch timeout with only "started" lines for whichever tests were in flight, which says nothing about what blocked the main actor. Sample the app host process before terminating xcodebuild and print the leading call graph in a log group, so the next hang names its stuck frames. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa * test: queue the first Codex prompt before asserting the second is rejected testCodexInputQueueBeforeThreadIsBounded spawned the first submit in a Task and immediately submitted a second prompt on the same main actor. The second call ran first, took the single pre-thread queue slot, and then awaited a thread the test only starts afterwards, so the test hung until the 30-minute app-host batch timeout (shard 2 on runs 34235694241 and 34245949340, and main's own shard 2). Yield to the spawned task before the second submit so the first prompt holds the slot and the second is rejected as intended. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa * test: gate the concurrent file-preview save on isSaving instead of a FIFO testSaveTextContentIgnoresConcurrentSaveRequest replaced the previewed file with a FIFO so the first write would stay in flight. Since the preview panel re-opens its watched path for change monitoring, a FIFO with no writer can block the app host's main thread, and this test was the unfinished XCTest in two hung app-host shards (shard 1 on run 34232451577, shard 5 on run 34245949340). saveTextContent() sets isSaving synchronously and completes on a later main-actor hop, so the second request is always observed while the first is still saving without any FIFO. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa * ci: match the app host for hang sampling even when launched with arguments Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa * test: wait for the Codex thread/start request before answering it; per-test CI allowance Every CodexAppServerSessionTests case fed the thread/start response after a single Task.yield() following the initialize response. The session sends `initialized` and then `thread/start` from a spawned main-actor task, so when that task had not reached the request yet the response was dropped as unknown, every later submit waited for a thread forever, and the reentrant turn test spun on Task.yield() until the batch timeout (shard 2 on runs 34245949340 and 34256455336; main shows the same). - CodexAppServerSession gains two read-only test seams (isAwaitingThreadStart, hasThread). - The tests wait for the request before feeding its response, and the pending-write spin loop is bounded. - CI passes XCTest's per-test execution allowance (300s by default) to the app-host batches so a single parked test fails with a spindump instead of taking the batch to its 30-minute timeout, and the hang sample prints more threads. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa * test: replace mobile lane timing wait with completion signal --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
* ci: persist nightly Xcode compilation caches * test(ci): require nightly caches to prune dead CAS generations before saving The nightly workflow test now demands that both nightly cache jobs prune dead Xcode CAS generations before measuring the 5 GiB bound, save on a miss from the bound step's verdict instead of rescanning with hashFiles, and keep the cache key prefix identical to the PR release build that restores it. It also adds a behavioural test for the pruning helper and wires it into the workflow guard job. Both fail until the next commit adds the helper and the workflow changes. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01R7fdh4BL4KeoUxNukaSdTs * ci: prune dead Xcode CAS generations so warm nightly caches persist Nightly never published a warm compilation cache. Xcode's CAS chains v1.N generations and only deletes the dead one at its next open, so a warm full build leaves two full generations behind: 6.3 GiB on main against the 5 GiB bound, which then removed the directory and the save step found nothing to upload ("Path Validation Error" in every warm run). Only cold builds, at about 3.1 GiB, ever saved, so main kept restoring a days-old entry. Prune the dead generations before measuring, the same way the CAS's own garbage collection would, in both nightly cache jobs and in the PR release build that restores the same cache. Save explicitly on a miss from the bound step's verdict instead of rescanning the directory with hashFiles, which the runner aborts after 120 seconds. Keep the cache key prefix as it was: PR release builds restore nightly's cache by that prefix, so the v2 key bump would have cut them off from the cache warmed from main. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_01R7fdh4BL4KeoUxNukaSdTs * fix(web): guard the devbox reachability entrypoint without Bun typings `import.meta.main` is Bun-only and the web typecheck program has no Bun typings, so `bun run typecheck` fails on main since manaflow-ai#12132: scripts/check-devbox-image-reachable.ts(148,17): error TS2339: Property 'main' does not exist on type 'ImportMeta'. That PR never ran the web typecheck because ci.yml only fires for a short path list. Guard the entrypoint by comparing import.meta.url with argv[1], the pattern the other scripts under web/scripts already use. Direct runs still execute main (`--print-key` prints the key) and importing the module from tests stays side-effect free. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(ci): require non-fatal cache pruning and pin the shared cache key prefix The nightly workflow test now demands that the pruner call in both nightly cache jobs and the PR release build cannot fail the build, and that every Xcode compilation cache key and restore-key in nightly.yml and ci.yml carries the one shared prefix, so renaming a key on either side (or a key but not its restore-keys) is caught. Mutation-tested: dropping the prune, moving it after the measurement, restoring the hashFiles gate, removing the miss gate, the bound step id, or the save= output, and every prefix rename now fail. Fails until the next commit makes the prune call non-fatal. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * ci: never fail a build because cache pruning failed Pruning is an optimisation. With `set -euo pipefail`, a pruner that cannot start (missing interpreter, syntax error) aborted the bound step and the whole nightly. Log a workflow warning and measure the directory unpruned instead, which at worst reproduces the old behaviour of skipping the save. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(ci): pruner must survive root-level generations and unlistable CAS dirs Pruning a generation that sits directly under the cache root removed a directory the candidate list still named, so the next iteration raised FileNotFoundError and the remaining CAS directories went unpruned for that run. A CAS directory the runner cannot list raised as well. Both now have to be skipped with a message while the other directories are still pruned. Fails until the next commit hardens the pruner. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * ci: keep pruning the remaining CAS dirs when one vanishes or cannot be listed Skip a candidate that an earlier root-level prune already removed, and treat an OSError while inspecting a CAS directory as "leave this one alone" instead of aborting the walk, so one odd directory cannot leave the others unpruned. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test(ci): nightly guard requires a build-only measurement lane and an oversize-cache warning A full branch dispatch of nightly.yml signs and notarizes under the release identity, so the only safe way to time the nightly build job from a branch is a run that stops at the unsigned universal build. The guard now requires `build_only` and `cold_cache` workflow_dispatch inputs that skip the helper, signing, notarization, dSYM upload and publication jobs, keep should_publish false, and use their own concurrency group so a measurement run can never cancel a pending publish. It also requires the bound step to report an oversize cache as a workflow warning, since that path silently freezes the cache at the last saved entry. This commit adds the test only; it fails until the workflow changes land. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018wsDzi4Q5uSdU9AYq75tmE * ci: add a build-only nightly measurement lane and warn when the cache bound skips a save nightly.yml gains two workflow_dispatch inputs. `build_only` runs decide and the unsigned universal app build and stops: the helper, signing, notarization, dSYM upload and publication jobs are gated off, should_publish is forced false, and the run gets its own concurrency group so it can never cancel a pending publish. `cold_cache` (only honoured with build_only) skips the compilation cache restore so the same commit can be measured as a cache miss; the missing restore output reads as a miss, so the cold build still saves. This is the only way to time the nightly build job from a branch without notarizing under the release identity. The bound step now reports an oversize cache as a workflow warning in nightly.yml and ci.yml: nothing is saved on that path, so every later build restores the same older entry and the cache silently freezes. The comments also state the real rotation rule from LLVM's UnifiedOnDiskCache: a new primary generation is started when the current one ends a build above half of COMPILATION_CACHE_LIMIT_SIZE, which is 1.5 GiB against a 3.3 GB working set, so every warm build rotates and leaves a dead generation behind. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018wsDzi4Q5uSdU9AYq75tmE * ci: prune-xcode-compilation-cache measures only the generations it removes The pruner walked every generation, including the two live ones, to print per-generation sizes, and the workflow then ran `du -sk` over the retained cache: two full traversals of up to 5 GiB for a log line (CodeRabbit on PR manaflow-ai#12039). Stale generations are now selected by name first and only those are measured before removal; the retained size comes from the workflow's single `du`. The speculative handling of generations placed directly under the cache root is gone: Xcode never writes that layout, and a layout the pruner does not recognise is left alone and measured unpruned rather than guessed at. An unlistable CAS directory is still skipped with a message while the remaining directories are pruned. Both behaviours keep their tests. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> Claude-Session: https://claude.ai/code/session_018wsDzi4Q5uSdU9AYq75tmE * test(ci): build_only must always build the universal app; pruner must leave unlocked directories alone Review findings on PR manaflow-ai#12039. The nightly guard now matches each job's complete job-level `if:` verbatim, so the build_only exclusion can only be a conjunctive clause, and requires build_only to imply should_build (a measurement dispatch on main must not depend on the nightly tag) and to disable the fast arm64 path (a measurement always builds the production universal workload). The pruner test requires a CAS directory without a `lock` file to be left alone, since nothing can be locked there, and only exercises the unlistable-directory path when permission bits actually took hold. Test-only commit; it fails until the fixes land. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * ci: build_only always builds the universal app; pruner leaves unlocked CAS directories alone A build_only dispatch is a measurement of the production nightly build, so it now implies should_build (on main it no longer depends on whether the nightly tag already matches HEAD) and ignores the fast arm64 input instead of quietly measuring a one-architecture build. The pruner only ever deletes generations while holding the CAS directory's `lock` exclusively. A directory without a `lock` file has never been opened by the toolchain, so there is nothing to lock; it is now reported and left alone rather than pruned unlocked. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * web(tests): carry the bun:test type shim fix so web-typecheck passes until main has it main at 8ba29ea fails `bun run typecheck`: the tests added by manaflow-ai#12257 (billing-alerts, cron-alerts) and manaflow-ai#12250 (billing-purchase `.mock.calls`, vm-devbox-image `import.meta.dir`) do not type against the repository's own `web/tests/bun-test.d.ts` shim, so `ci-status` is red for every PR that does not carry a fix. This takes the shim and the two one-line test edits verbatim from PR manaflow-ai#12105's branch (7df81ef), whose CI passes on the same base, so a later merge of that fix is conflict-free. No behaviour change: the shim is a `.d.ts` and the test edits keep their assertions. Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com> * test: reject non-keyboard events in file explorer shortcuts * fix: unblock app-host tests at event and async wait boundaries * ci: keep focused macOS tests on the required SDK 26 toolchain * test(ci): cap selected SDKs while preserving legacy runner fallback * ci: cap focused-test SDK selection without dropping older runners * test: install a valid shortcut before checking file explorer routing * test: fence detach command capture without blocking the main actor --------- Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>


The bake smokes the cmux-tui daemon once, at the moment an image is made. Nothing asks the question that matters later: can a client reach the daemon in the image we are shipping right now, using the cmux-tui build we are shipping right now?
Those two drift apart with no commit at all.
files.cmux.commoves to a new client on every cmux-tui release, while the manifest keeps pinning yesterday's snapshot, so the pair in production is one nobody ever tested together.bun scripts/check-devbox-image-reachable.tsboots the manifest's current default and runs a real round trip, not a port probe: enroll a device, connect overws://[::1]:1337/v1/link, open a workspace, run a process, read its output back. When the live client differs from the baked one, it installs the live binary in the guest and repeats the entire trip with it, so a protocol change that would break a freshly built app fails here instead of in front of a user.Measured on the current default: 17s with both legs, 9s when the live client is already the baked one, one
smVM, deleted on every path including failure.Only when necessary
What it needs
FREESTYLE_API_KEYas an Actions secret. Without it the job says so and skips rather than failing, so this can land before you decide on that. It is the same decision as the bake-on-main job we discussed; this one is a much smaller blast radius, since it only ever boots one small VM and deletes it.Need help on this PR? Tag
@codesmith-botwith what you need. Autofix is disabled.Note
Medium Risk
CI boots real Freestyle VMs with a production API key (mitigated by environment scoping, fork skips, and single small VM lifecycle), but failures could block merges on external client/image drift.
Overview
Adds cloud VM image reachability CI that boots the manifest’s default devbox snapshot and proves the shipped cmux-tui client can complete a full daemon round trip (not just a port check). When the live client from
files.cmux.comdiffers from what’s baked in the image, the check downloads the live binary in the guest and repeats the trip so image/client drift is caught before users hit it.A new
check-devbox-image-reachable.tsscript drives Freestyle VM create/exec/delete, exposesreachabilityKey(default image ids + live client sha256) for cache keys, and supports--print-keyso CI can skip unchanged pairs. The workflow runs on relevant path changes, daily (client can move without a repo commit), and workflow_dispatch; it gates onFREESTYLE_API_KEY(skip on fork PRs or missing secret), uses Actions cache so only new pairs boot a VM, and records passes only so failures re-run. Unit tests coverreachabilityKeystability and invalidation when images or client pins change.Reviewed by Cursor Bugbot for commit 95a83a5. Bugbot is set up for automated code reviews on this repo. Configure here.
Summary by cubic
Adds a CI job that boots the shipped devbox image and proves the shipped cmux-tui client can reach the daemon inside it, so the image/client pair in production is what actually gets tested.
ws://[::1]:1337/v1/link, open a workspace, run a process, read output), not a port probe.smVM, deleted on every path including failure; ~17s with both legs, ~9s when the client is already baked.FREESTYLE_API_KEY; without it the job skips rather than failing, and fork PRs skip without a credential.Written for commit 95a83a5. Summary will update on new commits.