Skip to content

Check that the shipped devbox image is reachable, only when the pair changed - #12132

Merged
lawrencecchen merged 3 commits into
mainfrom
feat-devbox-reachability-ci
Sep 8, 2026
Merged

lawrencecchen merged 3 commits into
mainfrom
feat-devbox-reachability-ci

Conversation

@lawrencecchen

@lawrencecchen lawrencecchen commented Sep 8, 2026 •

Copy link
Copy Markdown
Contributor

The bake smokes the cmux-tui daemon once, at the moment an image is made. Nothing asks the question that matters later: can a client reach the daemon in the image we are shipping right now, using the cmux-tui build we are shipping right now?

Those two drift apart with no commit at all. files.cmux.com moves to a new client on every cmux-tui release, while the manifest keeps pinning yesterday's snapshot, so the pair in production is one nobody ever tested together.

bun scripts/check-devbox-image-reachable.ts boots the manifest's current default and runs a real round trip, not a port probe: enroll a device, connect over ws://[::1]:1337/v1/link, open a workspace, run a process, read its output back. When the live client differs from the baked one, it installs the live binary in the guest and repeats the entire trip with it, so a protocol change that would break a freshly built app fails here instead of in front of a user.

Measured on the current default: 17s with both legs, 9s when the live client is already the baked one, one sm VM, deleted on every path including failure.

Only when necessary

  • Path-filtered to the manifest, the devbox source, the client resolver, and the check itself.
  • A daily schedule, because the client moves without a commit.
  • The run keys a cache on every default image id plus the live client's sha256 and skips a pair that has already passed, so the daily run costs nothing on a day when neither moved. Only a pass is recorded, so a failure re-runs.
  • Fork PRs skip cleanly: booting a machine for one would hand an untrusted branch the production account.

What it needs

FREESTYLE_API_KEY as an Actions secret. Without it the job says so and skips rather than failing, so this can land before you decide on that. It is the same decision as the bake-on-main job we discussed; this one is a much smaller blast radius, since it only ever boots one small VM and deletes it.


View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.


Note

Medium Risk
CI boots real Freestyle VMs with a production API key (mitigated by environment scoping, fork skips, and single small VM lifecycle), but failures could block merges on external client/image drift.

Overview
Adds cloud VM image reachability CI that boots the manifest’s default devbox snapshot and proves the shipped cmux-tui client can complete a full daemon round trip (not just a port check). When the live client from files.cmux.com differs from what’s baked in the image, the check downloads the live binary in the guest and repeats the trip so image/client drift is caught before users hit it.

A new check-devbox-image-reachable.ts script drives Freestyle VM create/exec/delete, exposes reachabilityKey (default image ids + live client sha256) for cache keys, and supports --print-key so CI can skip unchanged pairs. The workflow runs on relevant path changes, daily (client can move without a repo commit), and workflow_dispatch; it gates on FREESTYLE_API_KEY (skip on fork PRs or missing secret), uses Actions cache so only new pairs boot a VM, and records passes only so failures re-run. Unit tests cover reachabilityKey stability and invalidation when images or client pins change.

Reviewed by Cursor Bugbot for commit 95a83a5. Bugbot is set up for automated code reviews on this repo. Configure here.


Summary by cubic

Adds a CI job that boots the shipped devbox image and proves the shipped cmux-tui client can reach the daemon inside it, so the image/client pair in production is what actually gets tested.

  • Runs a real round trip (enroll, connect over ws://[::1]:1337/v1/link, open a workspace, run a process, read output), not a port probe.
  • When the live client differs from the baked one, installs the live binary in the guest and repeats the trip.
  • Triggered on path changes, a daily schedule, and manually; a cache key over image ids plus client sha256 skips pairs that already passed.
  • Costs one small sm VM, deleted on every path including failure; ~17s with both legs, ~9s when the client is already baked.
  • Requires FREESTYLE_API_KEY; without it the job skips rather than failing, and fork PRs skip without a credential.
  • A temporary override to a stock image proved the job fails cleanly and was reverted, so the check targets the manifest default.

Written for commit 95a83a5. Summary will update on new commits.

Review in cubic

…changed

The bake smokes the daemon once, when an image is made. Nothing asks whether
a client can reach the daemon in the image we are shipping now, using the
cmux-tui build we are shipping now. Those drift apart with no commit at all:
files.cmux.com moves to a new client on every cmux-tui release while the
manifest keeps pinning yesterday's snapshot.

The check boots the manifest's current default and runs a real round trip
(enroll a device, connect over ws://[::1]:1337/v1/link, open a workspace, run
a process, read its output). When the live client differs from the baked one
it repeats the whole trip with the live binary, so a protocol change that
would break a freshly built app fails here instead of in production.

It costs 17s with both legs, 9s when the client is the baked one, and one
small VM that is deleted on every path.

Running it when nothing changed is waste, so the job keys a cache on the
default image ids plus the client sha256 and skips a pair that has already
passed. That makes the daily schedule, which exists to catch the client
moving without a commit, free on a day when it did not.
@vercel

vercel Bot commented Sep 8, 2026 •

Copy link
Copy Markdown

The latest updates on your projects. Learn more about Vercel for GitHub.

Project Deployment Actions Updated
cmux166 Ready Ready Preview Sep 8, 2026 1:27pm UTC
cmux41 Ready Ready Preview Sep 8, 2026 1:27pm UTC

@coderabbitai

coderabbitai Bot commented Sep 8, 2026 •

Copy link
Copy Markdown

Warning

Review limit reached

Next included review available in 3 minutes.

Check out review usage here.

View limit details

Limit details: You’ve used all 10 included reviews currently available.

You've used all free OSS reviews for now. Wait for the free limit to reset to keep reviewing this public repository.

Learn how review limits work.

Review configuration:

⚙️ Run configuration

Configuration used: Path: .coderabbit.yaml

Review profile: ASSERTIVE

Plan: Advanced

Run ID: ab790a73-f6e0-4777-9c81-e69ee1f9f62e

📥 Commits

Reviewing files that changed from the base of the PR and between bf8ae03 and 95a83a5.

📒 Files selected for processing (3)
  • .github/workflows/cloud-vm-image-reachability.yml
  • web/scripts/check-devbox-image-reachable.ts
  • web/tests/vm-image-reachability.test.ts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@github-actions

github-actions Bot commented Sep 8, 2026

Copy link
Copy Markdown
Contributor

All contributors have signed the CLA ✍️ ✅
Posted by the CLA Assistant Lite bot.

Run 34210445406 booted freestyle/ubuntu-sm, waited out the 120s daemon
budget, reported "UNREACHABLE: the baked daemon never came up", deleted the
machine anyway, and exited 1, so the job is red when the image is broken and
leaves nothing running behind it.
@lawrencecchen
lawrencecchen temporarily deployed to cloud-vm-image-checks September 8, 2026 09:34 — with GitHub Actions Inactive
@lawrencecchen
lawrencecchen merged commit 31cdb0a into main Sep 8, 2026
26 of 30 checks passed
@lawrencecchen
lawrencecchen deleted the feat-devbox-reachability-ci branch September 8, 2026 09:38

@cursor cursor Bot left a comment

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cursor Bugbot has reviewed your changes using default effort and found 1 potential issue.

Fix All in Cursor

❌ Bugbot Autofix is OFF. To automatically fix reported issues with cloud agents, enable autofix in the Cursor dashboard.

Reviewed by Cursor Bugbot for commit 95a83a5. Configure here.


concurrency:
group: cloud-vm-image-reachability-${{ github.ref }}
cancel-in-progress: true

Copy link
Copy Markdown

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

Cancelled runs can leak VMs

Medium Severity

Concurrency cancels an in-progress reachability job when a new run starts on the same ref. The guest is only deleted in the script finally block, which does not run on SIGTERM, so the sm VM can stay billed after the job is cancelled.

Additional Locations (1)
Fix in Cursor Fix in Web

Reviewed by Cursor Bugbot for commit 95a83a5. Configure here.

rustybret pushed a commit to rustybret/bmux that referenced this pull request Sep 8, 2026
31cdb0a Check that the shipped devbox image is reachable, only when the pair changed (manaflow-ai#12132)
93cb5be Polish Find and Vault search spacing, corners, and hover (manaflow-ai#12109)
bf8ae03 cloud: give devbox panes the Ghostty terminal identity (manaflow-ai#12125)
f65fe42 web: run Cloud VM routes on one Effect runtime with a tag-keyed error table (manaflow-ai#12122)
austinywang added a commit that referenced this pull request Sep 8, 2026
`import.meta.main` (added on main in #12132) fails `tsgo --noEmit` because
the web tsconfig does not include bun-types, which turns the CI cheap layer
red and skips every macOS job. Cast the meta object so Bun's runtime flag
still gates main() while the typecheck passes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa
austinywang added a commit that referenced this pull request Sep 8, 2026
`import.meta.main` is Bun-only and the web typecheck program has no Bun
typings, so `bun run typecheck` fails on main since #12132:

  scripts/check-devbox-image-reachable.ts(148,17): error TS2339:
  Property 'main' does not exist on type 'ImportMeta'.

That PR never ran the web typecheck because ci.yml only fires for a short
path list. Guard the entrypoint by comparing import.meta.url with argv[1],
the pattern the other scripts under web/scripts already use. Direct runs
still execute main (`--print-key` prints the key) and importing the module
from tests stays side-effect free.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
austinywang added a commit that referenced this pull request Sep 8, 2026
)

`import.meta.main` is Bun-only and is not part of TypeScript's ImportMeta
type, so `bun run typecheck` has failed on main since #12132:

  scripts/check-devbox-image-reachable.ts(148,17): error TS2339:
  Property 'main' does not exist on type 'ImportMeta'.

Use the portable `fileURLToPath(import.meta.url)` vs `process.argv[1]`
check that the sibling devbox scripts already rely on. Behavior is
unchanged: `bun scripts/check-devbox-image-reachable.ts --print-key`
still runs main, and importing `reachabilityKey` from the test does not.

Closes #12147


Claude-Session: https://claude.ai/code/session_01Qiy38q5XFYQU3CeccDENty

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
austinywang added a commit that referenced this pull request Sep 9, 2026
* test: cover creation initial command plumbing

* test: isolate creation command regression

* feat: add initial commands to terminal creation

* test: require creation request dispatch

* test(web): preserve hosted config overrides

* test(cli): expect OMP restore path

* test(session): expect codex wrapper shim

* docs(cli): sync restore help contract

* test(cli): preserve interactive shell for creation command

* fix(cli): inject creation commands into interactive shells

* fix(cli): preserve terminal creation command text

* test(cli): reject invalid creation command input

* fix(cli): validate terminal creation input boundaries

* fix(cli): resolve creation validation in app context

* test(cli): cover review-found creation boundaries

* test(cli): import welcome setting contract

* test(cli): wait for terminal readiness before welcome assertion

* test(cli): exercise focused workspace welcome path

* fix(cli): close terminal creation review gaps

* test: cover app-host temp path aliases

* ci: accept validated macOS temp aliases

* Fix Xcode 26 warning regressions

* test: keep install command fixture inert

* test: repair app-host CLI fixtures

* fix: remove duplicate dock initial input plumbing

* test: bound workspace readiness regression wait

* fix: order dock creation arguments

* fix: order dock input after startup options

* fix: preserve null terminal creation types

* fix: keep readiness wait actor-safe

* test: isolate creation command socket environment

* test: keep cloud splits local for initial input

* fix: keep initial input out of cloud split routing

* ci: route browser skill contract through linux runner

* test: isolate cloud routing fixture from local surfaces

* refactor: keep terminal creation checks within file budgets

* fix: wire terminal creation helpers into app target

* fix: repair current main compile errors

* docs: document --command on terminal creation commands

Cover the new --command flag on new-split, new-pane, and new-surface (and
the existing one on new-workspace) in the cmux and cmux-workspace skills,
the CLI contract (table rows, an Initial terminal command section, and
--help probes), and the web API docs in English and Japanese.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa

* fix: silence unmutated payload warning in browser key replay

CI's Swift warning budget fails on main because the delivered-key payload
in TerminalController.swift is declared var but never mutated. Use let so
the tests-build-and-lag job can pass on this branch.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa

* fix(web): keep devbox reachability script typechecking without bun-types

`import.meta.main` (added on main in #12132) fails `tsgo --noEmit` because
the web tsconfig does not include bun-types, which turns the CI cheap layer
red and skips every macOS job. Cast the meta object so Bun's runtime flag
still gates main() while the typecheck passes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa

* fix: repair package test import and Swift warnings inherited from main

- FakeTerminalEngine.swift (from #10564) uses UUID without importing
  Foundation, which breaks `swift test` for CmuxTerminal in the
  swift-package-tests job.
- Parenthesize the two compactMap trailing closures in the pane memory
  guardrail guard and hop onto the main actor before reconcilePresentation
  in the session index table observer; both were new warnings over the
  cmux-owned Swift warning budget in tests-build-and-lag.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa

* fix: clear two more Swift warnings inherited from main

Drop the unreachable default branch in the cmux-tui snapshot parser's
exhaustive resource-kind switch and stop binding the unused rowID shadow in
SurfaceCatalogModel, so the cmux-owned warning budget passes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa

* test: bring MachinesPanelModelTests in line with the current cloud tree and catalog

Four tests in the app-host shard were stale against main:
- a sleeping or broken machine now keeps a Ports group whose status row
  explains how to discover ports (#12051), and an unregistered machine shows
  a Connecting placeholder instead of no children;
- display rows placed inside a remote workspace carry their tab id like
  every other placement;
- the catalog drops writes for a cloud machine with no registered provider,
  so the two workspace-group tests register the file's GroupFakeProvider
  before replacing resources.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa

* chore: normalize project.pbxproj after main merges

The auto-merged project file drifted from scripts/normalize-pbxproj.py
output, which fails the workflow-guard pbxproj check and gates every macOS
CI job behind it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa

* test: assert pane re-homing through the journal event for started turns

#11976 moved the pane-scoped notification clear for UserPromptSubmit and
PreToolUse out of the Claude hook and into the app's journal reconciler
(clearInvalidatedNotifications keys off the event's workspace and surface).
The two moved-pane tests still looked for the hook's old clear_notifications
send and turned the focused notification-routing step red on main.

They now assert the agent_journal_append event carries the re-homed
workspace and surface (and, for PreToolUse, the running phase), and keep
the guards against whole-workspace or stale-workspace clears.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa

* test: bound WebKit page-load waits in the design-mode screenshot suite

On CI app hosts the WebContent process sometimes disappears mid-batch
("Could not signal service com.apple.WebKit.WebContent"), after which a
loadHTMLString navigation never reports didFinish. The screenshot evaluator
tests awaited that signal without a bound, so one lost process wedged the
whole app-host batch until the 30-minute timeout (shard 4 on runs
34232451577 and 34235694241 both stalled in
smoothScrollingPageCapturesRequestedRegionAndRestoresOffset).

Race every real page-load wait against a 30-second budget and fail the
test fast instead, so the rest of the batch still runs and reports.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa

* test: tolerate an already-closed socket when bounding PTY bridge reads

testPTYBridgeDefersHalfCloseUntilAttachCompletes half-closes its bridge
socket, and the bridge answers and closes so quickly that the follow-up
SO_RCVTIMEO setsockopt sees a torn-down connection and fails with EINVAL.
That thrown error is counted as an unexpected failure and fails the whole
app-host batch (it reproduces on main's own shard 6 run 34221588627).

The timeout only bounds the read that follows, which returns EOF at once
on such a socket, so skip the option instead of throwing.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa

* ci: sample the wedged app host before killing a timed-out unit-test batch

App-host unit-test batches on main and on PRs intermittently sit at the
30-minute batch timeout with only "started" lines for whichever tests were
in flight, which says nothing about what blocked the main actor. Sample
the app host process before terminating xcodebuild and print the leading
call graph in a log group, so the next hang names its stuck frames.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa

* test: queue the first Codex prompt before asserting the second is rejected

testCodexInputQueueBeforeThreadIsBounded spawned the first submit in a Task
and immediately submitted a second prompt on the same main actor. The
second call ran first, took the single pre-thread queue slot, and then
awaited a thread the test only starts afterwards, so the test hung until
the 30-minute app-host batch timeout (shard 2 on runs 34235694241 and
34245949340, and main's own shard 2).

Yield to the spawned task before the second submit so the first prompt
holds the slot and the second is rejected as intended.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa

* test: gate the concurrent file-preview save on isSaving instead of a FIFO

testSaveTextContentIgnoresConcurrentSaveRequest replaced the previewed file
with a FIFO so the first write would stay in flight. Since the preview
panel re-opens its watched path for change monitoring, a FIFO with no
writer can block the app host's main thread, and this test was the
unfinished XCTest in two hung app-host shards (shard 1 on run 34232451577,
shard 5 on run 34245949340). saveTextContent() sets isSaving synchronously
and completes on a later main-actor hop, so the second request is always
observed while the first is still saving without any FIFO.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa

* ci: match the app host for hang sampling even when launched with arguments

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa

* test: wait for the Codex thread/start request before answering it; per-test CI allowance

Every CodexAppServerSessionTests case fed the thread/start response after a
single Task.yield() following the initialize response. The session sends
`initialized` and then `thread/start` from a spawned main-actor task, so
when that task had not reached the request yet the response was dropped as
unknown, every later submit waited for a thread forever, and the reentrant
turn test spun on Task.yield() until the batch timeout (shard 2 on runs
34245949340 and 34256455336; main shows the same).

- CodexAppServerSession gains two read-only test seams
  (isAwaitingThreadStart, hasThread).
- The tests wait for the request before feeding its response, and the
  pending-write spin loop is bounded.
- CI passes XCTest's per-test execution allowance (300s by default) to the
  app-host batches so a single parked test fails with a spindump instead of
  taking the batch to its 30-minute timeout, and the hang sample prints
  more threads.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa

* test: replace mobile lane timing wait with completion signal

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
austinywang added a commit that referenced this pull request Sep 10, 2026
* ci: persist nightly Xcode compilation caches

* test(ci): require nightly caches to prune dead CAS generations before saving

The nightly workflow test now demands that both nightly cache jobs prune
dead Xcode CAS generations before measuring the 5 GiB bound, save on a miss
from the bound step's verdict instead of rescanning with hashFiles, and keep
the cache key prefix identical to the PR release build that restores it.
It also adds a behavioural test for the pruning helper and wires it into the
workflow guard job.

Both fail until the next commit adds the helper and the workflow changes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R7fdh4BL4KeoUxNukaSdTs

* ci: prune dead Xcode CAS generations so warm nightly caches persist

Nightly never published a warm compilation cache. Xcode's CAS chains v1.N
generations and only deletes the dead one at its next open, so a warm full
build leaves two full generations behind: 6.3 GiB on main against the 5 GiB
bound, which then removed the directory and the save step found nothing to
upload ("Path Validation Error" in every warm run). Only cold builds, at
about 3.1 GiB, ever saved, so main kept restoring a days-old entry.

Prune the dead generations before measuring, the same way the CAS's own
garbage collection would, in both nightly cache jobs and in the PR release
build that restores the same cache. Save explicitly on a miss from the
bound step's verdict instead of rescanning the directory with hashFiles,
which the runner aborts after 120 seconds. Keep the cache key prefix as it
was: PR release builds restore nightly's cache by that prefix, so the v2
key bump would have cut them off from the cache warmed from main.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R7fdh4BL4KeoUxNukaSdTs

* fix(web): guard the devbox reachability entrypoint without Bun typings

`import.meta.main` is Bun-only and the web typecheck program has no Bun
typings, so `bun run typecheck` fails on main since #12132:

  scripts/check-devbox-image-reachable.ts(148,17): error TS2339:
  Property 'main' does not exist on type 'ImportMeta'.

That PR never ran the web typecheck because ci.yml only fires for a short
path list. Guard the entrypoint by comparing import.meta.url with argv[1],
the pattern the other scripts under web/scripts already use. Direct runs
still execute main (`--print-key` prints the key) and importing the module
from tests stays side-effect free.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(ci): require non-fatal cache pruning and pin the shared cache key prefix

The nightly workflow test now demands that the pruner call in both nightly
cache jobs and the PR release build cannot fail the build, and that every
Xcode compilation cache key and restore-key in nightly.yml and ci.yml carries
the one shared prefix, so renaming a key on either side (or a key but not
its restore-keys) is caught. Mutation-tested: dropping the prune, moving it
after the measurement, restoring the hashFiles gate, removing the miss gate,
the bound step id, or the save= output, and every prefix rename now fail.

Fails until the next commit makes the prune call non-fatal.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* ci: never fail a build because cache pruning failed

Pruning is an optimisation. With `set -euo pipefail`, a pruner that cannot
start (missing interpreter, syntax error) aborted the bound step and the
whole nightly. Log a workflow warning and measure the directory unpruned
instead, which at worst reproduces the old behaviour of skipping the save.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(ci): pruner must survive root-level generations and unlistable CAS dirs

Pruning a generation that sits directly under the cache root removed a
directory the candidate list still named, so the next iteration raised
FileNotFoundError and the remaining CAS directories went unpruned for that
run. A CAS directory the runner cannot list raised as well. Both now have to
be skipped with a message while the other directories are still pruned.

Fails until the next commit hardens the pruner.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* ci: keep pruning the remaining CAS dirs when one vanishes or cannot be listed

Skip a candidate that an earlier root-level prune already removed, and treat
an OSError while inspecting a CAS directory as "leave this one alone" instead
of aborting the walk, so one odd directory cannot leave the others unpruned.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(ci): nightly guard requires a build-only measurement lane and an oversize-cache warning

A full branch dispatch of nightly.yml signs and notarizes under the release
identity, so the only safe way to time the nightly build job from a branch is
a run that stops at the unsigned universal build. The guard now requires
`build_only` and `cold_cache` workflow_dispatch inputs that skip the helper,
signing, notarization, dSYM upload and publication jobs, keep should_publish
false, and use their own concurrency group so a measurement run can never
cancel a pending publish. It also requires the bound step to report an
oversize cache as a workflow warning, since that path silently freezes the
cache at the last saved entry. This commit adds the test only; it fails until
the workflow changes land.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018wsDzi4Q5uSdU9AYq75tmE

* ci: add a build-only nightly measurement lane and warn when the cache bound skips a save

nightly.yml gains two workflow_dispatch inputs. `build_only` runs decide and
the unsigned universal app build and stops: the helper, signing, notarization,
dSYM upload and publication jobs are gated off, should_publish is forced
false, and the run gets its own concurrency group so it can never cancel a
pending publish. `cold_cache` (only honoured with build_only) skips the
compilation cache restore so the same commit can be measured as a cache miss;
the missing restore output reads as a miss, so the cold build still saves.
This is the only way to time the nightly build job from a branch without
notarizing under the release identity.

The bound step now reports an oversize cache as a workflow warning in
nightly.yml and ci.yml: nothing is saved on that path, so every later build
restores the same older entry and the cache silently freezes. The comments
also state the real rotation rule from LLVM's UnifiedOnDiskCache: a new
primary generation is started when the current one ends a build above half
of COMPILATION_CACHE_LIMIT_SIZE, which is 1.5 GiB against a 3.3 GB working
set, so every warm build rotates and leaves a dead generation behind.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018wsDzi4Q5uSdU9AYq75tmE

* ci: prune-xcode-compilation-cache measures only the generations it removes

The pruner walked every generation, including the two live ones, to print
per-generation sizes, and the workflow then ran `du -sk` over the retained
cache: two full traversals of up to 5 GiB for a log line (CodeRabbit on
PR #12039). Stale generations are now selected by name first and only those
are measured before removal; the retained size comes from the workflow's
single `du`.

The speculative handling of generations placed directly under the cache root
is gone: Xcode never writes that layout, and a layout the pruner does not
recognise is left alone and measured unpruned rather than guessed at. An
unlistable CAS directory is still skipped with a message while the remaining
directories are pruned. Both behaviours keep their tests.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018wsDzi4Q5uSdU9AYq75tmE

* test(ci): build_only must always build the universal app; pruner must leave unlocked directories alone

Review findings on PR #12039. The nightly guard now matches each job's
complete job-level `if:` verbatim, so the build_only exclusion can only be a
conjunctive clause, and requires build_only to imply should_build (a
measurement dispatch on main must not depend on the nightly tag) and to
disable the fast arm64 path (a measurement always builds the production
universal workload). The pruner test requires a CAS directory without a
`lock` file to be left alone, since nothing can be locked there, and only
exercises the unlistable-directory path when permission bits actually took
hold. Test-only commit; it fails until the fixes land.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* ci: build_only always builds the universal app; pruner leaves unlocked CAS directories alone

A build_only dispatch is a measurement of the production nightly build, so it
now implies should_build (on main it no longer depends on whether the nightly
tag already matches HEAD) and ignores the fast arm64 input instead of quietly
measuring a one-architecture build.

The pruner only ever deletes generations while holding the CAS directory's
`lock` exclusively. A directory without a `lock` file has never been opened
by the toolchain, so there is nothing to lock; it is now reported and left
alone rather than pruned unlocked.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* web(tests): carry the bun:test type shim fix so web-typecheck passes until main has it

main at 8ba29ea fails `bun run typecheck`: the tests added by #12257
(billing-alerts, cron-alerts) and #12250 (billing-purchase `.mock.calls`,
vm-devbox-image `import.meta.dir`) do not type against the repository's own
`web/tests/bun-test.d.ts` shim, so `ci-status` is red for every PR that does
not carry a fix. This takes the shim and the two one-line test edits verbatim
from PR #12105's branch (7df81ef), whose CI passes on the same base, so a
later merge of that fix is conflict-free. No behaviour change: the shim is a
`.d.ts` and the test edits keep their assertions.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test: reject non-keyboard events in file explorer shortcuts

* fix: unblock app-host tests at event and async wait boundaries

* ci: keep focused macOS tests on the required SDK 26 toolchain

* test(ci): cap selected SDKs while preserving legacy runner fallback

* ci: cap focused-test SDK selection without dropping older runners

* test: install a valid shortcut before checking file explorer routing

* test: fence detach command capture without blocking the main actor

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
aerickson pushed a commit to aerickson/cmux that referenced this pull request Sep 13, 2026
…changed (manaflow-ai#12132)

* cloud: prove the shipped devbox image is reachable, and only when it changed

The bake smokes the daemon once, when an image is made. Nothing asks whether
a client can reach the daemon in the image we are shipping now, using the
cmux-tui build we are shipping now. Those drift apart with no commit at all:
files.cmux.com moves to a new client on every cmux-tui release while the
manifest keeps pinning yesterday's snapshot.

The check boots the manifest's current default and runs a real round trip
(enroll a device, connect over ws://[::1]:1337/v1/link, open a workspace, run
a process, read its output). When the live client differs from the baked one
it repeats the whole trip with the live binary, so a protocol change that
would break a freshly built app fails here instead of in production.

It costs 17s with both legs, 9s when the client is the baked one, and one
small VM that is deleted on every path.

Running it when nothing changed is waste, so the job keys a cache on the
default image ids plus the client sha256 and skips a pair that has already
passed. That makes the daily schedule, which exists to catch the client
moving without a commit, free on a day when it did not.

* TEMPORARY: point the reachability check at a stock image to prove it fails

* Restore the manifest default after proving the check fails

Run 34210445406 booted freestyle/ubuntu-sm, waited out the 120s daemon
budget, reported "UNREACHABLE: the baked daemon never came up", deleted the
machine anyway, and exited 1, so the job is red when the image is broken and
leaves nothing running behind it.
aerickson pushed a commit to aerickson/cmux that referenced this pull request Sep 13, 2026
…aflow-ai#12154)

`import.meta.main` is Bun-only and is not part of TypeScript's ImportMeta
type, so `bun run typecheck` has failed on main since manaflow-ai#12132:

  scripts/check-devbox-image-reachable.ts(148,17): error TS2339:
  Property 'main' does not exist on type 'ImportMeta'.

Use the portable `fileURLToPath(import.meta.url)` vs `process.argv[1]`
check that the sibling devbox scripts already rely on. Behavior is
unchanged: `bun scripts/check-devbox-image-reachable.ts --print-key`
still runs main, and importing `reachabilityKey` from the test does not.

Closes manaflow-ai#12147


Claude-Session: https://claude.ai/code/session_01Qiy38q5XFYQU3CeccDENty

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
aerickson pushed a commit to aerickson/cmux that referenced this pull request Sep 13, 2026
* test: cover creation initial command plumbing

* test: isolate creation command regression

* feat: add initial commands to terminal creation

* test: require creation request dispatch

* test(web): preserve hosted config overrides

* test(cli): expect OMP restore path

* test(session): expect codex wrapper shim

* docs(cli): sync restore help contract

* test(cli): preserve interactive shell for creation command

* fix(cli): inject creation commands into interactive shells

* fix(cli): preserve terminal creation command text

* test(cli): reject invalid creation command input

* fix(cli): validate terminal creation input boundaries

* fix(cli): resolve creation validation in app context

* test(cli): cover review-found creation boundaries

* test(cli): import welcome setting contract

* test(cli): wait for terminal readiness before welcome assertion

* test(cli): exercise focused workspace welcome path

* fix(cli): close terminal creation review gaps

* test: cover app-host temp path aliases

* ci: accept validated macOS temp aliases

* Fix Xcode 26 warning regressions

* test: keep install command fixture inert

* test: repair app-host CLI fixtures

* fix: remove duplicate dock initial input plumbing

* test: bound workspace readiness regression wait

* fix: order dock creation arguments

* fix: order dock input after startup options

* fix: preserve null terminal creation types

* fix: keep readiness wait actor-safe

* test: isolate creation command socket environment

* test: keep cloud splits local for initial input

* fix: keep initial input out of cloud split routing

* ci: route browser skill contract through linux runner

* test: isolate cloud routing fixture from local surfaces

* refactor: keep terminal creation checks within file budgets

* fix: wire terminal creation helpers into app target

* fix: repair current main compile errors

* docs: document --command on terminal creation commands

Cover the new --command flag on new-split, new-pane, and new-surface (and
the existing one on new-workspace) in the cmux and cmux-workspace skills,
the CLI contract (table rows, an Initial terminal command section, and
--help probes), and the web API docs in English and Japanese.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa

* fix: silence unmutated payload warning in browser key replay

CI's Swift warning budget fails on main because the delivered-key payload
in TerminalController.swift is declared var but never mutated. Use let so
the tests-build-and-lag job can pass on this branch.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa

* fix(web): keep devbox reachability script typechecking without bun-types

`import.meta.main` (added on main in manaflow-ai#12132) fails `tsgo --noEmit` because
the web tsconfig does not include bun-types, which turns the CI cheap layer
red and skips every macOS job. Cast the meta object so Bun's runtime flag
still gates main() while the typecheck passes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa

* fix: repair package test import and Swift warnings inherited from main

- FakeTerminalEngine.swift (from manaflow-ai#10564) uses UUID without importing
  Foundation, which breaks `swift test` for CmuxTerminal in the
  swift-package-tests job.
- Parenthesize the two compactMap trailing closures in the pane memory
  guardrail guard and hop onto the main actor before reconcilePresentation
  in the session index table observer; both were new warnings over the
  cmux-owned Swift warning budget in tests-build-and-lag.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa

* fix: clear two more Swift warnings inherited from main

Drop the unreachable default branch in the cmux-tui snapshot parser's
exhaustive resource-kind switch and stop binding the unused rowID shadow in
SurfaceCatalogModel, so the cmux-owned warning budget passes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa

* test: bring MachinesPanelModelTests in line with the current cloud tree and catalog

Four tests in the app-host shard were stale against main:
- a sleeping or broken machine now keeps a Ports group whose status row
  explains how to discover ports (manaflow-ai#12051), and an unregistered machine shows
  a Connecting placeholder instead of no children;
- display rows placed inside a remote workspace carry their tab id like
  every other placement;
- the catalog drops writes for a cloud machine with no registered provider,
  so the two workspace-group tests register the file's GroupFakeProvider
  before replacing resources.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa

* chore: normalize project.pbxproj after main merges

The auto-merged project file drifted from scripts/normalize-pbxproj.py
output, which fails the workflow-guard pbxproj check and gates every macOS
CI job behind it.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa

* test: assert pane re-homing through the journal event for started turns

manaflow-ai#11976 moved the pane-scoped notification clear for UserPromptSubmit and
PreToolUse out of the Claude hook and into the app's journal reconciler
(clearInvalidatedNotifications keys off the event's workspace and surface).
The two moved-pane tests still looked for the hook's old clear_notifications
send and turned the focused notification-routing step red on main.

They now assert the agent_journal_append event carries the re-homed
workspace and surface (and, for PreToolUse, the running phase), and keep
the guards against whole-workspace or stale-workspace clears.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa

* test: bound WebKit page-load waits in the design-mode screenshot suite

On CI app hosts the WebContent process sometimes disappears mid-batch
("Could not signal service com.apple.WebKit.WebContent"), after which a
loadHTMLString navigation never reports didFinish. The screenshot evaluator
tests awaited that signal without a bound, so one lost process wedged the
whole app-host batch until the 30-minute timeout (shard 4 on runs
34232451577 and 34235694241 both stalled in
smoothScrollingPageCapturesRequestedRegionAndRestoresOffset).

Race every real page-load wait against a 30-second budget and fail the
test fast instead, so the rest of the batch still runs and reports.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa

* test: tolerate an already-closed socket when bounding PTY bridge reads

testPTYBridgeDefersHalfCloseUntilAttachCompletes half-closes its bridge
socket, and the bridge answers and closes so quickly that the follow-up
SO_RCVTIMEO setsockopt sees a torn-down connection and fails with EINVAL.
That thrown error is counted as an unexpected failure and fails the whole
app-host batch (it reproduces on main's own shard 6 run 34221588627).

The timeout only bounds the read that follows, which returns EOF at once
on such a socket, so skip the option instead of throwing.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa

* ci: sample the wedged app host before killing a timed-out unit-test batch

App-host unit-test batches on main and on PRs intermittently sit at the
30-minute batch timeout with only "started" lines for whichever tests were
in flight, which says nothing about what blocked the main actor. Sample
the app host process before terminating xcodebuild and print the leading
call graph in a log group, so the next hang names its stuck frames.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa

* test: queue the first Codex prompt before asserting the second is rejected

testCodexInputQueueBeforeThreadIsBounded spawned the first submit in a Task
and immediately submitted a second prompt on the same main actor. The
second call ran first, took the single pre-thread queue slot, and then
awaited a thread the test only starts afterwards, so the test hung until
the 30-minute app-host batch timeout (shard 2 on runs 34235694241 and
34245949340, and main's own shard 2).

Yield to the spawned task before the second submit so the first prompt
holds the slot and the second is rejected as intended.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa

* test: gate the concurrent file-preview save on isSaving instead of a FIFO

testSaveTextContentIgnoresConcurrentSaveRequest replaced the previewed file
with a FIFO so the first write would stay in flight. Since the preview
panel re-opens its watched path for change monitoring, a FIFO with no
writer can block the app host's main thread, and this test was the
unfinished XCTest in two hung app-host shards (shard 1 on run 34232451577,
shard 5 on run 34245949340). saveTextContent() sets isSaving synchronously
and completes on a later main-actor hop, so the second request is always
observed while the first is still saving without any FIFO.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa

* ci: match the app host for hang sampling even when launched with arguments

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa

* test: wait for the Codex thread/start request before answering it; per-test CI allowance

Every CodexAppServerSessionTests case fed the thread/start response after a
single Task.yield() following the initialize response. The session sends
`initialized` and then `thread/start` from a spawned main-actor task, so
when that task had not reached the request yet the response was dropped as
unknown, every later submit waited for a thread forever, and the reentrant
turn test spun on Task.yield() until the batch timeout (shard 2 on runs
34245949340 and 34256455336; main shows the same).

- CodexAppServerSession gains two read-only test seams
  (isAwaitingThreadStart, hasThread).
- The tests wait for the request before feeding its response, and the
  pending-write spin loop is bounded.
- CI passes XCTest's per-test execution allowance (300s by default) to the
  app-host batches so a single parked test fails with a spindump instead of
  taking the batch to its 30-minute timeout, and the hang sample prints
  more threads.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01EXS7towgxy33eJ4ZTMeQHa

* test: replace mobile lane timing wait with completion signal

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>
aerickson pushed a commit to aerickson/cmux that referenced this pull request Sep 13, 2026
* ci: persist nightly Xcode compilation caches

* test(ci): require nightly caches to prune dead CAS generations before saving

The nightly workflow test now demands that both nightly cache jobs prune
dead Xcode CAS generations before measuring the 5 GiB bound, save on a miss
from the bound step's verdict instead of rescanning with hashFiles, and keep
the cache key prefix identical to the PR release build that restores it.
It also adds a behavioural test for the pruning helper and wires it into the
workflow guard job.

Both fail until the next commit adds the helper and the workflow changes.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R7fdh4BL4KeoUxNukaSdTs

* ci: prune dead Xcode CAS generations so warm nightly caches persist

Nightly never published a warm compilation cache. Xcode's CAS chains v1.N
generations and only deletes the dead one at its next open, so a warm full
build leaves two full generations behind: 6.3 GiB on main against the 5 GiB
bound, which then removed the directory and the save step found nothing to
upload ("Path Validation Error" in every warm run). Only cold builds, at
about 3.1 GiB, ever saved, so main kept restoring a days-old entry.

Prune the dead generations before measuring, the same way the CAS's own
garbage collection would, in both nightly cache jobs and in the PR release
build that restores the same cache. Save explicitly on a miss from the
bound step's verdict instead of rescanning the directory with hashFiles,
which the runner aborts after 120 seconds. Keep the cache key prefix as it
was: PR release builds restore nightly's cache by that prefix, so the v2
key bump would have cut them off from the cache warmed from main.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_01R7fdh4BL4KeoUxNukaSdTs

* fix(web): guard the devbox reachability entrypoint without Bun typings

`import.meta.main` is Bun-only and the web typecheck program has no Bun
typings, so `bun run typecheck` fails on main since manaflow-ai#12132:

  scripts/check-devbox-image-reachable.ts(148,17): error TS2339:
  Property 'main' does not exist on type 'ImportMeta'.

That PR never ran the web typecheck because ci.yml only fires for a short
path list. Guard the entrypoint by comparing import.meta.url with argv[1],
the pattern the other scripts under web/scripts already use. Direct runs
still execute main (`--print-key` prints the key) and importing the module
from tests stays side-effect free.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(ci): require non-fatal cache pruning and pin the shared cache key prefix

The nightly workflow test now demands that the pruner call in both nightly
cache jobs and the PR release build cannot fail the build, and that every
Xcode compilation cache key and restore-key in nightly.yml and ci.yml carries
the one shared prefix, so renaming a key on either side (or a key but not
its restore-keys) is caught. Mutation-tested: dropping the prune, moving it
after the measurement, restoring the hashFiles gate, removing the miss gate,
the bound step id, or the save= output, and every prefix rename now fail.

Fails until the next commit makes the prune call non-fatal.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* ci: never fail a build because cache pruning failed

Pruning is an optimisation. With `set -euo pipefail`, a pruner that cannot
start (missing interpreter, syntax error) aborted the bound step and the
whole nightly. Log a workflow warning and measure the directory unpruned
instead, which at worst reproduces the old behaviour of skipping the save.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(ci): pruner must survive root-level generations and unlistable CAS dirs

Pruning a generation that sits directly under the cache root removed a
directory the candidate list still named, so the next iteration raised
FileNotFoundError and the remaining CAS directories went unpruned for that
run. A CAS directory the runner cannot list raised as well. Both now have to
be skipped with a message while the other directories are still pruned.

Fails until the next commit hardens the pruner.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* ci: keep pruning the remaining CAS dirs when one vanishes or cannot be listed

Skip a candidate that an earlier root-level prune already removed, and treat
an OSError while inspecting a CAS directory as "leave this one alone" instead
of aborting the walk, so one odd directory cannot leave the others unpruned.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test(ci): nightly guard requires a build-only measurement lane and an oversize-cache warning

A full branch dispatch of nightly.yml signs and notarizes under the release
identity, so the only safe way to time the nightly build job from a branch is
a run that stops at the unsigned universal build. The guard now requires
`build_only` and `cold_cache` workflow_dispatch inputs that skip the helper,
signing, notarization, dSYM upload and publication jobs, keep should_publish
false, and use their own concurrency group so a measurement run can never
cancel a pending publish. It also requires the bound step to report an
oversize cache as a workflow warning, since that path silently freezes the
cache at the last saved entry. This commit adds the test only; it fails until
the workflow changes land.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018wsDzi4Q5uSdU9AYq75tmE

* ci: add a build-only nightly measurement lane and warn when the cache bound skips a save

nightly.yml gains two workflow_dispatch inputs. `build_only` runs decide and
the unsigned universal app build and stops: the helper, signing, notarization,
dSYM upload and publication jobs are gated off, should_publish is forced
false, and the run gets its own concurrency group so it can never cancel a
pending publish. `cold_cache` (only honoured with build_only) skips the
compilation cache restore so the same commit can be measured as a cache miss;
the missing restore output reads as a miss, so the cold build still saves.
This is the only way to time the nightly build job from a branch without
notarizing under the release identity.

The bound step now reports an oversize cache as a workflow warning in
nightly.yml and ci.yml: nothing is saved on that path, so every later build
restores the same older entry and the cache silently freezes. The comments
also state the real rotation rule from LLVM's UnifiedOnDiskCache: a new
primary generation is started when the current one ends a build above half
of COMPILATION_CACHE_LIMIT_SIZE, which is 1.5 GiB against a 3.3 GB working
set, so every warm build rotates and leaves a dead generation behind.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018wsDzi4Q5uSdU9AYq75tmE

* ci: prune-xcode-compilation-cache measures only the generations it removes

The pruner walked every generation, including the two live ones, to print
per-generation sizes, and the workflow then ran `du -sk` over the retained
cache: two full traversals of up to 5 GiB for a log line (CodeRabbit on
PR manaflow-ai#12039). Stale generations are now selected by name first and only those
are measured before removal; the retained size comes from the workflow's
single `du`.

The speculative handling of generations placed directly under the cache root
is gone: Xcode never writes that layout, and a layout the pruner does not
recognise is left alone and measured unpruned rather than guessed at. An
unlistable CAS directory is still skipped with a message while the remaining
directories are pruned. Both behaviours keep their tests.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>
Claude-Session: https://claude.ai/code/session_018wsDzi4Q5uSdU9AYq75tmE

* test(ci): build_only must always build the universal app; pruner must leave unlocked directories alone

Review findings on PR manaflow-ai#12039. The nightly guard now matches each job's
complete job-level `if:` verbatim, so the build_only exclusion can only be a
conjunctive clause, and requires build_only to imply should_build (a
measurement dispatch on main must not depend on the nightly tag) and to
disable the fast arm64 path (a measurement always builds the production
universal workload). The pruner test requires a CAS directory without a
`lock` file to be left alone, since nothing can be locked there, and only
exercises the unlistable-directory path when permission bits actually took
hold. Test-only commit; it fails until the fixes land.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* ci: build_only always builds the universal app; pruner leaves unlocked CAS directories alone

A build_only dispatch is a measurement of the production nightly build, so it
now implies should_build (on main it no longer depends on whether the nightly
tag already matches HEAD) and ignores the fast arm64 input instead of quietly
measuring a one-architecture build.

The pruner only ever deletes generations while holding the CAS directory's
`lock` exclusively. A directory without a `lock` file has never been opened
by the toolchain, so there is nothing to lock; it is now reported and left
alone rather than pruned unlocked.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* web(tests): carry the bun:test type shim fix so web-typecheck passes until main has it

main at 8ba29ea fails `bun run typecheck`: the tests added by manaflow-ai#12257
(billing-alerts, cron-alerts) and manaflow-ai#12250 (billing-purchase `.mock.calls`,
vm-devbox-image `import.meta.dir`) do not type against the repository's own
`web/tests/bun-test.d.ts` shim, so `ci-status` is red for every PR that does
not carry a fix. This takes the shim and the two one-line test edits verbatim
from PR manaflow-ai#12105's branch (7df81ef), whose CI passes on the same base, so a
later merge of that fix is conflict-free. No behaviour change: the shim is a
`.d.ts` and the test edits keep their assertions.

Co-Authored-By: Claude Fable 5.1 <noreply@anthropic.com>

* test: reject non-keyboard events in file explorer shortcuts

* fix: unblock app-host tests at event and async wait boundaries

* ci: keep focused macOS tests on the required SDK 26 toolchain

* test(ci): cap selected SDKs while preserving legacy runner fallback

* ci: cap focused-test SDK selection without dropping older runners

* test: install a valid shortcut before checking file explorer routing

* test: fence detach command capture without blocking the main actor

---------

Co-authored-by: Claude Fable 5.1 <noreply@anthropic.com>

This branch was successfully deployed

2 active and 1 inactive deployments
Preview – cmux166 — 95a83a50 Deployed Sep 8, 2026 by vercel[bot]
Preview – cmux41 — 95a83a50 Deployed Sep 8, 2026 by vercel[bot]
cloud-vm-image-checks — 95a83a50 Deployed Sep 8, 2026 by lawrencecchen via reachable #3
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

1 participant