Skip to content

cloud onboarding: introducing cmux cloud welcome, and cloud enablement screens - #17230

Merged
austinywang merged 42 commits into
cloud-gate-upgradefrom
cloud-onboarding
Oct 5, 2026
Merged

austinywang merged 42 commits into
cloud-gate-upgradefrom
cloud-onboarding

Conversation

@lucasr1b

@lucasr1b lucasr1b commented Oct 4, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

nothing in the app tells people cloud exists. you find it by opening the cloud tab, and that tab's setup screen didn't really sell it. this adds a one-time welcome and redoes the cloud tab's setup screen to match. stacked on #17127 (the enable / upgrade gate), so this diff is only the onboarding.

"introducing cmux cloud" welcome

  • a glass window with "introducing" over "cmux cloud" and a "new" badge, a preview of the cloud tree as the sidebar shows it mid-work (one machine open on its workspace and two terminals), the subtitle, three reasons (keeps running, any mac, shared with your team) and one button.
  • the button is the next step for the account: sign in (signed out), upgrade to pro (free), enable cloud (pro / max), or set up cloud (plan not known yet, opens the cloud tab whose gate decides). plus not now and the plan note.
  • shown once per mac (cmux.cloud.welcome.seen, so once per release channel): existing users see it on the first launch after updating, new users on their second launch (a fresh install has no cached remote flags yet when the first window opens, and cloud reads as off until they load). only when cloud is offered on this mac (cloud-machines-enabled-release) and still off, so it follows the cloud rollout and never shows to people already using cloud. considered once per launch at the first main window, so a window opened later never pops it up. skipped under xctest and ui tests.
  • macOS 26: the window content sits inside an NSGlassEffectView (real window glass, no second rim, nothing lensed). macOS 14/15: the window material the about window uses. glass apis are behind #if compiler(>=6.2) like the rest of the repo.
  • debug builds can reopen it from help > show cloud welcome….
  • new ProUpgradeSource.cloudWelcome (mac_cloud_welcome) so upgrades from here are attributed.

cloud tab setup screen

  • free / pro / plan unknown now show the dark app icon with a badge (cloud, or a lock while the plan needs pro), a title for the plan ("upgrade to use cmux cloud" for free, "use cmux cloud" otherwise), the same four reasons, one button and the plan note.
  • every other state (checking, failed, cancelled, unavailable) is a simpler centered status with the same type and buttons. buttons use the sidebar's 6pt radius.
  • CloudMachinesEnablementView now takes plain values (phase, plan flags, closures). the app wiring moved to a thin CloudMachinesEnablementPanel in MachinesPanelView+Activation.swift; the coordinator state maps through an exhaustive switch. same states, same actions and services, same accessibility identifiers.

for austin: copy direction. the new copy says "cmux cloud" in lowercase (the welcome title and the cloud tab titles). the rest of the app still says "Cloud" (the tab, settings, menus, "Enable Cloud"). if lowercase is the direction it should spread; if not, it's two strings to change back.

Preview (welcome (new))
Preview
Preview (cloud tab, every state (lab))
Preview

the cloud tab's "before" is the after shot on #17127 (free plan): https://github.com/user-attachments/assets/fc42d038-e1f6-42b0-998e-258a62423c44. the welcome is new, so it's after-only. the cloud tab recording is the lab rendering the real CloudMachinesEnablementView with fake data, since this build can't load a plan.

Testing

  • tagged build (onboarding, local backend mode) compiled and launched. dogfooded live: the welcome (via help, and the screenshot below), its buttons, and the cloud tab with cloud turned off. local mode can't load a plan, so the real app shows the plan-unknown version; the free version is from the lab below.
  • lab: iterated in a standalone swiftui window that compiles the exact CloudWelcome*.swift and CloudMachinesEnablementView.swift sources with fake data (all states x light/dark x 280-380pt sidebar widths). the cloud tab recording above is that lab.
  • tests added, not run locally (repo rule: no local xcodebuild test): cmuxTests/CloudWelcomeTests.swift covers when the welcome shows and which button it offers. CI runs them.
  • python3 scripts/localization_catalog.py check: 0 parity errors. ./scripts/check-pbxproj.sh: clean.
  • not verified yet: the free / upgrade path against a real plan (needs the dev backend), macOS 14/15 fallback (not run), and the xcode 16.2 compile (CI covers it).
  • localization audit: new strings (welcome title + eyebrow + badge + notes + set up cloud, cloud tab titles, four reason titles and the team line, the debug menu item) are translated in all 8 locales. unused strings from earlier iterations removed, plus cloud.enable.planNote which nothing uses anymore. localize-changes flags an interpolated defaultValue in AppDelegate.swift; it's an existing string in that file, this PR adds none there.

Changelog

Added: An "Introducing cmux cloud" window on first launch (and once after updating) explains Cloud and offers the next step, and the Cloud tab's setup screen shows why to use it

Review

correctness review (subagent) before opening, findings and what was done:

  1. glass apis broke the xcode 16.2 / macOS 15 sdk build (NSGlassEffectView, .glassEffect). fixed: wrapped in #if compiler(>=6.2) with the material fallback, like TextBoxInput.
  2. a fresh install could pop the welcome up mid-session on cmd+n once remote flags arrived. fixed: considered once per launch at the first main window; an unseen welcome waits for the next launch.
  3. footer could clip on macOS 26 (plausible). dogfooding then found a real crash on open: the hosting view inside NSGlassEffectView kept invalidating its titlebar safe area and its constraints until AppKit aborted (too many update constraints passes). fixed with hosting.safeAreaRegions = []: the content clears the traffic lights with its own top padding, so it never needs the titlebar inset, and the window is sized once from that layout. this also makes the window lay out exactly like the lab.
  4. a stray return could enable cloud as the window appears at launch. fixed: no default-action shortcut on the primary button (esc still dismisses).
  5. cloud tab differences vs before (intended, reviewed in the lab): cancelled no longer repeats the reasons, plan checking is a spinner and one line, failed (requires pro) shows retry before upgrade (upgrade stays the prominent one).
  6. smaller: stale comments, a dead parameter, the badge ring color, locale-aware lowercasing of "new", the orphaned cloud.enable.planNote. all fixed.

no issues found in: double presentation, window lifetime / retain cycles, main-actor use, test skips, removed symbols, pbxproj and catalog checks.

Checklist

  • Behavior changes have added or updated tests, or Testing says why not
  • UI, settings, menu, schema, help-text or user-facing docs change: localization audited, and the result is stated above

View with [code]smith Autofix with [code]smith
Need help on this PR? Tag @codesmith-bot with what you need. Autofix is disabled.


Summary by cubic

Adds a one-time "Introducing cmux cloud" welcome window and reworks the Cloud tab's setup screen so people learn Cloud exists and what their next step is.

  • The welcome appears once per Mac (once per release channel) when Cloud is offered and still off, and its one button follows the account: sign in, upgrade to Pro, enable Cloud, or open the Cloud tab. Its close shortcut routes to the welcome window, never to another window.
  • The Cloud tab setup screen now leads with a plan-aware intro — the app icon with a Cloud badge (a lock while the plan needs Pro) above "Upgrade to use cmux cloud" for free or "Use cmux cloud" otherwise — with four reasons; checking, failed, cancelled, and unavailable states are simpler centered statuses.
  • CloudMachinesEnablementView now takes plain values and closures; the app wiring moved to a thin CloudMachinesEnablementPanel.
  • The welcome uses window glass on macOS 26 and the About-window material on macOS 14/15.
  • Upgrades from the welcome are attributed as mac_cloud_welcome via the new ProUpgradeSource.cloudWelcome.

Written for commit 1aaf634. Summary will update on new commits.

Review in cubic

…ayouts, sidebar intro with banner and reasons, enablement view takes plain values (wip)
…sidebar intro titles by plan, 6pt buttons, drop the comparison layouts
…considered once per launch, size after hosting, no return shortcut, comments, orphaned string)
@coderabbitai

coderabbitai Bot commented Oct 4, 2026 •

Copy link
Copy Markdown

Important

Review skipped

Auto reviews are disabled on base/target branches other than the default branch.

Please check the settings in the CodeRabbit UI or the .coderabbit.yaml file in this repository. To trigger a single review, invoke the @coderabbitai review command.

⚙️ Run configuration
  • Configuration used: Repository: manaflow-ai/cmux/.coderabbit.yaml
  • Review profile: ASSERTIVE
  • Plan: Advanced
  • Run ID: 6b424248-beb7-4a6a-aa1e-53d371a29f16

You can disable this status message by setting the reviews.review_status to false in the CodeRabbit configuration file.

Use the checkbox below for a quick retry:

  • 🔍 Trigger review
  • Autopilot · Keep fixing CodeRabbit findings and required CI, and resolving merge conflicts

Thanks for using CodeRabbit! It's free for OSS, and your support helps us grow. If you like it, consider giving us a shout-out.

❤️ Share

Comment @coderabbitai help to get the list of available commands.

@lucasr1b lucasr1b changed the title cloud onboarding: introducing cmux cloud welcome, and a cloud tab intro that says why cloud onboarding: introducing cmux cloud welcome, and cloud enablement screens Oct 4, 2026
@cursor

cursor Bot commented Oct 4, 2026

Copy link
Copy Markdown

Bugbot is paused — on-demand spend limit reached

Bugbot uses usage-based billing for this team and has hit its on-demand spend limit.

A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue.

@cursor

cursor Bot commented Oct 4, 2026

Copy link
Copy Markdown

Bugbot is paused — on-demand spend limit reached

Bugbot uses usage-based billing for this team and has hit its on-demand spend limit.

A team admin can raise the spend limit in the Cursor dashboard, or wait for the next billing cycle to continue.

@github-actions

github-actions Bot commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

CI failure attribution

CI passes on 1aaf63492d (run 37247096921 attempt 1).

Written by scripts/ci/classify_failures.py (ci-failure-attribution.yml); signatures are its SIGNATURES table. A machine verdict is the runner's fault, not this PR's.

@github-actions

github-actions Bot commented Oct 4, 2026 •

Copy link
Copy Markdown
Contributor

Dogfood tours of 1aaf6349

browser-notifications-tour at 1aaf6349: not run

skipped: CI built this head on a runner pool whose products the UI test Macs cannot load, and media never compiles one; gh workflow run pr-media.yml -f pr=<n> -f allow_compile=true does

Tours are picked by the paths globs in dogfood/scenarios/*.json; a Dogfood-tours: a, b line in the description picks them instead (none turns this off). Look at every frame before merging: a green tour only means no step failed.

austinywang and others added 4 commits October 3, 2026 22:44
* fix(cloud): explain unavailable display actions

Show the existing Cloud failure message when New Display is unavailable, and explain the machine ownership restriction when a display is opened into another Cloud workspace.

— unregistered

* fix(cloud): explain unavailable display actions

Show the existing Cloud failure message when New Display is unavailable, and explain the machine ownership restriction when a display is opened into another Cloud workspace.

— unregistered

* fix(cloud): reject cross-machine display opens before projection

* fix(cloud): probe legacy desktop VMs for displays

* fix(cloud): initialize wallpapers for new displays

* fix(cloud): fail closed before display double-click opens

* fix(cloud): replay repeated display ownership hints

* test(cloud): cover display helper refresh

* fix(cloud): refresh installed display helper

* test(cloud): require display service readiness probe

* fix(cloud): wait for display helper readiness

* fix(cloud): keep display creation retryable

* fix(cloud): authorize dynamic display ports

* fix(cloud): let new display retry guest discovery

* Revert "fix(cloud): let new display retry guest discovery"

This reverts commit 0fe008e.

* test(cloud): sidebar visibility filter must keep display creation state

Moves applyingDeviceVisibility next to SurfaceCatalogSnapshot in
CmuxSurfaceCatalogModel (no behavior change) so it is testable, and adds
a failing regression test.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(cloud): keep display creation state through the sidebar filter

applyingDeviceVisibility rebuilt SurfaceCatalogSnapshot from scratch and
dropped displayCreationMachines, staleMachineIDs and display memberships,
so every desktop VM's New Display row reported additional displays as
unavailable. Filter a copy instead so all per-machine state survives.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(surfaces): a split with no room opens as a tab in the target pane

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(cloud): optimistic display row; fall back to a tab when a split has no room

New Display now shows a "Starting display…" row from the click until the
guest answers, driven by the catalog's in-flight creation set.

Opening a sidebar resource splits the focused pane; once split admission
had no room left, the open failed with "Could not create the pane:
noSpace". SurfacePaneFactory now opens it as a tab in the pane that would
have been split.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(surfaces): refused split is typed; sidebar gestures open a tab

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(surfaces): scope the no-room tab fallback to sidebar gestures

The factory now throws a typed noSpace instead of opening a tab, so the
layout replay and socket open verbs keep a truthful refused split. Sidebar
opens and New Display use openPreferringSplit, which retries as a tab in
the pane that would have been split. The empty Displays row no longer
shows beside the optimistic Starting display row.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(cloud): display creation needs no discovery round trip

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* perf(cloud): open the display pane at the click and skip pre-create discovery

New Display paid two VM exec round trips (list, then create, about 2s
each) before any pane appeared. The create reply is already the full guest
catalog, so the list is dropped. The pane now opens at the click with a
native Starting display state and is adopted by the projection once the
guest assigns the display; creation failure closes it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(cloud): harden the reserved display pane

Drop a second click before it opens a pane; bind the reservation to the
created display id so no other projection adopts it; keep creating when
the pane cannot open; leave the display in the pool when the person closed
the pane; show the failure on a pane the socket refuses to close.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(cloud): additional displays serve noVNC beside the primary desktop

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(cloud): make additional displays reachable on the VM private address

Displays 2+ ran websockify on 127.0.0.1 while display 1 listens on [::].
The client reaches every display through the private address, so each new
display's first noVNC connection was refused and only recovered after a
timeout. Bind like display 1 (Xvnc stays -localhost), replace proxies an
older helper left on loopback, and drop the open-port call added for
display ports: it changed no routing and cost a control-plane request,
plus a desktop heal exec for display 1, on every open.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(cloud): stale proxy match covers only this display's websockify

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(cloud): match stale display proxies by argv, not pgrep regex

pgrep's ERE has no (?:...) groups, so the stale websockify match failed,
the old loopback listener kept the port, and the rebound proxy exited.
Read /proc argv for this user's websockify on this display's ports.
Verified on a live VM: displays 2 and 3 now accept the noVNC websocket on
the private address, one proxy each, X sessions untouched.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(cloud): a failed first desktop connection retries before failing

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(cloud): retry a display's first connection; warm the carrier at create

noVNC stops at Connect after an initial connect failure; its reconnect only
follows a session that once connected. Restored and new displays could lose
that race while the proxy or guest listener started, leaving "Failed to
connect to server". Retry the route's first connection three times
(0.5s/1s/2s on the injected clock) before showing the failure card.

The first display open on a machine spent 12-18s starting its browser
carrier after the guest exec. Start it alongside the create request.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(cloud): an explicit Retry gets its own quiet first-connection retries

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(cloud): run the guest display script test on the main actor

CloudGuestDisplayScript is @mainactor; the test called it from a
nonisolated context and broke the CmuxCloud package test build.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…7239)

* test(ci): Blacksmith keeps the shared root; reruns keep a per-runner root (red)

Blacksmith macOS runners report RUNNER_ENVIRONMENT=self-hosted, so they got
per-runner roots: every build started cold and no seed could be adopted
(job 111334766861). take-product-canonical-root.sh also refused a product
built at a per-runner root on a fleet Mac without glaeda.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ci): per-runner roots only on fleet Macs without glaeda

canonical-build-root.sh derived /private/tmp/cmux-ci-<runner> for every
self-hosted runner without the glaeda helper, which included Blacksmith's
ephemeral macOS runners: their builds started cold and their seeds were
never adopted. Derive it only when the fleet directory exists. Let
take-product-canonical-root.sh accept such a root when glaeda isn't there
to hold it, so app-host test reruns build where the product was compiled.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
…ile (#17207)

* fix(ci): align owned-build-state guards with per-runner canonical roots

#17168 (496195b) derives a canonical root per self-hosted runner, exports
CMUX_CI_CANONICAL_ROOT from the build-slot step and drops the shared
/private/tmp/cmux-ci fallback in test-e2e's owned-state step. Three tests in
tests/test_ci_owned_build_state.py still assumed the old contract, turning
"CI fast guards" red on main.

Update them to the new contract: the slot step exports the root, a
per-runner root reads its own store, and a missing root reads nothing. The
owned-state step now fails with a clear message when the root is missing, and
rejects a root nested under a cmux-ci-* prefix (e.g. cmux-ci-a/../x), so the
looser glob cannot map the store outside the Mac's package directory.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(l10n): translate the New Machine plan loading strings

#17135 added machines.new.plan.loading, .retry and .error with the English
text copied into every locale, so `localization_catalog.py check` reports 24
parity errors and "Fast static checks" fails on every PR. Translate them for
all catalog locales; Retry reuses common.retry's wording.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(l10n): match the catalog's plan terminology in de, ko and zh-Hant

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(tests): repair cmuxTests compile after the sidebar reorder change

#17070 moved RightSidebarModeBarDragLayout into the CmuxSidebar package,
but RightSidebarTabCustomizationTests still relied on the app module
exporting it, and CloudMachineOrderingTests put `try` calls inside an `&&`
in #expect, which the macro rejects ("operator can throw but expression is
not marked with 'try'"). Import CmuxSidebar and hoist the throwing lookups.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(tests): hoist the second throwing lookup out of #expect in CloudMachineOrderingTests

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ci): satisfy dogfood-build runner guards and the Swift warning budget

#17206 turned dogfood-build into a macOS build, which tripped three guards:
no pinned Xcode, no fork branch in runs-on, and no static-preflight
dependency. Add `scripts/select-ci-xcode.sh`, the standard owner and fork-PR
branches ahead of the existing selector, and `needs: static-preflight`. The
job's `if` already excludes forks, so manaflow-ai PRs keep the same runner.

#17070's `RightSidebarModeBarDragController.coordinateSpace` is read from
a Sendable geometry closure and broke the zero-warning budget; it is a plain
String constant, so mark it `nonisolated`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(ci): a root the glaeda hook exported wins over the per-runner root

Fails on main: canonical-build-root.sh treats an exported /private/tmp/cmux-ci
as unset and derives /private/tmp/cmux-ci-<runner>, so main-compile-probe
refuses the root glaeda placed it in (runs 37157423969, 37155786981).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ci): honor the canonical root the glaeda hook exported

#17168 (496195b) derives /private/tmp/cmux-ci-<runner> on every self-hosted
runner unless CMUX_CI_CANONICAL_ROOT names a non-default root. glaeda's
runner hook holds root 1 (/private/tmp/cmux-ci) for a compile job and
exports exactly that, so the job then built somewhere glaeda does not hold,
and main-compile-probe refused it: "glaeda placed this job in
/private/tmp/cmux-ci, not /private/tmp/cmux-ci-cmux13s-mac-mini-glaeda".

Let any exported root win. A self-hosted job with no exported root still
gets its per-runner root, so #17168's isolation for unplaced runners stays.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(ci): an unplaced self-hosted job keeps the shared seeded root

Fails on the branch: canonical-build-root.sh derives /private/tmp/cmux-ci-<runner>,
whose fingerprint never matches main's seeds, so new aws runners
(aws-m4pro-9-glaeda-3, aws-m4pro-8-glaeda-4) compiled cold and timed out at
35 minutes in compile admission (run 37158561950, both attempts).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ci): build unplaced self-hosted jobs at the shared seeded root

#17168 gave every self-hosted job without an exported root its own
/private/tmp/cmux-ci-<runner>. That root is part of the cache fingerprint,
so main's DerivedData seeds and the owned build state never match it, and the
prepare step clears it each job: every new aws runner compiled cold and hit
the 35-minute admission limit (aws-m4pro-9-glaeda-3, aws-m4pro-8-glaeda-4 on
run 37158561950).

Go back to /private/tmp/cmux-ci when no root is exported. glaeda's hook still
exports root N for the jobs it places, so those stay isolated. Unplaced
concurrent runners on one Mac share the root again, as before #17168, until
glaeda places those jobs too.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(ci): the pool picker never picks an owned label it was not configured with

Fails on main: simple_pool_picker adds every glaeda-<class>-xcode-* label it
sees on a runner, and glaeda-std-xcode-26.3 (ten aws EC2 runners, five per
Mac) sorts before the minis' 26.6, so compile admission ran there and timed
out at 35 minutes while the minis were idle.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ci): pick only configured owned pools, not any owned-looking label

simple_pool_picker added every glaeda-<class>-xcode-* label it saw on an
online runner to the configured CI_OWNED_POOL_SLOTS pools, then tried them
in string order. The aws EC2 Macs carry glaeda-std-xcode-26.3 (ten runners,
five per Mac), which sorts before the minis' glaeda-std-xcode-26.6, so PR
compile admission landed on contended EC2 runners and timed out at 35
minutes while the minis sat idle. Use only the configured pools.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(ci): the pool picker reads every page of organization runners

Fails on main: LiveState.runners reads one page of 100, but manaflow-ai has
509 runners and the first page holds the aws Macs and no idle mini, so the
picker never saw the mini fleet (run 37166910798 fell back to Blacksmith with
29 idle minis online).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ci): read every page of runners in the pool picker

LiveState.runners read only the first page of 100 organization runners.
manaflow-ai has about 500, and the first page holds the aws Macs but no idle
mini, so the picker never saw the mini fleet: it took the aws 26.3 label
while that was discoverable, and fell back to Blacksmith once it was not.
Read up to ten pages. Against live data the picker now sees 38 online
glaeda-std-xcode-26.6 minis (33 free) and picks them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: clear the Cloud sidebar warnings that broke the Swift warning budget

Compile admission on a mini (run 37167089830) built cleanly but failed the
zero-warning budget on three warnings from the Cloud sidebar work:
`SurfaceCatalog.shared` used as a default argument of two @mainactor
CloudWorkspaceSidebarPresentation entry points (evaluated nonisolated), and an
implicit `self` in CloudTreeOutlineView's machine-lift reopen closure.
Resolve `.shared` inside the main-actor body and make the capture explicit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ci): preserve canonical root and translation wording

* test(ci): align canonical root guard with main behavior

* test: use isolated catalogs in Cloud sidebar fixtures

Pass each test fixture catalog through the sidebar snapshot factory and direct presentation assertion so Cloud machine metadata is read from the catalog the test populated.

Co-authored-by: Austin Wang <austinwang115@gmail.com>

* test: model loading Cloud sidebar state

Keep device projections visible while their catalog rows are restoring, and assert that a reserved loading card shows machine identity before its directory appears after adoption.

Co-authored-by: Austin Wang <austinwang115@gmail.com>

* test: use a terminal panel for cloud directory settings

The sidebar detail test supplies a reported directory, so give it a real terminal panel instead of a loading card that correctly suppresses directory metadata.

Co-authored-by: Austin Wang <austinwang115@gmail.com>

* fix(cloud): restore device directory fallback lost in main merge

The merge of main took #17107's presentation file and dropped the
fallback from 27a8d80, so a device projection whose catalog row is
still restoring rendered no directory and
SidebarCloudWorkspaceBadgeTests.deviceNameIsVisibleBesideItsDirectory
failed on b6ea663 (the regression run is that CI failure).

Co-authored-by: Austin Wang <austinwang115@gmail.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Austin Wang <austinwang115@gmail.com>
* ci: require merge-gate before merges

* Add API-only merge gate for pull requests

* test: cover merge-gate override decisions

* ci: wire merge-gate tests and push timing

* ci: pin merge gate workflow source

* Reject read-only override authors

* ci: harden merge gate runner and author checks

* ci: keep merge gate on hosted capacity

* ci: fail closed when collaborator lookup fails

* ci: reject stale successes during reruns

* ci: fail closed for untimestamped queued checks

* ci: fail closed for untimestamped pending statuses

* ci: verify matching override failure evidence

* test: cover mismatched override failure evidence

* ci: choose newest merge gate check run

* ci: require compile evidence for main-fix merges (#16993)

* test: reproduce main-fix merging without compile evidence

The installed helper bypasses all checks under --main-fix. The regression records that it merges with no compile checks present.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ci: require build evidence for main fixes

gh-merge-green --main-fix now requires successful Release, Debug and Swift test target build steps on the exact PR head. Existing Swift test failures are waivable only when their parsed issue records match the same Swift test step on the exact base SHA; the audit comment records each match and unrelated failures remain fatal.

CPU, memory and disk checks: the validator uses one bounded 90-second GitHub request per call, caps captured output at 32 MiB, writes audit bodies to temporary files, and does not retain logs, caches or stores.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* ci: harden merge gate freshness and evidence

* ci: require per-check override reasons

* ci: complete merge gate review fixes

* docs(ci): document merge-gate rollout order

* ci: make merge helper rollout repairable

* ci: point merge helper refusals to repair runbook

* ci: resolve merge helper symlink for main fixes

* docs: remove stale main-fix PR reference

* fix: avoid pipefail false negatives in merge helper

* fix: keep merge gate events fresh per pull request

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
lawrencecchen and others added 11 commits October 4, 2026 01:15
* test(vms): reproduce missing shared Cloud VM resource pool

* feat(vms): share one vCPU and memory pool across a plan's Cloud VMs

Pro, Team (per paid seat), and Founder's Edition get up to 5 active VMs
sharing 20 vCPUs and 40 GB RAM, with machines up to the 32 GB xl row.
Max gets 80 vCPUs and 160 GB RAM with machines up to the 64 GB 2xl row.
The repository enforces the pool next to the active-VM count, under the
same transaction and billing lock, on create, Base open/reset, paused
resume, CPU/memory resize, and fork. Pricing, docs, app, and iOS copy
now describe pooled resources.

* Pooled VMs: Pro overrides stop at 32 GB; drop unused iOS pool strings

* test(vms): reproduce leaked compute claim and stale-plan resume pool

* fix(vms): give back unused compute claims and resume against the caller's pool

* docs(pricing): describe the Team pool per paid seat and localize pool strings in every locale
* test(cloud-vm): snapshot create dedups by idempotency key (red)

A retry of POST /api/vm/:id/snapshot with the same Idempotency-Key must return
the first snapshot and take no second one; a retry while the first runs is
refused as in progress; the same key with another name is a conflict; a failed
attempt frees the key.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(cloud-vm): snapshot create dedups by Idempotency-Key

POST /api/vm/:id/snapshot now reads Idempotency-Key. A new ledger table,
cloud_vm_snapshot_requests (one row per machine and key, additive migration
20261004120000), records a pending attempt before the provider call and the
provider snapshot after it. A retry with the same key returns the first
snapshot and records no second usage event; a retry while the first runs gets
409 vm_snapshot_in_progress (retryable); the same key with another name gets
409 vm_snapshot_idempotency_conflict; a failed attempt frees the key; a pending
row older than 15 minutes (route budget 600 s) is taken over by a retry.
Requests without a key behave as before.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* chore(cloud-vm): smoke --snapshot-check proves snapshot idempotency

With --create, takes one snapshot of the throwaway smoke machine twice with the
same Idempotency-Key, requires the same snapshotId, deletes that snapshot,
then destroys the machine as before. No fixed VM id and no printed token: the
smoke mints its own throwaway user session.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Preserve selected tab when closing another surface

* Only select the closing pane's tab when that pane is focused

BonsplitController.selectTab also focuses the pane, so calling it for an
unfocused pane moved focus into the pane where a tab closed, the focus theft
this change is meant to stop. Bonsplit already keeps the selection when an
unselected tab closes; the shouldCloseTab change is the actual fix.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Leo Li <cheerleaderleo@outlook.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* test: guard managed contributor difficulty labels

* test: reject invalid difficulty descriptions

* test: exercise difficulty label sync requests

---------

Co-authored-by: Leo Li <cheerleaderleo@outlook.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
 (#17250)

* test(cloud): expect the display helper readiness loop

#17132 replaced the one-line `list || exit 1` probe with a bounded
retry loop and updated the package test, but cmuxTests still asserted
the old line, so CloudDisplayCatalogTests.guestCommandShape failed on
main (run 37186487936, shard 4/7).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(cloud): make the restored-display retry test a desktop

#17132's initialDesktopFailureRetries configured port 6902 without a
resource ID, so CloudBrowserAccessState.isDesktop was false and
desktopConnectionDidChange ignored the failure: no quiet retry, and the
following navigations[1] trapped and crashed the shard 5 app host on
main (run 37186487936). Give it the .display identity additional
displays carry and require the retry before indexing it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(cloud): a foreign display click shows one synchronous hint

#17103 made a display click on another machine's workspace reject
synchronously through showDisplayOpenHint, so no tree operation starts
and waitForOpen() waited out the 60 s suite limit on main (run
37186487936, shards 2 and 3). Assert no operation ran and that the
rejection is shown once; this stays red until the double-click stops
repeating the hint.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(cloud): show a foreign display's ownership hint once per double-click

AppKit delivers a double-click's first click to handleSingleClick,
which already shows the ownership hint, and #17103 also showed it from
handleDoubleClick, so the user got the same rejection twice.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(cloud): never show upstream tree failures in the Machines status row

#17103 started presenting treeError verbatim so its ownership hints
would survive, but treeError also carries raw upstream failures
(LocalizedError descriptions, create output), so a URL with query
parameters reached the toolbar text, hover help and copy menu.
MachinesCloudStatusTests.emptyStatusHasNoProgressPresentation caught it
on main (run 37186487936, shard 7).

Hints now travel through their own onHint sink (falling back to
onFailure for other callers), the panel records them as trusted copy,
and the status row shows the tree error verbatim only when it is that
hint; everything else gets the safe recovery message.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* fix: keep merge-gate diagnostics on the exact PR

* fix: print repair guidance for gate diagnostics

* fix: allow merge-gate PR comments when permitted

* fix: reject mismatched merge-gate event identities

* fix: keep merge-gate alive on comment errors
* test: codex wrapper must not put the cmux-cua credential in argv

The Codex wrapper passes the cmux-cua socket credential as
-c mcp_servers.cmux-cua.env.CMUX_CUA_SOCKET_AUTH_TOKEN=<value>, so any
local user can read it with ps. These assertions require the value to be
absent from argv, Codex to receive it in its environment, and the MCP
server to receive it through env_vars. The fake Codex now builds the MCP
server environment the way Codex does (allow-list + env_vars + env).

* fix: keep the cmux-cua credential out of the Codex argv

The wrapper passed the socket credential as a Codex -c env override, which
put the value in the codex process argv where any local user can read it
with ps. The wrapper now exports CMUX_CUA_SOCKET_AUTH_TOKEN in the parent
shell before exec and emits mcp_servers.cmux-cua.env_vars so Codex forwards
the variable by name to the MCP server (Codex env_vars, openai/codex#5246).
Codex's default shell_environment_policy excludes *TOKEN* names from agent
shell commands. Attachment stays fail-closed when no credential resolves.
* test: codex wrapper overrides must survive subcommand -c flags

Codex declares -c/--config, --enable, and --disable as clap global
arguments. When a subcommand such as exec or resume also receives one,
Codex 0.159.3 keeps only the subcommand-level values, so the wrapper's
cmux-cua MCP config, hook config, and --disable computer_use placed
before the subcommand are dropped. The fake Codex now applies that rule,
and new cases cover exec, resume, mixed root and subcommand -c, and a
literal prompt after --.

* fix: keep codex wrapper overrides when the subcommand gets -c

Codex declares -c/--config, --enable, and --disable as clap global
arguments and keeps only the subcommand-level values when a subcommand
also receives one. The wrapper put its cmux-cua MCP config, hook config,
and --disable computer_use before the user's argv, so codex exec -c ...
or codex resume ... -c ... dropped all of them. The wrapper now moves the
user's subcommand-level global arguments, in order, in front of the
subcommand. Codex reads one root-level list with the same precedence as
before. Tokens after -- stay in place, and interactive launches without a
subcommand are unchanged.
…17256)

`cmux browser repl` is a persistent JavaScript REPL that agents use to drive
cmux browser panes: Playwright page, locator, keyboard and mouse semantics with
native trusted input, budgeted accessibility snapshots with diffs, tabs,
cookie-bearing fetch, a sandboxed fs, named sessions, an MCP server mode and
site tools. Guards live outside agent code: fill-only secrets with redaction
and capture masking, a domain policy over every frame and fetch hop, a per-tab
clipboard for session-created tabs, private per-session temp directories and
bounded cells, fetches and timers. Agent work never moves the user's focus,
hibernated tabs wake on use, and crashed tabs report how to recover.

Squashed from #15570 (392 commits; the
CLA action cannot read more than 250 commits of one pull request). Same tree as
that branch's head.
* Cloud: VM file operation routes (port of #16936), missing path answers 404

Ports the web part of #16936
(feat-cmux-next) to main: /api/vm/[id]/fs/[operation] (list, read, stat,
write, mkdir, remove) with the Freestyle driver and gateway methods.

Includes the fix from feat-cmux-next: Freestyle removes a missing path with
success, so removeVmFile stats first and answers a missing file with
404 vm_file_not_found (any other stat failure, a missing VM included, stays a
provider failure).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(cloud): stat/read/dir of a missing VM path must answer vm_file_not_found

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(cloud): stat/read/dir of a missing VM path answer 404 vm_file_not_found

The staging rehearsal of #17254 showed stat of a removed file answering 502
vm_cloud_service_unavailable. Map the guest ENOENT on every file read, as
remove already did, and title the error 'File not found'.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
lawrencecchen and others added 18 commits October 4, 2026 05:20
* test: account deletion must not call the retired legacy Subrouter

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Stop calling the legacy Subrouter during account deletion

The legacy Subrouter at subrouter.cmux.dev is being retired. Account
deletion now skips the legacy revoke phase and needs no legacy env vars.
Local mapping rows are still deleted, and a legacy_delete_pending
tombstone from an older deployment resumes at the hosted checkpoint.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Split the account DELETE handler under the complexity limit

DELETE delegates to deleteAccount and named phase helpers that share one
progress record, so its complexity drops from 46 to under 20 and its
grandfathered baseline entry is removed. Behavior is unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
* Cloud: private network routes (port of #16948), missing firewall rule answers 404

Ports the web part of #16948
(feat-cmux-next) to main: /api/vm/firewall (list, get, create, delete),
/api/vm/network and /api/vm/tunnel/network/[operation], with the Freestyle
driver, gateway and private-network workflow pieces.

Includes the fix from feat-cmux-next: a firewall rule that is not in the
caller's network answers 404 vm_firewall_rule_not_found (a provider 404 race
on delete too); a missing VM endpoint stays vm_not_found.

Stacked on the file-routes port (#17254).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* refactor(cloud): move firewall endpoint parsing into services/vms/firewallEndpoint

No behavior change; makes the parser testable without the route.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(cloud): firewall must accept normal CIDR prefixes and team-owned VM endpoints

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(cloud): firewall accepts normal CIDR prefixes and team-owned VM endpoints

The staging rehearsal of #17255 found two defects. validCidr compared the
prefix with net.isIP(), which returns the family (4 or 6), so any IPv4
prefix above /4 was refused; it now uses canonicalCidr. The vmId ownership
check looked the VM up in the personal scope, but every new VM is
team-owned, so vmId endpoints were vm_not_found; the route now resolves the
account scope when a vmId is named, and the VM must be the caller's own
(the firewall edits the caller's network). The provider now gets only the
rule fields: the Freestyle driver spreads its input into the request body,
so userId and provider were sent to Freestyle.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(cloud): firewall rules must name a caller resource as destination; get/delete must find VM rules

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(cloud): decide firewall rule ownership on the shared provider account

The Freestyle account is shared by every cmux user, so the API decides
whose a rule is. Reading and deleting: the rule names at least one resource
and every resource it names is the caller's. get and delete now read the
rule by id; they searched only the network listing, so a rule that named a
VM and a CIDR was created (201) and then could not be found or deleted. The
list merges the network listing with one listing per caller VM. Creating:
the destination must be a caller resource (400 vm_invalid_firewall_rule),
because a destination of only an address range or the public Internet would
reach other tenants' machines. Every firewall call resolves the account
scope like the other VM routes.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(cloud): firewall refuses unknown endpoint fields as owned and sends canonical CIDRs

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(cloud): unknown firewall endpoint fields are never owned; send canonical CIDRs

Security review P2: the provider adds selectors as new optional fields, and
an unknown one could name another tenant's resource, so a rule with an
unknown endpoint field is not the caller's. Review P3: send the canonical
range so a valid non-canonical CIDR does not fail at the provider.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(cloud): title vm_firewall_rule_not_found 'Firewall rule not found'

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(cloud): firewall needs a 100-rule cap, a bounded list, and a per-user rate limit

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(cloud): cap firewall rules at 100, bound the unfiltered list, rate-limit mutations

These routes are new on a provider account shared by every cmux user.
Create refuses at 100 owned rules (409 vm_firewall_rule_limit). An
unfiltered list reads the network plus at most the 10 newest live VMs, one
provider call each; older VMs list with ?vmId. Create and delete are
throttled per user with the Vercel firewall rule CMUX_VM_FIREWALL_RATE_LIMIT_ID
(no other VM mutation route has a limiter, so this follows the team-invite
limiter: fail closed when the firewall is unavailable, fail open and report
when the rule is unset or removed).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>
…cted one (#16897)

* remote-tmux: failing test for a torn-down stream feeding the next one

A control stream torn down for a reconnect keeps delivering what its reader
had already buffered. The test holds the main actor while a first client
writes 560 KB, starts a reconnect, and expects none of those bytes to reach
the connection.

* remote-tmux: stop a torn-down stream from feeding the next one

Cancelling the task that reads a control client's stdout does not empty the
reader's buffer, so chunks the old client had already written were still
ingested after the teardown. Once the reconnect had respawned, a leftover
command result was taken for the new client's attach reply. The connection
then never asked for windows and the mirror stayed blank for good.

Each read loop now stops as soon as its process generation is no longer the
current one, and closes its reader.

---------

Co-authored-by: ejc3 <ejc3@users.noreply.github.com>
…e out (#17237)

* remote-tmux: failing test for a docked terminal closed when its window's mirrors move

* remote-tmux: keep a window whose Dock has panels when its mirrors move out

---------

Co-authored-by: ejc3 <ejc3@users.noreply.github.com>
Custom sidebars could not read a workspace's task-status lane. cmux already
resolves one per workspace and the control socket can pin it, but the
interpreter data context carried no field for it, so a sidebar had no way to
group or colour rows by whether a workspace needs attention.

`workspaces[i].status` now carries the resolved lane as its raw wire value:
todo, working, needs-attention, review or done. The snapshot takes it as a
required parameter so a dropped wiring breaks the build rather than reporting
a silent "todo".


Claude-Session: https://claude.ai/code/session_0113SqtxGQwjHjzFw8mkSgwU

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
* test: prune read notifications of agents that died with the previous app

An agent alive at quit dies without SessionEnd, so its last "Completed in"
notification is restored on every launch and shown as the workspace's
latest sidebar summary even though no agent is running. The stale-agent
sweep must drop it once the pane has no agent again.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: drop restored agent notifications once the agent does not return

Notifications persist across relaunch so an unseen agent result is not
lost, but an agent that was alive when cmux quit dies without SessionEnd.
Nothing clears its notification afterwards: agent PIDs are not restored,
so the 30s stale-PID sweep has no dead PID to catch.

Restore now records the notifications of local panes that hosted an agent
(resume binding or restorable agent snapshot). The stale-agent sweep
removes the read ones when the pane has no agent PID again; unread ones
survive until read, and a pane the agent resumes into is handed back to
its hooks. Remote terminals are skipped since their agent can outlive the
app.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: keep restored notifications while the resume is in flight

Covers a read notification posted after restore (never tracked) and a
pane whose restored resume has not reported an agent PID yet.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: defer restored notification prune while the resume is in flight

The 30-second sweep can run before an auto-resumed agent reports its PID.
Skip panes whose restored command is still in flight, using the
coordinator's existing ownsInFlightRestoredCommand contract. Also look up
tracked notifications by id instead of scanning the store per panel, and
mark the value-only snapshot helper nonisolated.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: keep read non-agent notifications on a restored agent pane

The restore tracks every notification persisted on a pane that hosted an
agent, so a read `cmux notify` banner on that pane is pruned with the
agent's result.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: track only agent-produced notifications on restored agent panes

Use the persisted `isAgentEvent` provenance instead of panel ownership so
a `cmux notify` banner on a pane that hosted an agent is not retired with
the agent's result. Unknown provenance restores as agent-produced, matching
TerminalNotificationStore.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* browser: apply external-open rules to sidebar links

The sidebar's pull-request and port links (SwiftUI and AppKit rows, and
the open-all-pull-requests action) opened in the embedded browser whenever
that preference was on, without consulting the URL rules that route a site
to the system browser. Sites listed there cannot work in the embedded web
view at all, so a rule now wins over the embedded preference on those
paths, the same way it does for a click inside a page.

* browser: require a user event before a link escapes to the system browser

WebKit reports a script calling click() on an anchor as .linkActivated,
the same as a real click, so the external-open rules on their own let a
page hand itself a system-browser open at a moment of its choosing. The
navigation-typed escape now also requires an AppKit event in flight (a
key, left-mouse, or middle-mouse event) and never intercepts a download,
on every path that consults the rules: the main navigation delegate, the
target=_blank UI delegate, and both popup delegates. The context menu's
Open Link in New Tab is a gesture by construction and says so.

The event check is a bound rather than a proof: NSApp.currentEvent says
an event is being dispatched, not that this navigation is the thing the
user asked for. Middle-clicks arrive as otherMouse events and count.

* browser: e2e coverage for external-open link routing

BrowserExternalOpenRoutingUITests drives real WebKit link activations
through the socket browser.click against a local fixture and asserts
routing at the delegate layer, where popup-vs-link-activation behavior
actually diverges and unit tests cannot reach. Escapes are captured to a
file through the existing DEBUG-only UI-test sink
(CMUX_UI_TEST_CAPTURE_EXTERNAL_OPEN_PATH) from the external-navigation
handler's default opener, so CI never opens Safari. Four cases: a matched
link click escapes while the embedded page stays put; an unmatched click
navigates embedded; a scripted window.open to a matched host never
escapes; a target=_blank form POST to a matched host stays embedded.

The shared BrowserFixtureSocketTestCase gains subclass hooks for launch
arguments and environment, falls back from the in-process socket client
to nc -U and then the bundled cmux CLI, disables hidden-webview discarding
for the backgrounded UI-test host, and polls browser.wait through the
cold-start content-process transient.

* browser: click links for real in the external-open UI tests

A link now leaves for the system browser only while a real input event
is in flight, and the socket browser.click runs JavaScript, so the matched
and unmatched link cases click through accessibility the way a person
does. The scripted popup and form cases keep the socket click, since a
scripted action is what they test.

* browser: keep only the sidebar links, drop the in-page activation change

The in-page half of the external-open rules has landed separately. What remains here is the sidebar: pull-request and port links follow the rules, with tests for the matcher.

* browser: use a rule the pattern safety check accepts in the port-link test

---------

Co-authored-by: ejc3 <ejc3@users.noreply.github.com>
Co-authored-by: Leo Li <cheerleaderleo@outlook.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* test: cover merging repos without aggregate CI workflow

* fix: make merge checks conditional on base workflows
* ci: use nearest ancestor for main-fix evidence

* ci: constrain ancestor evidence to path-filtered changes

* ci: inspect renamed paths in ancestor evidence

* ci: bound ancestor evidence traversal

* ci: reject incomplete ancestor path comparisons
* test(session): cover replay sweep and permissions

Signed-off-by: Alejandro Florez <soyeladice@gmail.com>

* fix(session): sweep stale scrollback replay files

Signed-off-by: Alejandro Florez <soyeladice@gmail.com>

* fix(session): sweep stale replay files before restore

Signed-off-by: Alejandro Florez <soyeladice@gmail.com>

* fix(session): remove synchronous replay sweep from app init

Signed-off-by: Alejandro Florez <soyeladice@gmail.com>

* fix(session): gate restore on off-main replay cleanup

Signed-off-by: Alejandro Florez <soyeladice@gmail.com>

* fix(app): keep replay sweep off startup critical path

Signed-off-by: Alejandro Florez <soyeladice@gmail.com>

* fix(session): skip replay sweep under XCTest

* fix(session): harden retained replay files during stale sweep

* docs(session): keep crash-recovery gate semantics accurate

* test(session): cover legacy replay permissions and sweep filters

---------

Signed-off-by: Alejandro Florez <soyeladice@gmail.com>
Co-authored-by: Leo Li <cheerleaderleo@outlook.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* docs: clarify Claude project dir decode asymmetry

* test: cover dotted Claude project dir in session directory scope

* test: scope Claude session roots per task instead of process env

Swift Testing runs suites in parallel, so setting CLAUDE_CONFIG_DIR process-wide could leak into other tests. A DEBUG-only TaskLocal override keeps the fixture root local to the test's task tree.

* test: cover Claude cwd-filter lookup without a DEBUG seam

Widen the Claude candidate enumerator and its two types to internal so the test reaches them through @testable import, per the no-test-debug-seam review rule. The test checks .claude/worktrees and .worktrees cwds and that dot-preserving and other project dirs are excluded.

---------

Co-authored-by: Leo Li <cheerleaderleo@outlook.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
* Answer the tmux session commands Claude Code calls

Claude Code's tmux backend tears down and reattaches its agent panes with
kill-session, switch-client, new-session -A and show-options -g prefix. The
compatibility layer rejected all four, so a Teams session failed with
"Unsupported tmux compatibility command" once it got past creating panes.

A tmux session is a cmux workspace, which has-session and new-session already
assume, so kill-session closes that workspace, switch-client selects it, and
new-session -A attaches to it when it exists instead of creating a duplicate.
show-options now answers from a table, and prefix reports the C-b that a
default tmux client would.

The sequence test drives all four through the real shim. Its fake socket also
now unwraps the capability envelope that the shell integration adds inside a
cmux terminal, so the test reports the behavior it checks rather than a JSON
decode error; that unwrapping moved into the shared helper.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Reject a socket payload that is not a request

The fake servers indexed request["method"] straight off json.loads, so a
payload that decoded to null, a list, or an object with a non-string method
raised inside the handler thread and surfaced as an unrelated CLI error. They
now answer those with an error line, which names the real problem.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test: reject empty fake socket request methods

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: reject a failed tmux session lookup and kill-session -a

A workspace.list error must not become a new session, and kill-session -a
must not close the caller.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Keep a failed tmux session lookup from creating a workspace

new-session -A treated every resolution error as a missing session, and
kill-session ignored -a and closed the caller. A missing session still
creates; a lookup failure and an unsupported flag now fail first.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Leo Li <cheerleaderleo@outlook.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
…7943) (#7974)

* Add failing test: sidebar nil-comparison yields nothing (#7943)

In a custom sidebar, `x != nil` / `x == nil` against a bound optional field
does not evaluate to true/false — it evaluates to nothing. Interpolation
renders empty, ternaries always take the else branch, and `if x != nil`
guards are never taken, so optional-guarded views never render.

This commit adds only the regression test (no fix) so CI shows it red.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* Fix sidebar nil-comparison to evaluate to a Bool (#7943)

The interpreter's value model had no null case and no `nil`-literal
evaluator branch, so `nil` evaluated to a host `SwiftValue?` of `nil` — the
same value that means "expression unsupported / no value". `evalInfix` then
bailed on the comparison, so `x != nil` / `x == nil` produced nothing:
interpolation rendered empty, ternaries always took the else branch, and
`if x != nil` guards were never taken. Optional-guarded views (including the
shipped status-board.swift / finder.swift examples) silently drew nothing.

- Add `SwiftValue.null` for the `nil` literal and for comparing an absent
  optional field against `nil` (distinct from host `nil` = "no value").
- Evaluate `NilLiteralExprSyntax` to `.null`.
- Handle `==` / `!=` before the operand guards, coalescing an absent operand
  (host `nil`) and the `nil` literal to `.null`, so the comparison yields a
  Bool that is true/false when present and false/true when absent.

Fixes #7943

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: preserve nil comparison evaluation failures

* fix: preserve nested nil comparison misses

* fix: preserve parenthesized nil comparison misses

* fix: handle nil optional binding and equality budget

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Leo Li <cheerleaderleo@outlook.com>
* test: cover neutral checks and helper checkout refresh

* test: tolerate absent git diagnostics

* fix: refresh clean main checkout before merging
Merge-main commit by scripts/merge-main.sh.
Merged by scripts/merge-main.sh: origin/main at 186cec7.

Resolved conflicts:
- Resources/Localizable.xcstrings: xcstrings key-level union
- cmux.xcodeproj/project.pbxproj: union of added entries, then normalize-pbxproj.py

Merge-main-previous-head: 3b26270
Merge-main-base: 186cec7
@austinywang

austinywang commented Oct 5, 2026 •

Copy link
Copy Markdown
Contributor

Fixed the Cloud welcome Cmd+W routing failure and added a behavioral regression for the shared auxiliary-window close decision. The regression is in 9186e34dcff; the fix is in a099d3ca958.

Final pushed head: 1aaf63492db33f6747c515f056021e3046546614. Merged main 186cec797812 through scripts/merge-main.sh, then merged the updated stacked base 5622f3656f5. Preserved main’s complete translations when resolving the overlapping catalog entries.

Verification:

  • CI and iOS tests passed. Current head has 61 successful checks, 19 skipped checks, and one neutral review check; no failing or pending checks.
  • App-host test log confirms all three Cloud welcome tests passed, including welcomeOwnsCloseShortcut. Swift Testing reports 5,368 tests passing; XCTest reports 2,210 tests, 26 skipped, zero failures.
  • python3 scripts/verify-local.py: all 9 selected checks passed. The auxiliary-window lint failed before the fix and passed afterward; its regression script and the package-conventions check also passed. The new Swift behavioral regression was executed only on the fixed head in hosted CI, not demonstrated red locally.
  • Localization audit: ./scripts/localize-changes --base a099d3ca958 passed; 10 catalogs, 9 required locales, zero parity errors. No new copy was introduced by the repair.

No review threads or remaining actionable comments were present on the final check. No new GUI dogfood was performed; skipped checks are not claimed as coverage. The PR remains open and unmerged.

@austinywang
austinywang merged commit 4588838 into cloud-gate-upgrade Oct 5, 2026
81 checks passed
@austinywang
austinywang deleted the cloud-onboarding branch October 5, 2026 01:43
austinywang added a commit that referenced this pull request Oct 5, 2026
…visible (#17127)

* cloud: enable cloud for pro, upgrade for free, keep the button when the panel rebuilds

the cloud tab and settings > cloud now show enable cloud only when the plan
includes cloud and upgrade for free accounts. the plan answer lives on the
account flow (per account) instead of the panel's view state, so switching
sidebar modes no longer drops the button back to "checking your cmux plan…".

* cloud: stop waiting on a slow plan check, title enable cloud machines, debug plan override

* cloud: cloud.fill for the pro-required screen, settings shows only the enable row until cloud is on, enable cloud machines title

* cloud: drop the debug plan override

* cloud: keep prominent cloud buttons visible in an inactive window, keep a known plan through failed or cancelled checks, recheck settings on account change

* cloud gate: newest plan check wins, drop answers for a switched account

* fix cloud billing plan concurrency and state ownership

* fix cloud plan localization parity

* fix cloud entitlement detection for team plans

* avoid stale cloud entitlements across team changes

* report billing refresh success to settings

* preserve account flow billing refresh contract

* cloud onboarding: introducing cmux cloud welcome, and cloud enablement screens (#17230)

* cloud welcome: introducing cmux cloud window, shown once, five layouts to compare from help (wip)

* cloud onboarding: glass welcome window, lowercase and machine focus layouts, sidebar intro with banner and reasons, enablement view takes plain values (wip)

* cloud onboarding: one welcome design (machine focus, all lowercase), sidebar intro titles by plan, 6pt buttons, drop the comparison layouts

* cloud onboarding: review fixes (glass behind compiler guard, welcome considered once per launch, size after hosting, no return shortcut, comments, orphaned string)

* cloud welcome: ignore the titlebar safe area (fixes a layout-loop crash on open), three reasons

* cloud tab intro: title first, no icon

* cloud tab intro: app icon banner back, lock badge while the plan needs pro

* cloud tab intro: dark app icon in the banner

* fix(cloud): wait for display helper readiness (#17132)

* fix(cloud): explain unavailable display actions

Show the existing Cloud failure message when New Display is unavailable, and explain the machine ownership restriction when a display is opened into another Cloud workspace.

— unregistered

* fix(cloud): explain unavailable display actions

Show the existing Cloud failure message when New Display is unavailable, and explain the machine ownership restriction when a display is opened into another Cloud workspace.

— unregistered

* fix(cloud): reject cross-machine display opens before projection

* fix(cloud): probe legacy desktop VMs for displays

* fix(cloud): initialize wallpapers for new displays

* fix(cloud): fail closed before display double-click opens

* fix(cloud): replay repeated display ownership hints

* test(cloud): cover display helper refresh

* fix(cloud): refresh installed display helper

* test(cloud): require display service readiness probe

* fix(cloud): wait for display helper readiness

* fix(cloud): keep display creation retryable

* fix(cloud): authorize dynamic display ports

* fix(cloud): let new display retry guest discovery

* Revert "fix(cloud): let new display retry guest discovery"

This reverts commit 0fe008ed3f6fd78338be9e9a2f8161d37984cc0f.

* test(cloud): sidebar visibility filter must keep display creation state

Moves applyingDeviceVisibility next to SurfaceCatalogSnapshot in
CmuxSurfaceCatalogModel (no behavior change) so it is testable, and adds
a failing regression test.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(cloud): keep display creation state through the sidebar filter

applyingDeviceVisibility rebuilt SurfaceCatalogSnapshot from scratch and
dropped displayCreationMachines, staleMachineIDs and display memberships,
so every desktop VM's New Display row reported additional displays as
unavailable. Filter a copy instead so all per-machine state survives.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(surfaces): a split with no room opens as a tab in the target pane

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(cloud): optimistic display row; fall back to a tab when a split has no room

New Display now shows a "Starting display…" row from the click until the
guest answers, driven by the catalog's in-flight creation set.

Opening a sidebar resource splits the focused pane; once split admission
had no room left, the open failed with "Could not create the pane:
noSpace". SurfacePaneFactory now opens it as a tab in the pane that would
have been split.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(surfaces): refused split is typed; sidebar gestures open a tab

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(surfaces): scope the no-room tab fallback to sidebar gestures

The factory now throws a typed noSpace instead of opening a tab, so the
layout replay and socket open verbs keep a truthful refused split. Sidebar
opens and New Display use openPreferringSplit, which retries as a tab in
the pane that would have been split. The empty Displays row no longer
shows beside the optimistic Starting display row.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(cloud): display creation needs no discovery round trip

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* perf(cloud): open the display pane at the click and skip pre-create discovery

New Display paid two VM exec round trips (list, then create, about 2s
each) before any pane appeared. The create reply is already the full guest
catalog, so the list is dropped. The pane now opens at the click with a
native Starting display state and is adopted by the projection once the
guest assigns the display; creation failure closes it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(cloud): harden the reserved display pane

Drop a second click before it opens a pane; bind the reservation to the
created display id so no other projection adopts it; keep creating when
the pane cannot open; leave the display in the pool when the person closed
the pane; show the failure on a pane the socket refuses to close.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(cloud): additional displays serve noVNC beside the primary desktop

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(cloud): make additional displays reachable on the VM private address

Displays 2+ ran websockify on 127.0.0.1 while display 1 listens on [::].
The client reaches every display through the private address, so each new
display's first noVNC connection was refused and only recovered after a
timeout. Bind like display 1 (Xvnc stays -localhost), replace proxies an
older helper left on loopback, and drop the open-port call added for
display ports: it changed no routing and cost a control-plane request,
plus a desktop heal exec for display 1, on every open.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(cloud): stale proxy match covers only this display's websockify

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(cloud): match stale display proxies by argv, not pgrep regex

pgrep's ERE has no (?:...) groups, so the stale websockify match failed,
the old loopback listener kept the port, and the rebound proxy exited.
Read /proc argv for this user's websockify on this display's ports.
Verified on a live VM: displays 2 and 3 now accept the noVNC websocket on
the private address, one proxy each, X sessions untouched.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(cloud): a failed first desktop connection retries before failing

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(cloud): retry a display's first connection; warm the carrier at create

noVNC stops at Connect after an initial connect failure; its reconnect only
follows a session that once connected. Restored and new displays could lose
that race while the proxy or guest listener started, leaving "Failed to
connect to server". Retry the route's first connection three times
(0.5s/1s/2s on the injected clock) before showing the failure card.

The first display open on a machine spent 12-18s starting its browser
carrier after the guest exec. Start it alongside the create request.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(cloud): an explicit Retry gets its own quiet first-connection retries

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(cloud): run the guest display script test on the main actor

CloudGuestDisplayScript is @MainActor; the test called it from a
nonisolated context and broke the CmuxCloud package test build.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ci): per-runner build roots only on fleet Macs without glaeda (#17239)

* test(ci): Blacksmith keeps the shared root; reruns keep a per-runner root (red)

Blacksmith macOS runners report RUNNER_ENVIRONMENT=self-hosted, so they got
per-runner roots: every build started cold and no seed could be adopted
(job 111334766861). take-product-canonical-root.sh also refused a product
built at a per-runner root on a fleet Mac without glaeda.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ci): per-runner roots only on fleet Macs without glaeda

canonical-build-root.sh derived /private/tmp/cmux-ci-<runner> for every
self-hosted runner without the glaeda helper, which included Blacksmith's
ephemeral macOS runners: their builds started cold and their seeds were
never adopted. Derive it only when the fleet directory exists. Let
take-product-canonical-root.sh accept such a root when glaeda isn't there
to hold it, so app-host test reruns build where the product was compiled.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ci): repair main's guards, localization parity and cmuxTests compile (#17207)

* fix(ci): align owned-build-state guards with per-runner canonical roots

#17168 (496195b) derives a canonical root per self-hosted runner, exports
CMUX_CI_CANONICAL_ROOT from the build-slot step and drops the shared
/private/tmp/cmux-ci fallback in test-e2e's owned-state step. Three tests in
tests/test_ci_owned_build_state.py still assumed the old contract, turning
"CI fast guards" red on main.

Update them to the new contract: the slot step exports the root, a
per-runner root reads its own store, and a missing root reads nothing. The
owned-state step now fails with a clear message when the root is missing, and
rejects a root nested under a cmux-ci-* prefix (e.g. cmux-ci-a/../x), so the
looser glob cannot map the store outside the Mac's package directory.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(l10n): translate the New Machine plan loading strings

#17135 added machines.new.plan.loading, .retry and .error with the English
text copied into every locale, so `localization_catalog.py check` reports 24
parity errors and "Fast static checks" fails on every PR. Translate them for
all catalog locales; Retry reuses common.retry's wording.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(l10n): match the catalog's plan terminology in de, ko and zh-Hant

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(tests): repair cmuxTests compile after the sidebar reorder change

#17070 moved RightSidebarModeBarDragLayout into the CmuxSidebar package,
but RightSidebarTabCustomizationTests still relied on the app module
exporting it, and CloudMachineOrderingTests put `try` calls inside an `&&`
in #expect, which the macro rejects ("operator can throw but expression is
not marked with 'try'"). Import CmuxSidebar and hoist the throwing lookups.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(tests): hoist the second throwing lookup out of #expect in CloudMachineOrderingTests

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ci): satisfy dogfood-build runner guards and the Swift warning budget

#17206 turned dogfood-build into a macOS build, which tripped three guards:
no pinned Xcode, no fork branch in runs-on, and no static-preflight
dependency. Add `scripts/select-ci-xcode.sh`, the standard owner and fork-PR
branches ahead of the existing selector, and `needs: static-preflight`. The
job's `if` already excludes forks, so manaflow-ai PRs keep the same runner.

#17070's `RightSidebarModeBarDragController.coordinateSpace` is read from
a Sendable geometry closure and broke the zero-warning budget; it is a plain
String constant, so mark it `nonisolated`.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(ci): a root the glaeda hook exported wins over the per-runner root

Fails on main: canonical-build-root.sh treats an exported /private/tmp/cmux-ci
as unset and derives /private/tmp/cmux-ci-<runner>, so main-compile-probe
refuses the root glaeda placed it in (runs 37157423969, 37155786981).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ci): honor the canonical root the glaeda hook exported

#17168 (496195b) derives /private/tmp/cmux-ci-<runner> on every self-hosted
runner unless CMUX_CI_CANONICAL_ROOT names a non-default root. glaeda's
runner hook holds root 1 (/private/tmp/cmux-ci) for a compile job and
exports exactly that, so the job then built somewhere glaeda does not hold,
and main-compile-probe refused it: "glaeda placed this job in
/private/tmp/cmux-ci, not /private/tmp/cmux-ci-cmux13s-mac-mini-glaeda".

Let any exported root win. A self-hosted job with no exported root still
gets its per-runner root, so #17168's isolation for unplaced runners stays.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(ci): an unplaced self-hosted job keeps the shared seeded root

Fails on the branch: canonical-build-root.sh derives /private/tmp/cmux-ci-<runner>,
whose fingerprint never matches main's seeds, so new aws runners
(aws-m4pro-9-glaeda-3, aws-m4pro-8-glaeda-4) compiled cold and timed out at
35 minutes in compile admission (run 37158561950, both attempts).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ci): build unplaced self-hosted jobs at the shared seeded root

#17168 gave every self-hosted job without an exported root its own
/private/tmp/cmux-ci-<runner>. That root is part of the cache fingerprint,
so main's DerivedData seeds and the owned build state never match it, and the
prepare step clears it each job: every new aws runner compiled cold and hit
the 35-minute admission limit (aws-m4pro-9-glaeda-3, aws-m4pro-8-glaeda-4 on
run 37158561950).

Go back to /private/tmp/cmux-ci when no root is exported. glaeda's hook still
exports root N for the jobs it places, so those stay isolated. Unplaced
concurrent runners on one Mac share the root again, as before #17168, until
glaeda places those jobs too.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(ci): the pool picker never picks an owned label it was not configured with

Fails on main: simple_pool_picker adds every glaeda-<class>-xcode-* label it
sees on a runner, and glaeda-std-xcode-26.3 (ten aws EC2 runners, five per
Mac) sorts before the minis' 26.6, so compile admission ran there and timed
out at 35 minutes while the minis were idle.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ci): pick only configured owned pools, not any owned-looking label

simple_pool_picker added every glaeda-<class>-xcode-* label it saw on an
online runner to the configured CI_OWNED_POOL_SLOTS pools, then tried them
in string order. The aws EC2 Macs carry glaeda-std-xcode-26.3 (ten runners,
five per Mac), which sorts before the minis' glaeda-std-xcode-26.6, so PR
compile admission landed on contended EC2 runners and timed out at 35
minutes while the minis sat idle. Use only the configured pools.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(ci): the pool picker reads every page of organization runners

Fails on main: LiveState.runners reads one page of 100, but manaflow-ai has
509 runners and the first page holds the aws Macs and no idle mini, so the
picker never saw the mini fleet (run 37166910798 fell back to Blacksmith with
29 idle minis online).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ci): read every page of runners in the pool picker

LiveState.runners read only the first page of 100 organization runners.
manaflow-ai has about 500, and the first page holds the aws Macs but no idle
mini, so the picker never saw the mini fleet: it took the aws 26.3 label
while that was discoverable, and fell back to Blacksmith once it was not.
Read up to ten pages. Against live data the picker now sees 38 online
glaeda-std-xcode-26.6 minis (33 free) and picks them.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: clear the Cloud sidebar warnings that broke the Swift warning budget

Compile admission on a mini (run 37167089830) built cleanly but failed the
zero-warning budget on three warnings from the Cloud sidebar work:
`SurfaceCatalog.shared` used as a default argument of two @MainActor
CloudWorkspaceSidebarPresentation entry points (evaluated nonisolated), and an
implicit `self` in CloudTreeOutlineView's machine-lift reopen closure.
Resolve `.shared` inside the main-actor body and make the capture explicit.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ci): preserve canonical root and translation wording

* test(ci): align canonical root guard with main behavior

* test: use isolated catalogs in Cloud sidebar fixtures

Pass each test fixture catalog through the sidebar snapshot factory and direct presentation assertion so Cloud machine metadata is read from the catalog the test populated.

Co-authored-by: Austin Wang <austinwang115@gmail.com>

* test: model loading Cloud sidebar state

Keep device projections visible while their catalog rows are restoring, and assert that a reserved loading card shows machine identity before its directory appears after adoption.

Co-authored-by: Austin Wang <austinwang115@gmail.com>

* test: use a terminal panel for cloud directory settings

The sidebar detail test supplies a reported directory, so give it a real terminal panel instead of a loading card that correctly suppresses directory metadata.

Co-authored-by: Austin Wang <austinwang115@gmail.com>

* fix(cloud): restore device directory fallback lost in main merge

The merge of main took #17107's presentation file and dropped the
fallback from 27a8d804fdf, so a device projection whose catalog row is
still restoring rendered no directory and
SidebarCloudWorkspaceBadgeTests.deviceNameIsVisibleBesideItsDirectory
failed on b6ea6632cb1 (the regression run is that CI failure).

Co-authored-by: Austin Wang <austinwang115@gmail.com>
Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Austin Wang <austinwang115@gmail.com>

* ci: require a written merge override for red CI (#17217)

* ci: require merge-gate before merges

* Add API-only merge gate for pull requests

* test: cover merge-gate override decisions

* ci: wire merge-gate tests and push timing

* ci: pin merge gate workflow source

* Reject read-only override authors

* ci: harden merge gate runner and author checks

* ci: keep merge gate on hosted capacity

* ci: fail closed when collaborator lookup fails

* ci: reject stale successes during reruns

* ci: fail closed for untimestamped queued checks

* ci: fail closed for untimestamped pending statuses

* ci: verify matching override failure evidence

* test: cover mismatched override failure evidence

* ci: choose newest merge gate check run

* ci: require compile evidence for main-fix merges (#16993)

* test: reproduce main-fix merging without compile evidence

The installed helper bypasses all checks under --main-fix. The regression records that it merges with no compile checks present.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* ci: require build evidence for main fixes

gh-merge-green --main-fix now requires successful Release, Debug and Swift test target build steps on the exact PR head. Existing Swift test failures are waivable only when their parsed issue records match the same Swift test step on the exact base SHA; the audit comment records each match and unrelated failures remain fatal.

CPU, memory and disk checks: the validator uses one bounded 90-second GitHub request per call, caps captured output at 32 MiB, writes audit bodies to temporary files, and does not retain logs, caches or stores.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* ci: harden merge gate freshness and evidence

* ci: require per-check override reasons

* ci: complete merge gate review fixes

* docs(ci): document merge-gate rollout order

* ci: make merge helper rollout repairable

* ci: point merge helper refusals to repair runbook

* ci: resolve merge helper symlink for main fixes

* docs: remove stale main-fix PR reference

* fix: avoid pipefail false negatives in merge helper

* fix: keep merge gate events fresh per pull request

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* Share one vCPU and memory pool across a plan's Cloud VMs (#17238)

* test(vms): reproduce missing shared Cloud VM resource pool

* feat(vms): share one vCPU and memory pool across a plan's Cloud VMs

Pro, Team (per paid seat), and Founder's Edition get up to 5 active VMs
sharing 20 vCPUs and 40 GB RAM, with machines up to the 32 GB xl row.
Max gets 80 vCPUs and 160 GB RAM with machines up to the 64 GB 2xl row.
The repository enforces the pool next to the active-VM count, under the
same transaction and billing lock, on create, Base open/reset, paused
resume, CPU/memory resize, and fork. Pricing, docs, app, and iOS copy
now describe pooled resources.

* Pooled VMs: Pro overrides stop at 32 GB; drop unused iOS pool strings

* test(vms): reproduce leaked compute claim and stale-plan resume pool

* fix(vms): give back unused compute claims and resume against the caller's pool

* docs(pricing): describe the Team pool per paid seat and localize pool strings in every locale

* Cloud VM: snapshot create honors Idempotency-Key (#17244)

* test(cloud-vm): snapshot create dedups by idempotency key (red)

A retry of POST /api/vm/:id/snapshot with the same Idempotency-Key must return
the first snapshot and take no second one; a retry while the first runs is
refused as in progress; the same key with another name is a conflict; a failed
attempt frees the key.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(cloud-vm): snapshot create dedups by Idempotency-Key

POST /api/vm/:id/snapshot now reads Idempotency-Key. A new ledger table,
cloud_vm_snapshot_requests (one row per machine and key, additive migration
20261004120000), records a pending attempt before the provider call and the
provider snapshot after it. A retry with the same key returns the first
snapshot and records no second usage event; a retry while the first runs gets
409 vm_snapshot_in_progress (retryable); the same key with another name gets
409 vm_snapshot_idempotency_conflict; a failed attempt frees the key; a pending
row older than 15 minutes (route budget 600 s) is taken over by a retry.
Requests without a key behave as before.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* chore(cloud-vm): smoke --snapshot-check proves snapshot idempotency

With --create, takes one snapshot of the throwaway smoke machine twice with the
same Idempotency-Key, requires the same snapshotId, deletes that snapshot,
then destroys the machine as before. No fixed VM id and no printed token: the
smoke mints its own throwaway user session.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Preserve selected tab when closing another surface (#16645)

* Preserve selected tab when closing another surface

* Only select the closing pane's tab when that pane is focused

BonsplitController.selectTab also focuses the pane, so calling it for an
unfocused pane moved focus into the pane where a tab closed, the focus theft
this change is meant to stop. Bonsplit already keeps the selection when an
unselected tab closes; the shouldCloseTab change is the actual fix.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* docs: detail tmux help options (#16780)

Co-authored-by: Leo Li <cheerleaderleo@outlook.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* test: guard managed contributor difficulty labels (#16448)

* test: guard managed contributor difficulty labels

* test: reject invalid difficulty descriptions

* test: exercise difficulty label sync requests

---------

Co-authored-by: Leo Li <cheerleaderleo@outlook.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* fix(cloud): repair main's Cloud app-host failures from #17132 and #17103 (#17250)

* test(cloud): expect the display helper readiness loop

#17132 replaced the one-line `list || exit 1` probe with a bounded
retry loop and updated the package test, but cmuxTests still asserted
the old line, so CloudDisplayCatalogTests.guestCommandShape failed on
main (run 37186487936, shard 4/7).

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(cloud): make the restored-display retry test a desktop

#17132's initialDesktopFailureRetries configured port 6902 without a
resource ID, so CloudBrowserAccessState.isDesktop was false and
desktopConnectionDidChange ignored the failure: no quiet retry, and the
following navigations[1] trapped and crashed the shard 5 app host on
main (run 37186487936). Give it the .display identity additional
displays carry and require the retry before indexing it.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test(cloud): a foreign display click shows one synchronous hint

#17103 made a display click on another machine's workspace reject
synchronously through showDisplayOpenHint, so no tree operation starts
and waitForOpen() waited out the 60 s suite limit on main (run
37186487936, shards 2 and 3). Assert no operation ran and that the
rejection is shown once; this stays red until the double-click stops
repeating the hint.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(cloud): show a foreign display's ownership hint once per double-click

AppKit delivers a double-click's first click to handleSingleClick,
which already shows the ownership hint, and #17103 also showed it from
handleDoubleClick, so the user got the same rejection twice.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix(cloud): never show upstream tree failures in the Machines status row

#17103 started presenting treeError verbatim so its ownership hints
would survive, but treeError also carries raw upstream failures
(LocalizedError descriptions, create output), so a URL with query
parameters reached the toolbar text, hover help and copy menu.
MachinesCloudStatusTests.emptyStatusHasNoProgressPresentation caught it
on main (run 37186487936, shard 7).

Hints now travel through their own onHint sink (falling back to
onFailure for other callers), the panel records them as trusted copy,
and the status row shows the tree error verbatim only when it is that
hint; everything else gets the safe recovery message.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* fix(ci): keep merge-gate diagnostics on the exact PR (#17248)

* fix: keep merge-gate diagnostics on the exact PR

* fix: print repair guidance for gate diagnostics

* fix: allow merge-gate PR comments when permitted

* fix: reject mismatched merge-gate event identities

* fix: keep merge-gate alive on comment errors

* fix: keep the cmux-cua credential out of the Codex argv (#17252)

* test: codex wrapper must not put the cmux-cua credential in argv

The Codex wrapper passes the cmux-cua socket credential as
-c mcp_servers.cmux-cua.env.CMUX_CUA_SOCKET_AUTH_TOKEN=<value>, so any
local user can read it with ps. These assertions require the value to be
absent from argv, Codex to receive it in its environment, and the MCP
server to receive it through env_vars. The fake Codex now builds the MCP
server environment the way Codex does (allow-list + env_vars + env).

* fix: keep the cmux-cua credential out of the Codex argv

The wrapper passed the socket credential as a Codex -c env override, which
put the value in the codex process argv where any local user can read it
with ps. The wrapper now exports CMUX_CUA_SOCKET_AUTH_TOKEN in the parent
shell before exec and emits mcp_servers.cmux-cua.env_vars so Codex forwards
the variable by name to the MCP server (Codex env_vars, openai/codex#5246).
Codex's default shell_environment_policy excludes *TOKEN* names from agent
shell commands. Attachment stays fail-closed when no credential resolves.

* fix: keep codex wrapper overrides when the subcommand gets -c (#17257)

* test: codex wrapper overrides must survive subcommand -c flags

Codex declares -c/--config, --enable, and --disable as clap global
arguments. When a subcommand such as exec or resume also receives one,
Codex 0.159.3 keeps only the subcommand-level values, so the wrapper's
cmux-cua MCP config, hook config, and --disable computer_use placed
before the subcommand are dropped. The fake Codex now applies that rule,
and new cases cover exec, resume, mixed root and subcommand -c, and a
literal prompt after --.

* fix: keep codex wrapper overrides when the subcommand gets -c

Codex declares -c/--config, --enable, and --disable as clap global
arguments and keeps only the subcommand-level values when a subcommand
also receives one. The wrapper put its cmux-cua MCP config, hook config,
and --disable computer_use before the user's argv, so codex exec -c ...
or codex resume ... -c ... dropped all of them. The wrapper now moves the
user's subcommand-level global arguments, in order, in front of the
subcommand. Codex reads one root-level list with the same precedence as
before. Tokens after -- stay in place, and interactive launches without a
subcommand are unchanged.

* Add cmux browser repl: a Playwright-shaped browser REPL for agents (#17256)

`cmux browser repl` is a persistent JavaScript REPL that agents use to drive
cmux browser panes: Playwright page, locator, keyboard and mouse semantics with
native trusted input, budgeted accessibility snapshots with diffs, tabs,
cookie-bearing fetch, a sandboxed fs, named sessions, an MCP server mode and
site tools. Guards live outside agent code: fill-only secrets with redaction
and capture masking, a domain policy over every frame and fetch hop, a per-tab
clipboard for session-created tabs, private per-session temp directories and
bounded cells, fetches and timers. Agent work never moves the user's focus,
hibernated tabs wake on use, and crashed tabs report how to recover.

Squashed from https://github.com/manaflow-ai/cmux/pull/15570 (392 commits; the
CLA action cannot read more than 250 commits of one pull request). Same tree as
that branch's head.

* Cloud: VM file operation routes (port of #16936) (#17254)

* Cloud: VM file operation routes (port of #16936), missing path answers 404

Ports the web part of https://github.com/manaflow-ai/cmux/pull/16936
(feat-cmux-next) to main: /api/vm/[id]/fs/[operation] (list, read, stat,
write, mkdir, remove) with the Freestyle driver and gateway methods.

Includes the fix from feat-cmux-next: Freestyle removes a missing path with
success, so removeVmFile stats first and answers a missing file with
404 vm_file_not_found (any other stat failure, a missing VM included, stays a
provider failure).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(cloud): stat/read/dir of a missing VM path must answer vm_file_not_found

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(cloud): stat/read/dir of a missing VM path answer 404 vm_file_not_found

The staging rehearsal of #17254 showed stat of a removed file answering 502
vm_cloud_service_unavailable. Map the guest ENOENT on every file read, as
remove already did, and title the error 'File not found'.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Stop calling the legacy Subrouter during account deletion (#17273)

* test: account deletion must not call the retired legacy Subrouter

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Stop calling the legacy Subrouter during account deletion

The legacy Subrouter at subrouter.cmux.dev is being retired. Account
deletion now skips the legacy revoke phase and needs no legacy env vars.
Local mapping rows are still deleted, and a legacy_delete_pending
tombstone from an older deployment resumes at the hosted checkpoint.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* Split the account DELETE handler under the complexity limit

DELETE delegates to deleteAccount and named phase helpers that share one
progress record, so its complexity drops from 46 to under 20 and its
grandfathered baseline entry is removed. Behavior is unchanged.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* ci: remove merge gate and restore exact-head merging (#17275)

* Cloud: private network routes (port of #16948) (#17255)

* Cloud: private network routes (port of #16948), missing firewall rule answers 404

Ports the web part of https://github.com/manaflow-ai/cmux/pull/16948
(feat-cmux-next) to main: /api/vm/firewall (list, get, create, delete),
/api/vm/network and /api/vm/tunnel/network/[operation], with the Freestyle
driver, gateway and private-network workflow pieces.

Includes the fix from feat-cmux-next: a firewall rule that is not in the
caller's network answers 404 vm_firewall_rule_not_found (a provider 404 race
on delete too); a missing VM endpoint stays vm_not_found.

Stacked on the file-routes port (#17254).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* refactor(cloud): move firewall endpoint parsing into services/vms/firewallEndpoint

No behavior change; makes the parser testable without the route.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(cloud): firewall must accept normal CIDR prefixes and team-owned VM endpoints

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(cloud): firewall accepts normal CIDR prefixes and team-owned VM endpoints

The staging rehearsal of #17255 found two defects. validCidr compared the
prefix with net.isIP(), which returns the family (4 or 6), so any IPv4
prefix above /4 was refused; it now uses canonicalCidr. The vmId ownership
check looked the VM up in the personal scope, but every new VM is
team-owned, so vmId endpoints were vm_not_found; the route now resolves the
account scope when a vmId is named, and the VM must be the caller's own
(the firewall edits the caller's network). The provider now gets only the
rule fields: the Freestyle driver spreads its input into the request body,
so userId and provider were sent to Freestyle.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(cloud): firewall rules must name a caller resource as destination; get/delete must find VM rules

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(cloud): decide firewall rule ownership on the shared provider account

The Freestyle account is shared by every cmux user, so the API decides
whose a rule is. Reading and deleting: the rule names at least one resource
and every resource it names is the caller's. get and delete now read the
rule by id; they searched only the network listing, so a rule that named a
VM and a CIDR was created (201) and then could not be found or deleted. The
list merges the network listing with one listing per caller VM. Creating:
the destination must be a caller resource (400 vm_invalid_firewall_rule),
because a destination of only an address range or the public Internet would
reach other tenants' machines. Every firewall call resolves the account
scope like the other VM routes.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(cloud): firewall refuses unknown endpoint fields as owned and sends canonical CIDRs

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(cloud): unknown firewall endpoint fields are never owned; send canonical CIDRs

Security review P2: the provider adds selectors as new optional fields, and
an unknown one could name another tenant's resource, so a rule with an
unknown endpoint field is not the caller's. Review P3: send the canonical
range so a valid non-canonical CIDR does not fail at the provider.

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(cloud): title vm_firewall_rule_not_found 'Firewall rule not found'

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* test(cloud): firewall needs a 100-rule cap, a bounded list, and a per-user rate limit

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* fix(cloud): cap firewall rules at 100, bound the unfiltered list, rate-limit mutations

These routes are new on a provider account shared by every cmux user.
Create refuses at 100 owned rules (409 vm_firewall_rule_limit). An
unfiltered list reads the network plus at most the 10 newest live VMs, one
provider call each; older VMs list with ?vmId. Create and delete are
throttled per user with the Vercel firewall rule CMUX_VM_FIREWALL_RATE_LIMIT_ID
(no other VM mutation route has a limiter, so this follows the team-invite
limiter: fail closed when the firewall is unavailable, fail open and report
when the rule is unset or removed).

Co-Authored-By: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 (1M context) <noreply@anthropic.com>

* remote-tmux: stop a torn-down control stream from feeding the reconnected one (#16897)

* remote-tmux: failing test for a torn-down stream feeding the next one

A control stream torn down for a reconnect keeps delivering what its reader
had already buffered. The test holds the main actor while a first client
writes 560 KB, starts a reconnect, and expects none of those bytes to reach
the connection.

* remote-tmux: stop a torn-down stream from feeding the next one

Cancelling the task that reads a control client's stdout does not empty the
reader's buffer, so chunks the old client had already written were still
ingested after the teardown. Once the reconnect had respawned, a leftover
command result was taken for the new client's attach reply. The connection
then never asked for windows and the mirror stayed blank for good.

Each read loop now stops as soon as its process generation is no longer the
current one, and closes its reader.

---------

Co-authored-by: ejc3 <ejc3@users.noreply.github.com>

* remote-tmux: keep a window whose Dock has panels when its mirrors move out (#17237)

* remote-tmux: failing test for a docked terminal closed when its window's mirrors move

* remote-tmux: keep a window whose Dock has panels when its mirrors move out

---------

Co-authored-by: ejc3 <ejc3@users.noreply.github.com>

* Expose the workspace task-status lane to custom sidebars (#17245)

Custom sidebars could not read a workspace's task-status lane. cmux already
resolves one per workspace and the control socket can pin it, but the
interpreter data context carried no field for it, so a sidebar had no way to
group or colour rows by whether a workspace needs attention.

`workspaces[i].status` now carries the resolved lane as its raw wire value:
todo, working, needs-attention, review or done. The snapshot takes it as a
required parameter so a dropped wiring breaks the build rather than reporting
a silent "todo".


Claude-Session: https://claude.ai/code/session_0113SqtxGQwjHjzFw8mkSgwU

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Clear restored agent notifications once the agent is gone (#17067)

* test: prune read notifications of agents that died with the previous app

An agent alive at quit dies without SessionEnd, so its last "Completed in"
notification is restored on every launch and shown as the workspace's
latest sidebar summary even though no agent is running. The stale-agent
sweep must drop it once the pane has no agent again.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: drop restored agent notifications once the agent does not return

Notifications persist across relaunch so an unseen agent result is not
lost, but an agent that was alive when cmux quit dies without SessionEnd.
Nothing clears its notification afterwards: agent PIDs are not restored,
so the 30s stale-PID sweep has no dead PID to catch.

Restore now records the notifications of local panes that hosted an agent
(resume binding or restorable agent snapshot). The stale-agent sweep
removes the read ones when the pane has no agent PID again; unread ones
survive until read, and a pane the agent resumes into is handed back to
its hooks. Remote terminals are skipped since their agent can outlive the
app.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: keep restored notifications while the resume is in flight

Covers a read notification posted after restore (never tracked) and a
pane whose restored resume has not reported an agent PID yet.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: defer restored notification prune while the resume is in flight

The 30-second sweep can run before an auto-resumed agent reports its PID.
Skip panes whose restored command is still in flight, using the
coordinator's existing ownsInFlightRestoredCommand contract. Also look up
tracked notifications by id instead of scanning the store per panel, and
mark the value-only snapshot helper nonisolated.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: keep read non-agent notifications on a restored agent pane

The restore tracks every notification persisted on a pane that hosted an
agent, so a read `cmux notify` banner on that pane is pruned with the
agent's result.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* fix: track only agent-produced notifications on restored agent panes

Use the persisted `isAgentEvent` provenance instead of panel ownership so
a `cmux notify` banner on a pane that hosted an agent is not retired with
the agent's result. Unknown provenance restores as agent-produced, matching
TerminalNotificationStore.

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

---------

Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* Send sidebar links through the external-open rules (#7397)

* browser: apply external-open rules to sidebar links

The sidebar's pull-request and port links (SwiftUI and AppKit rows, and
the open-all-pull-requests action) opened in the embedded browser whenever
that preference was on, without consulting the URL rules that route a site
to the system browser. Sites listed there cannot work in the embedded web
view at all, so a rule now wins over the embedded preference on those
paths, the same way it does for a click inside a page.

* browser: require a user event before a link escapes to the system browser

WebKit reports a script calling click() on an anchor as .linkActivated,
the same as a real click, so the external-open rules on their own let a
page hand itself a system-browser open at a moment of its choosing. The
navigation-typed escape now also requires an AppKit event in flight (a
key, left-mouse, or middle-mouse event) and never intercepts a download,
on every path that consults the rules: the main navigation delegate, the
target=_blank UI delegate, and both popup delegates. The context menu's
Open Link in New Tab is a gesture by construction and says so.

The event check is a bound rather than a proof: NSApp.currentEvent says
an event is being dispatched, not that this navigation is the thing the
user asked for. Middle-clicks arrive as otherMouse events and count.

* browser: e2e coverage for external-open link routing

BrowserExternalOpenRoutingUITests drives real WebKit link activations
through the socket browser.click against a local fixture and asserts
routing at the delegate layer, where popup-vs-link-activation behavior
actually diverges and unit tests cannot reach. Escapes are captured to a
file through the existing DEBUG-only UI-test sink
(CMUX_UI_TEST_CAPTURE_EXTERNAL_OPEN_PATH) from the external-navigation
handler's default opener, so CI never opens Safari. Four cases: a matched
link click escapes while the embedded page stays put; an unmatched click
navigates embedded; a scripted window.open to a matched host never
escapes; a target=_blank form POST to a matched host stays embedded.

The shared BrowserFixtureSocketTestCase gains subclass hooks for launch
arguments and environment, falls back from the in-process socket client
to nc -U and then the bundled cmux CLI, disables hidden-webview discarding
for the backgrounded UI-test host, and polls browser.wait through the
cold-start content-process transient.

* browser: click links for real in the external-open UI tests

A link now leaves for the system browser only while a real input event
is in flight, and the socket browser.click runs JavaScript, so the matched
and unmatched link cases click through accessibility the way a person
does. The scripted popup and form cases keep the socket click, since a
scripted action is what they test.

* browser: keep only the sidebar links, drop the in-page activation change

The in-page half of the external-open rules has landed separately. What remains here is the sidebar: pull-request and port links follow the rules, with tests for the matcher.

* browser: use a rule the pattern safety check accepts in the port-link test

---------

Co-authored-by: ejc3 <ejc3@users.noreply.github.com>
Co-authored-by: Leo Li <cheerleaderleo@outlook.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* ci: require merge checks only when their workflows exist (#17284)

* test: cover merging repos without aggregate CI workflow

* fix: make merge checks conditional on base workflows

* ci: fall back to ancestor evidence for --main-fix (#17277)

* ci: use nearest ancestor for main-fix evidence

* ci: constrain ancestor evidence to path-filtered changes

* ci: inspect renamed paths in ancestor evidence

* ci: bound ancestor evidence traversal

* ci: reject incomplete ancestor path comparisons

* fix(session): sweep stale scrollback replay files (#16056)

* test(session): cover replay sweep and permissions

Signed-off-by: Alejandro Florez <soyeladice@gmail.com>

* fix(session): sweep stale scrollback replay files

Signed-off-by: Alejandro Florez <soyeladice@gmail.com>

* fix(session): sweep stale replay files before restore

Signed-off-by: Alejandro Florez <soyeladice@gmail.com>

* fix(session): remove synchronous replay sweep from app init

Signed-off-by: Alejandro Florez <soyeladice@gmail.com>

* fix(session): gate restore on off-main replay cleanup

Signed-off-by: Alejandro Florez <soyeladice@gmail.com>

* fix(app): keep replay sweep off startup critical path

Signed-off-by: Alejandro Florez <soyeladice@gmail.com>

* fix(session): skip replay sweep under XCTest

* fix(session): harden retained replay files during stale sweep

* docs(session): keep crash-recovery gate semantics accurate

* test(session): cover legacy replay permissions and sweep filters

---------

Signed-off-by: Alejandro Florez <soyeladice@gmail.com>
Co-authored-by: Leo Li <cheerleaderleo@outlook.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* Cover dotted Claude project dir in session directory search (#4939)

* docs: clarify Claude project dir decode asymmetry

* test: cover dotted Claude project dir in session directory scope

* test: scope Claude session roots per task instead of process env

Swift Testing runs suites in parallel, so setting CLAUDE_CONFIG_DIR process-wide could leak into other tests. A DEBUG-only TaskLocal override keeps the fixture root local to the test's task tree.

* test: cover Claude cwd-filter lookup without a DEBUG seam

Widen the Claude candidate enumerator and its two types to internal so the test reaches them through @testable import, per the no-test-debug-seam review rule. The test checks .claude/worktrees and .worktrees cwds and that dot-preserving and other project dirs are excluded.

---------

Co-authored-by: Leo Li <cheerleaderleo@outlook.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>

* Answer the tmux session commands Claude Code calls (#13632)

* Answer the tmux session commands Claude Code calls

Claude Code's tmux backend tears down and reattaches its agent panes with
kill-session, switch-client, new-session -A and show-options -g prefix. The
compatibility layer rejected all four, so a Teams session failed with
"Unsupported tmux compatibility command" once it got past creating panes.

A tmux session is a cmux workspace, which has-session and new-session already
assume, so kill-session closes that workspace, switch-client selects it, and
new-session -A attaches to it when it exists instead of creating a duplicate.
show-options now answers from a table, and prefix reports the C-b that a
default tmux client would.

The sequence test drives all four through the real shim. Its fake socket also
now unwraps the capability envelope that the shell integration adds inside a
cmux terminal, so the test reports the behavior it checks rather than a JSON
decode error; that unwrapping moved into the shared helper.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* Reject a socket payload that is not a request

The fake servers indexed request["method"] straight off json.loads, so a
payload that decoded to null, a list, or an object with a non-string method
raised inside the handler thread and surfaced as an unrelated CLI error. They
now answer those with an error line, which names the real problem.

Co-Authored-By: Claude Opus 5 (1M context) <noreply@anthropic.com>

* test: reject empty fake socket request methods

Co-Authored-By: Claude Opus 5.5 <noreply@anthropic.com>

* test: reject a failed tmux session lookup and kill-session -a

A workspace.list error must not become a new session, and kill-session -a
must not close the caller.

Co-authored-by: Cursor <cursoragent@cursor.com>

* Keep a failed tmux session lookup from creating a workspace

new-session -A treated every resolution error as a missing session, and
kill-session ignored -a and closed the caller. A missing session still
creates; a lookup failure and an unsupported flag now fail first.

Co-authored-by: Cursor <cursoragent@cursor.com>

---------

Co-authored-by: Claude Opus 5 (1M context) <noreply@anthropic.com>
Co-authored-by: Leo Li <cheerleaderleo@outlook.com>
Co-authored-by: Cursor <cursoragent@cursor.com>

* Fix custom-sidebar nil-comparison so optional-guarded views render (#7943) (#7974)

* Add failing test: sidebar nil-comparison yields nothing (#7943)

In a custom sidebar, `x != nil` / `x == nil` against a bound optional field
does not evaluate to true/false — it evaluates to nothing. Interpolation
renders empty, ternaries always take the else branch, and `if x != nil`
guards are never taken, so optional-guarded views never render.

This commit adds only the regression test (no fix) so CI shows it red.

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* Fix sidebar nil-comparison to evaluate to a Bool (#7943)

The interpreter's value model had no null case and no `nil`-literal
evaluator branch, so `nil` evaluated to a host `SwiftValue?` of `nil` — the
same value that means "expression unsupported / no value". `evalInfix` then
bailed on the comparison, so `x != nil` / `x == nil` produced nothing:
interpolation rendered empty, ternaries always took the else branch, and
`if x != nil` guards were never taken. Optional-guarded views (including the
shipped status-board.swift / finder.swift examples) silently drew nothing.

- Add `SwiftValue.null` for the `nil` literal and for comparing an absent
  optional field against `nil` (distinct from host `nil` = "no value").
- Evaluate `NilLiteralExprSyntax` to `.null`.
- Handle `==` / `!=` before the operand guards, coalescing an absent operand
  (host `nil`) and the `nil` literal to `.null`, so the comparison yields a
  Bool that is true/false when present and false/true when absent.

Fixes #7943

Co-Authored-By: Claude Opus 4.8 <noreply@anthropic.com>

* fix: preserve nil comparison evaluation failures

* fix: preserve nested nil comparison misses

* fix: preserve parenthesized nil comparison misses

* fix: handle nil optional binding and equality budget

---------

Co-authored-by: Claude Opus 4.8 <noreply@anthropic.com>
Co-authored-by: Leo Li <cheerleaderleo@outlook.com>

* fix: refresh merge helper and honor neutral checks (#17294)

* test: cover neutral checks and helper checkout refresh

* test: tolerate absent git diagnostics

* fix: refresh clean main checkout before merging

* test: cover cloud welcome close shortcut ownership

* fix: route cloud welcome close shortcut to its window

---------

Signed-off-by: Alejandro Florez <soyeladice@gmail.com>
Co-authored-by: Austin Wang <austinwang115@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Leo <cheerleaderleo@outlook.com>
Co-authored-by: Lawrence Chen <54008264+lawrencecchen@users.noreply.github.com>
Co-authored-by: BlueRaddish <jeeholife2@gmail.com>
Co-authored-by: EJ <ej@campbell.name>
Co-authored-by: ejc3 <ejc3@users.noreply.github.com>
Co-authored-by: Philipp Mochine <philipp@mochine.de>
Co-authored-by: mys <wowpotato@naver.com>
Co-authored-by: Alejandro Florez <soyeladice@gmail.com>
Co-authored-by: Sungho Park <relilau00@gmail.com>
Co-authored-by: Darío Kondratiuk <dariokondratiuk@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Mark Xian <mark-xian@foxmail.com>
Co-authored-by: Austin Wang <38676809+austinywang@users.noreply.github.com>

---------

Signed-off-by: Alejandro Florez <soyeladice@gmail.com>
Co-authored-by: Austin Wang <38676809+austinywang@users.noreply.github.com>
Co-authored-by: Austin Wang <austinwang115@gmail.com>
Co-authored-by: Claude Opus 5.5 <noreply@anthropic.com>
Co-authored-by: Leo <cheerleaderleo@outlook.com>
Co-authored-by: Lawrence Chen <54008264+lawrencecchen@users.noreply.github.com>
Co-authored-by: BlueRaddish <jeeholife2@gmail.com>
Co-authored-by: EJ <ej@campbell.name>
Co-authored-by: ejc3 <ejc3@users.noreply.github.com>
Co-authored-by: Philipp Mochine <philipp@mochine.de>
Co-authored-by: mys <wowpotato@naver.com>
Co-authored-by: Alejandro Florez <soyeladice@gmail.com>
Co-authored-by: Sungho Park <relilau00@gmail.com>
Co-authored-by: Darío Kondratiuk <dariokondratiuk@gmail.com>
Co-authored-by: Cursor <cursoragent@cursor.com>
Co-authored-by: Mark Xian <mark-xian@foxmail.com>
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.